Spoken Words and Their Timings

This is a guest blog post from Mark Boas, Co-founder of Hyperaudio Inc. Mark will be speaking at the upcoming BBC Sound: Now and Next event on the future of broadcast audio.
The history of communication is fascinating, it was revolutionised 500,000 years ago with speech. Cave paintings appeared around 40,000 years ago, symbols 30,000 years ago and writing 5,000 years ago.Communication technology has evolved from prehistoric fires, beacons, smoke signals and horns, through post or mail in the 6th century, the Printing Press in 1440, Telegraph, Telephone and Radio in the eighteen-hundreds and Television, Internet, Email and the World Wide Web in the last century.
Fast forward to the present day and we're still very much in the midst of a communication revolution.Over the last five years I have become interested in the idea of coupling two forms of communication - the spoken word and its textual representation. A great example of this is subtitles - we see the words on the screen at roughly the same time as they are spoken. Another example is Karaoke.
It's interesting to note that one of the first uses of subtitles was to caption silent movies where words were uttered but never heard. In 1938 the BBC was one of the first to broadcast a 'talkie' with subtitles.These days subtitles are rarely burned onto the actual footage, instead they're have become decoupled from video, thanks in part to free software like SubRip which extracted subtitles and their timings from videos. The resulting .srt file was a simple text file that could be easily edited and thus translated.
Services like Amara.org (formerly Universal Subtitles) further facilitated this effort. Interestingly the format for subtitles supported within browsers, WebVTT is based on the .srt format. Nowadays we are finding that many countries are legally obliged to provide subtitles and captions to make them accessible to the deaf and hard of hearing.
But what about other written representations of the spoken word? Well we've had transcripts for a while now and in today's web-connected world they have been an important way of archiving audio and video for later retrieval. Unlike the media they represent, transcripts can be easily picked up by search engines.But we can go further and you can already find interactive transcripts in sites or services like YouTube.
My efforts have centred around tightly coupling transcripts with audio. Tight coupling in this case means storing timings for each word and using these to highlight words exactly as they are spoken. A bit like karaoke's famous bouncing ball. Conversely you can click on the words to jump around the audio or video. It makes navigating a video very intuitive. The advantage of this particular style of media representation means that you can get a feeling for the pace of spoken material without actually having to hear it, the pregnant pauses, the bursts of excited and flowing speech - they're all laid bare for anyone to experience.
Coupled with the audio, seeing words light up as they are spoken can be a great reinforcement for people learning languages. Additionally the ability to click on words to hear them repeated can be a tremendous help to the budding multi-linguist.Word timings also allow precise sharing of audio and text on platforms such as social media. In addition to letting you share the text of a piece it provides a link to the exact part of the corresponding audio or video. But wait, there's more!
If we represent audiovisual content as a word-timed transcript we can use that transcript as a basis for editing - you could edit media by dragging and dropping text as if you were using a text-editor. That to me was a very exciting thought and this is what the Hyperaudio project is all about. Hyperaud.io was conceived as an opensource non-profit initiative supported by the Knight Foundation and Mozilla. Hyperaud.io aims to provide the workflow for people to transcribe and remix media, one of its core tools is the Hyperaudio Pad which allows drag and drop editing and embedding.
It's fantastic to see other people using the open source code behind Hyperaudio to create their own interpretations. BBC's Chris Baume has enhanced the Hyperaudio Pad to create something new in the form of his speech editor prototype. Building media-literacy amongst young people is something I feel strongly about and I had the opportunity to work with the Mozilla Foundation and Chattanooga Public Library to try the Hyperaudio Pad out on some of the kids there.
We created a version of Hyperaud.io called chattanooga.hyperaud.io that contains a subset of content from Chattanooga.It’s not too difficult to imagine giving students transcribed resources from which they can create their own video-essays on certain subjects, they could even intersperse them with their own transcribed talking heads.
But how do we transcribe things with word timings in the first place? Well there's a number of ways: If you have the subtitle file we can guess them from those loose timings - which can be good enough for some. People requiring more accuracy can use professional transcription services, happily most of these will offer the word timings.
However services that employ people can be expensive and take time. Automatic transcription can be cheaper, quicker and surprisingly accurate, however you nearly always need to edit the results. Lately I've been concentrating my efforts on a transcript editing tool that aims to provide a low-cost way for people to edit and create word timed interactive transcripts. It feels like an exciting new age of communication is upon us!