System and method for remote audio caption visualizations
Claim Score by NHIP
Abstract
A system and method for remote audio caption visualizations is presented. A user uses a personal device during an event to display an enhanced captioning stream corresponding to the event. A media-playing device provides a media stream corresponding to the enhanced captioning stream. The media-playing device provides a synchronization signal to the personal device which instructs the personal device to start playing the enhanced captioning stream on the personal device's display. The user views text on the personal display while the media stream plays. The user is able to adjust the timing of the enhanced captioning stream in order to fine-tune the synchronization between the enhanced captioning stream and the media stream.

Term
Term ended
Projected expiry passed 3 September 2022, 4.1 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
30 claims: 9 independent, 21 dependent
- 1Broadest claimClaim Score 85, broad(NHIP)A method for providing a user with an audio caption, said method comprising:receiving a media stream from a first source;receiving an enhanced captioning stream from a second source;and displaying the enhanced captioning stream that corresponds to the media stream on an enhanced captioning device.
- 12An information handling system comprising:one or more processors;a memory accessible by the processors;one or more nonvolatile storage devices accessible by the processors;a display accessible by the processors;and an audio captioning tool for processing audio captions, the audio captioning tool including: receiving logic for receiving a media stream from a first source;receiving logic for receiving an enhanced captioning stream from a second source;and display logic for displaying the enhanced captioning stream that corresponds to the media stream on an enhanced captioning device.
- 17A computer program product stored in a computer operable media for providing audio captions, said computer program product comprising:means for receiving a media stream from a first source;means for receiving an enhanced captioning stream from a second source;and means for displaying the enhanced captioning stream that corresponds to the media stream on an enhanced captioning device.
- 25A method for providing a user with an audio caption, said method comprising:receiving a media stream from a first source;receiving an enhanced captioning stream from a second source;synchronizing the enhanced captioning stream with the media stream based on the comparing;displaying the enhanced captioning stream that corresponds to the media stream on an enhanced captioning device;and displaying the enhanced captioning stream on an enhanced captioning device in response to the synchronization.
- 26A method for providing a user with an audio caption, said method comprising:receiving a media stream;determining whether the media stream matches one or more words included in an enhanced captioning stream wherein the enhanced captioning stream is external to the media stream;changing an adjustment time in response to the determination;storing the adjustment time in a non-volatile storage area;and displaying the enhanced captioning stream on an enhanced captioning device using the adjustment time.
- 27An information handling system comprising:one or more processors;a memory accessible by the processors;one or more nonvolatile storage devices accessible by the processors;a display accessible by the processors;and an audio captioning tool for processing audio captions, the audio captioning tool including: receiving logic for receiving a media stream;detection logic for detecting an audible signal, the audible signal corresponding to the media stream;comparison logic for comparing the audible signal with an enhanced captioning stream wherein the enhanced captioning stream is external to the media stream;synchronization logic for synchronizing the enhanced captioning stream with the media stream based on the comparing;and display logic for displaying the enhanced captioning stream on an enhanced captioning device in response to the synchronization.
- 28An information handling system comprising:one or more processors;a memory accessible by the processors;one or more nonvolatile storage devices accessible by the processors;a display accessible by the processors;and an audio captioning tool for processing audio captions, the audio captioning tool including: receiving logic for receiving a media stream;determination logic for determining whether the media stream matches one or more words included in an enhanced captioning stream wherein the enhanced captioning stream is external to the media stream;alteration logic for changing an adjustment time in response to the determination;storage logic for storing the adjustment time in a non-volatile storage area;and display logic for displaying the enhanced captioning stream on an enhanced captioning device using the adjustment time.
- 29A computer program product stored in a computer operable media for providing audio captions, said computer program product comprising:means for receiving a media stream from a first source;means for receiving an enhanced captioning stream from a second source;and means for detecting an audible signal, the audible signal corresponding to the media stream;means for comparing the audible signal with an enhanced captioning stream wherein the enhanced captioning stream is external to the media stream;means for synchronizing the enhanced captioning stream with the media stream based on the comparing;and means for displaying the enhanced captioning stream on an enhanced captioning device in response to the synchronization.
- 30A computer program product stored in a computer operable media for providing audio captions, said computer program product comprising:means for receiving a media stream from a first source;means for determining whether the media stream matches one or more words included in an enhanced captioning stream wherein the enhanced captioning stream is from a second source;means for changing an adjustment time in response to the determination;means for storing the adjustment time in a non-volatile storage area;and means for displaying the enhanced captioning stream on an enhanced captioning device using the adjustment time.
Independent claims9
72 paragraphs in 4 sections, as filed
BACKGROUND OF THE INVENTION
[0001] 1. Technical Field
[0002] The present invention relates in general to a system and method for remote audio caption visualization. More particularly, the present invention relates to a system and method for playing an enhanced captioning stream on a personal device while playing a media stream on a media-playing device wherein the enhanced captioning stream is external to the media stream.
[0003] 2. Description of the Related Art
[0004] Many individuals use captioning to comprehend audio or video media. Hearing impaired individuals depend on captioning for everyday activities, such as watching television shows. Individuals with normal hearing may also use captioning to comprehend a television program in public areas with high noise levels, such as in an exercise facility. Two types of captioning transmission methods are open captioning and closed captioning. Open captioning places text on a screen at all times, often in a black reader box. Closed captioning does not automatically place the text on a screen but rather uses a decoder unit to decode the captioning and place the text on the screen at a user's discretion.
[0005] A content provider (i.e. television network) may use online captioning or offline captioning to generate a captioning text stream. Online captioning is generated as an event occurs. For example, television news shows, live seminars, and sports events may use online captioning. Online captions may be generated from a script (live display), or generated in real-time. Someone listening to an event with the script loaded on a computer system generates live display captioning. The person presses a “next caption” button to show a viewer the next line of captioning. Alternatively, the script may come from a prompter in which the viewer sees the same text that the speaker is seeing. Live display typically scrolls text up one line at a time on a television screen.
[0006] A challenge of live-display is that the content provider only captions what is scripted, and if the speaker deviates from the script, the captions are incorrect. For example, a newscast using live-display may have clean, high-quality captions as an anchorperson reads the stories off of a prompter. As soon as the newscast performs a live interview, the captions stop. Typically, content providers that use prompter-based captions leave a third to a half of each newscast uncaptioned.
[0007] On the other hand, real-time captioning uses stenocaptioners to caption an entire broadcast. Stenocaptioners listen to a live broadcast and type what they hear on a shorthand keyboard. Special computer software translates the stenocaptioner's phonetic shorthand into English. A closed-caption encoder receives the phonetic shorthand and places it on the broadcast signal for a viewer to see. Stenocaptioning costs more than live-display captioning, but it allows the entire broadcast to be captioned. However, stenocaptioning is more prone to errors than live-display captioning.
[0008] Many newscasts use a combination of captioning techniques to try to achieve both the accuracy of live-display captioning and the complete coverage of real-time stenocaptioning. To accomplish this, the stenocaptioner dials in to the newsroom computer system about an hour before the broadcast, and copies all of the scripts into the captioning system. The captioner then sorts and cleans up the scripts, names the segments, and marks which ones will require live stenocaptioning.
[0009] During the broadcast, the stenocaptioner may move many times between sending script lines and writing real-time. A casual viewer may notice a difference in that real-time captions appear one word at a time wherein live display captions appear one line at a time.
[0010] Alternatively, offline captioning is performed “after the fact” in a studio. Examples of offline captioning include television game shows, videotapes of movies, and corporate videotapes (e.g., training videos). The text of the captions is created on a computer, and synchronized to the video using time codes. The captions are then transferred to the videotape before it is broadcast or distributed.
[0011] A challenge found with captioning is that limited funds and resources limit the amount of audio and video media that a content provider captions. Typically, mainstream television shows are captioned, while other less popular shows are not. While the FCC is requiring captioning for audio and video media, exceptions do apply. For example, a video programmer is not required to spend more than 2% of its annual gross revenues on captioning. Additionally, programs aired between 2:00 am and 6:00 am are not required to be captioned. Furthermore, programming from “new networks” is not required to be captioned.
[0012] Another challenge found with captioning is that movies in a movie theater rarely have captioning capability. In many cases, a hearing impaired person waits for a movie to be available in video rental stores before the person is able to view the movie with captions.
[0013] Finally, caption information in the media stream is lost during conversion to web stream formats. A challenge found is that a hearing impaired person may not be able to understand a web cast event without first downloading a corresponding transcript and follow the transcript as the speaker talks. This process may become cumbersome to the user.
[0014] What is needed, therefore, is a way for a person to view an enhanced captioning stream on an individual basis for situations when captioning is not available for a particular event.
SUMMARY
[0015] It has been discovered that the aforementioned challenges are resolved by using a personal device to display an enhanced captioning stream that is synchronized with a corresponding media stream. The personal device uses synchronization signals from a media-playing device to synchronize the enhanced captioning stream with the media stream. A user may use the personal device to understand events that do not have captioning.
[0016] The user attends an event that includes the media-playing device that plays the media stream. For example, the user may wish to see a movie that is played on a movie projector. The user instructs the personal device to download the enhanced captioning streams. The personal device downloads the enhanced captioning stream using a variety of methods. Using the example described above, the user may download a script corresponding to the movie using a wireless connection when the user enters the movie theater. In one embodiment, the enhanced captioning stream may include graphic information to support the text.
[0017] After the personal device downloads the enhanced captioning stream, the personal device waits for the media-playing device to provide the synchronization signal. The personal device uses the synchronization signal to synchronize the enhanced captioning stream with the media stream. Using the example described above, the personal device uses the synchronization signal to align the script with the movie displayed on a screen. The synchronization signal may be an audible signal, a wireless signal, or a manual signal. An audible signal is a signal, such as a speech pattern, that the personal device detects, and matches the detected audible signal with the enhanced captioning stream. When the personal device finds a match, the personal device displays the corresponding enhanced captioning stream relative to the location point of the detected signal. A wireless signal may be an RF signal, such as Bluetooth, in which the media-playing device transmits. The wireless signal informs processing when to display the enhanced captioning stream. A manual signal may be a queue to the user as to when to push a “start” button on the personal device. For example, the user may attend a movie and push the “start” button when a particular movie scene is displayed on the movie screen.
[0018] The media playing device may provide one or more resynchronization signals throughout the duration of the media stream. For example, the user may enter the movie theater after the movie has started and miss the first audible signal. In this example, the personal device “listens” to the movie's audio and compares it with the enhanced captioning stream. When the personal device detects a match, the personal device displays the corresponding enhanced caption text relative to the movie scene. The user is also able to adjust the timing of the enhanced captioning stream on the personal device. For example, the user may change an adjustment time by selecting soft keys on a PDA to increase the speed of the enhanced captioning stream.
[0019] The foregoing is a summary and thus contains, by necessity, simplifications, generalizations, and omissions of detail; consequently, those skilled in the art will appreciate that the summary is illustrative only and is not intended to be in any way limiting. Other aspects, inventive features, and advantages of the present invention, as defined solely by the claims, will become apparent in the non-limiting detailed description set forth below.
BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present invention may be better understood, and its numerous objects, features, and advantages made apparent to those skilled in the art by referencing the accompanying drawings. The use of the same reference symbols in different drawings indicates similar or identical items.
[0021]FIG. 1 is a diagram showing a personal device synchronizing with a media-playing device to display an enhanced captioning stream corresponding to a media stream;
[0022]FIG. 2 is a flowchart showing steps taken in a personal device displaying an enhanced captioning stream and adjusting the enhanced captioning stream to correlate with a media stream;
[0023]FIG. 3 is a diagram showing a personal device receiving a synchronization signal from a media playing device and displaying an enhanced captioning stream;
[0024]FIG. 4 is a detail flowchart showing steps taken in playing an enhanced captioning stream on a personal device;
[0025]FIG. 5 is a flowchart showing steps taken in generating an enhanced captioning stream corresponding to an audio stream;
[0026]FIG. 6A is a user interface window on an enhanced captioning device showing an enhanced captioning stream corresponding to a conversation;
[0027]FIG. 6B is a user interface window on an enhanced captioning device showing an enhanced captioning stream corresponding to a musical event; and
[0028]FIG. 7 is a block diagram of an information handling system capable of implementing the present invention.
DETAILED DESCRIPTION
[0029] The following is intended to provide a detailed description of an example of the invention and should not be taken to be limiting of the invention itself. Rather, any number of variations may fall within the scope of the invention which is defined in the claims following the description.
[0030]FIG. 1 is a diagram showing a personal device synchronizing with a media playing device to display an enhanced captioning stream corresponding to a media stream. User <b>175</b> may be a hearing impaired individual that uses personal device <b>100</b> to display captioned text corresponding to an event. For example, user <b>175</b> may wish to view a movie at a movie theater in which the movie does not have captioning. In another example, users may video overlay caption visualizations on their television during television shows that do not provide captioning. Personal device <b>100</b> is an electronic device that includes a display, such as a personal digital assistant (PDA), a mobile telephone, or a computer.
[0031] User <b>175</b> attends an event that includes media playing device <b>120</b>. Media playing device <b>120</b> retrieves media stream <b>140</b> from media content store <b>130</b>. Using the-example described above, media stream <b>140</b> may be a movie stored on digital media or the movie may be stored on a film reel. Media content store <b>130</b> may be stored on a non-volatile storage area, such as non-volatile memory. Media content store <b>130</b> may also be a storage area to store film reels.
[0032] Personal device <b>100</b> includes captioned text area <b>110</b> where processing displays enhanced caption text. In one embodiment, personal device <b>100</b> may include queue area <b>105</b> where processing displays manual synchronization queues, such as movie scenes.
[0033] Personal device <b>100</b> downloads enhanced captioning stream <b>160</b> from enhanced captioning stream store <b>150</b>. Enhanced captioning stream <b>160</b> includes text and related timing information corresponding to media stream <b>140</b>. In one embodiment, enhanced captioning stream <b>160</b> may include graphic information to support the text in which the graphic information may be compiled into a binary file and stored in non-volatile memory. For example, graphical information may include “bouncing ball” emoticons to display word delivery and attitude or “bouncing ball musical bar charts” may be displayed to support media streams that include music. In another embodiment, enhanced captioning stream <b>160</b> may include text in a different language than media stream <b>160</b>. For example, enhanced captioning stream <b>160</b> may include text in English whereas media stream <b>140</b> may be a German movie. In yet another embodiment, personal device <b>100</b> may project perspective corrected visuals from enhanced captioning stream <b>160</b> onto a beamsplitter glass as to not disturb nearby patrons. In this example, the enhanced captioning stream is visible only to a person that is directly in front of the glass, such as with what a speaker may use on a podium while reading his speech.
[0034] In yet another embodiment, enhanced captioning stream <b>160</b> may include audio descriptions that are non-spoken words that describe what is occurring, such as an emotion of an actor (i.e. angry, sad, etc.). Personal device <b>100</b> may download enhanced captioning stream <b>160</b> using a variety of methods, such as using a global computer network (i.e. the Internet) or by using a wireless network. Using the example described above, user <b>175</b> may download a script corresponding to the movie using a wireless connection when user <b>175</b> enters the movie theater.
[0035] After personal device <b>100</b> downloads enhanced captioning stream <b>160</b>, personal device <b>100</b> waits for media playing device <b>120</b> to provide synchronization signal <b>165</b>. Personal device <b>100</b> uses synchronization signal <b>165</b> to synchronize the enhanced captioning stream with the media stream. Using the example described above, personal device <b>100</b> uses the synchronization signal to align the script with the movie displayed on a screen. The synchronization signal may be an audible signal, a wireless signal, or a manual signal. An audible signal is a signal, such as a speech pattern, that personal device <b>100</b> detects, and matches the detected audible signal with enhanced captioning stream <b>160</b>. When processing finds a match, processing displays the enhanced captioning stream on caption text area <b>110</b> at a point corresponding to the location point of the detected signal. A wireless signal may be an RF signal, such as Bluetooth, in which media playing device <b>120</b> transmits. The wireless signal informs processing when to display the enhanced captioning stream. A manual signal may be a queue to the user as to when to push a “start” button on the personal device. For example, a user may attend a movie and push the “start” button when a particular movie scene is displayed on the movie screen.
[0036] Media playing device <b>120</b> may provide one or more re-synchronization signals throughout the duration of playing media stream <b>140</b>. For example, user <b>175</b> may enter the movie theater after the movie has started and miss the first audible signal. In this example, personal device <b>100</b> “listens” to the movie's audio and compares it with enhanced captioning stream <b>160</b>. When personal device <b>100</b> detects a match, personal device <b>100</b> displays enhanced caption text on caption text area <b>110</b> corresponding to the movie scene being played. In addition to personal device synchronizing on re-synchronization signals, personal device <b>100</b> may frequently re-synchronize using the media stream (i.e. audio) and matching the media stream with enhanced captioning stream <b>160</b>.
[0037] User <b>175</b> is also able to adjust the timing of the enhanced captioning stream by sending timing adjust <b>180</b> to personal device <b>100</b>. For example, user <b>175</b> may change an adjustment time by selecting soft keys on a PDA to increase the speed of enhanced captioning stream <b>160</b> (see FIGS. 2 through 4 for further details regarding timing adjustment).
[0038]FIG. 2 is a flowchart showing steps taken in a personal device displaying an enhanced captioning stream and adjusting the enhanced captioning stream to correlate with a media stream. Media processing commences at <b>200</b>, whereupon processing downloads a media stream file from media store <b>208</b>. For example, the media stream may be a movie. Media store <b>208</b> may be stored on a non-volatile storage area, such as non-volatile memory. Media store <b>208</b> may also be a storage area to store movie film reels.
[0039] Processing provides synchronization signal <b>212</b> to the personal device which notifies the personal device to start playing the enhanced captioning stream (step <b>210</b>). The synchronization signal may be an audible signal, a wireless signal, or a manual signal. An audible signal may be a speech pattern from the media stream. For example, the media-playing device may be a movie projector and the audible signal may be an actor's speech. A wireless signal may be an RF signal, such as Bluetooth, in which the media-playing device transmits to instruct the personal device to start playing the enhanced captioning stream. A manual signal may be a queue to the user as to when to push a “start” button on the personal device.
[0040] Processing plays the media stream at step <b>215</b>. A determination is made as to whether to provide a re-synchronization signal to the personal device (decision <b>220</b>). Using the example described above, the actor's speech may be a continuous re-synchronization signal to the personal device. Another example is a wireless signal may be sent every five minutes to inform the personal device as to what point in the movie the movie is being shown. If a re-synchronization signal should be sent to the personal device, decision <b>220</b> branches to “Yes” branch <b>222</b> which sends synchronization signal <b>223</b> to the personal device. On the other hand, if a re-synchronization signal should not be sent, decision <b>220</b> branches to “No” branch <b>224</b> bypassing re-synchronization steps.
[0041] A determination is made as to whether the media stream is finished (decision <b>225</b>). If the media stream is not finished, decision <b>225</b> branches to “No” branch <b>227</b> which loops back to continue processing the media stream. This looping continues until the media stream is finished, at which point decision <b>225</b> branches to “Yes” branch <b>229</b> whereupon processing stops the media stream (step <b>230</b>). Media processing commences at <b>235</b>.
[0042] Personal device processing commences at <b>240</b>, whereupon processing downloads an enhanced captioning stream from enhanced captioning stream store <b>248</b> (step <b>245</b>). The enhanced captioning stream includes text information and timing information corresponding to a media stream. In one embodiment, the enhanced captioning stream may include graphic enhancement information corresponding to the timing information. The graphic enhancement information may be compiled into a binary file and stored in a non-volatile storage area, such as non-volatile memory. For example, processing may display a bouncing ball emoticon over each word at the word's corresponding timestamp. Processing may download the enhanced captioning stream using a global computer network, such as the Internet. In another embodiment, processing may download the enhanced captioning stream using a wireless network, such as Bluetooth. Using the example described above, a hearing impaired user may wish to view a movie at a movie theater in which the particular movie does not have captioned text. In this example, the user enters the movie theater and downloads enhanced captioning stream using a wireless network.
[0043] A determination is made as to whether the personal device receives synchronization signal <b>212</b> which informs the personal device to start playing the enhanced captioning stream (decision <b>250</b>). The synchronization signal may be an audible signal, a wireless signal, or a manual signal. An audible signal is a signal, such as a speech pattern, that processing detects, and matches the detected audible signal with the enhanced captioning stream. When processing finds a match, processing displays the enhanced captioning stream at a point corresponding to the location point of the detected signal. A wireless signal may be an RF signal, such as Bluetooth, that the media-playing device transmits. The wireless signal informs processing when to display the enhanced captioning stream. A manual signal may be a queue to the user as to when to push a “start” button on the personal device. For example, a user may attend a movie and pushing the “start” button when a particular movie scene is displayed on the movie screen. If the personal device has not received a synchronization signal, decision <b>250</b> branches to “No” branch <b>252</b> which loops back to wait for the synchronization signal. This looping continues until the personal device receives synchronization signal <b>212</b>, at which point decision <b>250</b> branches to “Yes” branch <b>258</b>.
[0044] Processing starts playing the enhanced captioning stream at step <b>260</b>. Processing uses the timing information included in the enhanced captioning stream to display words (i.e. script) on the personal device's screen in correlation with the corresponding media stream (i.e. movie) that the user is viewing. A determination is made as to whether the user wishes to adjust the timing of the displayed captioning (decision <b>265</b>). Using the example described above, the user may wish to have the words displayed slightly before or after they are actually spoken. Users may also wish to display several sentences of dialogue in the past and/or future relative to when the words are spoken as a default. Another example is that the user may wish to increase the enhanced captioning stream display rate for a short time in order to “catch-up” the enhanced captioning stream to the media stream. If the user wishes to adjust the timing, decision <b>265</b> branches to “Yes” branch <b>266</b> whereupon the user changes an adjustment time (step <b>268</b>). On the other hand, if the user does not wish to adjust the enhanced captioning stream timing, decision <b>265</b> branches to “No” branch <b>269</b> bypassing timing adjustment steps.
[0045] A determination is made as to whether processing wishes to re-synchronize (decision <b>270</b>). Using the example described above, the user's enhanced captioning stream may be a few minutes behind the media stream and the user may wish to re-synchronize at the next scene in the movie. If the user does not wish to re-synchronize, decision <b>270</b> branches to “No” branch <b>272</b> bypassing re-synchronization steps. On the other hand, if processing wishes to resynchronize, decision <b>270</b> branches to “Yes” branch <b>274</b>.
[0046] A determination is made as to whether processing has received synchronization signal <b>223</b> (decision <b>275</b>). If processing has not received synchronization signal <b>223</b>, decision <b>275</b> branches to “No” branch <b>277</b> to wait for synchronization signal <b>223</b>. This looping continues until processing receives synchronization signal <b>223</b>, at which point decision <b>275</b> branches to “Yes” branch <b>279</b> whereupon processing re-synchronizes the enhanced captioning stream (step <b>280</b>).
[0047] A determination is made as to whether the enhanced captioning stream is finished (decision <b>285</b>). If the enhanced captioning stream is not finished, decision <b>285</b> branches to “No” branch <b>287</b> to continue processing the enhanced captioning stream. This looping continues, until the enhanced captioning stream is finished, at which point decision <b>285</b> braches to “Yes” branch <b>289</b>. Personal device processing ends at <b>290</b>.
[0048]FIG. 3 is a diagram showing a personal device receiving a synchronization signal from a media playing device and displaying an enhanced captioning stream. Personal device <b>300</b> is an electronic device with a display, such as a personal digital assistant (PDA), a mobile telephone, or a computer.
[0049] Personal device <b>300</b> includes caption generator <b>340</b> which retrieves an enhanced captioning stream and displays the enhanced captioning stream on display <b>360</b>. Caption generator <b>340</b> retrieves enhanced captioning stream <b>320</b> from enhanced captioning stream store <b>330</b>. Enhanced captioning stream <b>320</b> includes text <b>322</b> and timing <b>328</b> which correspond to a media stream. For example, text <b>322</b> may include a movie script and timing <b>328</b> includes corresponding time-stamp information that correlates the movie script to movie scenes. Enhanced captioning stream store <b>330</b> may be stored on a non-volatile storage area, such as non-volatile memory. In one embodiment, personal device <b>300</b> may download enhanced captioning stream <b>320</b> from an external source using a global computer network or wireless network and store enhanced captioning stream <b>320</b> in its local memory (i.e. enhanced captioning stream store <b>330</b>).
[0050] Personal device <b>300</b> uses audible monitor <b>380</b> to detect a synchronization signal (i.e. speech pattern) from media playing device <b>310</b>. Audible monitor <b>380</b> may be a “voice engine” that is capable of detecting audio, such as speech. Audible monitor <b>380</b> matches speech patterns transmitted from media playing device <b>310</b> with locations in the enhanced captioning stream. When audible monitor <b>380</b> identifies a match, audible monitor <b>380</b> informs timer <b>350</b> at what point to display the enhanced captioning stream based upon the match location. For example, media playing device <b>310</b> may be playing a movie and audible monitor <b>380</b> is listening to the actor speaking. In this example, audible monitor searches the enhanced captioning stream for a speech pattern similar to the actor's speech.
[0051] As audible monitor <b>380</b> detects speech patterns and instructs timer <b>350</b>, timer <b>350</b> may send adjusted timing <b>390</b> to enhanced captioning stream store <b>330</b>. Adjusted timing <b>390</b> includes new timing information to replace timing <b>328</b> the next time enhanced captioning stream is played.
[0052]FIG. 4 is a detail flowchart showing steps taken in playing an enhanced captioning stream on a personal device. The personal device is an electronic device with a display such as a computer, a personal digital assistant (PDA), or a mobile phone. Enhanced captioning stream processing commences at <b>400</b>, whereupon processing retrieves the enhanced captioning stream from enhanced captioning stream store <b>415</b> (step <b>410</b>). The enhanced captioning stream includes text and timing information that correlates the text with a corresponding media stream. In one embodiment, the enhanced captioning stream may include graphic enhancement information corresponding to the timing information. The graphic enhancement information may be compiled into a binary file and stored in a non-volatile storage area, such as non-volatile memory. For example, processing may display a bouncing ball over each word at the word's corresponding timestamp. Enhanced captioning stream store <b>415</b> may be stored on a non-volatile storage area, such as non-volatile memory.
[0053] A determination is made as to whether processing receives a synchronization signal from media playing device <b>425</b> (decision <b>420</b>). The synchronization signal may be an audible signal, a wireless signal, or a manual signal. An audible signal is a signal, such as a speech pattern, that processing detects, and matches the detected audible signal with the enhanced captioning stream. When processing finds a match, processing displays the enhanced captioning stream at a point corresponding to the location point of the detected signal. A wireless signal may be an RF signal, such as Bluetooth, in which media playing device <b>425</b> transmits. The wireless signal informs processing when to display the enhanced captioning stream. A manual signal may be a queue to the user as to when to push a “start” button on the personal device. For example, a user may attend a movie and push the “start” button when a particular movie scene is displayed on the movie screen.
[0054] The synchronization signal may be an automated signal or a manual signal. An automated signal example may be a movie theater sending an RF signal (i.e. Bluetooth) to the personal device which instructs the personal) device to start the enhanced captioning stream. A manual signal example may be the beginning of a movie and the user depresses a “start” button on the personal device to start the enhanced captioning stream. If processing has not received the synchronization signal, decision <b>420</b> branches to “No” branch <b>422</b> which loops back to wait for the synchronization signal. On the other hand, if processing received the synchronization signal, decision <b>420</b> branches to “Yes” branch <b>428</b>.
[0055] Processing starts timer <b>435</b> which uses the timing information to instruct processing as to when to display a particular word (step <b>430</b>). The first word in the enhanced captioning stream is displayed on display <b>445</b> at step <b>440</b>. In one embodiment, processing may display one sentence at a time, and then highlight the first word using a different color or place a bouncing ball over the first word.
[0056] A determination is made as to whether processing should adjust the time which correlates the enhanced captioning stream text with the media stream (decision <b>450</b>). For example, the media stream may be playing at a faster rate than the enhanced captioning stream and the user may wish to “speed-up” the enhanced captioning stream. If processing should adjust the timing, decision <b>450</b> branches to “Yes” branch <b>452</b> whereupon processing adjusts the timing at step <b>460</b>. In one embodiment, processing may frequently detect an audible signal to synchronize the enhanced captioning stream. On the other hand, if the user does not wish to adjust the timing, decision <b>450</b> branches to “No” branch <b>458</b> bypassing timing adjustment steps.
[0057] A determination is made, as to whether there are more words to display in the enhanced captioning stream (decision <b>470</b>). If there are more words in the enhanced captioning stream, decision <b>470</b> branches to “Yes” branch <b>472</b> whereupon a determination is made as to whether timer <b>435</b> has reached the next time stamp which instructs processing to display the next word (decision <b>480</b>). Time stamps are included in the timing information and correspond to when each word should be displayed. If timer <b>435</b> has not reached the next time stamp, decision <b>480</b> branches to “No” branch <b>482</b> which loops back to wait for timer <b>435</b> to reach the next time stamp. This looping continues until timer <b>435</b> reaches the next time stamp, at which point decision <b>480</b> branches to “Yes” branch <b>488</b> to display the next word.
[0058] This looping continues until there are no more words to display in the enhanced captioning stream, at which point decision <b>470</b> branches to “No” branch <b>478</b>. Processing ends at <b>490</b>.
[0059]FIG. 5 is a flowchart showing steps taken in generating an enhanced captioning stream corresponding to an audio stream. Enhanced captioning stream generation commences at <b>500</b>, whereupon processing retrieves a text file from text store <b>520</b> (step <b>510</b>). The text file includes words corresponding to an audio stream, such as lyrics to a song or a script to a movie. Processing retrieves the corresponding audio stream from audio store <b>535</b>. Processing plays the audio stream on audio player <b>545</b> at step <b>540</b>. Audio player <b>545</b> may be an electronic device capable of playing an audio source or an audio/video source, such as a stereo or a television.
[0060] Processing selects the first word in the text file at step <b>550</b>. A determination is made as to whether audio player <b>545</b> has played the first word in the audio file (decision <b>560</b>). If the first word has not been played, decision <b>560</b> branches to “No” branch <b>562</b> which loops back to wait for audio player <b>545</b>, to play the first word. This looping continues until the first word is played, at which point decision <b>560</b> branches to “Yes” branch <b>568</b> whereupon processing time-stamps the first word. For example, processing may time-stamp the first word at “t=0”.
[0061] A determination is made as to whether there are more words in the text file (decision <b>580</b>). If there are more words in the text file, decision <b>580</b> branches to “Yes” branch <b>582</b> which loops back to select (step <b>585</b>), and process the next word. This looping continues until there are no more words in the text file, at which point decision <b>580</b> branches to “No” branch <b>588</b>.
[0062] A determination is made as to whether the user wishes to manually adjust the timing corresponding to time stamps in timing store <b>575</b> (decision <b>590</b>). For example, the user may wish to increase the rate at which words are displayed on a personal device. If the user wishes to manually adjust the timing, decision <b>590</b> branches to “Yes” branch <b>591</b> whereupon the user adjusts the timing (step <b>592</b>). On the other hand, if the user does not wish to manually adjust the timing, decision <b>590</b> branches to “No” branch <b>594</b> bypassing manual adjustment steps.
[0063] Processing generates an enhanced captioning stream using text information located in text store <b>520</b> and time-stamp information in timing store <b>575</b> and stores the enhanced captioning stream in enhanced captioning stream store <b>598</b> (step <b>595</b>). In one embodiment, processing may add graphic enhancements corresponding to the timestamps. The graphic enhancement information may be compiled into a binary file and stored in a non-volatile storage area, such as non-volatile memory. For example, a bouncing ball may be positioned over a word when the word's corresponding timestamp comes. Processing ends at <b>599</b>.
[0064]FIG. 6A is a user interface window on an enhanced captioning device showing an enhanced captioning stream corresponding to a conversation. Window <b>600</b> shows two individuals, Bryan and Mike, having a conversation. For example, Bryan and Mike may be actors in a movie in which a user is viewing. Highlight <b>610</b> shows that Bryan has already spoken his sentence. Emoticon (emotion icon) <b>620</b> shows that Bryan spoke his sentence in a pleasant tone.
[0065] Text <b>640</b> informs the user of Mike's voice tone when speaking his sentence. In this example, text <b>640</b> indicates that Mike is shouting while speaking his sentence. The speakers lines are indented to indicate the time that each speaker says his sentence as indicated at point <b>630</b>. Highlight <b>660</b> indicates that Mike has spoken the first three words of his sentence, and is ready to speak the fourth word as indicated by point <b>665</b> where highlight <b>660</b> ends. Emoticon <b>650</b> shows that Mike is shouting his sentence.
[0066] Text <b>670</b> shows descriptive audio that is occurring during the conversation. In addition to descriptive audio describing sounds other than speech, descriptive audio may describe an action being performed, such as “Bryan is walking towards the fence”. Descriptive audio information may be input into devices for the visually impaired, such as a portable Braille device. Descriptive audio is stored in the enhanced captioning stream along with the time at which the enhanced captioning device should display the descriptive audio.
[0067]FIG. 6B is a user interface window on an enhanced captioning device showing an enhanced captioning stream corresponding to a musical event. Window <b>680</b> shows musical notes corresponding to a media stream, such as a song. For example, a user may be listening to a song and the enhanced captioning device is displaying notes of the song and the timing at which each note is played. Highlight <b>690</b> indicates that the first five notes of the song have been played, and the sixth note is about to be played as indicated by point <b>695</b>. The, enhanced captioning device may synchronize to musical notes in the same manner in which the enhanced captioning device synchronizes to speech. The enhanced captioning device may listen to one or more notes, and compare the notes with the enhanced captioning stream. Once the enhanced captioning device detects a match between the notes and a location point within the enhanced captioning stream, the enhanced captioning device synchronizes the enhanced captioning stream with the notes, and displays, highlight <b>690</b> accordingly. The enhanced captioning device may, also receive a manual synchronization signal from the user, or a wireless synchronization signal from a media-playing device.
[0068]FIG. 7 illustrates information handling system <b>701</b> which is a simplified example of a computer system capable of performing the invention described herein. Computer system <b>701</b> includes processor <b>700</b> which is coupled to host bus <b>705</b>. A level two (L2) cache memory <b>710</b> is also coupled to the host bus <b>705</b>. Host-to-PCI bridge <b>715</b> is coupled to main memory <b>720</b>, includes cache memory and main memory control functions, and provides bus control to handle transfers among PCI bus <b>725</b>, processor <b>700</b>, L2 cache <b>710</b>, main memory <b>720</b>, and host bus <b>705</b>. PCI bus <b>725</b> provides an interface for a variety of devices including, for example, LAN card <b>730</b>. PCI-to-ISA bridge <b>735</b> provides bus control to handle transfers between PCI bus <b>725</b> and ISA bus <b>740</b>, universal serial bus (USB) functionality <b>745</b>, IDE device functionality <b>750</b>, power management functionality <b>755</b>, and can include other functional elements not shown, such as a real-time clock (RTC), DMA control, interrupt support, and system management bus support. Peripheral devices and input/output (I/O) devices can be attached to various interfaces <b>760</b> (e.g., parallel interface <b>762</b>, serial interface <b>764</b>, infrared (IR) interface <b>766</b>, keyboard interface <b>768</b>, mouse interface <b>770</b>, and fixed disk (HDD) <b>772</b>) coupled to ISA bus <b>740</b>. Alternatively, many I/O devices can be accommodated by a super I/O controller (not shown) attached to ISA bus <b>740</b>.
[0069] BIOS <b>780</b> is coupled to ISA bus <b>740</b>, and incorporates the necessary processor executable code for a variety of low-level system functions and system boot functions. BIOS <b>780</b> can be stored in any computer readable medium, including magnetic storage media, optical storage media, flash memory, random access memory, read only memory, and communications media conveying signals encoding the instructions (e.g., signals from a network). In order to attach computer system <b>701</b> to another computer system to copy files over a network, LAN card <b>730</b> is coupled to PCI bus <b>725</b> and to PCI-to-ISA bridge <b>735</b>. Similarly, to connect computer system <b>701</b> to an ISP to connect to the Internet using a telephone line connection, modem <b>775</b> is connected to serial port <b>764</b> and PCI-to-ISA Bridge <b>735</b>.
[0070] While the computer system described in FIG. 7 is capable of executing the invention described herein, this computer system is simply one example of a computer system. Those skilled in the art will appreciate that many other computer system designs are capable of performing the invention described herein.
[0071] One of the preferred implementations of the invention is an application, namely, a set of instructions (program code) in a code module which may, for example, be resident in the random access memory of the computer. Until required by the computer, the set of instructions may be stored in another computer memory, for example, on a hard disk drive, or in removable storage such as an optical disk (for eventual use in a CD ROM) or floppy disk (for eventual use in a floppy disk drive), or downloaded via the Internet or other computer network. Thus, the present invention may be implemented as a computer program product for use in a computer. In addition, although the various methods described are conveniently implemented in a general purpose computer selectively activated or reconfigured by software, one of ordinary skill in the art would also recognize that such methods may be carried out in hardware, in firmware, or in more specialized apparatus constructed to perform the required method steps.
[0072] While particular embodiments of the present invention have been shown and described, it will be obvious to those skilled in the art that, based upon the teachings herein, changes and modifications may be made without departing from this invention and its broader aspects and, therefore, the appended claims are to encompass within their scope all such changes and modifications as are within the true spirit and scope of this invention. Furthermore, it is to be understood that the invention is solely defined by the appended claims. It will be understood by those with skill in the art that if a specific number of an introduced claim element is intended, such intent will, be explicitly recited in the claim, and in the absence of such recitation no such limitation is present. For a non-limiting example, as an aid to understanding the following appended claims contain usage of the introductory phrases “at least one” and “one or more” to introduce claim elements. However, the use of such phrases should not be construed to imply that the introduction of a claim element by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim element to inventions containing only one such element, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an”; the same holds true for the use in the claims of definite articles.
Contents4
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| CN104980790A | Cited by | China | Search report |
| US2014359079A1 | Cited by | United States of America | Pre-grant |
| US2016322080A1 | Cited by | United States of America | Search report |
| US11375347B2 | Cited by | United States of America | Search report |
| US2008123636A1 | Cited by | United States of America | Pre-grant |
| US7983307B2 | Cited by | United States of America | Search report |
| US8761568B2 | Cited by | United States of America | Search report |
| US8634030B2 | Cited by | United States of America | Search report |
| KR101360316B1 | Cited by | Republic of Korea | Examiner |
| US8550334B2 | Cited by | United States of America | Applicant |
| US9686584B2 | Cited by | United States of America | Applicant |
| US2023215067A1 | Cited by | United States of America | Search report |
| US8640956B2 | Cited by | United States of America | Applicant |
| US2007188657A1 | Cited by | United States of America | Pre-grant |
| US10015483B2 | Cited by | United States of America | Applicant |
| US9027015B2 | Cited by | United States of America | Applicant |
| US8511540B2 | Cited by | United States of America | Applicant |
| US8386339B2 | Cited by | United States of America | Applicant |
| US2023030342A1 | Cited by | United States of America | Search report |
| US2009232473A1 | Cited by | United States of America | Pre-grant |
| WO2012115859A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US9148686B2 | Cited by | United States of America | Applicant |
| US8534540B2 | Cited by | United States of America | Applicant |
| US8782262B2 | Cited by | United States of America | Applicant |
| WO2007142648A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US8439257B2 | Cited by | United States of America | Applicant |
| US9571888B2 | Cited by | United States of America | Applicant |
| US2010141834A1 | Cited by | United States of America | Pre-grant |
| US8827150B2 | Cited by | United States of America | Applicant |
| US10506295B2 | Cited by | United States of America | Applicant |
| US8408466B2 | Cited by | United States of America | Applicant |
| US9652108B2 | Cited by | United States of America | Applicant |
| US8781824B2 | Cited by | United States of America | Search report |
| CN105323648A | Cited by | China | Search report |
| US8443407B2 | Cited by | United States of America | Applicant |
| US9066046B2 | Cited by | United States of America | Search report |
| US9788048B2 | Cited by | United States of America | Applicant |
| US2009146965A1 | Cited by | United States of America | Pre-grant |
| US2007011012A1 | Cited by | United States of America | Pre-grant |
| US10382807B2 | Cited by | United States of America | Applicant |
| US2010275132A1 | Cited by | United States of America | Pre-grant |
| US10015550B2 | Cited by | United States of America | Applicant |
| US11954778B2 | Cited by | United States of America | Search report |
| US8786410B2 | Cited by | United States of America | Applicant |
| US2005280636A1 | Cited by | United States of America | Pre-grant |
| US7750892B2 | Cited by | United States of America | Applicant |
| US10820061B2 | Cited by | United States of America | Applicant |
| US9843613B2 | Cited by | United States of America | Search report |
| US2015295986A1 | Cited by | United States of America | Pre-grant |
| US2008064326A1 | Cited by | United States of America | Pre-grant |
| US8468610B2 | Cited by | United States of America | Applicant |
| US8856853B2 | Cited by | United States of America | Applicant |
| US8723815B2 | Cited by | United States of America | Applicant |
| US2004139482A1 | Cited by | United States of America | Pre-grant |
| US10027939B2 | Cited by | United States of America | Applicant |
| US10055693B2 | Cited by | United States of America | Applicant |
| US7913155B2 | Cited by | United States of America | Search report |
| US9280515B2 | Cited by | United States of America | Applicant |
| FR3006525A1 | Cited by | France | Search report |
| US9781465B2 | Cited by | United States of America | Applicant |
| EP2232365A4 | Cited by | European Patent Office (EPO) | Search report |
| US2012173235A1 | Cited by | United States of America | Pre-grant |
| WO2010068388A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US2010232762A1 | Cited by | United States of America | Pre-grant |
| US2008259210A1 | Cited by | United States of America | Pre-grant |
| US9792612B2 | Cited by | United States of America | Applicant |
| CN100452874C | Cited by | China | Search report |
| US2007106516A1 | Cited by | United States of America | Pre-grant |
| US2019096407A1 | Cited by | United States of America | Search report |
| US2007140656A1 | Cited by | United States of America | Pre-grant |
| US8875173B2 | Cited by | United States of America | Applicant |
| US9736469B2 | Cited by | United States of America | Applicant |
| US8430302B2 | Cited by | United States of America | Applicant |
| EP2232365A2 | Cited by | European Patent Office (EPO) | Search report |
| US8833640B2 | Cited by | United States of America | Applicant |
| EP2045810A1 | Cited by | European Patent Office (EPO) | Search report |
| US8746554B2 | Cited by | United States of America | Applicant |
| CN103986940A | Cited by | China | Search report |
| US11936958B2 | Cited by | United States of America | Search report |
| US9654737B2 | Cited by | United States of America | Applicant |
| US2009204404A1 | Cited by | United States of America | Pre-grant |
| US8931031B2 | Cited by | United States of America | Applicant |
| US12010393B2 | Cited by | United States of America | Search report |
| US10726842B2 | Cited by | United States of America | Search report |
| US2015295986A1 | Cited by | United States of America | Search report |
| US10165321B2 | Cited by | United States of America | Applicant |
| US8497939B2 | Cited by | United States of America | Search report |
| US8553146B2 | Cited by | United States of America | Applicant |
| US2009226046A1 | Cited by | United States of America | Pre-grant |
| US9092830B2 | Cited by | United States of America | Applicant |
| US8736761B2 | Cited by | United States of America | Search report |
| US8886172B2 | Cited by | United States of America | Applicant |
| US9329966B2 | Cited by | United States of America | Applicant |
| US2008244676A1 | Cited by | United States of America | Pre-grant |
| US2010157151A1 | Cited by | United States of America | Pre-grant |
| US2010293598A1 | Cited by | United States of America | Pre-grant |
| US9280905B2 | Cited by | United States of America | Search report |
| US9367669B2 | Cited by | United States of America | Applicant |
| EP2811749A1 | Cited by | European Patent Office (EPO) | Search report |
| US2013150990A1 | Cited by | United States of America | Pre-grant |
1 member in 1 office
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 23397302 | United States of America | A | |
| US20020233973 | – | – | – |
Members1
| Document | Office | Kind | |
|---|---|---|---|
| US2004044532A1 | United States of America | A1 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Information on status: application discontinuationEXPRESSLY ABANDONED -- DURING EXAMINATIONSTCB | STCB | |
| AssignmentAS | AS |
Numbers
- Publication, DOCDB
- 2004044532
- Publication, EPODOC
- US2004044532
- Application
- 10233973
- Application, DOCDB
- 23397302
- Application, EPODOC
- US20020233973
Titles
- English
- System and method for remote audio caption visualizations
Classification
- CPC, 7
- H04N21/41415
- G11B27/102
- H04N21/4126
- H04N21/4622
- H04N21/4884
- H04N21/8133
- H04N21/43079
- IPC, 1
- G11B27 10
- USPC, 2
- 704271000
- G9B027018