Recognition of speech in editable audio streams
Abstract
This record has no abstract on file.
Term
Projected expiry 23 November 2027.
- Priority
- Filed
- Granted
- Today
- Projected expiry
23 claims: 7 independent, 16 dependent
- 1コンピュータで実現される方法であって、(A)話者の第1の音声を表す第1の部分的なオーディオストリームを生成するステップと、(B)前記第1の部分的なオーディオストリームを、前記第1の部分的なオーディオストリームが一部分であるディクテーションストリームの開始時間に相対する第1の開始時間と関連付けるステップと、(C)前記第1の部分的なオーディオストリームの終了後に続く前記話者の第2の音声を表す第2の部分的なオーディオストリームを生成するステップと、(D)前記第2の部分的なオーディオストリームを、前記第2の部分的なオーディオストリームが一部分である前記ディクテーションストリームの開始時間に相対する第2の開始時間と関連付けるステップであって、 巻き戻しの時間に対応する量によって、 前記ディクテーションストリームの開始時間に相対する、前記第1の部分的なオーディオストリームの第1の終了時間 よりも早くなるように前記第2の開始時間を設定することが可能である、 ステップと、(E)自動音声認識機器で、(1)前記第1の部分的なオーディオストリームを受信するステップと、(2)前記第1の開始時間に基づいた場所で、前記第1の部分的なオーディオストリームを有効ディクテーションストリームに書き込むステップと、(3)前記第2の部分的なオーディオストリームを受信するステップと、(4)前記第2の開始時間に基づいた場所で、前記第2の部分的なオーディオストリームを前記有効ディクテーションストリームに書き込むステップと、(5)前記(E)(4)のステップを完了する前に前記有効ディクテーションのトランスクリプトを生成するために、前記有効ディクテーションの少なくとも一部分に自動音声認識処理を適用するステップと、を含むことを特徴とする方法。
- 2前記(E)(5)のステップは、前記(E)(3)のステップの完了の前に出力を作るために、前記有効ディクテーションの少なくとも一部分に自動音声認識処理を適用するステップを含むことを特徴とする請求項1に記載の方法。
- 3前記(E)(2)のステップは、前記(C)のステップが完了する前に完了することを特徴とする請求項1に記載の方法。
- 4前記(E)(1)のステップは、前記(A)のステップが完了する前に、開始されることを特徴とする請求項1に記載の方法。
- 5前記(D)のステップは、挿入される他の部分的なオーディオストリームの時間に対応する量によって、前記第1の終了時間から前記第2の開始時間を増加させるように設定すること がさらに 可能であることを特徴とする請求項1に記載の方法。
- 6前記(E)(1)のステップは、ネットワーク上の前記第1の部分的なオーディオストリームを受信するステップを含むことを特徴とする請求項1に記載の方法。
- 7(F)前記(C)のステップの前に、前記ディクテーションストリーム上で編集操作を指定する前記話者からの入力を受信するステップと、(G)前記編集操作入力に応答して前記第1の部分的なオーディオストリームを終了し、前記第2の部分的なオーディオストリームを開始するステップと、をさらに含むことを特徴とする請求項1に記載の方法。
- 8前記(F)のステップは、前記ディクテーションストリームの相対的開始時間は、新規相対的開始時間に変更されるということを指定する前記話者からの第1の入力を受信するステップと、前記ディクテーションストリームが前記新規相対的開始時間で再開されるということを指定する前記話者からの第2の入力を受信するステップと、を含み、 前記第2の部分的なオーディオストリームの前記第2の開始時間は、前記第1の部分的なオーディオストリームの前記第1の開始時間よりも、前記ディクテーションストリームにおける開始時間に相対して早いことを特徴とする請求項7に記載の方法。
- 9前記(E)(5)のステップは、前記有効ディクテーションの少なくとも一部分を再生するステップを含むことを特徴とする請求項1に記載の方法。
- 10前記(E)(5)のステップは、前記(E)(4)のステップが完了した後にのみ、前記トランスクリプトをユーザへ表示するステップをさらに含むことを特徴とする請求項1に記載の方法。
- 11前記(E)(4)のステップは、(E)(4)(a)前記第2の部分的なオーディオストリームの前記第2の開始時間の既定の閾値内である時間で、前記有効ディクテーション内の沈黙の間としての言葉の一時停止を識別するステップと、(E)(4)(b)前記(E)(4)(a)のステップで識別された時間で、前記第2の部分的なオーディオストリームを前記有効ディクテーションに書き込むステップと、を含むことを特徴とする請求項10に記載の方法。
- 12(F)前記第1の部分的なオーディオストリームと関連付けられる文脈情報を識別するステップと、(G)前記第1の部分的なオーディオストリームの前記第1の開始時間を前記文脈情報と関連付けるステップと、(H)前記自動音声認識機器で、前記第1の部分的なオーディオストリームの前記第1の開始時間と関連する前記文脈情報を受信するステップと、をさらに含むことを特徴とする請求項1に記載の方法。
- 13前記(E)(5)のステップは、前記第1の部分的なオーディオストリームおよび前記文脈情報を反映する出力を作るために、前記第1の部分的なオーディオストリームおよび前記文脈情報に自動音声認識処理を適用するステップを含むことを特徴とする請求項12に記載の方法。
- 14前記(F)のステップは、前記文脈情報を識別する前記話者からの入力を受信するステップを含むことを特徴とする請求項12に記載の方法。
- 15前記文脈情報は、画像を含むことを特徴とする請求項12に記載の方法。
- 16装置であって、 話者の第1の音声を表す第1の部分的なオーディオストリームを生成するための第1の部分的なオーディオストリーム生成手段と、 前記第1の部分的なオーディオストリームを、前記第1の部分的なオーディオストリームが一部分であるディクテーションストリームにおける開始時間に相対する第1の開始時間と関連付けるための、第1の相対的時間手段と、 前記第1の部分的なオーディオストリームの終了後に続く前記話者の第2の音声を表す第2の部分的なオーディオストリームを生成するための第2の部分的なオーディオストリーム生成手段と、 前記第2の部分的なオーディオストリームを、前記第2の部分的なオーディオストリームが一部分である前記ディクテーションストリームにおける開始時間に相対する第2の開始時間と関連付け、 巻き戻しの時間に対応する量によって、 前記ディクテーションストリームの開始時間に相対する、前記第1の部分的なオーディオストリームの終了時間 よりも早くなるように前記第2の開始時間を 設定することが可能である、第2の相対的時間手段と、 自動音声認識機器であって、 前記第1の部分的なオーディオストリームを受信するための第1の受信手段と、 前記第1の開始時間に基づいた場所で、前記第1の部分的なオーディオストリームを有効ディクテーションストリームに書き込むための第1の書き込み手段と、 前記第2の部分的なオーディオストリームを受信するための第2の受信手段と、 前記第2の開始時間に基づいた場所で、前記第2の部分的なオーディオストリームを前記有効ディクテーションストリームに書き込むための第2の書き込み手段と、 前記第2の部分的なオーディオストリームの書き込みが完了する前に、前記有効ディクテーションのトランスクリプトを生成するために、前記有効ディクテーションの少なくとも一部分に自動音声認識処理を適用するための自動音声認識処理手段と、を含む自動音声認識機器と、を含むことを特徴とする装置。
- 17前記音声認識手段は、前記第2の部分的なオーディオストリームの受信が完了する前に、前記有効ディクテーションのトランスクリプトを生成するために、前記有効ディクテーションの少なくとも一部分に自動音声認識処理を適用するための手段を含むことを特徴とする請求項16に記載の装置。
- 18前記第1の書き込み手段は、前記第2の部分的なオーディオストリームの生成が完了する前に、前記第1の部分的なオーディオストリームを書き込むための手段を含むことを特徴とする請求項16に記載の装置。
- 19前記第1の受信手段は、前記第1の部分的なオーディオストリームの生成が完了する前に、前記第1の部分的なオーディオストリームの受信を開始するための手段を含むことを特徴とする請求項16に記載の装置。
- 20コンピュータで実行される方法であって、(A)話者の第1の音声を表す第1の部分的なオーディオストリームを生成するステップと、(B)前記第1の部分的なオーディオストリームを、前記第1の部分的なオーディオストリームが一部分であるディクテーションストリームにおける開始時間に相対する第1の開始時間と関連付けるステップと、(C)前記第1の部分的なオーディオストリームの終了後に続く前記話者の第2の音声を表す第2の部分的なオーディオストリームを生成するステップと、(D)前記第2の部分的なオーディオストリームを、前記第2の部分的なオーディオストリームが一部分である前記ディクテーションストリームにおける開始時間に相対する第2の開始時間と関連付ける ステップであって、巻き戻しの時間に対応する量によって、前記ディクテーションストリームの開始時間に相対する、前記第1の部分的なオーディオストリームの第1の終了時間よりも早くなるように前記第2の開始時間を設定することが可能である、 ステップと、(E)自動音声認識機器で、(1)ネットワーク上の前記第1の部分的なオーディオストリームを受信するステップと、(2)前記第1の開始時間に基づいた場所で、前記第1の部分的なオーディオストリームを有効ディクテーションストリームに書き込むステップと、(3)前記ネットワーク上で前記第2の部分的なオーディオストリームを受信するステップと、(4)前記第2の開始時間に基づいた場所で、前記第2の部分的なオーディオストリームを前記有効ディクテーションストリームに書き込むステップと、(5)前記(E)(4)のステップの完了の前に前記有効ディクテーションのトランスクリプトを生成するために、前記有効ディクテーションの少なくとも一部分に自動音声認識処理を適用するステップと、を含むことを特徴とする方法。
- 21(F)前記(C)のステップの前に、前記ディクテーションストリームの一時停止を指定する前記話者からの第1の入力を受信するステップと、前記ディクテーションストリームの再開を指定する前記話者からの第2の入力を受信するステップと、をさらに含むことを特徴とする請求項20に記載の方法。
- 22装置であって、 話者の第1の音声を表す第1の部分的なオーディオストリームを生成するための第1の生成手段と、 前記第1の部分的なオーディオストリームを、前記第1の部分的なオーディオストリームが一部分であるディクテーションストリームにおける開始時間に相対する第1の開始時間と関連付けるための、第1の関連付け手段と、 前記第1の部分的なオーディオストリームの終了後に続く前記話者の第2の音声を表す第2の部分的なオーディオストリームを生成するための第2の生成手段と、 前記第2の部分的なオーディオストリームを、前記第2の部分的なオーディオストリームが一部分であるディクテーションストリームにおける開始時間に相対する第2の開始時間と関連付け、 巻き戻しの時間に対応する量によって、前記ディクテーションストリームの開始時間に相対する、前記第1の部分的なオーディオストリームの終了時間よりも早くなるように前記第2の開始時間を設定することが可能である 、第2の関連付け手段と、 自動音声認識機器であって、 ネットワーク上の前記第1の部分的なオーディオストリームを受信するための第1の受信手段と、 前記第1の開始時間に基づいた場所で、前記第1の部分的なオーディオストリームを有効ディクテーションストリームに書き込むための第1の書き込み手段と、 前記ネットワーク上の前記第2の部分的なオーディオストリームを受信するための第2の受信手段と、 前記第2の開始時間に基づいた場所で、前記第2の部分的なオーディオストリームを前記有効ディクテーションストリームに書き込むための第2の書き込み手段と、 前記第2の部分的なオーディオストリームの書き込みの完了の前に、前記有効ディクテーションのトランスクリプトを生成するために、前記有効ディクテーションの少なくとも一部分に自動音声認識処理を適用するための自動音声認識処理手段と、を含む自動音声認識機器と、を含むことを特徴とする装置。
- 23前記第2の部分的なオーディオストリームの生成の前に、前記ディクテーションストリームの一時停止を指定する、前記話者からの第1の入力を受信するための第3の受信手段と、 前記ディクテーションストリームの再開を指定する前記話者からの第2の入力を受信するための第4の受信手段と、をさらに含むことを特徴とする請求項22に記載の装置。
Independent claims23
66 paragraphs, as filed
The present invention relates to speech recognition in an editable audio stream.
Various automatic speech recognition devices exist to transcribe speech. Such systems can generally be operated in "word-for-word transcript" mode, where all spoken words are transcribed in the order in which they are spoken. However, it is not desirable for the speaker to create a verbatim transcript when performing an edit operation that disables previously dictated speech.
For example, consider a speaker who dictates to a portable digital recorder. The speaker speaks a few sentences and then notices that he made a mistake. The speaker wants to re-record (replace) the previous 10-second audio, and he rewinds the 10-second recording (perhaps by pressing the rewind button on the recording device), and then again. Start speaking and correct the last 10 seconds of voice.
The verbatim transcript of such audio is therefore not only the audio that the speaker intends to be part of the final transcript, but also other audio (eg, redictated 10 seconds). Will be replaced by audio), resulting in additional audio that should not be part of the final transcript. Some existing speech recognition devices can create transcripts that reflect such changes made to the spoken audio stream before the entire audio stream is dictated, but such systems Doing so by requiring each part of the audio stream to recognize a delay of a certain amount of time after that part is spoken, the resulting transcript of that part of the audio stream follows. Ensure that it is not disabled by voice (or at least increase its likelihood).
<p> Audio processing systems divide the spoken audio stream into partial audio streams called "snippets." The system may split a portion of the audio stream into two snippets where the speaker performs an editing operation, such as pausing, then starting recording, or rewinding, then starting recording, and so on. It is possible. Once the snippet is generated, the snippet may be continuously transmitted to a consumer such as an automatic speech recognition device or a playback device. When the snippets are received, the consumer can process (eg, recognize or play) them. Consumers can modify their output in response to edit operations reflected in the snippet. Consumers can process an audio stream while it is being generated and transmitted, even if the audio stream contains an edit operation that invalidates a previously transmitted partial audio stream. This allows for shorter turnaround times between the dictation and consumption of the complete audio stream.</p><p> (Cross-reference of related applications) This application claims the benefit of US Provisional Patent Application No. 60 / 867,105, entitled "Recognition of Speech in Editable Audio Streams," filed November 22, 2006.</p><p> This application relates to US Patent Application No. 10 / 923,517, entitled "Automated Extraction of Semantic content and Generation of a Structured Document from Speech," filed on August 10, 2004, the content of which is by reference. Incorporated herein.</p>
<figref num="1">FIG. 5 is a data flow diagram of a system for processing (eg, transcribing or reproducing) audio according to an embodiment of the present invention.</figref><figref num="2">FIG. 5 is a diagram of a data structure for storing a partial audio stream (snippet) of audio according to an embodiment of the present invention.</figref><figref num="3A">It is a figure which shows the flowchart of the method performed by the system of FIG. 1 for processing voice according to one Embodiment of this invention.</figref><figref num="3B">It is a figure which shows the flowchart of the method performed by the system of FIG. 1 for processing voice according to one Embodiment of this invention.</figref><figref num="3C">FIG. 5 illustrates a flow chart of a method used by a voice consumer in response to an invalidation of previously processed voice by an editing operation according to an embodiment of the invention.</figref><figref num="3D">It is a figure which shows the flowchart of the method for completing the generation of the voice transcript and allowing the user to edit the transcript according to one embodiment of the present invention.</figref><figref num="4">It is a figure which shows the flowchart of the method for initializing the system of FIG. 1 by one Embodiment of this invention.</figref><figref num="5">It is a figure which shows the data flow of the system for displaying and editing a transcript according to one Embodiment of this invention.</figref><figref num="6">FIG. 5 illustrates a flow chart of how a snippet is written to a dictation stream, thereby adjusting where the snippet begins during a pause in words.</figref><figref num="7">It is a figure which shows the data flow of the system for storing contextual information in the dictation stream of FIG.</figref>
Embodiments of the present invention allow speech to be transcribed automatically and in real time (ie, before the speaker is speaking and the speech is complete). Such transcription can be performed even when the speaker naturally speaks and performs editing operations such as changing the recording location while speaking by rewinding or transferring. Rewinding and then starting dictation is an example of the term "editing operation" used herein. Another example of an "editing operation" is to pause the recording, then start the recording, and then continue the dictation.
A portion of the audio (referred to herein as a "snippet") can be transcribed without delay. In other words, the first snippet can be transcribed while being spoken or without waiting for the delay time to end, even if subsequent snippets modify or delete the first snippet. Is.
In addition, the speaker can dictate while speaking without a system that presents the speaker with a draft transcript. Rather, the draft document can only be displayed to the speaker after the dictation is complete. This allows, for example, a radiologist who dictates a report to focus on an overview and description of the radiologist image while dictating, rather than editing the text. The speaker can only provide the opportunity to edit the draft transcript when the dictation is complete. This is different from traditional speech recognition systems, which typically display a draft document to the user while the user is speaking and require the user to change the dictation by changing the text on the screen. ..
Embodiments of the present invention are described in more detail herein. With reference to FIG. 1, a data flow diagram shows a system 100 for processing (eg, transcribing or reproducing) speech according to an embodiment of the present invention. With reference to FIGS. 3A-3B, the flowchart shows a method 300 that can be performed by the system 100 of FIG. 1 for transcribing audio according to one embodiment of the invention.
In general, a speaker 102, such as a doctor, begins speaking to a device 106, such as a digital recording device, a personal computer to which a microphone is connected, a personal digital auxiliary device, or a telephone (step 302). The speaker's voice is shown in FIG. 1 as "dictation" 104, which shows the entire spoken audio stream that is desired to be transcribed by the time Method 300 shown in FIGS. 3A-3B is completed. ..
As described in more detail below, the recording device 106 can divide the dictation 104 into multiple partial audio streams, referred to herein as "snippets." While the recording device 106 records each snippet, the recording device faces the start of the dictation 104 (or to any other reference point in the dictation 104), the snippet start time 130, and the snippet 202 fruit (absolute). It is possible to record a start time 132 (to keep the snippet responding to other forms of user input, such as clicking a button in the GUI). When the speaker 102 begins speaking, the recording device 106 can initialize the relative start time 130 and the absolute start time 132 (shown in method 400, steps 402 and 404 of FIG. 4, respectively).
The recording device 106 can initialize, create a new snippet (step 304), and start recording the currently spoken portion of the dictation 104 into the snippet (step 306). An exemplary data structure 200 for storing such snippets is shown in FIG. The snippet 200 may include, for example, (1) a time-continuous audio stream 202 representing a portion of the dictation 104 associated with the snippet 200, (2) a start time 204 of the audio stream 202 relative to the start of the dictation 104, and (3) a portion. The real (absolute) start time 206 of the typical audio stream 202 and (4) the editing operation 208 associated with the snippet 200 (if any) can be included or associated with them. The recording device 106 can copy the values of the relative start time 130 and the absolute start time 132 to the relative start time 204 and the absolute start time 206, respectively, when the snippet 200 is initialized. is there.
The recording device 106 is, for example, when the speaker 102 uses the recording device 106 to perform an editing operation such as pausing, rewinding, or transferring the recording during recording (step 308). , It is possible to exit the current snippet 200. To exit the snippet 200, the recording device 106 interrupts recording of additional audio into the snippet 200's audio stream 202 and provides information about the editing operations performed by the speaker 102 in the snippet 200 field 208. It is possible to record to (step 310). The recording device 106 can then send the current snippet 200 on the network 112 to the consumer 114, such as a human transcriptionist, an automatic speech recognition device, or an audio playback device (step 312). .. An example of how a consumer 114 can consume a snippet is described below.
It should be noted that in the example shown in FIG. 3A, the current snippet 200 is sent to the consumer 114 after the snippet 200 is finished. However, this is merely an example and does not limit the present invention. For example, the recording device 106 can stream the current snippet 200 to the consumer 114 before the snippet 200 is finished. For example, the recording device 106 may start streaming the current snippet 200 as soon as the recording device 106 begins storing the audio stream 202 in the snippet 200, as more audio streams 202 are stored in the snippet 200. , It is possible to keep streaming the snippet 200. As a result, the consumer 114 begins to process (eg, recognize or play) the faster part of the snippet 200, even when the speaker 102 speaks and the recording device 106 records and transmits the slower part of the snippet 200. Is possible.
When speaker 102 continues to dictate after the end of the current snippet (step 302), recording device 106 stores the new snippet in relative start times 130 and absolute start times, respectively, in fields 204 and 206. It can be initialized with a current value of 132, as well as an empty audio stream 202 (step 304). In other words, the recording device 106 divides the dictation 104 into consecutive snippets 102a-n, and when the snippet 102a-n is created, the recording device 106 continuously transmits to the consumer 114, so that the speaker 102 is natural. It is possible to continue dictating. The snippet 102a-n thereby forms a dictation stream 108 that the recording device 106 transmits to the consumer 114. The dictation stream 108 can be formatted as a single sequential stream of bytes, for example in a socket, HTTP connection, or streaming object via API.
In parallel with such continuous recording of dictation by speaker 102 and dictation 104 by recording device 106, consumer 114 is capable of receiving each snippet (step 314). For example, if the speaker 102 creates a dictation 104 using the client-side speech recording device 106, the consumer 114 can be a server-side automatic speech recognition device.
Consumer 114 is capable of processing each of the snippets 110a-n as they are received, in other words, without causing any delay before initiating such processing. Further, the consumer 114 is capable of processing one snippet while the recording device 106 continues to record and transfer subsequent snippets in the dictation stream 108. For example, if the consumer 114 is an automatic speech recognition device, the automatic speech recognition device can transfer each of the snippets 110a-n as they are received, thereby configuring the dictation 104. When the snippet 110a-n to be received is received by the consumer 114, it creates a working transcript 116 of the dictation 104.
Consumer 114 is capable of combining received snippets 110a-n into a single combined audio stream referred to herein as "valid dictation" 120 on the consumer (eg, server) side. .. In general, the purpose for effective dictation 120 is to represent the speaker's intent towards the transcribed speech. For example, the original dictation 104 contains 10 seconds of voice, and if speaker 102 rewinds to these 10 seconds of voice and dictates on them, they are continuously disabled and Even if the audio appears in the stream of snippet 110a-n sent to the original dictation 104 and consumer 114, the deleted (disabled) 10 second audio should not appear in the valid dictation 120. .. Consumer 114 repeatedly updates the valid dictation 120 when it receives the snippet 110a-n.
More specifically, the consumer 114 can include a "reader" component 122 and a "processor" component 124. At some point before receiving the first snippet, the reader 122 initializes the valid dictation 120 to the empty audio stream (Figure 4, step 406) and indicates the write time 134 to the start of the valid dictation 120. Initialize as follows (step 408). The write time 134 displays the time within the valid dictation 120 in which the reader 122 writes the next snippet.
Next, when the reader 122 receives the snippet 110a-n (step 314), the reader 122 updates the valid dictation 120 based on the contents of the snippet 110a-n. The reader 122 can begin updating the valid dictation 120 as soon as it begins receiving snippets 110a-n, and therefore all snippets 110a-n will be received. As a result, the reader 122 can update the effective dictation 120 based on the reception of the first half snippet at the same time that the reader 122 receives the subsequent snippet.
When the reader 122 receives the snippet, the reader can identify the relative start time of the snippet from the field 204 of the snippet (step 320). The reader 122 can then update the valid dictation 120 by using the snippet and writing the contents of the snippet's audio stream 202 to the valid dictation 120 at the identified start time (step 322). ).
The reader 122 can "write" the audio stream 202 to the effective dictation 120 in various ways. For example, the reader 122 can write the audio stream 202 to the active dictation 120 in "overwrite" mode, in which mode the reader 122 is now stored at the identified start time (step 320). Overwrite the valid dictation 120 with the data from the new snippet. As another example, the reader 122 can write the audio stream 202 to the active dictation 120 in "insert" mode, in which the reader 122 (1) activates the current snippet dictation 120. Insert into and start at the start time identified in step 320, and (2) increase the relative start time of subsequent snippets already stored in effective dictation 120 by an amount equal to the time of the newly inserted snippet. .. As yet another example, the reader 122 can write the audio stream 202 to the enabled dictation 120 in "truncate" mode, in which the reader 122 is (1) at the identified start time. Overwrite the stored (step 320) current data in the valid dictation 120 with the data from the snippet, and (2) erase any data in the valid dictation 120 after the newly written snippet.
The reader 122 can use overwrite mode, insert mode, or truncate mode in any various way to determine whether to write the current snippet to the valid dictation 120. For example, the reader 122 can be configured to write all snippets 110a-n to a particular dictation stream 108 using a similar mode (eg, overwrite or insert). As another example, the edit operation field 208 for each snippet can specify which mode should be used to use for writing that snippet.
If the relative start time 204 of the current snippet points to or exceeds the end of the valid dictation 120, then the reader 122, regardless of whether the reader 122 is operating in overwrite mode or insert mode. It is possible to add the audio stream 202 of the current snippet to the effective dictation 120.
Consider how the operation of the reader 122 described above affects the effective dictation 120 in the case of two specific types of editing operations, "pause recording" and "pause and rewind". In the case of a recording pause, the speaker 102 pauses the recording on the recording device 106 and then restarts the recording at the subsequent "real" (absolute) time. In response, the recording device 106 terminates the current snippet and creates a new snippet when speaker 102 resumes recording, as described above with respect to FIG. 3A. The resulting two snippets contain an audio stream representing their respective audio before and after the pause. In this case, the recording device 106 can be set so that the second relative start time of the two snippets is equal to the relative end time of the first snippet.
If the reader 122 receives the first and second snippets, the reader 122 will take steps 320-322 because the relative end time of the first snippet corresponds to the relative start time of the second snippet. It is possible to effectively combine both snippets into a single long audio stream. This reflects the possible intention of speaker 102 to create a single continuous audio stream from two snippets.
In the case of "pause and rewind", the speaker 102 pauses, rewinds, and resumes recording on the recording device 106. In this case, the recording device 106 is capable of producing two snippets within the dictation stream 108, one for the audio spoken before the pause / rewind is performed. The other is for the audio spoken after the pause / rewind has been performed. The relative start time of the second recorded snippet can be set to be earlier than the relative end time of the first recorded snippet, depending on the amount corresponding to the rewind time. Thereby, the effect of rewinding is reflected. As a result, the first and second recorded snippets can be discontinuous with respect to the start time of dictation 104 (or the reference point in dictation 104).
When the reader 122 receives the first of these two snippets, the reader will first write the first snippet to the valid dictation by performing steps 320-322. The reader 122 then receives the second of these two snippets, and the reader 122 receives that snippet earlier in the effective dictation, which corresponds to the earlier relative start time of the second snippet. Is inserted, thereby reflecting the effect of the rewind operation.
The techniques described above differ from those employed by existing transfer systems, which are combined into a single combination audio stream as soon as a partial audio stream is created. In other words, existing systems do not retain partial audio streams (because they are within the dictation streams 108 described herein) and a single audio stream is processed (eg, transcribed or played). ) To be transferred to the consumer. To allow rewinding, the combined audio stream is typically transferred to the consumer after a sufficient delay, and the partial audio stream that has already been transferred to the consumer is not modified by subsequent edit operations. Make sure that, or at least reduce its chances.
One disadvantage of such a system is that there is no absolute guarantee that subsequent editing operations will not modify the previous audio, even after a long delay. For example, even in a system with a delay of 5 minutes, the speaker may speak for 10 minutes from scratch before deciding to resume dictation. Another disadvantage of such systems is that the delays they cause delay the generation of transcripts.
In an embodiment of the invention, in contrast, an audio stream that reflects the application of an editing operation is not transferred to a consumer 114 (eg, a speech recognition device). Instead, a series of partial audio streams (snippets 110a-n) are transferred that further include audio streams that are modified or deleted by subsequent audio streams.
An example of how processor 124 is capable of handling valid dictation 120 is described here (step 324). In general, the processor 124 can operate in parallel with other elements of the system 100, such as the recording device 106 and the reader 122. After initializing the system (Figure 4), processor 124 is able to initialize read time 138 to zero (step 410). Read time 138 points to a position in the valid dictation where processor 124 reads the next. Processor 124 is even more capable of initializing transfer location 140 to zero (step 412). The location of the transcription points to a location in transcript 116 where the processor then writes the text.
Once the reader 122 begins to store the audio data in the valid dictation 120, the processor 124 can begin reading the data starting at a position within the valid dictation 120 specified by read time 138 (step 326). .. In other words, processor 124 does not have to wait at all before starting to read and process the data from the valid dictation 120. The processor 124 updates (increases) the read time 138 as the processor 124 reads the audio data from the valid dictation (step 328).
Processor 124 transcribes a portion of the valid dictation 120 read in step 326, creates the transcribed text, and writes such text to transcript 116 at the current transcription location 140 (step 330). Processor 124 updates the current transcription location 140 and points to the end of the text transcribed in step 330 (step 332). Processor 124 returns to step 326 and continues the process of reading audio from the valid dictation 120.
It should be noted that processor 124 is capable of performing functions other than transcription and / or in addition to it. For example, processor 124 may perform playback of audio at effective dictation 120, instead of or in addition to transcribing the effective dictation.
For the reasons mentioned above, there is no guarantee that any data read by processor 124 from effective dictation 120, and the process will be part of the final recording. For example, after processor 124 has transcribed a portion of the audio in the effective dictation 120, that portion of the audio can be deleted or overwritten within the effective dictation 120 by continuously received snippets. is there.
A flow chart is provided here showing how 350, with reference to FIG. 3C, is capable of responding to such invalidation of previously processed audio using consumer 114. The reader 122 is capable of accessing the current read time 138 of processor 124. The reader 122 is capable of reading the processor read time 138 (step 352) (after identifying the relative start time of the snippet in step 320 of FIG. 3B, etc.), thereby causing the reader 122 to read the reader. Allows the snippet currently being processed by 122 to detect whether processor 124 invalidates a portion of valid dictation 120 that has already been processed. More specifically, the reader 122 is capable of comparing the relative start time 204 of the snippet currently being processed by the reader 122 with the read time 138 of the processor 124. If its relative start time 204 is earlier than read time 128 (step 354), then reader 122 can provide update event 136 to processor 124 (step 356), and the data has already been processed. Indicates that it is no longer valid.
Update event 136 can include information such as the relative start time of the snippet being processed by reader 122. In response to receiving update event 136, processor 124 changes its read time 138 to the relative start time indicated by update event 136 (step 358), and then processes the valid dictation 120, starting with new read time 138. Can be resumed (step 362).
Method 350, shown in Figure 3C, is just an example of how consumer 114 can respond to the reception of a snippet that invalidates a previously processed snippet. The appropriate response to update event 136 depends on the consumer 114. For example, if the consumer 114 is an audio player, the audio player may ignore event 136 because it is not possible to "non-play" the audio. However, if the consumer 114 is an automatic speech recognition device, the speech recognition device discards the partial recognition result (text and / or partial hypothesis, etc.) corresponding to the currently invalid portion of the valid dictation 120 (text and / or partial hypothesis, etc.). In step 360), it is possible to restart processing (recognition) at a new load time 138 within the valid dictation 120 (step 362). The steps to discard the partial recognition result in step 360 are to remove the text from the current version of the transcript 116 for speech that is no longer part of the valid dictation 120 and for the transcript corresponding to the new load time 138. It is possible to include a step of updating the transfer location 140 to correspond to a location within 116.
Referring to FIG. 3D, the flowchart shows method 370 performed by system 100 upon completion of dictation 104. If the recording device 106 detects that the speaker 102 has finished dictating the dictation 104 (step 372), the recording device 106 can send a dictation completion indication 142 to the consumer 114 (step 372). 374) In response, the consumer 114 can terminate the processing of the dictation stream 108 to make the final version of the transcript 116, which is any editing operation performed by the speaker 102. (Steps 376 and 378).
Once the final transcript 116 is complete, a text editor 502 (Figure 5) or other component can display the rendering 504 of transcript 116 to speaker 102 for reconsideration (step 380). ). Speaker 102 can issue edit command 506 to the text editor 502 to edit transcript 116 to correct errors in transcript 116 or change the format of transcript 116 (step 382). .. Any person other than speaker 102 may perform such reconsideration and editing. In addition, one or more persons may perform such reconsideration and editing. For example, a medical record transcriptionist can reconsider and edit transcript 116 for linguistic accuracy, while a physician can reconsider and edit transcript 116 for factual accuracy. It is possible to do.
Speaker 102 generally finds it difficult to precisely rewind to the moment the speaker wants a dictate, and even a one-hundredth of a second difference affects the output of a speech recognition device. It should be noted that rewind events are generally very uncertain because of the potential. As a result, if speaker 102 rewinds and redicts, speaker 102 may rewind slightly longer or may not rewind long enough, and is intended by the user. If not, a small amount of words will be overwritten, or if the user's intention is to redict them, a small amount of words will remain.
One way to tackle this problem is shown by method 600 in FIG. 6, in which the reader 122 automatically adjusts the write time 134 when the speaker 102 rewinds, thereby making it new. Snippets are written to effective dictation 120 during silence (word pause). Method 600 can be performed, for example, after step 320 and before step 322 in FIG. 3B.
For example, when speaker 102 rewinds to a particular new relative start time, the reader 122 can search between word pauses within the valid dictation 120 near that new start time (). Step 602). If such a word pause is found within a shorter time frame than typical words (eg, hundreds of thousands of seconds) or within some other default threshold time (step 604). It is possible to infer that the duplication is an error. In such a case, the reader 122 can adjust the new write time 134 to be equal to the pause position of the word (step 606). Intelligent auto-relocation can improve recognition results by eliminating recognition errors caused by inaccurate rewind placement by speaker 102.
The advantages of the embodiments of the present invention are one or more of the following. Embodiments of the present invention are capable of performing the transfer in real time, i.e., when the audio 104 is being spoken or played, even when transcribing an audio stream, including editing operations. is there. No delay needs to be introduced after the partial audio stream has been spoken or played and before it has been transcribed or processed. As a result, transcription of voice 104 can be produced more quickly.
In addition to the benefits that allow the transcript to be used more quickly, the increased transcription speed makes it easier for the speaker 102 himself to edit the transcript 116, rather than by a third party. It is possible to reduce the transfer cost. In addition, the increased transcription rate increases the quality of transcription by allowing the speaker 102 to correct the error when it is new to the speaker's memory.
The techniques disclosed herein can incorporate any editing operations performed during dictation into the final transcript 116. As a result, the increased rate obtained by real-time treatment does not require any sacrifice in the transcript.
Moreover, the techniques disclosed herein can be applied to audio streams produced by speaking naturally. For example, speaker 102 can rewind, transfer, or pause the recording while dictating, and such editing operations can be reflected in the final transcript 116. As a result, the benefits of the technology disclosed herein can be obtained without the speaker having to change his dictation behavior.
Further, the technique disclosed in the present invention is different from various conventional systems in which the speaker 102 is required to make edits by editing the text of the draft transcript produced by the system. It is possible to execute without displaying the recognition result to the speaker 102. The ability to avoid the need for such text editing is for use in mobile recording / transmitting devices (such as mobile voice recorders and mobile phones) and in situations where speaker 102 cannot access a computer with a display. The techniques disclosed herein are specifically adapted for use. Eliminating the need for text display, even when a display is available, allows speaker 102 to focus freely on dictating and visual tasks (such as reading radiographic images) other than editing text. to enable.
Although the present invention has been described above with respect to specific embodiments, it should be understood that the aforementioned embodiments are provided by way of example and do not limit or define the scope of the invention. Various other embodiments include, are not limited to, and are within the scope of the claims. For example, the elements and components described herein can be subdivided or connected together to form fewer components to perform similar functions.
The recording device 106 can be any type of device. The recording device 106 may be software that runs on the computer, or may include it. Although only the transmitted dictation stream 108 is shown in FIG. 1, the recording device 106 may further store the dictation stream 108 or its equivalent within the recording device or in another storage means. Some or all dictations 108 can be removed from the recording device 106 at any time after being transferred to the consumer 114.
Further, although the recording device 106 and the consumer 114 are shown in FIG. 1 as different devices that communicate on the network 112, this is merely an example and does not constitute a limitation of the present invention. The recording device 106 and the consumer 114 can be implemented, for example, in a single device. For example, the recording device 106 and the consumer 114 can both be implemented in software running on a similar computer.
The network 112 can be any mechanism for transmitting the dictation stream 108. For example, network 112 can be the public internet or LAN.
The execution of the editing operation is described herein as a trigger for splitting the dictation 104 into snippets 110a-n, but the dictation 104 can be split into snippets 110a-n in other ways. For example, the recording device 106 can end the current snippet and create a new snippet periodically, for example, every 5 seconds, even if the speaker 102 does not perform the editing operation. As another example, recording device 106 can terminate the current snippet and create a new snippet after each long pause in dictation 104, or after a predetermined number of shorter pauses. ..
The recording device 106 is capable of recording data in addition to the audio data, as shown in FIG. 7, exemplifying the modification 700 to the system 100 of FIG. The specific elements of FIG. 1 are omitted from FIG. 7 solely for the purpose of facilitating illustration.
Consider an example in which speaker 102 is a doctor dictating these radiographic images while viewing the radiographic images as displayed by the radiological software on a monitor. When a physician dictates a comment about a particular such image, the recording device 106 records PACS (Picture Archiving and Communication System) information about the image and records that information (including the image itself) within the dictation stream 108. It is possible to send.
Such image information is merely an example of information 702a-m about the context of the dictation stream audio transmitted within or in connection with the dictation stream 108 itself. As a result, the dictation stream 108 can be more than just an audio stream, but a more general multimedia stream that results from the multimodal inputs provided by speaker 102 (eg, voice and keyboard inputs). is there.
As a result, the audio (snippet 110a-n) in the dictation stream 108 can correlate with any additional contextual information 702a-m associated with the audio 110a-n. Such correlations can be performed in any variety of ways. For example, an image can correlate with one or more snippets 110a-n by marking the image with the absolute start time of the snippet. As a result, the consumer 114 is able to match the images it receives or other contextual information 702a-m to the snippets they correspond to.
As a result, the consumer 114 can be a more general multimedia processor rather than just a speech recognizer, audio player, or other speech processor. For example, if processor 124 plays dictation stream 108, it is possible for processor 124 to play the snippet at the same time that processor 124 further displays the image or contextual information 702a-m associated with the snippet. Allows the reviewer / editor to understand or outline the contextual information associated with dictation stream 108 at the appropriate time.
The recording device 106 can determine whether to add the contextual information 702a-m to the dictation stream 108 in any variety of ways. For example, if speaker 102 is viewing an image as described above, the recording device 106 will automatically provide information about each image associated with a portion of the dictated dictation stream 108 while the image is being viewed. Can be added to. As another example, the recording device 106 may, by default, not transmit the image information of the dictation stream 108, but rather only the information about the image identified by the speaker 102. For example, if the speaker 102 considers a particular image to be important, the speaker 102 can hit a default hotkey or provide another input 704, which the recording device 106 identifies. In response to which recording device 106 does so, it is instructed to add information about the image to the dictation stream 108.
Instead, for example, if the consumer 114 is an automatic speech recognition device and the consumer receives the dictation stream 108, the processor 124 may store an image or other contextual information 708 recorded within the transcript 116. It is possible. The transcript 116 can be, for example, the type of structured document described in the referenced patent application, entitled "Automated Extraction of Semantic content and Generation of a Structured Document from Speech". The contextual information 708 in the transcript 116 can be tied to the text corresponding to the speech dictated by the speaker 102 when the contextual information was created. As a result, the image seen by the speaker 102 can be displayed when the text representing the image is then displayed by the text editor 502.
In the particular examples described herein, speech recognition is performed by an automated speech recognition device that operates on a server, but this is merely an example and is not a limitation of the present invention. Rather, speech recognition and other processing can be performed anywhere and need not occur within the client-server environment.
The art of art can be performed, for example, in hardware, software, firmware, or any combination thereof. The techniques described above include a processor, a processor-readable storage medium (eg, including volatile memory and non-volatile memory and / or storage elements), at least one input device, and at least one output device. It can be run in one or more computer programs that run on possible computers. The program code can be applied to the input input using the input device to perform the described functions and to produce the output. The output can be provided to one or more output devices.
Each computer program within the scope of the following claims can be run in any program language, such as an assembly language, a machine language, a high-level procedural program language, or an object-oriented program language. The program language may be, for example, a compiled or interpreted program language.
Each such computer program can be executed in a computer program product that is clearly embodied in a machine-readable storage device for execution by a computer processor. The steps of the methods of the invention can be performed on a computer processor that is explicitly embodied in a computer-readable medium for performing the functions of the invention by manipulating the inputs and producing the outputs. Suitable processors include, for example, both general purpose and special purpose microprocessors. In general, the processor receives instructions and data from read-only memory and / or random access memory. Suitable storage devices for explicitly embodying computer program instructions are all types of non-volatile memory, internal hard disks and removable disks, such as semiconductor memory devices, including, for example, EPROM, EEPROM and flash memory devices. Etc., including magnetic disks, magneto-optical disks, and CD-ROMs. Any of the above may be complemented or incorporated into a specially designed ASIC (application specific integrated circuit) or FPGA (field programmable gate array). Computers can also generally receive programs and data from storage media such as internal disks (not shown) or removable disks. These elements are also combined with any printing or marking engine, display screen, or other brown tube device capable of creating color or gray scale on paper, film, display screens, or other output media. It will be found on conventional desktop or workstation computers as well as other computers suitable for running computer programs that perform the methods described herein.
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2001228897A | Cites | Japan |
| JP2005079821A | Cites | Japan |
10 members in 5 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 60867105 | United States of America | – | |
| 86710506 | United States of America | P | |
| 2007085472 | United States of America | W | |
| 2006867105 | – | – | – |
| 2007085472 | – | – | – |
| US20060867105P | – | – | – |
| WO2007US85472 | – | – | – |
Members10
| Document | Office | Kind | |
|---|---|---|---|
| CA2662564A1 | Canada | A1 | |
| WO2008064358A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008064358A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US2008221881A1 | United States of America | A1 | |
| EP2095363A2 | European Patent Office (EPO) | A2 | |
| JP2010510556A | Japan | A | |
| US7869996B2 | United States of America | B2 | |
| CA2662564C | Canada | C | |
| EP2095363A4 | European Patent Office (EPO) | A4 | |
| JP4875752B2This record | Japan | B2 |
32 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313111S111 | S111 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written measure of declining of transfer procedureJAPANESE INTERMEDIATE CODE: R370R370 | R370 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification for declining of transfer of rightsJAPANESE INTERMEDIATE CODE: R360R360 | R360 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313111S111 | S111 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on accelerated examinationJAPANESE INTERMEDIATE CODE: A971005A975 | A975 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Explanation of circumstances concerning accelerated examinationJAPANESE INTERMEDIATE CODE: A871A871 | A871 |
Numbers
- Publication
- 4875752
- Publication, DOCDB
- 4875752
- Publication, EPODOC
- JP4875752B
- Application
- 2009538525
- Application, DOCDB
- 2009538525
- Application, EPODOC
- JP20090538525
Titles2
- Japanese
- 編集可能なオーディオストリームにおける音声の認識
- English
- Speech recognition in an editable audio stream
Classification
- CPC, 3
- G10L15/22
- G11B27/036
- G11B27/105
- IPC, 3
- G10L15 00
- G10L15 22
- G06F3 16