Speech end-pointer
Abstract
Rule-based end pointers separate spoken speech contained within the speech stream from background noise and non-spoken transients. Rule-based end pointers include multiple rules for determining the start and end of spoken speech based on various utterance characteristics. A rule may analyze an audio stream or part of an audio stream based on an event, a combination of events, a continuation of an event, or a continuation of an event. The rules can be manually or dynamically customized depending on the characteristics of the audio stream itself, the expected response contained within the audio stream, or factors that may include ambient conditions.
Term
Term ended
Projected expiry passed 3 April 2026, 0.5 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
39 claims: 5 independent, 34 dependent
- 1音声発話セグメントの開始および終了のうちの少なくとも一方を決定するエンドポインタであって、該エンドポインタは、 発話事象を含む音声ストリームの一部分を識別する音声トリガーモジュールと、 該音声トリガーモジュールと通信するルールモジュールであって、該ルールモジュールは、該音声ストリームの少なくとも一部を分析することによって、発話事象に関する音声発話セグメントが音声エンドポイント内にあるかを決定する複数の継続時間ルールを含む、ルールモジュールと を備える、エンドポインタ。
- 2前記音声トリガーモジュールが母音を識別する、請求項1に記載のエンドポインタ。
- 3前記音声トリガーモジュールがS音またはX音を識別する、請求項1に記載のエンドポインタ。
- 4前記音声ストリームの前記一部分がフレームを有する、請求項1に記載のエンドポインタ。
- 5前記ルールモジュールが前記音声ストリームの前記一部分におけるエネルギーの不足を分析する、請求項1に記載のエンドポインタ。
- 6前記ルールモジュールが前記音声ストリームの前記一部分におけるエネルギーを分析する、請求項1に記載のエンドポインタ。
- 7前記ルールモジュールが前記音声ストリームの前記一部分における経過時間を分析する、請求項1に記載のエンドポインタ。
- 8前記ルールモジュールが前記音声ストリームの前記一部分における所定の数の破裂音を分析する、請求項1に記載のエンドポインタ。
- 9前記ルールモジュールが前記音声発話セグメントの前記開始と終了とを検出する、請求項1に記載のエンドポインタ。
- 10エネルギー検出器モジュールをさらに備える、請求項1に記載のエンドポインタ。
- 11マイクロフォン出力部、処理ユニットおよびメモリと通信する、処理環境をさらに備え、前記ルールモジュールは該メモリ内に存在する、請求項1に記載のエンドポインタ。
- 12複数の決定ルールを有するエンドポインタを用いて音声発話セグメントの開始および終了のうちの少なくとも一方を決定する方法であって、該方法は、 音声ストリームの一部分を受信することと、 該音声ストリームの該一部分がトリガー特性を含むかを決定することと、 少なくとも1つの継続時間決定ルールを該トリガー特性に関する該音声ストリームの一部分に対して適用し、該音声ストリームの該一部分が音声エンドポイント内にあるかを決定することと を包含する、方法。
- 13前記決定ルールが、前記トリガー特性を含む前記音声ストリームの前記一部分に対して適用される、請求項12に記載の方法。
- 14前記決定ルールが、前記音声ストリームのうちの前記トリガー特性を含む前記一部分とは異なる一部分に対して適用される、請求項12に記載の方法。
- 15前記トリガー特性が母音である、請求項12に記載の方法。
- 16前記トリガー特性がS音またはX音である、請求項12に記載の方法。
- 17前記音声ストリームの前記一部分がフレームである、請求項12に記載の方法。
- 18前記ルールモジュールが前記音声ストリームの前記一部分におけるエネルギーの不足を分析する、請求項12に記載の方法。
- 19前記ルールモジュールが前記音声ストリームの前記一部分におけるエネルギーを分析する、請求項12に記載の方法。
- 20前記ルールモジュールが前記音声ストリームの前記一部分における経過時間を分析する、請求項12に記載の方法。
- 21前記ルールモジュールが前記音声ストリームの前記一部分における所定の数の破裂音を分析する、請求項12に記載の方法。
- 22前記ルールモジュールが潜在的な発話セグメントの開始および終了を検出する、請求項12に記載の方法。
- 23音声ストリームにおける音声発話セグメントの開始および終了のうちの少なくとも一方を決定するエンドポインタであって、該エンドポインタは、 該音声ストリームのうちの少なくとも1つのダイナミックな局面を分析することによって該音声発話セグメントが音声エンドポイント内にあるかを決定する複数の継続時間ルールを含む、エンドポインタモジュールと、 該エンドポインタモジュールと通信するメモリであって、該複数のルールのうちの1つ以上の継続時間を変更するプロファイル情報を保存するように構成されている、メモリと を備える、エンドポインタ。
- 24前記音声ストリームの前記ダイナミックな局面が話者の少なくとも1つの特徴を含む、請求項23に記載のエンドポインタ。
- 25前記話者の前記特徴が話者の話すペースを含む、請求項24に記載のエンドポインタ。
- 26前記音声ストリームの前記ダイナミックな局面が前記音声ストリームにおけるバックグラウンドノイズを含む、請求項23に記載のエンドポインタ。
- 27前記音声ストリームの前記ダイナミックな局面が、該音声ストリームにおいて予測された音を含む、請求項23に記載のエンドポインタ。
- 28前記予測された音が、話者に対して与えられた質問に対する少なくとも1つの予測された回答を含む、請求項27に記載のエンドポインタ。
- 29マイクロフォン入力部、処理ユニットおよびメモリと通信する、処理環境をさらに備え、前記エンドポインタモジュールは該メモリ内に存在する、請求項23に記載のエンドポインタ。
- 30音声ストリームにおける音声発話セグメントの開始および終了のうちの少なくとも一方を決定するエンドポインタであって、該エンドポインタは、 周期的な音声信号を含む音声ストリームの一部分を識別する音声トリガーモジュールと、 複数のルールに基づいて認識装置へ入力された該音声ストリームの量を変動させる、エンドポインタモジュールと を備え、 該複数のルールは、周期的な音声信号に関する音声ストリームの一部分が音声エンドポイント内にあるかを決定するための継続時間ルールを含む、エンドポインタ。
- 31前記認識装置が自動音声認識装置である、請求項30に記載のエンドポインタ。
- 32音声発話セグメントの開始および終了のうちの少なくとも一方を決定するための命令のセットを含む、コンピュータ可読記憶媒体であって、該命令のセットは、 音波を電気信号に変換することと、 該電気信号の周期性を識別することと、 該識別された周期性に関する該電気信号の可変部分を分析することによって、該電気信号が音声エンドポイント内にあるかを決定することと を含む、コンピュータ可読記憶媒体。
- 33前記電気信号の可変部分を分析することが、有声発話音の前の継続時間を分析することを含む、請求項32に記載のコンピュータ可読記憶媒体。
- 34前記電気信号の可変部分を分析することが、有声発話音の後の継続時間を分析することを含む、請求項32に記載のコンピュータ可読記憶媒体。
- 35前記電気信号の可変部分を分析することが、有声発話音の前または後の推移の数を分析することを含む、請求項32に記載のコンピュータ可読記憶媒体。
- 36前記電気信号の可変部分を分析することが、有声発話音の前の連続した沈黙の継続を分析することを含む、請求項32に記載のコンピュータ可読記憶媒体。
- 37前記電気信号の可変部分を分析することが、有声発話音の後の連続した沈黙の継続を分析することを含む、請求項32に記載のコンピュータ可読記憶媒体。
- 38前記コンピュータ可読媒体が車両のオンボードコンピュータ内に格納されている、請求項32に記載のコンピュータ可読記憶媒体。
- 39前記コンピュータ可読媒体が音声システムと通信する、請求項32に記載のコンピュータ可読記憶媒体。
Independent claims39
41 paragraphs, as filed
The present invention relates to automatic speech recognition, and more particularly to a system that separates spoken speech from background noise and non-speech transients.
In a vehicle environment, an automatic speech recognition (ASR) system can be used to give navigation instructions to passengers based on voice input. This functionality reduces safety considerations in that the driver's attention is not distracted from the road while manually entering or reading information from the screen. In addition, the ASR system can also be used to control voice systems, air conditioning controls, or other vehicle functions.
The ASR system allows the user to speak to the microphone. The ASR system also translates the signal into commands that are recognized by the computer. Upon recognizing the command, the computer can execute the application. One factor in running an ASR system is recognizing exactly what is said. To do this, it is necessary to find the beginning and / or end of the statement (end pointing).
Some systems search for energy in the audio frame. When the energy is detected, the system subtracts a predetermined time from the point where the energy is detected (to determine the start time of the speech), or adds a predetermined time from the point where the energy is detected (the end time of the speech). Predict the endpoint of the statement by (to determine). This selected portion of the audio stream is then passed to the ASR to determine the spoken remark.
The energy in an acoustic signal can come from many sources. In a vehicle environment, for example, acoustic signal energy can be derived from transient noise such as road bumps, door bangs, bumps, bangs, engine noises, motives, and the like. The above systems, which focus on the presence of energy, may misunderstand these transient noises as spoken statements and send them to the ASR system to process the periphery of the signal. As a result, the ASR system unnecessarily attempts to recognize temporary noise as an uttered command, which can generate false positive signals or delay the response to the actual command.
<p> Therefore, there is a need for an intelligent end-pointer system that can identify spoken statements in temporary noise situations.</p>
<p> A rule-based end pointer contains one or more rules that determine the start, end, or both start and end of a speech utterance segment in a speech stream. Rules can be based on a variety of factors, such as the occurrence or combination of events, or the continuation of the presence / absence of speech characteristics. Further, the rule may include analyzing a period of silence, a voiced audio event, an unvoiced audio event or any combination of such events, a continuation of an event, or a continuation of an event. Depending on the rules applied or the content of the audio stream being analyzed, the amount of audio stream that the rule-based end pointer sends to the ASR can vary.</p><p> The dynamic end pointer analyzes one or more dynamic aspects of the speech stream and may determine the start, end, or both start and end of the speech utterance based on the analyzed dynamic aspects. Dynamic aspects that can be analyzed include (1) the audio stream itself, such as the pace of the speaker's speech, the pitch of the speaker's speech, and (2) the expected answer to the question given to the speaker (eg,). , "Yes" or "No"), or (3) ambient noise levels, echoes, and other ambient conditions, but are not limited to these. The rule may utilize one or more dynamic aspects to endpoint the voice utterance segment.</p><p> Other systems, methods, features and advantages of the present invention will be apparent (or will be apparent) to those skilled in the art by considering the drawings and detailed description below. All such additional systems, methods, features and advantages are included herein, within the scope of the present invention, and are intended to be protected by the claims described below.</p><p> The present invention can be better understood with reference to subsequent drawings and description. The elements in the figure are not necessarily the actual size, and are emphasized in order to illustrate the principles of the present invention. Moreover, throughout the various figures, the same reference numbers in the figures indicate corresponding parts.</p>
Rule-based end pointers can consider one or more characteristics of the audio stream to obtain trigger characteristics. Trigger characteristics can include voiced or unvoiced sounds. Voiced speech segments (eg, vowels) generated when the vocal cords vibrate give off a nearly periodic time domain signal. The unvoiced speech sound produced when the vocal cords do not vibrate (such as when speaking the letter "f" in English) has no periodicity and has a time domain signal similar to a noise-like structure. By identifying trigger characteristics in the speech stream and adopting a set of rules that act on the original characteristics of the utterance sound, the end pointer can improve the decision to start and / or end the utterance.
Alternatively, the end pointer can analyze at least one dynamic aspect of the audio stream. The dynamic aspects of the audio stream that can be analyzed include (1) the pace of speech of the speaker, the pitch of the speaker's speech, etc., the audio stream itself, and (2) the expected questions given to the speaker. Possible answers in the audio stream, such as answers (eg, "yes" or "no"), or (3) ambient noise levels, echoes, and other ambient conditions, but are not limited to these. .. Dynamic end pointers can be rule-based. The dynamic nature of the end pointer allows for improved determination of the start and / or end of the utterance segment.
FIG. 1 is a block diagram of a device 100 for performing utterance endpointing based on voice. The end pointing device 100 may include hardware or software that may run on one or more processors associated with one or more operating systems. The end pointing device 100 may include a processing environment 102 such as a computer. The processing environment 102 may include a processing unit 104 and a memory 106. The processing unit 104 can perform calculations and logics and / or control operations by accessing memory 106 via a bidirectional bus. Memory 106 may store the input audio stream. Memory 106 may include a rule module 108 used to detect the start and / or end of a voice utterance segment. The memory 106 may further include a vocal analysis module 116 used to discover the triggering characteristics of the speech segment, and / or an ASR unit 118 that may be used to recognize the speech input. Further, the memory device 106 may store the buffered audio information obtained during the operation of the end pointer. The processing unit 104 communicates with the input / output (I / O) unit 110. The I / O unit 110 receives the input audio stream from the device that converts the sound wave into the electric signal 114, and transmits the output signal to the device that converts the electric signal into the audio sound 112. The I / O unit 110 can act as an interface between the processing unit and a device that converts an electrical signal into an audio sound 112 and a sound wave into an electrical signal 114. The I / O unit 110 may convert an input audio stream received via a device that converts sound waves into an electrical signal 114 from an acoustic waveform into a computer-understandable format. Similarly, the I / O unit 110 may convert the signal transmitted from the processing environment 102 into an electrical signal for output via a device that converts the electrical signal into audio sound 112. The processing unit 104 displays the flowcharts of FIGS. 3 and 4.
FIG. 2 shows an end pointer device 100 incorporated in the vehicle 200. The vehicle 200 may include a driver seat 202, a passenger seat 204 and a rear seat 206. Further, the vehicle 200 may include an end pointer device 100. The processing environment 102 can be incorporated into the vehicle 200's onboard computer, such as an electronic controller, electronic control module, body control module, or the existing circuit of the vehicle 200 using one or more acceptable protocols. It can be a separate unit built in after manufacture that can communicate with. Some of the protocols are under J1850VPW, J1850PWM, ISO, ISO9141-2, ISO14230, CAN, High Speed CAN, MOST, LIN, IDB-1394, IDB-C, D2B, Bluetooth, TTCAN, TTP, or FlexRay . May include protocols traded in. One or more devices that convert electrical signals to sound 112 may be located in the passenger cavity of the vehicle 200, such as the front passenger cavity. A device that converts sound waves to an electrical signal 114, but not limited to this configuration, may be connected to an I / O unit 110 to receive an input audio stream. Alternatively or additionally, an additional device that converts an electrical signal into audio sound 212 to receive an audio stream from passengers in the back seat and output information to these same passengers, and an electrical signal for sound waves. A device to convert to 214 may be placed in the rear passenger cavity of the vehicle 200.
FIG. 3 is a flowchart of the utterance end pointer system. The system can operate by dividing the input audio stream into separate sections, such as frames, so that the input audio stream can be analyzed frame by frame. Each frame can contain any position from about 10 ms to about 100 ms of the entire input audio stream. The system may buffer a predetermined amount of input audio data, such as about 350 ms to about 500 ms, before starting to process the data. As shown in block 302, an energy detector can be used to determine if energy is present in addition to noise. The energy detector examines a portion of the audio stream, such as a frame, to determine the amount of energy present and compares the amount with a noise energy assessment. The evaluation of noise energy may be constant or dynamically determined. The difference in decibels (dB), or power ratio, can be the instantaneous signal-to-noise ratio (SNR). Prior to analysis, the frame can be assumed to be non-utterance, and as a result, if the energy detector determines that energy is present within the frame, the frame is marked as non-utterance, as shown in block 304. .. After the energy is detected, the frame, as shown in block 306<sub>n</sub>A vocal analysis of the current frame, shown as, can be performed. Vocal analysis can be performed as described in US Patent Application No. 11 / 131,150 filed May 17, 2005. The specification of the application is incorporated herein by reference. Vocal analysis is a frame<sub>n</sub>You can check any trigger characteristics that can be present in. In vocal analysis, voice "S" or "X" is the frame<sub>n</sub>You can check if it exists in. Alternatively, vocal analysis can check for the presence of vowels. For the purpose of explanation rather than limitation, the rest of FIG. 3 will be described as using vowels as the triggering characteristic of vocal analysis.
There are various ways in which vocal analysis can identify the presence of vowels in a frame. One method involves the use of a pitch estimator. The pitch estimator indicates that vowels can be present and can search for periodic signals in the frame. Alternatively, the pitch estimator can search the frame for a given level of natural frequency. The pitch estimator may indicate the presence of vowels.
Vowel is the frame<sub>n</sub>If the vocal analysis determines that it is within the frame<sub>n</sub>Is marked as an utterance, as shown in block 310. The system can then consider one or more earlier frames. As shown in block 312, the system is the preceding frame, the frame<sub>n-1</sub>Can be examined. The system may determine if the previous frame was previously marked as containing an utterance, as shown in block 314. If the previous frame was already marked as an utterance (ie the "yes" answer to block 314), the system has already determined that the utterance is contained within the frame, as shown in block 304. , Move on to the analysis of new audio frames. If the previous frame was not marked as an utterance (ie, the "no" answer to block 314), the system may use one or more rules to determine if the frame is marked as an utterance.
As shown in FIG. 3, block 316, shown as the decision block "external endpoint", may use a routine to determine if a frame is marked as an utterance using one or more rules. One or more rules can be applied to any part of the audio stream, such as a frame or group of frames. The rules may determine if the current frame under investigation contains utterances. Rules can indicate whether there is an utterance in a frame or group of frames or not. If the utterance is present, the frame can be designated as being within the endpoint.
If the rule indicates that there is no utterance, the frame can be specified as outside the endpoint. flame<sub>n-1</sub>If block 316 indicates that is outside the endpoint (eg, there is no utterance), then a new audio frame (frame), as shown in block 304.<sub>n + 1</sub>) Is entered into the system and marked as not spoken. flame<sub>n-1</sub>If block 316 indicates that is in the endpoint (eg, an utterance exists), then the frame, as shown in block 318.<sub>n-1</sub>Is marked as an utterance. As shown in block 320, the previous audio stream can be analyzed frame by frame until the last frame in memory is analyzed.
FIG. 4 is a more detailed flowchart of the block 316 shown in FIG. As mentioned earlier, block 316 can contain one or more rules. The rule may relate to any aspect of the presence and / or absence of utterances. In this way, the rules can be used to determine the start and / or end of the spoken remark.
The rule is an event (eg voiced energy, unvoiced energy, presence and / or absence of silence) or any combination of events (eg, followed by silence followed by voiced energy, unvoiced energy, followed by unvoiced energy). It can be based on the analysis of silence followed by silence, silence, etc.). Specifically, the rule may consider the transition from a period of silence to an energy event or the transition from a period of silence to an energy event. By the rule that an utterance can contain one or less transitions from an unvoiced event or silence before a vowel, the rule can analyze the number of transitions before a vowel. Alternatively, the rule can analyze the number of transitions after a vowel by the rule that an utterance can contain no more than two transitions from an unvoiced event or silence after a vowel.
One or more rules can examine different durations. Specifically, the rule may examine the continuation of an event (eg, voiced energy, unvoiced energy, presence and / or absence of silence). By the rule that an utterance can include a duration in the range of about 300 ms to 400 ms before a vowel and can be about 350 ms, the rule can analyze the duration before a vowel. Alternatively, the rule can analyze the duration after a vowel by the rule that the utterance can include a duration in the range of about 400 ms to 800 ms after the vowel and can be about 600 ms.
One or more rules can examine the duration of an event. Specifically, the rules can look for certain energy periods or energy shortages. Unvoiced energy is a type of energy that can be analyzed. By the rule that an utterance can include a continuation of continuous unvoiced energy in the range of about 150 ms to 300 ms and can be about 200 ms, the rule can analyze the continuation of continuous unvoiced energy. .. Alternatively, continuous silence can be analyzed as a lack of energy. By the rule that an utterance can include a continuation of continuous silence in the range of about 50 ms to 80 ms before a vowel and can be about 70 ms, the rule is a continuous silence before a vowel. The continuation of can be analyzed. Alternatively, by the rule that an utterance can include a continuation of continuous silence in the range of about 200 ms to 300 ms after a vowel and can be about 250 ms, the rule is a continuous silence after a vowel. The continuation of can be analyzed.
At block 402, a check is made to determine if the frame or group of frames being analyzed has energy above the background noise level. Frames or groups of frames with energies above background noise levels can be further analyzed based on certain energy continuations or event-related continuations. If the frame or group of frames under analysis does not have energy above the background noise level, then the frame or group of frames is a continuation of continuous silence, a transition from a period of silence to an energy event, or a silence. It can be further analyzed based on the transition from period to energy event.
In block 404, the "energy" counter is incremented if energy is present in the frame or group of frames being analyzed. The "energy" counter counts the amount of time. The amount of time increases by the frame length. If the frame size is about 32 ms, the block 404 counts "energy" as about 32 ms. In decision 406, the check is made to see if the "energy" counter value exceeds the time threshold. The threshold evaluated in decision block 406 corresponds to a continuous non-spoken energy rule that can be used to determine the presence and / or absence of speech. In decision block 406, a threshold can be evaluated for the maximum continuation of continuous unspoken energy. If the determination 406 determines that the "energy" counter value exceeds the threshold setting, then the frame or group of frames being analyzed is considered to be outside the endpoint in block 408 (eg, no utterances). It is specified. As a result, referring to Figure 3 again, the system jumps to block 304, where it is a new frame, the frame.<sub>n + 1</sub>Is entered into the system and marked as not spoken. Alternatively, multiple thresholds may be evaluated in block 406.
If the "energy" counter value does not exceed the time threshold in block 406, a check is made in block 410 to determine if the "no energy" counter exceeds the separation threshold. Like the "energy" counter 404, the "no energy" counter 418 counts the time and is incremented by the frame length if the frame or group of frames being analyzed does not have energy above the noise level. The separation threshold is a time threshold that defines the amount of time between two plosive events. Plosives are literally consonants lined up from the speaker's mouth. The momentary blockage of air creates pressure to make a plosive sound. Examples of the plosive sound include sounds "P", "T", "B", "D", and "K". This threshold can be in the range of about 10 ms to about 50 ms and can be about 25 ms. If the separation threshold is exceeded, a separated unvoiced energy event, i.e. a plosive surrounded by silence (eg, P of STOP), is identified and the "separation event" counter 412 is incremented. The "separation event" counter 412 is incremented by an integer value. After incrementing the "separation event" counter 412, the "no energy" counter 418 is reset at block 414. This counter is reset because energy was found in the frame or group of frames being analyzed. If the "no energy" counter 418 does not exceed the separation threshold, the "no energy" counter 418 is reset at block 414 without incrementing the "separation event" counter 412. Again, the "no energy" counter 418 is reset because energy was found in the frame or group of frames being analyzed. By returning a value of "No" in block 416 after resetting the "No Energy" counter 418, an analysis outside the endpoint has a frame or group of frames being analyzed within the endpoint (eg, utterances). (Exists). As a result, see Figure 3.
Alternatively, if the determination 402 determines that there is no energy above the noise level within the frame or group of frames being analyzed, then the frame or group of frames being analyzed contains silence or background noise. In this case, the "no energy" counter 418 is incremented. In determination 420, a check is made to see if the "no energy" counter value exceeds the time threshold. The threshold evaluated in decision block 420 corresponds to a continuous silent energy rule threshold that can be used to determine the presence and / or absence of utterances. In decision block 420, the threshold for continuation of continuous silence may be evaluated. If the determination 420 determines that the "no energy" counter value exceeds the threshold setting, then the frame or group of frames being analyzed is considered to be outside the endpoint in block 408 (eg, no utterances). It is specified. As a result, referring to Figure 3 again, the system jumps to block 304, where it is a new frame, the frame.<sub>n + 1</sub>Is entered into the system and marked as not spoken. Alternatively, a number of thresholds may be evaluated in block 420.
If the "no energy" counter 418 does not exceed the time threshold, a check is made in decision block 422 to determine if the maximum number of separation events allowed has occurred. The "separation event" counter provides the information needed to answer this check. The maximum number of separation events allowed is a configurable parameter. If the grammar is expected (eg, "yes" or "no" answer), the maximum number of segregation events allowed can be set accordingly to "narrow down" the end pointer results. If the maximum number of segregation events allowed is exceeded, then the frame or group of frames being analyzed is designated in block 408 as being out of the endpoint (eg, no utterance). As a result, referring to Figure 3 again, the system jumps to block 304, where it is a new frame, the frame.<sub>n + 1</sub>Is entered into the system and marked as not spoken.
If the maximum number of separation events allowed has not been reached, the "energy" counter 404 is reset at block 424. The "energy" counter 404 may be reset if a frame in which no energy is present is identified. By returning a value of "No" in block 416 after resetting the "Energy" counter 404, an analysis outside the endpoint has a frame or group of frames being analyzed within the endpoint (eg, utterances are present). ). As a result, referring to FIG. 3, the system marks the analyzed frame as an utterance at 318 or 322.
Figures 5-9 show some actual time series of simulated speech streams, various characteristic plots of these signals and a spectrograph of the corresponding actual signals. In FIG. 5, block 502 shows the actual time series of the simulated audio stream. The simulated audio stream includes spoken statements "No" 504, "Yes" 506, "No" 504, "YES" 506, "NO" 504, "YESSSSS" 508, "NO" 504 and many "NO" 504. Includes a click 510. These clicks may represent the sounds produced when the vehicle's turn signals are used. Block 512 shows various characteristic plots for the actual time series audio stream. Block 512 displays the number of samples along the X axis. Plot 514 is one representation of the end pointer analysis. If plot 514 is at level 0, the end pointer has not determined the presence of spoken remarks. If plot 514 is at a non-zero level, the end pointer indicates the boundary between the start and / or end of the spoken speech. Plot 516 represents energy that exceeds the background energy. Plot 518 represents a statement spoken in the time domain. Block 520 shows a spectral representation of the corresponding audio stream identified in block 502.
Block 512 shows how the end pointer can respond to the input audio stream. As shown in FIG. 5, the end pointer plot 514 accurately captures the "NO" 504 and "YES" 506 signals. When the "YESSSSS" 508 is analyzed, the end pointer plot 514 captures an extended "S" for some time, but when it finds that it exceeds the maximum time after a vowel or the maximum duration of continuous unvoiced energy, The end pointer is cut. The rule-based end pointer sends a portion of the audio stream bounded by the end pointer plot 514 to the ASR. As shown in block 512 and FIGS. 6-9, the portion of the audio stream sent to the ASR varies depending on the rules applied. The "click" 510 was detected as having energy. This is represented by the background energy plot 516 on the far right of block 512. However, no vowels were detected in the "click" 510, so the end pointer excludes these vowels.
Figure 6 is a close-up of one end-pointed "NO" 504. Due to time smearing, the spoken speech plot 518 is delayed by one or two frames. Plot 518 continues for the duration of the energy detection and is represented by the energy plot 516 above. As the spoken speech plot 518 rises, it level off and continues to the background energy plot 516 above. The end pointer plot 514 starts when speech energy is detected. During the period represented by plot 518, neither end pointer rule is violated and the audio stream is recognized as spoken speech. The end pointer breaks at the far right if either the maximum continuation rule for continuous silence after a vowel or the maximum time rule after a vowel may have been violated. As shown, a portion of the audio stream sent to the ASR contains approximately 3150 samples.
Figure 7 is a close-up of one end-pointed "YES" 506. Again, due to time smearing, the spoken speech plot 518 is delayed by one or two frames. The end pointer plot 514 starts when energy is detected. The end pointer plot 514 continues until the energy drops to noise, i.e., until the maximum continuation rule or maximum time rule of continuous silence after the vowel is violated. As shown, a portion of the audio stream sent to the ASR contains approximately 5550 samples. The difference in the amount of audio streams sent to the ASR in Figures 6 and 7 is caused by the end pointers that provide different rules.
Figure 8 is a close-up of one end-pointed "YESSSSS" 508. The end pointer recognizes the energy after the vowel as a possible consonant, but only for a reasonable amount of time. After a reasonable amount of time, the maximum continuation or maximum time rule of continuous unvoiced energy after a vowel may have been broken, and the pointer diminishes by limiting the data passed to the ASR. As shown, a portion of the audio stream sent to the ASR contains approximately 5750 samples. The spoken remarks continue for the 6500 samples to be burned, but the amount of audio stream sent to the ASR is different from that sent in Figures 6 and 7 because the end pointer is interrupted after a reasonable amount of time. different.
Figure 9 is a close-up of one "NO" 504, end-pointed, followed by several "clicks" 510. Similar to Figures 6-8, time smearing delays the spoken speech plot 518 by one or two frames. The end pointer plot 514 starts when energy is detected. The first click is contained within the endpoint plot 514 because there is energy above the background noise energy level, and this energy can be a consonant (ie, an extended "T"). However, there is about 300 milliseconds of silence between the first click and the next click. According to the threshold used in this example, this period of silence breaks the maximum continuation of continuous silence after a vowel. Therefore, the end pointer excluded the energy after the first click.
The end pointer can also be configured to determine the start and / or end of a speech speech segment by analyzing at least one dynamic aspect of the speech stream. FIG. 10 is a partial flow chart of an end pointer system that analyzes at least one dynamic aspect of an audio stream. Initialization of the global phase can be done at 1002. The global aspect may include the characteristics of the audio stream itself. For the purpose of explanation rather than limitation, these global aspects include the pace of speaker utterances or the pitch of speaker utterances. Initialization of the local aspect can be done at 1004. For the purpose of explanation rather than limitation, these local aspects are expected speaker answers (eg "yes" or "no") ambient conditions (echo or feedback in the system). An open or closed environment that affects the presence of the), or an assessment of background noise.
Global and local initialization can occur many times throughout the operation of the system. Background noise assessment (initialization of local aspects) can be done each time the system is started and / or after a given time. Determining the pace or pitch of a speaker's utterance (global initialization) can be initialized at a lower rate. Similarly, the local aspect where a particular response is expected is initialized at a lower rate. Similarly, this initialization can occur when the ASR communicates with an end pointer where an answer is expected. Local aspects of ambient conditions can be configured to initialize only once per power cycle.
During the initialization periods 1002 and 1004, the end pointer may operate with its default threshold settings as described above with respect to FIGS. 3 and 4. If any of the initial settings require threshold setting or timer changes, the system may dynamically change the appropriate limits. Alternatively, the system may call the profile of a particular user or general user previously stored in the system's memory based on the default values. This profile can change all or specific threshold settings or timers. If, during the initialization process, the system decides that the user speaks at a fast pace, the maximum duration of a rule can be at the level stored in the profile. In addition, it may be possible to operate the system in training mode, where the system performs initialization to create a user profile and store it for later use. One or more profiles may be stored in system memory for later use.
A dynamic end pointer similar to the end pointer described in FIG. 1 may be configured. In addition, the dynamic end pointer can include a bidirectional bus between the processing environment and the ASR. The bidirectional bus can send data and control information between the processing environment and the ASR. The information passed from the ASR to the processing environment may include data indicating a certain response that is expected in response to the question given to the speaker. The information passed from the ASR to the processing environment can be used to dynamically analyze aspects of the audio stream.
The behavior of dynamic end pointers has been described with respect to FIGS. 3 and 4, except that one or more thresholds of one or more rules of the "out-of-endpoint" routine (block 316) can be set dynamically. Can resemble an end pointer. In the presence of large amounts of background noise, the threshold for energy beyond the noise determination (block 402) can be dynamically increased to account for this condition. When doing this reset, the dynamic end pointer can reject more transient and non-spoken sounds, thereby reducing the number of false positive signals. The dynamically set threshold is not limited to the background noise level. Any threshold used by the dynamic end pointer can be set dynamically.
The methods shown in FIGS. 3, 4 and 10 can be encoded on a computer-readable medium such as a signal bearing medium, memory, programmed into a device such as one or more integrated circuits, or processed by a controller or computer. When the method is performed by software, the software resides in memory residing in rule module 108 or is interfaced through any type of communication interface. The memory may contain an ordered list of executable instructions for implementing a logical function. Logical functions can be implemented via analog sources, such as via digital circuits, via source code, via analog circuits, or via electrical, audio or video signals. The software may be embodied in any computer-readable or signal bearing medium for use by, or in combination with, a system, device or device capable of executing instructions. Such systems may include computer-based systems, systems that include processors, systems that can execute instructions, or other systems that can also execute instructions and selectively derive instructions from a device or device.
"Computer-readable media", "machine-readable media", "propagation signal" media, and / or "signal bearing media" are used by or in combination with instruction-executable systems, devices or equipment. It may include any means of including, storing, communicating, disseminating, or transferring software. Machine-readable media can optionally be, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, equipment, or propagation media. To enumerate non-limiting examples of machine-readable media examples, "electronic" electrical connections with one or more wires, portable magnetic disks or optical disks, random access memory "RAM" (electronic), Examples include read-only memory "ROM" (electronic), erasable programmable read-only memory (EPROM or flash memory (electronic)), or volatile memory such as optical fiber (optical). Machine-readable media are machine-readable because the software is stored electronically, compiled, and / or interpreted or processed as an image or in other format (via optical scanning). It can also include tangible media that can be made. The processed medium can then be stored in computer and / or machine memory.
Although various embodiments of the present invention have been described, it will be apparent to those skilled in the art that more embodiments and implementations are possible within the scope of the present invention. Therefore, the present invention can only be limited in consideration of the accompanying claims and their equivalents.
<figref num="1">FIG. 1 is a block diagram of a speech end pointing system.</figref><figref num="2">FIG. 2 is a partial illustration of a speech end pointing system built into a vehicle.</figref><figref num="3">FIG. 3 is a flowchart of the speech end pointer.</figref><figref num="4">FIG. 4 is a more detailed flowchart of a portion of FIG.</figref><figref num="5">Figure 5 shows the end-pointing of the simulated utterance sound.</figref><figref num="6">FIG. 6 is a detailed end-pointing of some of the simulated utterance sounds of FIG.</figref><figref num="7">FIG. 7 is a second detailed end-pointing of some of the simulated utterance sounds of FIG.</figref><figref num="8">FIG. 8 is a third detailed end-pointing of some of the simulated utterance sounds of FIG.</figref><figref num="9">FIG. 9 is a fourth detailed end-pointing of some of the simulated utterance sounds of FIG.</figref><figref num="10">FIG. 10 is a partial flow chart of a dynamic speech end pointing system based on voice.</figref>
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2017078869A | Cited by | Japan | Search report |
| JP2017078869A | Cited by | Japan | Search report |
| US11551709B2 | Cited by | United States of America | Applicant |
| US10929754B2 | Cited by | United States of America | Applicant |
| US10269341B2 | Cited by | United States of America | Applicant |
| US11062696B2 | Cited by | United States of America | Applicant |
| US10593352B2 | Cited by | United States of America | Applicant |
| JP2013545133A | Cited by | Japan | Examiner |
| JP2017078869A | Cited by | Japan | Search report |
| US11676625B2 | Cited by | United States of America | Applicant |
| US11710477B2 | Cited by | United States of America | Applicant |
| JP2017078848A | Cited by | Japan | Search report |
| US9330667B2 | Cited by | United States of America | Applicant |
| WO2004111996A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
21 members in 7 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 11152922 | United States of America | – | |
| 15292205 | United States of America | A | |
| 15292205 | United States of America | A | |
| 2006000512 | Canada | W | |
| 2006000512 | Canada | W | |
| 2005152922 | – | – | – |
| 2006000512 | – | – | – |
| US20050152922 | – | – | – |
| WO2006CA00512 | – | – | – |
Members21
| Document | Office | Kind | |
|---|---|---|---|
| CA2575632A1 | Canada | A1 | |
| US2006287859A1 | United States of America | A1 | |
| WO2006133537A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1771840A1 | European Patent Office (EPO) | A1 | |
| KR20070088469A | Republic of Korea | A | |
| CN101031958A | China | A | |
| EP1771840A4 | European Patent Office (EPO) | A4 | |
| US2007288238A1 | United States of America | A1 | |
| JP2008508564AThis record | Japan | A | |
| US2008228478A1 | United States of America | A1 | |
| JP2011107715A | Japan | A | |
| US8165880B2 | United States of America | B2 | |
| US8170875B2 | United States of America | B2 | |
| CN101031958B | China | B | |
| US2012265530A1 | United States of America | A1 | |
| US8311819B2 | United States of America | B2 | |
| US2012303366A1 | United States of America | A1 | |
| CA2575632C | Canada | C | |
| US8457961B2 | United States of America | B2 | |
| US8554564B2 | United States of America | B2 | |
| JP5331784B2 | Japan | B2 |
8 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Notification of change in applicantJAPANESE INTERMEDIATE CODE: A711A711 | A711 | |
| Re-examination (zenchi) completed and case transferred to appeal boardAppealJAPANESE INTERMEDIATE CODE: A912A912 | A912 | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of acceptance of power of attorneyJAPANESE INTERMEDIATE CODE: A7422RD02 | RD02 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 |
Numbers
- Publication
- 2008508564
- Publication, DOCDB
- 2008508564
- Publication, EPODOC
- JP2008508564
- Application
- 2007524151
- Application, DOCDB
- 2007524151
- Application, EPODOC
- JP20070524151
Titles2
- Japanese
- スピーチエンドポインタ
- English
- Speech end pointer
Classification
- CPC, 1
- G10L25/87
- IPC, 2
- G10L15 04
- G10L25 93
Designated states4
- Regional, 4
- Zimbabwe
- Turkmenistan
- Türkiye
- Togo