Audio/video synchronization processing device, terminal, audio/video synchronization processing method and program
Abstract
Problem to be solved.To appropriately switch a function for synchronizing an audio signal and a video signal according to a situation of a conference so that the audio signal and the video signal can be optimally controlled according to a scene. An audio-video synchronization processing device according to the present invention is an audio-video synchronization processing means that synchronizes a video signal and an audio signal based on a received video signal and time information included in the received audio signal, and a transmitted audio signal. The first voice determination means for determining the voice detection section or the voice non-detection section in the above, the second voice determination means for determining the voice detection section or the voice non-detection section in the received voice signal, the first voice detection means, and the second voice above. It is provided with a synchronization control means that dynamically controls synchronization by the audio / video synchronization processing means according to the audio detection result by the audio detection means. [Selection diagram] Fig. 1

Term
Projected expiry 27 August 2035.
- Priority and filed
- Published
- Today
- Projected expiry
8 claims: 3 independent, 5 dependent
- 1受信映像信号及び受信音声信号に含まれる時間情報に基づいて、映像信号及び音声信号の同期を行う音声映像同期処理手段と、 送信音声信号における音声検出区間又は音声非検出区間を判定する第1音声判定手段と、 上記受信音声信号における音声検出区間又は音声非検出区間を判定する第2音声判定手段と、 上記第1音声検出手段及び上記第2音声検出手段による音声検出結果に応じて、上記音声映像同期処理手段による同期を動的に制御する同期制御手段と を備えることを特徴とする音声映像同期処理装置。
- 2上記同期制御手段が、上記送信音声信号の音声検出区間と、上記受信音声信号の音声検出区間との非重複のときに、上記同期を有効にすることを特徴とする請求項1に記載の音声映像同期処理装置。
- 3上記同期制御手段が、上記送信音声信号の音声検出区間と上記受信音声信号の音声検出区間との重複区間で、上記同期を無効にすることを特徴とする請求項1に記載の音声映像同期処理装置。
- 4上記同期制御手段が、上記送信音声信号の音声検出区間と上記受信音声信号の音声検出区間との重複時点で上記同期を無効にし、所定の有効復帰条件に従って、上記同期を有効にすることを特徴とする請求項1に記載の音声映像同期処理装置。
- 5上記同期制御手段は、 上記受信音声信号の音声非検出区間長が所定時間を経過しているときに、次の受信音声信号の音声検出時に上記同期を有効にすること または、上記送信音声信号の音声非検出区間長が所定時間を経過しているときであり、上記受信音声信号の音声非検出区間長が所定時間を経過しているときに、次の受信音声信号の音声検出時に上記同期を有効にすること を特徴とする請求項4に記載の音声映像同期処理装置。
- 6映像信号及び音声信号を含むメディア情報を授受して、映像及び音声を出力する端末において、 請求項1~5のいずれかに記載の音声映像同期処理装置を備えることを特徴とする端末。
- 7音声映像同期処理装置の音声映像同期処理方法であって、 上記音声映像同期処理装置は、 受信映像信号及び受信音声信号に含まれる時間情報に基づいて、映像信号及び音声信号の同期を行う音声映像同期処理ステップと、 送信音声信号における音声検出区間又は音声非検出区間を判定する第1音声判定ステップと、 上記受信音声信号における音声検出区間又は音声非検出区間を判定する第2音声判定ステップと、 上記第1音声検出手段及び上記第2音声検出手段による音声検出結果に応じて、上記音声映像同期処理手段による同期を動的に制御する同期制御ステップと を備えることを特徴とする音声映像同期処理方法。
- 8コンピュータを、 受信映像信号及び受信音声信号に含まれる時間情報に基づいて、映像信号及び音声信号の同期を行う音声映像同期処理手段と、 送信音声信号における音声検出区間又は音声非検出区間を判定する第1音声判定手段と、 上記受信音声信号における音声検出区間又は音声非検出区間を判定する第2音声判定手段と、 上記第1音声検出手段及び上記第2音声検出手段による音声検出結果に応じて、上記音声映像同期処理手段による同期を動的に制御する同期制御手段と して機能させることを特徴とする音声映像同期処理プログラム。
Independent claims8
91 paragraphs, as filed
0001The present invention relates to an audio / video synchronization processing device, a terminal, an audio / video synchronization processing method and a program, and is applied to, for example, an audio / video synchronization processing device that synchronizes an audio signal and a video signal in an H.323 TV conference system. What you get.
0002In a video conferencing system that holds a conference through a network, there is a lip-sync function as a function to synchronize video and audio.
0003For example, in an H.323-compliant video conferencing system, the video signal is generally delayed as compared with the audio signal because the delay time associated with the coding of the audio signal and the video signal is different. Therefore, there is a discrepancy between the video and audio of the speaker. In order to solve such a problem, the conference terminal of the video conference system (in Patent Document 1 described later, the main body and the voice terminal are supported) is enabled by enabling the lip sync function to be the partner of the conference of the conference terminal. Regarding the video signal and audio signal acquired by the other party's conference terminal (corresponding to the remote terminal in Patent Document 1 described later), the audio signal is delayed at the time of playback on the main unit and the audio terminal on the receiving side. And output. Therefore, the deviation between the video and the audio is reduced.
<p num="0004"><patcit num="1"><text>JP-A-2002-290938</text></patcit></p>
<p num="0005"> However, the lip-sync function in the conventional video conferencing system is to enable or disable the function prior to the start of the call, and it is not possible to switch the enable or disable during the call.</p><p num="0006"> When the lip sync function is enabled, the audio signal is also delayed and reproduced as described above, so that the audio signal and the video signal related to the conference are always delayed.</p><p num="0007"> On the other hand, consider the effects of the above delays on the conference, assuming the following two situations.</p><p num="0008"> One is when someone keeps speaking and the other participants are listening to it, as in a presentation.</p><p num="0009"> The other is a scene where multiple participants speak to each other or at the same time, as in a discussion.</p><p num="0010"> In these two scenes, the difference between whether lip-sync is effective or not depending on the situation, such as the scene where synchronization between the audio signal and the video signal is important and the scene where the small delay of the audio signal is important. There is.</p><p num="0011"> At present, it is not possible to judge the status of the meeting and dynamically switch the lip-sync function. Therefore, the user had to decide in advance whether or not to use the lip-sync function in consideration of the nature of the conference.</p><p num="0012"> Therefore, the audio / video synchronization processing device and terminal capable of appropriately switching the function for synchronizing the audio signal and the video signal according to the situation of the conference and optimally controlling the audio signal and the video signal according to the scene. , Audio-video synchronization processing programs and information processing terminals are required.</p>
<p num="0013"> In order to solve such a problem, the audio / video synchronization processing device according to the first aspect of the present invention is (1) between the video signal and the audio signal based on the received video signal and the time information contained in the received audio signal. Audio-video synchronization processing means for synchronization, (2) first audio determination means for determining audio detection section or audio non-detection section in transmitted audio signal, and (3) audio detection section or audio non-detection section in received audio signal. A second audio determination means for determining the above, and (4) a synchronization control means for dynamically controlling synchronization by the audio / video synchronization processing means according to the audio detection result by the first audio detection means and the second audio detection means. It is characterized by having.</p><p num="0014"> The second terminal according to the present invention is a terminal that receives and receives media information including video signals and audio signals and outputs video and audio, and includes the audio-video synchronization processing device according to the first invention. It is characterized by.</p><p num="0015"> The third audio-video synchronization processing method according to the present invention is the audio-video synchronization processing method of the audio-video synchronization processing device, and the audio-video synchronization processing device is included in (1) the received video signal and the received audio signal. An audio-video synchronization processing step that synchronizes between an audio signal and an audio signal based on time information, and (2) a first audio determination step that determines an audio detection section or an audio non-detection section in the transmitted audio signal. (3) The second voice determination step for determining the voice detection section or the voice non-detection section in the received voice signal, and (4) the voice according to the voice detection results by the first voice detection step and the second voice detection step. It is characterized by including a synchronization control step that dynamically controls synchronization by the video synchronization processing step.</p><p num="0016"> The fourth audio-video synchronization processing program according to the present invention synchronizes a computer between a video signal and an audio signal based on (1) the received video signal and the time information contained in the received audio signal. Synchronous processing means, (2) a first voice determination means for determining a voice detection section or a voice non-detection section in a transmitted voice signal, and (3) a second voice determination section for determining a voice detection section or a voice non-detection section in a received voice signal. It is characterized by functioning as a voice determination means and (4) a synchronization control means for dynamically controlling synchronization by a voice-video synchronization processing means according to a voice detection result by the first voice detection means and the second voice detection means. And.</p>
<p num="0017"> According to the present invention, it is possible to appropriately switch the function of synchronizing the audio signal and the video signal according to the situation of the conference, and optimally control the audio signal and the video signal according to the scene.</p>
0018<figref num="1">It is an internal block diagram which shows the internal structure of the conference terminal which concerns on embodiment.</figref><figref num="2">It is a flowchart which shows the voice determination process which concerns on embodiment.</figref><figref num="3">It is explanatory drawing explaining that the phoneme for the voice signal which concerns on embodiment is obtained.</figref><figref num="4">It is a state transition diagram which shows the operation of the synchronization control processing by the synchronization control unit which concerns on embodiment.</figref><figref num="5">It is a timing chart which shows the synchronization control processing by the synchronization control part which concerns on embodiment (the 1).</figref><figref num="6">It is a timing chart which shows the synchronization control processing by the synchronization control part which concerns on embodiment (the 2).</figref><figref num="7">It is a timing chart which shows the synchronization control processing by the synchronization control part which concerns on embodiment (the 3).</figref>
0019(A) Main embodiment Hereinafter, embodiments of an audio / video synchronization processing device, a terminal, an audio / video synchronization processing method, and a program according to the present invention will be described in detail with reference to the drawings.
0020For example, a case where the present invention is applied to a conference terminal of an H.323 compliant video conferencing system as a communication protocol will be described as an example.
0021(A-1) Configuration of Embodiment FIG. 1 is an internal configuration diagram showing an internal configuration of a conference terminal according to an embodiment. The hardware of the conference terminal 1 has, for example, circuits such as a CPU, ROM, RAM, EEPROM, an input / output interface unit, and a communication unit. Further, the function as the conference terminal 1 is realized by executing the processing program (audio / video synchronization processing program) stored in the ROM by the CPU. It should be noted that the processing program (audio / video synchronization processing program) may be installed to be constructed, and even in that case, the processing program to be executed is shown as shown in FIG.
0022In FIG. 1, the conference terminal 1 collects voice signals (hereinafter referred to as transmission voice signals [conference terminal 1 side]) emitted by conference participants who use the conference terminal 1 via their respective connection interface units. The voice signal of the other party's conference participant input by the other party's conference terminal, which is the other party of the conference of the voice input device 2 such as a microphone and the conference terminal 1 (hereinafter referred to as the received voice signal [the other party]]. ), A video output device such as a display that outputs the video signal of the other party's conference participant (hereinafter referred to as the received video signal [the other party]) input by the other party's conference terminal. It can be connected to 4. Although not shown in FIG. 1, the conference terminal 1 outputs an imaging device that images conference participants and gives video data to the conference terminal 1, table data, moving image data, and the like used in the conference. It may be possible to connect to an information processing device such as a personal computer.
0023Further, in FIG. 1, the conference terminal 1 according to the embodiment is roughly classified into an audio / video synchronization control unit 10, a transmission audio signal processing unit 18, a received audio signal processing unit 19, and a received video signal processing unit 20.
0024The transmission audio signal processing unit 18 receives and transmits the audio signal input from the audio input device 2 connected to the conference terminal 1 as the transmission audio signal [conference terminal 1 side] via the audio / video synchronization control unit 10. Sends a packet containing an audio signal [conference terminal 1 side] to the network.
0025The received audio signal processing unit 19 extracts the audio signal received by the other party's conference terminal included in the packet received via the network as the received audio signal [other party], and synchronizes the received audio signal [other party] with the audio / video. Send to control unit 10.
0026The received video signal processing unit 20 extracts the video signal received by the other party's conference terminal included in the packet received via the network as the received video signal [other party], and synchronizes the received video signal [other party] with audio / video. Send to control unit 10.
0027The audio / video synchronization control unit 10 receives the audio signal input from the audio input device 2 as the transmitted audio signal [conference terminal 1 side], receives the received audio signal [other party] from the received audio signal processing unit 19, and receives the received video. Receives the received video signal [other party] from the signal processing unit 20. The audio / video synchronization control unit 10 receives the received audio signal [other party] and the received video signal [other party] output from the conference terminal 1 based on the transmitted audio signal [conference terminal 1 side] and the received audio signal [other party]. ], The received audio signal [other party] and the received video signal [other party] are synchronized, and the output timing of the received audio signal [other party] and the received video signal [other party] is controlled (or synchronized). And do not control the output timing).
0028When the audio / video synchronization control unit 10 detects only the received audio signal [other party] based on the input audio signal [conference terminal 1 side] and the received audio signal [other party], the audio / video as a lip sync function Controls the output timing of the received audio signal [other party] and the received video signal [other party] with synchronization processing enabled, and detects the transmitted audio signal [conference terminal 1 side] and the received audio signal [other party] at about the same time. If this happens, the audio / video synchronization processing is disabled and the output timing of the received audio signal [other party] and the received video signal [other party] is not controlled. That is, when the audio-visual synchronization control unit 10 determines the discussion status in a multipoint conference based on the speech status of the conference participants, and it seems that only the conference participants of the other conference terminal are speaking. It is possible to synchronize the received audio signal [other party] and the received video signal [other party], and conversely, the conference participants of the other party's conference terminal and the conference participants of the conference terminal 1 are discussing at almost the same time. Occasionally, synchronization between the received audio signal [other party] and the received video signal [other party] is not performed (or the accuracy of synchronization between the received audio signal [other party] and the received video signal [other party]]. Drop). The audio / video synchronization control unit 10 outputs the received audio signal [other party] processed based on the enable / disable of the audio / video synchronization processing to the audio output device 3, and outputs the received video signal [other party] to the video output device 4. Output to. As a result, when only the conference participants of the other party's conference terminal speak, for example, in a presentation or a conference report, the sound can be output according to the movement of the speaker's mouth projected on the video, and the conference terminal can be output. When exchanging discussions between one conference participant and the conference participant of the other party's conference terminal, priority can be given to audio output with less delay.
0029As shown in FIG. 1, the audio / video synchronization control unit 10 includes an audio / video synchronization processing unit 11, a first audio distribution unit 12, a first audio / non-audio determination unit 13, a synchronization control unit 14, and a second audio distribution unit 15. , The second voice / non-voice judgment unit 16, and the phoneme database 17.
0030The first audio distribution unit 12 receives the transmission audio signal input from the audio input device 2 as the transmission audio signal [conference terminal 1 side], and duplicates and distributes the transmission audio signal [conference terminal 1 side]. In order to transmit the transmitted audio signal [conference terminal 1 side] to the other party conference terminal, the first audio distribution unit 12 sends one of the transmitted audio signals [conference terminal 1 side] to the transmission audio signal processing unit 18 and also conferes. The other transmitted voice signal [conference terminal 1 side] is sent to the first voice / non-voice determination unit 13 in order to judge the discussion status of.
0031The first voice / non-voice determination unit 13 receives a transmitted voice signal [conference terminal 1 side] from the first voice distribution unit 12 in order to detect the speech status of the conference participants on the conference terminal 1 side, and the phoneme database 17 Refers to to detect whether or not the transmitted voice signal [conference terminal 1 side] contains voice (phonemes). The first voice / non-voice determination unit 13 sends a voice determination result based on the transmitted voice signal [conference terminal 1 side] to the synchronization control unit 14. That is, the first voice / non-voice determination unit 13 refers to the phoneme database 17 and determines whether or not the transmitted voice signal [conference terminal 1 side] contains phonemes as words. As the method of voice judgment processing by the first voice / non-voice judgment unit 13, various methods can be widely applied as long as it can detect whether or not the conference participant on the conference terminal 1 side is speaking. it can.
0032The phoneme database 17 is a database that holds a large amount of phoneme data. The phoneme database 17 may be registered in advance in the conference terminal 1, or the phoneme database 17 is provided on the network so that the conference terminal 1 can acquire phoneme data through the network. There may be.
0033The second voice distribution unit 15 duplicates and distributes the received voice signal [other party] received from the received voice signal processing unit 19. The second audio distribution unit 15 sends one received audio signal [other party] to the audio / video synchronization processing unit 11 in order to output the received audio signal [other party], and also determines the discussion status of the conference. The other received voice signal [other party] is sent to the second voice / non-voice determination unit 16.
0034The second voice / non-voice determination unit 16 receives the received voice signal [other party] from the second voice distribution unit 15 in order to detect the speech status of the conference participant of the other party's conference terminal, and refers to the phoneme database 17. Then, it is detected whether or not the received voice signal [other party] contains voice (phoneme). The second voice / non-voice determination unit 16 gives a voice determination result based on the received voice signal [other party] to the synchronization control unit 14. As the method of voice determination processing by the second voice / non-voice determination unit 16, the same method as the processing method by the first voice / non-voice determination unit 13 can be applied.
0035The synchronization control unit 14 sends the voice determination result based on the transmitted voice signal [conference terminal 1 side] from the first voice / non-voice judgment unit 13 and the received voice signal [other party] from the second voice / non-voice judgment unit 16. Based on the audio determination result based on, the discussion status in the call is determined from the transmitted audio system and the received audio system, and the audio / video synchronization processing unit 11 is instructed to enable or disable the audio / video synchronization processing. That is, the current state of speech of the conference participants is determined from the presence / absence of audio of the conference terminal 1 and the received audio signal [other party] based on the transmitted audio signal [conference terminal 1 side] and the received audio signal [other party]. Check and enable audio / video synchronization processing when the received audio signal [other party] contains audio only. On the other hand, when the transmitted audio signal [conference terminal 1 side] and the received audio signal [other party] contain audio, the audio / video synchronization processing is disabled.
0036The audio / video synchronization processing unit 11 enables or disables the lip-sync function for synchronizing the received audio signal [other party] and the received video signal [other party] according to the instruction of the synchronization control unit 14, and the received audio signal [other party]. The audio based on the [side] is output to the audio output device 3, and the video based on the received video signal [other side] is output to the video output device 4.
0037The audio / video synchronization processing unit 11 has a synchronization switching unit 111 and a synchronization processing unit 112.
0038The synchronization switching unit 111 dynamically switches between enabling and disabling the audio-video synchronization processing as a lip-sync function according to the instruction of the synchronization control unit 14, and instructs the synchronization processing unit 112.
0039The synchronization processing unit 112 is included in the received audio signal [other side] from the second audio distribution unit 15 and the received video signal [other side] from the received video signal processing unit 20 when the audio / video synchronization processing is enabled. The received audio signal [other party] and the received video signal are synchronized with the output timing of the received audio signal [other party] and the output timing of the received video signal [other party] based on the time information (synchronization information). It outputs [the other party]. More specifically, the synchronization processing unit 112 sets the time of both of the received audio signal [other party] based on the time information of the received audio signal [other party] and the time information of the received video signal [other party], and sets the received audio signal [ Outputs [other party] and received video signal [other party]. When the audio / video synchronization processing is disabled, the audio / video synchronization processing unit 11 outputs the received audio signal [other party] and the received video signal [other party] without synchronizing.
0040Here, the synchronization processing unit 112 synchronizes the output timing of the received audio signal [other party] with the output timing of the received video signal [other party], and synchronizes the received audio signal [other party] with the received video signal [other party]. ], The time information of the next received audio signal [other party] and the time information of the received video signal [other party] may be used. For example, the time information of the received audio signal [other party] is received from the received audio signal processing unit 19 via the second audio distribution unit 15, and is related to the received audio signal [other party] extracted from the packet. It is time information. For example, the time information of the received video signal [other party] is received from the received video signal processing unit 20, and is the time information related to the received video signal [other party] extracted from the packet.
0041(A-2) Operation of the embodiment Next, the operation of the audio-video synchronization processing in the conference terminal 1 according to the embodiment will be described in detail with reference to the drawings.
0042When the conference participant speaks at the place where the conference terminal 1 is installed, the voice signal is collected and input by the voice input device 2. The input audio signal is received by the conference terminal 1 as a transmission audio signal [conference terminal 1 side]. The transmitted audio signal [conference terminal 1 side] received by the conference terminal 1 is duplicated by the first audio distribution unit 12 into audio signals for two systems. One of the transmitted audio signals [conference terminal 1 side] is sent to the transmitted audio signal processing unit 18, and the transmitted audio signal processing unit 18 generates a packet related to the conference system and transmits it to the other party's conference terminal via the network. To. The other transmitted voice signal [conference terminal 1 side] is sent to the first voice / non-voice determination unit 13.
0043Further, the audio signal and the video signal of the conference participant received by the other party's conference terminal are received by the received audio signal processing unit 19 and the received video signal processing unit 20 as the received audio signal [other party] and the received video signal [other party]. Be extracted. Here, the received audio signal [other party] and the received video signal [other party] are extracted from the packet received from the other party conference terminal.
0044The received voice signal [other party] extracted by the received voice signal processing unit 19 is duplicated by the second voice distribution unit 15 into voice signals for two systems. One received audio signal [other party] is sent to the audio / video synchronization processing unit 11. The other received voice signal [other party] is sent to the second voice / non-voice determination unit 16.
0045Here, an example of the voice determination process in the first voice / non-voice determination unit 13 and the second voice / non-voice determination unit 16 will be described. In this embodiment, an acoustic model based on a hidden Markov model used as part of a voice recognition technique is adopted to illustrate a case where it is determined whether or not the received voice signal contains voice.
0046FIG. 2 is a flowchart showing the voice determination process according to the embodiment.
0047The first voice / non-voice judgment unit 13 and the second voice / non-voice judgment unit 16 are voice signals (the first voice / non-voice judgment unit 13 is a transmission voice signal [conference terminal 1 side], the second voice / non-voice judgment. When the received audio signal [other party]) is received in unit 16 (S101), the audio signal is converted into a digital signal by a predetermined AD conversion (S102). Since S101 and S102 are processes when an analog signal is input, the processes of S101 and S102 may be omitted when a digital signal is input.
0048The first voice / non-voice judgment unit 13 and the second voice / non-voice judgment unit 16 divide the input digital signal into frames at predetermined intervals (for example, 20 msec) in order to make the input digital signal into a predetermined processing unit (S103). .. Then, the signal in frame units is multiplied by a time window function such as a humming window or a hanning window while sliding in the time axis direction to generate an audio signal in which high frequency noise is reduced by frame division (S104).
0049The first audio / non-audio determination unit 13 and the second audio / non-audio determination unit 16 calculate the spectrum of the audio signal by Fourier transforming the frame-based audio signal multiplied by the time window function (one example is the discrete Fourier transform). (S105), the audio feature is extracted using the spectrum of the audio signal for each frame (S106). Various methods may be applied to the method for extracting the voice feature amount, but in this embodiment, the voice feature amount is indexed based on the Mel frequency Kepratram (MFCC) value. More specifically, the spectrum of the audio signal for each frame is subjected to, for example, a melscale band filter or the like to calculate the power of the frequency component for each frequency band. The mel frequency Kepratram coefficient (MFCC) is obtained by calculating the logarithmic value of the power of each frequency component and performing the discrete cosine transform.
0050Here, in order to determine the audio section or the non-audio section, what indicates the shape of the mouth, not the frequency of the sound source, is indicated by the envelope of the power of the frequency component. This is represented by a number called the vocal tract parameter. Vocal tract parameters vary widely, for example, depending on male and female, young and elderly, individuals, and the like. However, these varieties need to be eliminated and indexed. There is MFCC as a typical method of this index, and quantification is performed by the MFCC method. The reason for using MFCC in this embodiment is to utilize the characteristic that the generated voice has a similar numerical value. In addition, there are already data calculated from a large number of samples as to what the MFCC value will be for phoneme data. Therefore, a large amount of phoneme data is held in the phoneme database 17, and the first voice / non-speech determination unit 13 and the second voice / non-speech determination unit 16 refer to the phoneme database 17.
0051The first voice / non-voice judgment unit 13 and the second voice / non-voice judgment unit 16 collate the voice feature amount (for example, MFCC value) obtained for each frame with the phoneme data stored in the phoneme database 17. (S107). Then, for example, Fig. 3 (Source: Shogo Ando, Rei Oguro, "Try with Raspberry Pi! Voice recognition <6> Important data used for recognition processing! Find voice feature MFCC", Inteface April 2014 No., CQ Publishing Co., Ltd., Vol. 40, No. 4, Vol. 442, Issued April 1, 2014, p. 150, Fig. 11) Corresponds to the received voice signal "Good morning" You can get the phonemes "o", "h", "a", "y", "o".
0052When it is detected that there is voice by matching with the phoneme data of the phoneme database 17 (S108), the first voice / non-voice judgment unit 13 and the second voice / non-voice judgment unit 16 use the voice as the voice detection section. The determination result is output to the synchronization control unit 14. Further, when it is determined that there is no voice (S109), the first voice / non-voice judgment unit 13 and the second voice / non-voice judgment unit 16 send the voice judgment result of the voice non-detection section to the synchronization control unit 14. Send (S110).
0053When performing voice recognition, if the MFCC value is assigned to the phoneme as it is in fragmented frame units, the sound will be more than words. However, if voice / non-voice can be detected, if sufficient, the determination may be made in a state in which this fragmented phoneme is detected and in a state in which it cannot be detected.
0054In addition, the first voice / non-speech determination unit 13 and the second voice / non-voice determination unit 16 are intended to determine whether or not the conference participants on the conference terminal 1 side and the other side are speaking words. To do. That is, the voice input device 2 connected to the conference terminal 1 and the reception voice signal processing unit 19 from the voice signal received by the other party terminal capture sounds and noise other than the voice of the conference participants, but the first voice -In the non-sound judgment unit 13 and the second voice / non-sound judgment unit 16, the state of discussion in the conference is, for example, talking to each other, or one of the conference participants gives a presentation or explanation. Like, it speaks in one direction and judges whether it is in a state where other conference participants are listening to the remark. Here, as an example, the first voice / non-voice judgment unit 13 and the second voice / non-voice judgment unit 16 use the voice judgment process to transmit a voice signal [conference terminal 1 side] and receive a voice signal [ It is determined whether or not the other party contains phonemes.
0055Therefore, the first voice / non-voice judgment unit 13 and the second voice / non-voice judgment unit 16 do not need to process which phoneme the voice included in the received voice signal is, and are input. It suffices if it can be determined whether or not the speech signal contains phonemes as words rather than noise or non-speech. Therefore, here, the case of obtaining a phoneme for the input voice signal as in the example of FIG. 3 is illustrated, but when the phoneme is detected in the voice signal for each frame, the first voice / non-voice determination unit 13, The second voice / non-voice determination unit 16 may output a voice determination result indicating that there is voice. Further, for example, the first voice / non-voice determination unit 13 and the second voice / non-voice determination unit 16 perform voice determination processing for a predetermined time in the time series of the received voice signal to determine the presence or absence of a voice section. You may.
0056Next, the synchronization control unit 14 determines the voice between the other party and the conference terminal 1 based on the voice determination results from the first voice / non-voice determination unit 13 and the second voice / non-voice determination unit 16. The non-audio state is determined, and the audio / video synchronization processing unit 11 is instructed when to perform the audio / video synchronization processing and when not to perform the audio / video synchronization processing.
0057More specifically, the synchronization control unit 14 transmits and receives a voice detection section of the transmission system related to the transmission voice signal [conference terminal 1 side], a voice detection section of the reception system related to the reception voice signal [other side], and transmission / reception. Based on both audio non-detection sections, an instruction is given to enable or disable the audio / video synchronization processing.
0058The time when the situation where the audio and video are not synchronized is desired is very short when the audio detection section of transmission and reception is detected at the same time (that is, when the remarks collide with both the conference terminal 1 side and the other side). This is when the detection of the voice detection section transmitted / received in time is switched.
0059FIG. 4 is a state transition diagram showing the operation of the synchronization control process by the synchronization control unit 14 according to the embodiment.
0060As shown in FIG. 4, there are three states to be managed by the synchronization control unit 14, the synchronization control enabled state 141, the synchronization control disabled state 142, and the synchronization control return waiting state 143.
0061The events that generate the transition states of the synchronous control enabled state 141, the synchronous control disabled state 142, and the synchronous control return waiting state 143 are "TX voice (ON)" indicating voice detection of the transmission system and voice non-detection of the transmission system. "TX voice (OFF)" indicating, "TX non-voice Timeout" indicating the timeout of the duration of voice non-detection of the transmitting system, "RX voice (ON)" indicating voice detection of the receiving system, voice non-detection of the receiving system "RX voice (OFF)" indicating, and "RX non-voice Timeout" indicating the timeout of the duration of voice non-detection of the receiving system.
0062The synchronous control enabled state 141 will be described. When "TX voice (ON)", the RX voice state is determined (S201), and when "RX voice (OFF)", the state transitions to the synchronization control enabled state 141. Further, when "RX voice (ON)", the TX voice state is determined (S202), and when "TX voice (OFF)", the state transitions to the synchronization control enabled state 141. Further, at any of "TX voice (OFF)" and "RX voice (OFF)", the state transitions to the synchronization control enabled state 141.
0063As will be described later, the trigger for the transition from the synchronous control invalid state 142 or the synchronous control return wait state 143 to the synchronous control valid state 141 is when "RX non-voice TimerTimeout".
0064Next, the synchronization control invalid state 142 will be described. When the "TX voice (ON)" event occurs in the synchronization control enabled state 141, the RX voice state is determined (S201), and if it is "RX voice (ON)", the synchronization control disabled state 142 is entered. Further, when the "RX voice (ON)" event occurs in the synchronization control enabled state 141, the TX voice state is determined (S202), and if it is "TX voice (ON)", the synchronization control disabled state 142 is entered. To do. This is intended to be a state in which the voice detection sections overlap (voices collide) in both the transmission / reception systems.
0065That is, the trigger from the synchronous control valid state 141 to the synchronous control invalid state 142 transitions when the voice detection sections of the transmission system and the reception system partially overlap.
0066When the "RX voice (OFF)" event occurs after the transition to the synchronization control invalid state 142, the synchronization control unit 14 performs "RX non-voice Timer Start" to clock the Time Out time of the voice non-detection section of the receiving system ( S203). After the Time Out time of the voice non-detection section of the receiving system is timed, when "RX voice (ON)" is reached, the Time Out time is stopped, so "RX non-voice Timer Stop" is set (S204). When it becomes "RX non-voice TimerTimeOut" (S205), it shifts to S209. In S209, it becomes "RX non-voice TimerTimeOut" (S205), and when "TX non-voice TimerStop" is performed, it transitions to the synchronization control enabled state 141.
0067Further, when the "TX voice (OFF)" event occurs after the transition to the synchronization control invalid state 142, the synchronization control unit 14 clocks the Time Out time of the voice non-detection section of the transmission system, so that "TX non-voice Timer Start" is performed. (S206). After the Time Out time of the voice non-detection section of the transmission system is timed, when "TX voice (ON)" is reached, the Time Out time is stopped, so "TX non-voice Timer Stop" is performed (S207).
0068When "TX non-voice TimerTimeOut" is set (S208), the state transitions to the synchronous control return waiting state 143.
0069S203 to S208 are in a state where audio is not detected in a temporary short period of time in either or both of the transmission system and the reception system, or a predetermined timeout time in either the transmission system or the reception system. With the above, it is determined whether the voice is in the non-detected state. That is, in a conference, either the conference terminal 1 side and the other party's conference participants are in a state of mutual speech, or only one of the conference terminal 1 side or the other party's conference participant speaks. Determine if the other is not speaking.
0070Further, the S209 transitions to the synchronization control enabled state 141 after the other party's conference participant no longer speaks after the predetermined TimeOut time is exceeded.
0071Next, the synchronous control return waiting state 143 will be described. It becomes "TX non-voice Timer Time Out" (S208), and when the "RX voice (OFF)" event occurs, "RX non-voice Timer Start" is performed to time the Time Out time of the voice non-detection section of the receiving system (S210). After the Time Out time of the voice non-detection section of the receiving system is timed, when the "RX voice (ON)" event occurs, the Time Out time is stopped, so "RX non-voice Timer Stop" is set (S211). At this time, since the voices are duplicated in the transmission system and the reception system, the synchronization control return waiting state 143 is maintained instead of transitioning to the synchronization control enabled state 141. Further, when "RX non-voice TimerTimeOut" is set (S212), both the transmission system and the reception system are in the voice non-detection section, so that the synchronization control enabled state 141 is entered.
0072Further, when the "TX non-voice TimerTimeOut" is set (S208) and the "TX voice (ON)" event occurs, the state transitions to the synchronization control invalid state 142.
0073Here, the setting of the TimeOut time related to the non-voice TimeOut detection in the synchronization control invalid state 142 and the synchronization control return waiting state 143 will be described.
0074First, the timing of transitioning from the synchronous control return waiting state 143 to the synchronous control enabled state 141 and restarting the audio / video synchronous processing is an absolute condition that the mutual audio detection sections of the transmitting system and the receiving system do not overlap.
0075In addition, it is considered that even if the delay time due to the audio / video synchronization processing is delayed, the mutual audio detection sections of the transmission system and the reception system do not overlap. In addition, since mutual conversation is performed via the network, the round-trip delay time of the signal is also taken into consideration. In addition, consider the response time (eg, 300 msec) of the person before responding to the action (utterance) of the conference participant watching the video.
0076Considering the above points, the non-voice TimeOut time can be set as follows. Non-voice TimeOut time (msec) = system delay time (msec) + human reaction delay time (msec) ... (1)
0077The system delay time is the delay time related to the conference system. More specifically, as described above, the delay time related to the audio-video synchronization processing and the round-trip delay time of the signal are included. Therefore, Eq. (1) can be converted as Eq. (2). Non-voice TimeOut time (msec) = Lip sync delay (msec) + Network round trip delay (msec) + Human reaction delay (msec) ... (2)
0078By using the above TimeOut time, by managing the transition from the synchronization control invalid state 142 to the synchronization control enabled state 141, by waiting for the timing when the mutual conversation has settled down, the audio / video synchronization processing is frequently turned ON / OFF. Unnaturalness due to various switching can be reduced.
00795 to 7 are timing charts showing the synchronization control processing by the synchronization control unit 14 according to the embodiment.
00805 (A) to 7 (A) show the output timing of the first voice / non-voice determination unit 13 which is the transmission system related to the transmitted voice signal [conference terminal 1 side], and are shown in FIGS. 5 (B) to 7 (A). 7 (B) shows the output timing of the second voice / non-voice determination unit 16 which is the reception system related to the received voice signal [other party], and FIGS. 5 (C) to 7 (C) show the synchronization control unit. The output timing of 14 is shown. The horizontal axes of FIGS. 5 to 7 indicate time, and numbers are added for convenience in explaining the timing.
0081FIG. 5 is a timing chart when the voice detection sections of the transmission system and the reception system do not overlap. That is, it is a case where the statements of the conference participants on the conference terminal 1 side and the other party do not collide.
0082As shown in FIGS. 5 (A) and 5 (B), a voice detection section detected by the first voice / non-voice determination unit 13 and a voice detection section detected by the second voice / non-voice determination unit 16. Does not overlap in time. In such a case, the synchronization control unit 14 does not give an instruction to invalidate the audio / video synchronization processing. In the case of this example, since it is considered that the conference terminal 1 side and the other side make a statement with a certain time interval, the synchronization control unit 14 gives an instruction to enable the audio-video synchronization processing. , Instruct the audio / video synchronization processing unit 11.
0083FIG. 6 is a timing chart when the voice detection sections of the transmission system and the reception system partially overlap. FIG. 6 is a timing chart of the operation of transitioning from S205 to S209 in FIG. 4 and transitioning to the synchronous control enabled state 141.
0084As shown in FIGS. 6 (A) and 6 (B), the voice detection section of the first voice / non-voice determination unit 13 and the voice detection section of the second voice / non-voice determination unit 16 are temporally separated. Some overlap. In such a case, the synchronization control unit 14 maintains an instruction to invalidate the audio-video synchronization processing according to the following specific condition 1.
0085Here, the specific condition is a condition in which the synchronization control unit 14 maintains an instruction to invalidate the audio-video synchronization processing.
0086In the case of FIG. 6, the specific condition 1 starts the invalid instruction when a part of the transmission / reception voice detection section overlaps, and after both transmission / reception transition to the non-voice section, the next received voice is detected. The time point is the end of the invalidation instruction.
0087In the case of the example of FIG. 6, since the audio detection sections for transmission and reception overlap at the time of "time information: 3.5", the synchronization control unit 14 starts the invalidation instruction of the audio / video synchronization processing at this point. At the time of "time information: 10", the receiving system becomes voice non-detection, at the time of "time information: 10.7", the transmitting system becomes voice non-detection, and at the time of "time information: 11.0", the receiving system Voice non-detection exceeds non-voice TimeOut time. At this time, the synchronization control unit 14 maintains the invalidity of the audio / video synchronization processing, and the synchronization control unit 14 effectively switches the audio / video synchronization processing at a time near time information: 11.9.
0088In addition, for example, from "time information: 5.5" to "time information: 5.9", there is no audio detection section of the receiving system, and only the transmitting system has an audio detection section. Since the detection does not exceed the non-audio TimeOut time, the synchronization control unit 14 maintains the invalid instruction of the audio-video synchronization processing according to the specific condition 1. In addition, only the voice detection section of the transmission system is available, such as from "time information: 5.5" to the vicinity of "time information: 5.8", but in this case as well, the voice non-detection of the transmission system exceeds the non-voice Time Out time. Therefore, the synchronization control unit 14 maintains the invalid instruction of the audio / video synchronization processing according to the specific condition 1.
0089FIG. 7 is a timing chart when the voice detection sections of the transmission system and the reception system partially overlap. FIG. 7 is a timing chart of the operation of transitioning from S208 to S210 to S212 in FIG. 4 and transitioning to the synchronous control enabled state 141.
0090As shown in FIGS. 7 (A) and 7 (B), the voice detection section of the first voice / non-voice determination unit 13 and the voice detection section of the second voice / non-voice determination unit 16 are temporally separated. Some overlap. In such a case, the synchronization control unit 14 maintains an instruction to invalidate the audio-video synchronization processing according to the following specific condition 2.
0091In FIG. 7, a particular condition 2, the time when the overlap portion of the transmission and reception of voice detection section occurs as invalid instruction start, feeding on signal systems of audio undetected exceeds the non-speech TimeOut time, the receiving system When the non-voice detection exceeds the non-voice TimeOut time and the transmission system detects voice, the invalid instruction ends when the next received voice is detected.
0092In the case of the example of FIG. 7, since the transmission / reception audio detection sections overlap at the time of "time information: 3.5", the synchronization control unit 14 starts the invalidation instruction of the audio / video synchronization processing at this point. At the time near "Time information: 8.5", the voice of the transmission system is not detected, and at the time near "Time information: 9.5", the voice of the transmission system exceeds the non-voice TimeOut time. At around 10.7 , the audio of the receiving system is not detected, and at around time information: 10.9 , the audio of the receiving system exceeds the non-audio TimeOut time. After that, the synchronization control unit 14 keeps the audio / video synchronization processing disabled until the audio is detected in the receiving system, and switches the audio / video synchronization processing to valid at around "time information: 11.1". Here, the non-voice TimeOut time related to the voice of the receiving system in the synchronous control return waiting state 143 detects that the conference participant on the other side does not speak once while the conference participant on the conference terminal 1 side does not speak. As long as possible, the time is set to be shorter than the non-voice TimeOut time related to the voice of the receiving system in the synchronization control disabled state 142.
0093In this case as well, for example, from "time information: 4" to "time information: 4.5", there is no audio detection section of the transmission system, and only the reception system has an audio detection section. In this case, transmission is performed. Since the audio non-detection of the system does not exceed the non-audio TimeOut time, the synchronization control unit 14 maintains the invalid instruction of the audio-video synchronization processing. In addition, only the voice detection section of the transmission system is available, such as from "time information: 5.5" to the vicinity of "time information: 5.8", but in this case as well, the voice non-detection of the reception system exceeds the non-voice Time Out time. Therefore, the synchronization control unit 14 maintains the invalid instruction of the audio / video synchronization processing.
0094(A-3) Effect of embodiment As described above, according to the embodiment, it is detected that conversations collide with each other, and a state in which more real-time performance is required is recognized. As a result, even during a conference, the lip clip function for synchronizing the audio signal and the video signal can be temporarily suppressed, and the mutual audio delay can be minimized for output. Therefore, when many conference participants speak at almost the same time, it can be performed as if it were a normal conversational state.
0095Further, according to the embodiment, after disabling the audio-video synchronization processing (lip sync), the synchronization is restarted at the timing according to the specific condition of the synchronization control return waiting state. This makes it possible to prevent continuous switching of the lip sync and prevent unnatural switching.
0096Furthermore, just by recognizing each other's voice states, lip-sync is frequently enabled / disabled, and the delay that occurs when lip-sync is enabled may make the call feel uncomfortable. However, according to the embodiment, by performing timing control, it is possible to suppress a sense of discomfort and transition to the lip-sync effective state.
0097Further, according to the embodiment, the timing that does not give a sense of discomfort even when the lip sync is returned is when the non-speech time continues longer than the delay time required for the lip sync. After that state, by making the lip sync delay time wait and synchronizing between the audio signal and the video signal, the video signal is synchronized at the timing of the next audio signal. As a result, the person watching the image can see the image smoothly.
0098(B) Other embodiments Although various modified embodiments have been mentioned in the above-described embodiments, the present invention can also be applied to the following embodiments.
0099(B-1) In the above-described embodiment, a case where it is applied to a conference terminal used in a conference system is illustrated. However, the present invention is not limited to a conference, and may be applied to, for example, a videophone terminal that makes a one-to-one call. It can also be applied to television terminals used in conference systems at two or more points. Even in a multipoint conference system with three or more points, audio signals picked up at three or more points and captured video signals will be received. However, according to the present invention, it is possible to control whether to enable or disable the audio / video synchronization processing (lip sync function) based on the audio detection of the transmission system and the reception system in the conference terminal. , Applicable to multi-point conference system with 3 or more points.
0100(B-2) In the above-described embodiment, the video signal [other party] of the other party's conference participant acquired by the other party's conference terminal and the audio signal [other party] spoken by the other party's conference participant. The case of controlling the synchronization with and is illustrated. However, for example, even when communicating a moving image such as a music clip or a movie clip in a presentation or the like, even if the conference terminal controls the synchronization between the video signal and the audio signal received via the network. good.
01011 ... conference terminal, 2 ... audio input device, 3 ... audio output device, 4 ... video output device, 10 ... audio / video synchronization control unit, 11 ... audio / video synchronization processing unit , 12 ... 1st audio distribution unit, 13 ... 1st audio / non-audio judgment unit, 14 ... synchronous control unit, 15 ... 2nd audio distribution unit, 16 ... 2nd audio / Non-audio judgment unit, 17 ... phonetic database, 18 ... transmission audio signal processing unit, 19 ... received audio signal processing unit, 20 ... received video signal processing unit.
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| WO2022269788A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| JP2000124809A | Cites | Japan | A | Search report | – |
| US2005237378A1 | Cites | United States of America | A | Search report | – |
| JP2009294537A | Cites | Japan | A | Search report | – |
| JP2012217172A | Cites | Japan | A | Search report | – |
| JPH08317362A | Cites | Japan | X | Search report | 1-8 |
| JPH0879252A | Cites | Japan | A | Search report | – |
| JPS6318788A | Cites | Japan | A | Search report | – |
1 member in 1 office
Members1
| Document | Office | Kind | |
|---|---|---|---|
| JP2017046235AThis record | Japan | A |
4 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2017046235
- Application
- 168177
Titles2
- Japanese
- 音声映像同期処理装置、端末、音声映像同期処理方法及びプログラム
- English
- Audio-video synchronization processing device, terminal, audio-video synchronization processing method and program
Classification
- IPC, 3
- H04N7 15
- H04N21 43
- H04M3 56