Device and method for synchronizing received audio data with video data
Abstract
This record has no abstract on file.
Term
Projected expiry 8 December 2028.
- Priority and filed
- Granted
- Today
- Projected expiry
18 claims: 4 independent, 14 dependent
- 1通信装置(12)により受信された、複数のビデオデータセグメントを含むビデオデータと、前記通信装置により受信された、複数のオーディオデータセグメントを含むオーディオデータとを同期させる方法であって、 オーディオデータの第1セグメントを前記通信装置(12)で受信するステップと、 前記オーディオデータの第1セグメントと同時又はそれより遅い時点で、該 オーディオデータの 第1 セグメント に論理的に関連するビデオデータの第1セグメントを前記通信装置(12)で受信するステップと、 前記オーディオデータの第1セグメントの始まりに関連する事前通知を前記通信装置で発生するとともに、前記オーディオデータの第1セグメントの始まりを再生する前に、前記事前通知を示す視覚情報を表示し、又は、前記事前通知を示すオーディオ情報を再生する ステップと、 前記事前通知を示す視覚情報を表示し、又は、前記事前通知を示すオーディオ情報を再生した後に、 前記オーディオデータの第1セグメントと前記ビデオデータの第1セグメントとの間に同期処理を適用するステップと、 を有することを特徴とする方法。
- 2前記オーディオデータの最終セグメントの終わりに関連する事前通知を前記通信装置で発生するステップと、 前記オーディオデータの最終セグメントの終わりを再生する前に、前記事前通知を示す視覚情報を表示し、又は、前記事前通知を示すオーディオ情報を再生するステップと、 を更に有することを特徴とする請求項1に記載の方法。
- 3前記視覚情報を画像又は文字として前記通信装置の画面に表示し、又は、光源により発生される光として表示し、又は、音声を放射するために前記オーディオデータを処理するステップを更に有することを特徴とする請求項1に記載の方法。
- 4顔の有無を検出するために前記受信されたビデオデータの第1セグメントを分析するステップを更に有することを特徴とする請求項1に記載の方法。
- 5前記ビデオデータの第1セグメント中に前記顔が存在する確率を判定するために前記受信されたビデオデータの第1セグメントをフィルタリングするステップを更に有することを特徴とする請求項4に記載の方法。
- 6前記判定された確率が所定の閾値確率より高い場合に前記同期処理をオンするステップを更に有することを特徴とする請求項5に記載の方法。
- 7対応する後続ビデオデータセグメントにおいて次に判定された顔が存在する確率が前記所定の閾値確率より低い場合は前記同期処理をオフするステップを更に有することを特徴とする請求項6に記載の方法。
- 8前記同期処理を適用するステップは、 前記ビデオデータの第1セグメントを受信又は復号化する前に、前記オーディオデータの第1セグメントの始まりを再生するステップと、 所定の速度より遅い低再生速度で少なくとも前記オーディオデータの第1セグメントを再生するステップと、 前記ビデオデータの第1セグメントを再生のために利用可能である場合、前記オーディオデータの第1セグメント又はそれに続くセグメントの前記低再生速度を前記所定の速度まで上げるステップと、 を含むことを特徴とする請求項1に記載の方法。
- 9前記オーディオデータの最終セグメントを前記通信装置で受信するステップと、 前記オーディオデータの最終セグメントと同時又はそれより遅い時点で、前記オーディオデータの最終セグメントに論理的に関連する前記ビデオデータの最終セグメントを前記通信装置で受信するステップと、 前記オーディオデータの最終セグメントと前記ビデオデータの最終セグメントとが同期しないように、所定の速度より速い高再生速度で前記オーディオデータの最終セグメントを再生するステップと、 前記ビデオデータの最終セグメントの再生を終了する前に、前記オーディオデータの最終セグメントの再生を終了するステップと、 を更に有することを特徴とする請求項1に記載の方法。
- 10複数のビデオデータセグメントを含む受信したビデオデータと、複数のオーディオデータセグメントを含む受信したオーディオデータとを同期させる通信装置(12)であって、 オーディオデータの第1セグメントを受信し、該オーディオデータの第1セグメントと同時又はそれより遅い時点で、該オーディオデータの第1セグメントに論理的に関連するビデオデータの第1セグメントを受信する入出力ユニット(18)と、 前記オーディオデータの第1セグメントの始まりに関連する事前通知を発生するとともに、前記オーディオデータの第1セグメントの始まりを再生する前に、前記事前通知を示す視覚情報を表示し、又は、前記事前通知を示すオーディオ情報を再生 し、前記事前通知を示す視覚情報を表示し、又は、前記事前通知を示すオーディオ情報を再生した後に、 前記オーディオデータの第1セグメントと前記ビデオデータの第1セグメントとの間に同期処理を適用するプロセッサ(24)と、 を有することを特徴とする通信装置。
- 11前記プロセッサは、更に、 前記オーディオデータの最終セグメントの終わりに関連する事前通知を発生し、 前記オーディオデータの最終セグメントの終わりを再生する前に、前記事前通知を示す視覚情報を表示し、又は、前記事前通知を示すオーディオ情報を再生する ことを特徴とする請求項10に記載の通信装置。
- 12前記視覚情報を前記通信装置の画面に画像又は文字として表示し、又は、光源により発生される光として表示する表示ユニット、又は、 前記オーディオデータを音声として再生する音声再生ユニット、 を更に有することを特徴とする請求項10に記載の通信装置。
- 13前記プロセッサは、更に、 顔の有無を検出するために前記受信されたビデオデータの第1セグメントを分析することを特徴とする請求項10に記載の通信装置。
- 14前記プロセッサは、更に、 前記ビデオデータの第1セグメント中に前記顔が存在する確率を判定するために前記受信されたビデオデータの第1セグメントをフィルタリングすることを特徴とする請求項13に記載の通信装置。
- 15前記プロセッサは、更に、 前記判定された確率が所定の閾値確率より高い場合に前記同期処理をオンすることを特徴とする請求項14に記載の通信装置。
- 16前記プロセッサは、更に、 対応する後続ビデオデータセグメントにおいて次に判定された顔が存在する確率が前記所定の閾値確率より低い場合は前記同期処理をオフすることを特徴とする請求項14に記載の通信装置。
- 17前記プロセッサは、更に、 前記ビデオデータの第1セグメントを受信又は復号化する前に、少なくとも前記オーディオデータの第1セグメントを再生し、 所定の速度より遅い低再生速度で前記オーディオデータの第1セグメントを再生し、 前記ビデオデータの第1セグメントを再生のために利用可能である場合、前記オーディオデータの第1セグメント又はそれに続くオーディオデータセグメントの前記低再生速度を前記所定の速度まで上げる ことを特徴とする請求項10に記載の通信装置。
- 18前記入出力ユニットは、更に、前記オーディオデータの最終セグメントを受信し、該オーディオデータの最終セグメントと同時又はそれより遅い時点で、前記オーディオデータの最終セグメントに論理的に関連する前記ビデオデータの最終セグメントを受信し、 前記プロセッサは、更に、前記オーディオデータの最終セグメントと前記ビデオデータの最終セグメントとが同期しないように、所定の速度より速い高再生速度で前記オーディオデータの最終セグメントを再生し、 前記ビデオデータの再生の終了前に、前記オーディオデータの再生を終了する ことを特徴とする請求項10に記載の通信装置。
Independent claims18
60 paragraphs, as filed
The present invention relates to methods, devices and systems capable of processing audio data and video data, and more particularly to techniques and methods for compensating for time delays associated with reproduction of video data transmitted with the audio data.
Communication using multimedia, including both video data and related audio data, is becoming more important in the field of communication technology for both fixed and mobile access. As improvements are being made to include video components (ie, video data) in addition to traditional voice-only calls, users are more likely to communicate using so-called "video telephony". ..
Video data associated with a videophone call is typically created by the video camera of the transmitting device. The transmitting device may be a portable device such as a mobile phone. In some cases, the user orients the transmitting device to position the camera to show the speaker's face. However, the camera may also be used to show other things that the user considers to be related to the conversation, such as the scenery that the user wants to show to the person in the conversation. Therefore, what is shown during a communication session can change. In this regard, video and audio data are usually generated with logical relevance. For example, the user's voice is associated with the image of the user's face corresponding to the user emitting the voice.
If the speaker user is also displayed on the viewer user's screen, it is desirable that the audio and video data be synchronized so that the user experiences the proper combination of audio and video. In order to achieve the proper combination, the movement of the user's lips should normally be synchronized with the sound from the device's speakerphone. This will provide the same relationship between the movement of the lips and the words heard, as if two people were in close proximity to each other in a normal conversation. In the present specification, this is referred to as lip synchronization or the logical relationship between audio data and video data.
Therefore, refer to 3G circuit-switched videophones (eg, 3GPP TS26.111 (according to the 3GPP Standards Group of route des Lucioles 06921 Sophia-Antipolis Codex, ETSI Mobile Competence Center 650, France). Support for inter-media synchronization is provided for existing equipment such as (embedded) and IP multimedia services such as the new technology IMS Multimedia Telephony (see, eg, 3GPP TS 22.173 and ETSI TS181002 by ETSI). It is desired. Hereinafter, a conventional method for synchronizing audio and video will be described. For Circuit Switched Multimedia, you can tell how much the audio is delayed to synchronize the audio with the video (see ITU-T H.324). Real-time transport protocol (RTP, IETF) For services transported via RFC 3550), RTP timestamps can be used with the RTP Control Protocol (RTCP) Transmitter Report as input to achieve synchronization (see IETF RFC 3550). However, some existing multimedia communication services do not perform media synchronization, so users can only make inadequate calls when lip synchronization is required.
Systems that synchronize audio with video typically delay audio data for a certain length of time before the video data is decoded in order to achieve the desired lip synchronization, and then two data. Are played at the same time. However, this synchronization method is not preferred for the user because of the increased delay, resulting in longer response times and conversational problems. For example, video data typically has a longer delay time from the camera to the screen than the delay time from the microphone to the speakerphone. The long delays in video data are due to the long algorithm delays associated with coding and decoding (compared to audio data), slow frame rates, and in some cases high bitrates that result in high transfer delays. It is to be long. Assuming that the receiving device synchronizes audio and video, the device must delay the audio data flow before playing the audio data. This, of course, causes the user to be dissatisfied with the conversation, resulting in poor conversation quality. For example, when audio data delay exceeds a certain limit (about 200ms), conversation quality begins to be affected. First, the other speaker feels slow and in some cases both speakers start speaking at the same time (because the two speakers only notice this problem after some time delay). The user will be somewhat frustrated, as it may be). If the delay is long (eg, over 500ms), it will start to be difficult to continue a normal conversation. Therefore, one of the causes of dissatisfaction among speakers using videophones is that the response time of the other speaker is too long, unlike normal face-to-face conversations or voice-only telephone conversations. Is.
Therefore, it is desired to provide devices, systems and methods for audio and video communication that avoid the above problems and drawbacks.
According to an exemplary embodiment, a method of synchronizing video data including a plurality of video data segments received by a communication device with audio data including a plurality of audio data segments received by the communication device. Is provided. In the method, video data logically related to the first segment of the audio data at the same time as or later than the step of receiving the first segment of the audio data by the communication device and the first segment of the audio data. It has a step of receiving the first segment of the above in a communication device, and a step of applying a synchronization process between the first segment of the audio data and the first segment of the video data based on a predetermined index.
According to another exemplary embodiment, a communication device is provided that synchronizes received video data including a plurality of video data segments with received audio data including a plurality of audio data segments. The communication device receives the first segment of the audio data, and at the same time as or later than the first segment of the audio data, the first segment of the video data logically related to the first segment of the audio data. It has an input / output unit that receives the data, and a processor that applies a synchronization process between the first segment of the audio data and the first segment of the video data based on a predetermined index.
<figref num="1">It is a figure which shows the communication system including the transmitting side apparatus, receiving side apparatus and communication network which concerns on one exemplary Embodiment.</figref><figref num="2">It is a figure which shows the transmitting side apparatus or receiving side apparatus which concerns on one exemplary Embodiment.</figref><figref num="3">It is a figure which shows the timing of the audio data and the video data exchanged between a transmitting side apparatus and a receiving side apparatus.</figref><figref num="4">FIG. 5 is a diagram showing the timing of audio data and video data exchanged between a transmitting side device and a receiving side device using advance notification according to an exemplary embodiment.</figref><figref num="5">It is a flowchart which shows the step which is executed to send the advance notice which concerns on one exemplary Embodiment.</figref><figref num="6">It is a figure which shows the process which turns on and off a synchronization function based on the face detection process which concerns on one exemplary Embodiment.</figref><figref num="7">It is a flowchart which shows the step for turning on and off a synchronization function based on the face detection process which concerns on one exemplary Embodiment.</figref><figref num="8">It is a figure which shows the process of turning on and off a synchronization function based on the user input which concerns on one exemplary Embodiment.</figref><figref num="9">It is a flowchart which shows the step for turning on and off a synchronization function based on the user input which concerns on one exemplary Embodiment.</figref><figref num="10">It is a figure which shows the timing of the audio data and the video data exchanged between the transmitting side apparatus and the receiving side apparatus with time scaling.</figref><figref num="11">It is a flowchart which shows the step for applying the time scaling to the 1st segment of the audio data which concerns on one exemplary Embodiment.</figref><figref num="12">It is a flowchart which shows the step for applying the time scaling to the final segment of the audio data which concerns on one exemplary Embodiment.</figref><figref num="13">FIG. 13 is a flowchart showing steps of a method of synchronizing video data with audio data.</figref>
(Abbreviation) RTP: Real-Time Transport Protocol: Real-time Transport Protocol RTCP: Real-Time Control Protocol: Real-time control protocol AVS: Audio-video signal: Audio-video signal LED: Light Emitting Diode: Light Emitting Diode UDP: User Datagram Protocol: User Datagram Protocol IP: Internet Protocol: Internet Protocol AMR: Adaptive Multi-Rate: Adaptive Multi-Rate DVD: Digital Versatile Disc: Digital Versatile Disc ASIC: Application Specific Integrated Circuit: Integrated circuit for specific applications DSP: Digital Signal Processor: Digital Signal Processor FPGA: Field Programmable Gate Array: Field Programmable Gate Array IC: Integrated Circuit: Integrated Circuit FM: Frequency Modulated: Frequency modulation LCD: Liquid Crystal Display: Liquid crystal display OLED: Organic Light-Emitting Diode: Organic Light-Emitting Diode WLAN: Wireless Local Area Network: Wireless Local Area Network
(Detailed explanation) In the description of the following exemplary embodiments, reference is made to the accompanying drawings. In the drawings, the same reference numerals indicate the same element or similar elements. The following detailed description is not limited to the present invention. The scope of the present invention is defined by the appended claims. For brevity, the following embodiments describe a user who also uses the mobile phone to communicate with another user who uses the mobile phone. However, the embodiments described below are not limited to this system and may be applied to other existing audio and video transmission systems.
Throughout this specification, the term "one embodiment" means that one particular feature, structure or property described in connection with one embodiment is included in at least one embodiment of the present invention. .. Therefore, the expression "in one embodiment" found in various places herein does not always refer to the same embodiment. Moreover, specific features, structures or properties can be combined in any suitable manner in one or more embodiments.
As shown in FIG. 1, according to an exemplary embodiment, the system 10 includes a first communication device 12 and a second communication device 14 connected to each other via a communication network 16. Devices 12 and 14 may be desktops, laptops, mobile phones, conventional phones, PDAs, digital cameras, video cameras and the like. The two devices can be connected to each other via a wired or wireless interface. The two devices may be connected directly to each other, or may be connected via one or more base stations (not shown) that are part of a communication network. As used herein, the term "base station" is used as a general term to refer to any device that facilitates the exchange of data between connected devices, such as modems, stations in telecommunications systems, and entire networks. used.
As shown in FIG. 2, the structure of the device 12 or 14 includes an input / output port 18 configured to send and receive the audio-video signal AVS. Audio-video signal AVS may include audio data and video data. Each of the audio or video data can contain multiple segments. A segment may contain multiple frames corresponding to a particular time. However, the definition of this segment may be further limited depending on the particular environmental conditions. Examples of specific embodiments will be given later. The plurality of audio data segments and / or video data segments may include a first segment and a final segment, and may further include a plurality of other segments between the first segment and the final segment. One audio data segment may correspond to one segment of video data, as in the case of a user recording an audio message while video recording a face, for example.
To receive the video signal AVS, the I / O port 18 may be connected to the antenna 22 or wireline (not shown) via bus 20. The antenna 22 may be a single antenna or multiple antennas and may be configured to receive the audio-video signal AVS via infrared, radio frequency or other well-known radio interfaces. The input / output port 18 is further connected to a processor 24 that receives and processes the audio-video signal AVS. Processor 24 may be connected to memory 26 via bus 20. The memory 26 may store the audio-video signal AVS and other data required by the processor 24.
In one exemplary embodiment, the device 12 may have a display 28 configured to display an image corresponding to the received audio-video signal AVS. The display 28 may be a screen and may further include one or more LEDs or any other well-known light source emitting device. The display 28 may be a combination of a screen and an LED. In another exemplary embodiment, the device 12 may have an input / output interface 30, such as a keyboard, mouse, microphone, video camera, etc., capable of inputting instructions and / or data from the user.
The device 12 can measure various indicators of the audio-video signal AVS connected to and received by the bus 20, or can analyze the video data of the AVS to extract the user's face, or at different speeds. It may have a processing unit 32 capable of reproducing AVS audio data at a speed (faster or slower than the recording speed). The device 12 may have a voice unit 34 configured to generate voice based on audio data received by the device. Further, the voice unit 34 may emit voice or record voice according to the instruction of the processor 24. In one exemplary embodiment, the audio unit may include a speakerphone and a microphone. The device 14 shown in FIG. 1 may have the same structure as the device 12 shown in FIG.
In the following description, for the sake of brevity, device 12 (see Figure 1) is considered to be the transmitter and device 14 (see Figure 1 as well) is considered to be the receiver. However, devices 12 and 14 may both operate as transmitters and / or receivers. When user 1 of device 12 transmits video data and audio data to user 2 of device 14, the operation shown in FIG. 3 occurs in device 12 first, and then the operation shown in FIG. 3 in device 14 occurs. Behavior occurs. More specifically, at time t1, the device 12 receives audio data S1 from user 1 or another source and video data V1 from user 1 or another source. Both the audio data S and the video data V1 are encoded by the device 12 and then sent to the user 2 via the input / output unit 18 or the antenna 22. The encoded audio data S2 is transmitted at a time t2 that is slower than t1 but earlier than the time t3 when the encoded video data V2 is transmitted. Figure 3 shows that the video data is already delayed by t3-t2 from the audio data. The reason why the transmission of the coded video data V2 is delayed in this way is that the video data coding process requires a longer time than the audio data process.
The coded audio data S2 is received by the device 14 of the user 2 at time t4, and the coded video data V2 is received by the device 14 at a later time t6. Due to the delay of the coded video data V2, the receiving device 14 starts decoding the coded audio data S2 at time t5, which is later than the time t4 but earlier than the time t6 when the coded video data V2 is received by the device 14. Things can happen. However, in one exemplary embodiment, time t5 may be later than time t6. The device 14 decodes the encoded video data V2 at a time t7 later than the time t6.
The earliest time when the apparatus 14 can synchronize and reproduce both the decoded video data V3 and the decoded audio data S3 is time t8. Therefore, in the conventional device, the device 14 delays the audio data from the time t5 to the time t8, and starts the reproduction of both the decoded audio data S3 and the decoded video data V3 at the time t8. This delay from t5 to t8 causes the problems described in the "Background Techniques" section for conventional equipment. FIG. 3 further shows the timing and coding / decoding data when the user 2 responds to the user 1 and that the user 1 experiences the reaction time T1 of the user 2 as the experience reaction time T.
According to one exemplary embodiment, the receiving device of the receiving user may notify the receiving user that the transmitting user has stopped speaking. By obtaining this information, the receiving user can avoid starting talking while the receiving device is still processing the received data. Regarding this point, in the conventional device, due to the internal processing of the receiving device, (i) the time when the receiving device receives the final part of the audio data from the transmitting user and (ii) the receiving user There is a delay with the time to notice the fact of reception. However, according to this embodiment, this delay is reduced or eliminated. According to another exemplary embodiment, the receiving device may provide user 2 with an instruction that the talk of user 1 is about to end. As a result, the reaction time T1 is shortened because the user 2 can start talking earlier than when such an instruction is not given. This instruction may be a visual signal (eg, LED lighting or a symbol on the screen of the device) that persists for the duration of the conversation. The signal may be another visual or audible signal.
According to one exemplary embodiment shown in FIG. 4, the video data and audio data flow can include voice advance notification "a" to user 2. More specifically, the user 2 can receive (generate) the advance notification "a" that the user 1 has stopped transmitting the audio data. The advance notice "a" can occur at or immediately after t4, and the receiving user will notice the reception of the audio data at t4, not at t7, where the audio data is played. This advance notification shortens the reaction time T1 of user 2. The effect (ie, reducing the time delay of the audio data) is shown as an "A" in Figure 4. In this regard, the timings and symbols used in FIG. 4 are similar to the timings and symbols used in FIG. 3, and will not be repeated here.
According to another exemplary embodiment, the receiving device may generate advance notice notifying user 2 that audio data from user 1 is not detected. This advance notification may be generated and displayed at t8, which is earlier than the time t9 in which the user 2 has conventionally determined that the voice from the user 1 has not been received. Therefore, this time difference t9-t8 can be another effect of user 2. In this exemplary embodiment, the end of the final segment of audio data is determined and advance notice is generated based on the end of the final segment.
In another exemplary embodiment, when the start of audio data from user 2 is received, the instruction "b" that user 2 has started sending audio data is given to user 1 on the device of user 1. Occurs. From this voice advance notification "b", the user 1 determines that the talk is not resumed until the information from the user 2 is presented in a state of being synchronized with other media. The voice advance notification "b" can be realized in the same manner as the advance notification "a". By using the voice advance notification in this way, the risk of both talking parties speaking at the same time is considerably reduced, and the reaction time of both parties is also shortened.
In one exemplary embodiment, the advance notices "a" and "b" may both be implemented in communication devices 12 or 14, respectively. In the present embodiment, the user is warned by his / her device that audio data from another user has started, and that the audio data has stopped is also warned at the time of stop before the audio data is played. To.
The total effect obtained by using the two advance notices "a" and "b" in this exemplary embodiment (ie, reducing the time delay of audio data) is that the user 2 is notified of the conversation interval. The round trip delay is actually reduced, resulting in a shorter reaction time for user 2, which has the advantage of reducing the risk of crosstalk due to user 1 being notified of audio data reception from user 2. Combined. This total effect is shown as "B" in Figure 4. Therefore, according to an exemplary embodiment described herein, a device configured to generate voice advance notification reduces the risk of crosstalk (simultaneous speech of user 1 and user 2) and / or ( Reduce user response time (because it can more accurately determine when the user should start speaking). Another advantage of one or all of the exemplary embodiments described above is realized because the device uses information already available on the terminal, ie, no extra-terminal signaling is required. Is easy.
FIG. 5 shows a method of synchronizing the video data received by the communication device and the audio data received by the communication device by an exemplary method for realizing the exemplary embodiment described above. The video data includes a plurality of video data segments and the audio data includes a plurality of audio data segments. The method consists of step 50 of receiving the first segment of audio data by a communication device and the first of video data logically related to the first segment of audio data at the same time as or later than the first segment of audio data. Step 52 to receive the segment on the communication device, step 54 to generate advance notification on the communication device related to the first segment of audio data, and process the advance notification to generate visual or audio information indicating the advance notification. Including step 56 and.
More specifically, step 54 gives advance notice related to the beginning of the first segment of the audio data in step 54-1, which causes the communication device to give advance notice before playing the beginning of the first segment of the audio data. It may include step 54-2 of displaying the visual information to be shown or playing back the audio information showing the advance notice. Alternatively, step 54 is a visual information indicating step 54-3 in which the communication device generates advance notification related to the end of the first segment of the audio data and advance notification before playing the end of the first segment of the audio data. May include steps 54-4 and playback of audio information indicating the display or advance notice. In yet another embodiment, step 54 may include all of steps 54-1 to 54-4.
In a further exemplary embodiment, the receiving device does not generate or receive the aforementioned advance notice, but has an image analysis function (eg, an image analysis function) that detects whether or not a face is present in the received video data. Face detection function) may be included. When the receiving device detects a face, the receiving device may activate the synchronization function. If the receiving device does not detect a face, the receiving device does not activate the synchronization function. This optimizes the quality of conversations, including both audio and video. This exemplary technique will be described in more detail below.
Communication between device 12 and device 14 may be set up using a conventional session setup protocol. For brevity, exemplary embodiments relating to techniques including face analysis are described based on the RTP / User Datagram Protocol (UDP) / Internet Protocol (IP) communication system using RTCP as the activation synchronization protocol. Will be done. However, exemplary embodiments may apply to other systems and protocols.
The receiving device is configured to apply the synchronization function as needed. The synchronization function may include a time delay of audio data with respect to video data, novel techniques described herein, or a combination thereof. The synchronization function may be realized by the processor 24 or the processing unit 32 shown in FIG. At least one of the communication devices 12 and 14 includes an exemplary synchronization function according to the present embodiment. In another exemplary embodiment, communication devices 12 and 14 both include a synchronization function.
The communication device may be initially configured to have the synchronization function turned on or off. During communication, the transmitting device continues to transmit audio and video data along with standard protocol tools to enable synchronization. In one exemplary embodiment, the receiving device continues to analyze the received video data and uses the face detection function to detect whether or not a face is present in the received video data. In another exemplary embodiment, the receiving device analyzes the received video data at predetermined intervals in order to detect the face. The face detection function generates a "with face" value or a "without face" value as an output. An example of face detection is Polar Rose, MINC (Anckargripsgatan 321119, Sweden) Available from Malmo). Other face detection products may be used in the communication device as will be appreciated by those skilled in the art. In another exemplary embodiment, the face detection feature provides binary output with / without face, such as certainty, such as a percentage indicating the probability of the presence or absence of a face. Further soft output may be generated. This soft output may be used to filter the information as described below.
If the face / faceless output switches at high speed, the result is that the synchronization function is turned on / off too often in a later step, for example, when moving the camera to the user's face. Therefore, in order to avoid this, a low-pass filter for face detection information may be applied. Such frequent switching will adversely affect voice quality. The filter function produces a filtered detection output that avoids frequent switching, along with a "faced" or "faceless" value. Advanced face detection features "soft" confidence information in the output that contains a 0-100% confidence value that represents the confidence when the detection algorithm concludes whether the video data being analyzed contains a face. It may occur. When the face detection function generates soft certainty information as described, this information can be used in the filtering function. For example, if the detection confidence is low, long-term filtering is applied to provide a more reliable basis for determining the detection state change between "faced" and "faceless".
If the output value of the detection after filtering is "with face", the synchronization function is applied to synchronize the audio data and the video data. If the sync feature was not previously used (ie turned off), it will be turned on based on the "faced" output. This on may be performed immediately after the "faced" output, but in that case there will be a gap in the audio. Alternatively, the on may be performed in a more sophisticated way that eliminates audio gaps, such as using time scaling (discussed below) or waiting for a break during a conversation to achieve synchronization functionality. Good.
When the output value of the detection after filtering is "faceless", the synchronization between the audio data and the video data is not applied because the sound is not accompanied by the movement of the lips. Therefore, according to the present embodiment, the audio data is reproduced at the time of decoding, so that the voice quality is improved. In this case, if the synchronization function is on and the output value is "no face", the synchronization function is turned off. The off may be performed immediately, but then clipping of the audio segment will occur. Alternatively, the off may be achieved in a more sophisticated way that eliminates clipping of the audio segment, for example by using time scaling or waiting for a break in the audio data.
Therefore, the receiving device is in a position to turn on / off the synchronization function as needed. In an RTP / UDP / IP system that uses RTCP as the activation protocol, according to an exemplary embodiment, the communicator turns the synchronization feature on and off at any time by monitoring and keeping track of the RTCP transmitter report. It is possible.
Figure 6 shows the process of turning the synchronization function on and off. Data is transmitted from the transmitting device to the receiving device. Therefore, in step 60, the data is received by the receiving device. In step 62, the receiving device determines whether or not a face is present in the video data of the transmitted data. If it is determined in this step that a face is present, the synchronization function is turned on in step 64, and the process proceeds to step 66. In step 66, it is determined whether or not there is an end to the data. When the end of the transmitted data is determined, the receiving device stops the switching process. However, if it is determined that the face does not exist in the data received in step 62, the process proceeds to step 68. At step 68, the synchronization feature is turned off. The process then proceeds to step 66, which is as described above. If the end of the data is not determined, the process returns to step 62.
An embodiment of the method following the process described above is shown in FIG. The method of this embodiment is a method of synchronizing video data received by a communication device and including a plurality of video data segments with audio data received by a receiving device and containing a plurality of audio data segments. The method consists of step 70 of receiving the first segment of audio data in a communication device and the first of video data logically related to the first segment of audio data at the same time as or later than the first segment of audio data. It includes step 72 of receiving the segment on the communication device, step 74 of analyzing the first segment of the received video data to detect the face, and step 76 of turning on the synchronization function when the face is detected. ..
According to another exemplary embodiment, the receiving device does not use the face recognition function to turn the synchronization process on and off. In this exemplary embodiment, the user determines when the synchronization process should be turned on / off. In other words, synchronization is not applied at the receiving device at the start of communication between the transmitting device and the receiving device. In this exemplary embodiment, for the sake of brevity, it is considered that the synchronization function is in the receiving device and not in the transmitting device. However, the synchronization function may be applied to either or both devices. If the user of the receiving device receives media that requires synchronization, the user may press the software key of the receiving device. As a result, the receiving device begins to apply synchronization, that is, audio data delay, time scaling or other methods. Therefore, the user may choose whether or not to apply synchronization according to preference and current communication conditions. The synchronization function may be a function described above for an exemplary embodiment or a function well known to those skilled in the art. The user may further configure the default choice of receiving device call processing for synchronization through the setting of the receiving device option menu. For example, the synchronization function may be turned on by default, and the synchronization function may be turned off when the video data is significantly delayed.
As described above, the process for applying synchronization is shown in FIG. In an exemplary embodiment, the user's communication device is configured to start when the synchronization function is turned off. In another exemplary embodiment, the communication device may be started when the synchronization function is turned on. During communication between the transmitting device and the receiving device, the transmitting device transmits audio data and video data to the receiving device based on standard protocol tools in order to enable the synchronization function. In step 80, the receiving device receives the data. The user will have a menu option to turn the sync feature on and off, for example, "audio and video sync" if the sync feature is currently off and "minimum audio delay" if the sync feature is currently on. You may have a software key that indicates "Synchronization". In another exemplary embodiment, the user may have a non-software key (ie, a dedicated hardware button) on the communication device to perform the above selection. If the user selects "audio and video synchronization" in step 82, the synchronization function is turned on in step 84. Next, the process proceeds to step 86, and the communication device determines whether or not the end of data has been received. If the end of data has been received, processing is stopped. If no end of data has been received, processing returns to step 82. If the user selects "minimize audio delay" in step 82, the synchronization between the audio data and the video data is turned off in step 88, and the process proceeds to step 86 described above.
An embodiment of the method following the process described above is shown in FIG. The method of this embodiment is a method of synchronizing video data received by a communication device and including a plurality of video data segments with audio data received by a communication device and containing a plurality of audio data segments. The method consists of step 90 of receiving the first segment of audio data on a communication device and the first of video data logically related to the first segment of audio data at the same time as or later than the first segment of audio data. It includes step 92 of receiving the segment on the communication device and step 94 of receiving a user input instruction to turn on or off the audio data and video data synchronization function.
Therefore, according to these exemplary embodiments, the user determines when the audio and video data should be synchronized. Lip synchronization is used when the speaker's lips are present in the image and the user wants to perform synchronization. Otherwise, synchronization is not used and the quality of voice conversations is optimized by minimizing audio data delays. An exemplary embodiment may be implemented with only the receiver, in which case signaling from the network or signal exchange with the transmitting device is not required.
According to the exemplary embodiments shown below, the audio data may be synchronized with the video data based on a novel method described later. Prior notification, face detection or user input is not required in the following exemplary embodiments. The synchronization process at the start of the conversation is as described with reference to FIG. The synchronization process is not desirable because it causes a large time delay in the audio data. However, according to the novel method of the present invention, the delay is shortened so that the time delay of the audio data does not bother the user of the communication system. According to one exemplary embodiment, the reduction in time delay may be achieved by time scaling the audio data.
More specifically, during the first part of the conversation, one or more segments of audio data play at a different speed than during the second part of the conversation, which is temporally slower than the first part. Will be done. The first part of the conversation may include a first segment and one or more subsequent segments. In this regard, in this exemplary embodiment, the first segment, which is generally defined as time-fast, has the normal speed of audio data with less delay before video data with greater delay. It may be further defined to last from the time indicating the beginning of a conversational interval that may be played at a slower speed to the time the audio data catches up with the video data, i.e. One way to monitor and determine when the audio and video data are synchronized is to monitor the time stamps of the audio and video data frames. The final segment of the first part of the conversation is associated with the end of the conversation interval and may last from the playback time of the conversation interval to the current time when the beginning of the silence period is detected. According to one exemplary embodiment, each part of the conversation may correspond to one conversation section.
Therefore, the audio data may start with a reduced time delay, and then during the first part of the conversation, time scaling the segment of the audio data to synchronize the audio data with the video data ("slow"). Playing the audio in "motion") adds more delay. There are various ways to time scale audio data so that the cognitive quality of the audio data does not deteriorate significantly. For example, Appendix 1 of ITU-T's Recommendation G.711 (which is incorporated herein by reference in its entirety) is this type of method, Waveform Shift Overlap. Includes a description of Add (WSOLA). When synchronization is performed, the audio and video data are played back at normal speed until just before the end. In other words, the audio data will be played faster than the video data, and since the two types of data will initially have the same length, by playing the first segment of the audio data at a slower speed than normal. , At least the first segment of audio data will be "extended". According to one exemplary embodiment, more segments (first segment of audio data and multiple subsequent segments) may be played at a slower rate in order to achieve synchronization of audio and video data. ..
At the end of the audio data received from the user, by using the time scaling of the received audio data again (to speed up at least the last segment of the audio data, i.e. to play the audio data in "fast motion"". Therefore, the reaction delay of the other user can be shortened. The audio data is out of sync with the video data, but the user can respond to the other user with a shorter reaction time and a shorter delay. Scaling at the end of the conversation will be described in more detail below. Scaling may be achieved on a device that does not achieve scaling at the beginning of the conversation. However, in one exemplary embodiment, both scaling methods are realized by at least one user. These new processes allow the user to interact more smoothly while keeping the video and audio data in sync for most of the conversation duration.
According to an exemplary embodiment, FIG. 10 shows that user 1 sends audio and video data to user 2 and user 2 also responds to the received audio and video data to user 1. Indicates that audio data and video data are being transmitted. The input of the audio data A1 and the video data V1, the coding and reception of the audio data A2 and the video data V2, and the decoding of the audio data A3 and the video data V3 are as described above with reference to FIG. The reproduction of the decoded audio data A3 and the decoded video data V3 is a novel format different from the reproduction described with reference to FIG. This will be described in detail next in relation to FIG.
According to one exemplary embodiment, the decrypted audio data A3 is played after being decoded, rather than delaying the decrypted audio data A3 until the decrypted video data V3 is available. Therefore, as shown in FIG. 10, the start time t of the decrypted audio data A3<sub>Astart</sub>Is the start time of the decrypted video data V3 t<sub>Vstart</sub>Faster. Therefore, the conversation time delay is shortened as compared with the conventional delay method. However, in order to achieve synchronization between the decrypted audio data A3 and the decrypted video data V3, at least the first part "A" of the audio data is at normal speed until the decrypted video data V3 is available. Plays at a slower speed than (predetermined speed). Even when the decrypted video data V3 becomes available, the audio data may be played at a slower rate in order to "catch up" the video data with the audio data. In order to synchronize the audio data and the video data, the video data is at time t<sub>Vstart</sub>After starting to play at a slow speed in, the audio data may be played after a period of "a". In one exemplary embodiment, the period "a" is a predetermined value, eg 2 seconds. According to another exemplary embodiment, the audio and video data are at a particular time t<sub>catch-up</sub>May be synchronized after (eg 1s) and "a" is t<sub>catch-up</sub>-Defined to be "A". At the end of period "a", time t<sub>s</sub>The audio speed is increased to the normal speed so that the audio data and the video data are synchronized with each other. In one exemplary embodiment, during time "a" the audio data rate is slowly (continuously and / or monotonically) increased to normal rate. In another embodiment, the audio data rate is rapidly (gradually) increased from low speed to normal speed.
In some methods, it is necessary to detect in advance that a period of silence is about to begin in order to increase speed at the end of the conversation section. One method may be to find the packet as soon as possible at the end of the sound buffer to enable speedup. During periods of silence, certain audio codecs (eg, AMR) can detect silence from differences in frame size and speed. At the end of the audio data, applying time scaling to at least the last segment of the audio data (increasing the speed of the audio data, i.e. "first motion" of the audio data) may reduce the reaction delay of the other user. .. As shown in FIG. 10, the playback speed of the audio data exceeds the normal speed in order to prompt the user to present the audio data A3 earlier than the decoded video data V3 immediately before the end of the decrypted video data V3. Will be increased to. In one exemplary embodiment, the audio data ends earlier by the time interval B than the decoded video data V3. The audio will not be synchronized with the video, but the user will be able to respond to the other user with a shorter reaction time and a shorter delay T1 (short by B). Scaling at the end of the audio (B or D) may be achieved by a device that does not achieve scaling at the beginning of the audio (A or C).
Similarly, user 1 may start audio earlier than without time scaling. This is because the user 2 starts sending information earlier and the audio data starts earlier in the communication device. Since the audio data is played at a low speed at the start, synchronization of the audio data and the video data is also realized after a certain time (after the partial C is played). This method prevents the user 1 from starting to send information while the information is being received from the user 2, for example, the user 1 from starting to speak. Further, this method shortens the experiential reaction time as compared with the conventional treatment, so that the interference level is lowered.
From the already received voice frame, for example, the end of the conversation section can be determined when a silent portion is detected. Based on this detection, in order to shorten the response time of the other user, the end of the audio data is such that the video data is played without audio data for the time interval D (because the audio data has already been played). It may be played at high speed. Therefore, according to the exemplary embodiments described above, most of the time (eg, periods A, B, C and), with minimal impact on conversation quality by reducing voice delays. Audio (except D) is synced with video.
An embodiment of the method of scaling the first segment is shown in FIG. The method of this embodiment is a method of synchronizing video data received by a communication device and including a plurality of video data segments with audio data received by a communication device and containing a plurality of audio data segments. The method consists of step 110 of receiving the first segment of audio data on a communication device and the first of video data that is logically related to the first segment of audio data at the same time as or later than the first segment of audio data. Step 112 to receive the segment on the communication device, step 114 to scale the first segment of the audio data, and play back the first segment of the scaled audio data before receiving or decoding the first segment of the video data. Includes step 116 and.
An embodiment of another method of scaling the final segment of audio data will be described with reference to FIG. The method includes step 120 of receiving the final segment of the audio data by the communication device, step 122 of receiving the final segment of the video data by the communication device at the same time as or later than the final segment of the audio data, and the final of the audio data. It includes step 124 of scaling the segment and step 126 of playing back the final segment of the scaled audio data before receiving or decoding the final segment of the video data. The steps shown in FIG. 12 may be performed in association with the steps shown in FIG. 11, but may be performed independently of the steps shown in FIG.
FIG. 13 is a flowchart showing a step of a method of synchronizing video data received by a communication device and including a plurality of video data segments with audio data received by a communication device and containing a plurality of audio data segments. The method consists of step 130 of receiving the first segment of audio data in a communication device and the first of video data logically related to the first segment of audio data at the same time as or later than the first segment of audio data. It includes step 132 of receiving the segment on the communication device and step 134 of applying a synchronization process between the first segment of audio data and the first segment of video data based on a predetermined index. The synchronization process may be one of the new synchronization processes described above.
In the above, various exemplary embodiments have been described individually. However, it will be appreciated by those skilled in the art that any combination of those exemplary embodiments can be used.
An exemplary embodiment disclosed is a communication device, system, method and computer program that sends audio and video data from a transmitting device to a receiving device and synchronizes the audio and video data on the receiving device. provide. It should be understood that the above description is not intended to limit the present invention. The exemplary embodiments described above are intended to include alternative configurations, modifications and equivalent structures within the spirit and scope of the invention as defined by the appended claims. Further, in the detailed description of the exemplary embodiment, a number of specific detailed matters are described in order to comprehensively understand the invention described in the claims. However, it will be appreciated by those skilled in the art that various embodiments may be implemented without including such specific details.
Similarly, those skilled in the art will appreciate that exemplary embodiments may be implemented in wireless communication devices, wired communication devices or telecommunications networks, or as methods or as computer programs. Therefore, the exemplary embodiment may take the form of a completely hardware embodiment, but may also be in the form of an embodiment that combines a hardware aspect and a software aspect. Further, the exemplary embodiment may take the form of a computer program stored in a computer-readable storage medium in which computer-readable instructions are realized. Any suitable computer-readable medium may be utilized, including hard disks, CD-ROMs, digital versatile discs (DVDs), optical storage devices, or magnetic storage devices such as floppy disks or magnetic tapes. Examples of other computer-readable media include, but are not limited to, flash memory or other well-known memory.
Although the features and elements of the exemplary embodiments of the invention have been described in particular combinations in embodiments, each feature or element can be used alone without the other features and elements of the embodiment. Alternatively, it can be used in various combinations with or without other disclosed features and elements. The methods or flowcharts presented in this application may be implemented as computer programs, software or firmware that are tangibly implemented in a computer-readable storage medium for execution by a general purpose computer or general purpose processor.
An exemplary embodiment may be implemented in an application specific integrated circuit (ASIC) or digital signal processor. Suitable processors include, for example, general purpose processors, dedicated processors, traditional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with DSP cores, controllers, microcontrollers, application specific integrated circuits. Includes (ASIC), field programmable gate array (FPGA) circuits, any other type of integrated circuit (IC) and / or state transition machine. Software-related processors may be used to implement radio frequency transceivers for use in user terminals, base stations or any host computer. User terminals include cameras, video camera modules, videophones, speakerphones, vibrating devices, speakers, microphones, television transceivers, hands-free headphones, keyboards, Bluetooth modules, frequency modulation (FM) wireless units, and liquid crystal display (LCD) displays. Related to units implemented in hardware and / or software such as units, organic light emitting diode (OLED) display units, digital music players, video game player modules, internet browsers and / or any wireless local area network (WLAN). May be used.
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP04137987A | Cites | Japan |
| JP2009508386A | Cites | Japan |
| WO2007031918A2 | Cites | World Intellectual Property Organization (WIPO) |
8 members in 4 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 2008051420 | Sweden | W | |
| 2008051420 | Sweden | W | |
| 2008051420 | – | – | – |
| WO2008SE51420 | – | – | – |
Members8
| Document | Office | Kind | |
|---|---|---|---|
| WO2010068151A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2356817A1 | European Patent Office (EPO) | A1 | |
| JP2012511279A | Japan | A | |
| US2012169837A1 | United States of America | A1 | |
| JP5363588B2This record | Japan | B2 | |
| EP2356817A4 | European Patent Office (EPO) | A4 | |
| US9392220B2 | United States of America | B2 | |
| EP2356817B1 | European Patent Office (EPO) | B1 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 5363588
- Publication, DOCDB
- 5363588
- Publication, EPODOC
- JP5363588B
- Application
- 2011539473
- Application, DOCDB
- 2011539473
- Application, EPODOC
- JP20110539473
Titles2
- Japanese
- 受信オーディオデータをビデオデータと同期させるための装置及び方法
- English
- Devices and methods for synchronizing received audio data with video data
Classification
- CPC, 6
- H04N7/147
- H04N21/2368
- H04N21/41407
- H04N21/4341
- H04N21/4788
- H04N21/43072
- IPC, 5
- H04N7 14
- H04N7 26
- H04N7 52
- H04N19 00
- H04N19 70