Lip-sync correcting device and lip-sync correcting method
13 claims: 7 independent, 6 dependent
- 1映像信号と音声信号の再生時間のずれを補正するリップシンク補正システムであって、 映像信号及び音声信号をそれぞれの伝送経路において、ソース装置からシンク装置へ、それぞれ所定のインタフェースに準拠して送信し、前記所定のインタフェースには双方向インタフェースが含まれ、 前記双方向インタフェースに準拠した伝送経路に接続された機器 における映像信号 の 遅延時間の総和 を示す遅延情報を取得し、前記取得した遅延情報を用いて前記再生時間のずれを補正するコントローラを備え、 前記コントローラは、前記再生時間のずれを補正するために前記音声信号の伝送経路上の機器 において音声信号 に追加遅延を付与する、リップシンク補正システム。
- 2前記コントローラは、前記追加遅延を前記音声信号の伝送経路上の最下流の機器 において音声信号 に付与する、請求項1記載のリップシンク補正システム。
- 3前記遅延情報を前記音声信号の付加情報として伝送する、請求項1又は2に記載のリップシンク補正システム。
- 4前記双方向インタフェースを介して受信される映像信号に関する遅延時間は、映像信号の伝送経路に接続する機器の、映像信号に関する遅延時間を累積的に加算して得られた情報である、 請求項1から3のいずれかに記載のリップシンク補正システム。
- 5前記双方向インタフェースを介して受信される音声信号に関する遅延時間は、音声信号の伝送経路に接続する機器の、音声信号に関する遅延時間を累積的に加算して得られた情報である、 請求項1から4のいずれかに記載のリップシンク補正システム。
- 6前記双方向インタフェースは、HDMI(High Definition Multimedia Interface)またはIEEE1394である、請求項1から3のいずれかに記載のリップシンク補正システム。
- 7映像信号と音声信号の再生時間のずれを補正するリップシンク補正装置であって、 前記映像信号及び音声信号は、それぞれの伝送経路において、ソース装置からシンク装置へ、それぞれ所定のインタフェースに準拠して送信され、前記所定のインタフェースには双方向インタフェースが含まれており、 前記リップシンク補正装置は、前記双方向インタフェースに準拠した伝送経路に接続された機器 における映像信号の遅延時間の総和 を示す遅延情報を取得し、前記取得した遅延情報を用いて前記再生時間のずれを補正するコントローラを備え、 前記コントローラは、前記再生時間のずれを補正するために前記音声信号の伝送経路上の機器 において音声信号 に追加遅延を付与する、リップシンク補正装置。
- 8前記コントローラは、前記追加遅延を前記音声信号の伝送経路上の最下流の機器 において音声信号 に付与する、請求項7記載のリップシンク補正装置。
- 9前記遅延情報を前記音声信号の付加情報として伝送する、請求項7又は8に記載のリップシンク補正装置。
- 10前記双方向インタフェースを介して受信される映像信号に関する遅延時間は、映像信号の伝送経路に接続する機器の、映像信号に関する遅延時間を累積的に加算して得られた情報である、 請求項7から9のいずれかに記載のリップシンク補正装置。
- 11前記双方向インタフェースを介して受信される音声信号に関する遅延時間は、音声信号の伝送経路に接続する機器の、音声信号に関する遅延時間を累積的に加算して得られた情報である、 請求項7から10のいずれかに記載のリップシンク補正装置。
- 12前記双方向インタフェースは、HDMI(High Definition Multimedia Interface)またはIEEE1394である、請求項7から10のいずれかに記載のリップシンク補正装置。
- 13映像信号と音声信号の再生時間のずれを補正するリップシンク補正方法であって、 a)映像信号及び音声信号を、ソース装置からシンク装置へ、それぞれの伝送経路において、それぞれ所定のインタフェースに準拠して送信し、 前記所定のインタフェースは双方向インタフェースを含み、 前記シンク装置は、 b)前記双方向インタフェースに準拠した伝送経路上の機器の総遅延時間を示す遅延情報を取得し、 c)前記取得した遅延情報を用いて映像信号と音声信号のずれを補正し、 d)前記ずれを補正するために前記音声信号の伝送経路上の機器に追加遅延を付与する、リップシンク補正方法。
Independent claims13
134 paragraphs, as filed
The present invention relates to a device and a method for correcting a time difference between a video signal and an audio signal, that is, a lip-sync shift.
Lip-sync is the synchronization of video and audio signals during playback, and various techniques for realizing lip-sync have been conventionally devised. FIG. 24 shows a configuration example of a conventional video / audio synchronization system. In the source device 210, the same time stamp (time code) is multiplexed for each of the video signal and the audio signal. The video signal and the audio signal in which the time stamps are multiplexed follow various paths such as repeaters 214 and 216, and both of them are finally input to the sink device 220. The sink device 220 takes out the time stamp of the video signal and the time stamp of the audio signal, and the time code comparator 222 detects the difference. Based on this difference, the video delayer 224 and the audio delayer 226 provided in each final path are controlled. For example, if the video signal is delayed more than the audio signal, an additional delay is added to the audio signal to correct the total delay time of both.
In Patent Document 1, the time stamp of the video signal and the time stamp of the audio signal are compared, and one of the video signal and the audio signal is delayed so that the two match, so that the audio time shift is automatically performed. The technique for correcting the above is disclosed.<patcit num="1"><text>Japanese Unexamined Patent Publication No. 2003-259314</text></patcit>
<p> In the conventional system, when the time stamps of the video signal and the audio signal can be detected by pairwise comparison at the same time, that is, when both time stamps are input to the same device at the same time, the deviation can be detected and the lip sync correction can be performed. However, in a topology in which the transmission paths of the video signal and the audio signal are separated, only one time stamp can be recognized, so the time difference between the video signal and the audio signal cannot be detected, and lip sync correction cannot be performed. is there.</p><p> The present invention has been made to solve the above problems, and an object of the present invention is a limp sync correction capable of reliably realizing lip sync correction even in a configuration in which the paths of a video signal and an audio signal are separated. To provide the device.</p>
<p> In the first aspect of the present invention, there is provided a lip-sync correction system that corrects a difference in reproduction time between a video signal and an audio signal. The lip-sync correction system transmits a video signal and an audio signal from the source device to the sink device in accordance with a predetermined interface in each transmission path. Certain interfaces include bidirectional interfaces. The lip-sync correction system is a device on a transmission path that conforms to a bidirectional interface.<u style="single">Sum of the delay times of the video signals in</u>It is provided with a controller which acquires the delay information indicating the above and corrects the deviation of the reproduction time by using the acquired delay information.</p><p> According to the above configuration, the controller acquires the delay information of the device through the transmission path of the bidirectional interface, and uses the acquired information to adjust the difference in the reproduction time between the video signal and the audio signal. According to this method, the controller can adjust the reproduction time difference between the video signal and the audio signal on the other transmission path based on the delay time information obtained from one transmission path, so that the video signal and the audio signal are transmitted. Even if the routes are different, lip sync correction can be reliably performed. In addition, the controller is a device on the audio signal transmission path to correct for playback time lag.<u style="single">Audio signal in</u>Gives an additional delay. Since the video signal is generally delayed with respect to the audio signal, the device on the audio signal transmission path<u style="single">Audio signal in</u>Lip-sync correction can be reliably realized by adding an additional delay to.</p><p> In the lip-sync correction system, the video signal may be transmitted by the bidirectional interface and the audio signal may be transmitted by the unidirectional interface. By transmitting the video signal via a bidirectional interface, the delay information of the device on the transmission path of the video signal can be acquired from the downstream of the transmission path of the video signal, and the acquired delay information can be obtained from the audio signal. It can be transmitted from the upstream of the transmission path. As a result, lip-sync correction becomes possible in the device on the transmission path of the audio signal.</p><p> The controller may add an additional delay to the most downstream device on the voice signal transmission path. By adjusting the additional delay time with the most downstream device, lip-sync correction can be realized more reliably and accurately.</p><p> The delay information may be transmitted as additional information of the audio signal. Thereby, even if the transmission path of the audio signal is a unidirectional interface, the delay information of the video signal can be acquired through the transmission path of the audio signal.</p><p> In the second aspect of the present invention, there is provided a lip-sync correction device that corrects a difference in reproduction time between a video signal and an audio signal. In the lip-sync correction device, the video signal and the audio signal are transmitted from the source device to the sink device in each transmission path according to a predetermined interface. A given interface includes a bidirectional interface. A lip-sync compensator is a device connected to a transmission path that conforms to a bidirectional interface.<u style="single">Sum of the delay times of the video signals in</u>It is provided with a controller that acquires the delay information indicating the above and corrects the deviation of the reproduction time by using the acquired delay information. The controller is a device on the transmission path of the audio signal to correct the deviation of the playback time.<u style="single">Audio signal in</u>Gives an additional delay.</p><p> A third aspect of the present invention provides a lip-sync correction method for correcting a difference in reproduction time between a video signal and an audio signal. The lip-sync correction method transmits a video signal and an audio signal from a source device to a sink device in accordance with a predetermined interface in each transmission path. Certain interfaces include bidirectional interfaces. The lip-sync correction method is a device on a transmission path that conforms to a bidirectional interface.<u style="single">Video signal in</u>Delay time<u style="single">Sum</u>The delay information indicating the above is acquired, and the deviation between the video signal and the audio signal is corrected by using the acquired delay information. Equipment on the audio signal transmission path to correct this deviation<u style="single">Audio signal in</u>Gives an additional delay.</p>
<p> According to the present invention, the controller acquires the delay information of the device through the transmission path of the bidirectional interface, and uses the acquired information to adjust the difference in the reproduction time between the video signal and the audio signal. In this way, the controller can adjust the reproduction time difference between the video signal and the audio signal on the other transmission path based on the delay time information obtained from one transmission path, so that the transmission path of the video signal and the audio signal can be adjusted. Even if they are different, lip sync correction can be reliably performed.</p>
Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In each figure, the same components or components having the same function are designated by the same reference numerals.
(Basic concept) First, the basic concept of the present invention regarding lip-sync correction will be described with reference to FIG.
The present invention is a lip-synching system in which a source device 400 that generates an audio signal and a video signal and a sink device 410 that reproduces and outputs the audio signal and the video signal are connected directly or via a repeater. It corrects the deviation of the sink. Further, the present invention includes a controller 420 as a function of adding an additional delay to correct the deviation of the lip sync.
In the present invention, video transmission is performed by a video transmission path using a bidirectional data interface such as IEEE1394 or HDMI. The source device 400 uses a command sequence for IEEE1394 and an EDID (Extended Display Identification Data) system for HDMI (High-Definition Multimedia Interface) via a video transmission path for the total latency time of the video signal (device on the video transmission path). Inquire about the total latency of. When the source device 400 acquires the information of the total latency time through the uplink, the source device 400 transmits the information of the total latency time to another device on the voice transmission path. Other devices on the audio transmission path transmit the audio latency information integrated for each device together with the total latency time information of the video. The last device on the audio transmission path outputs the total value of the audio latency as sDa together with the total latency time (TLv) of the video through the audio transmission path.
Controller 420 detects sDa and TLv only from the final last device. If TLv> sDa, the controller 420 adds an extra delay (TLv-sDa) to the last device on the audio transmission path.
As described above, the present invention acquires the total latent time information (delay information) of the video signal in advance and transmits the information to the device on the audio transmission path through the audio transmission path. Then, by adjusting the output time of the audio signal in the last device based on the latent time information of the finally obtained video signal and audio signal, the deviation between the video signal and the audio signal, that is, the deviation of the lip sync is eliminated. to correct.
The function of the controller 420 is typically included in the sink device, but if the function of adding an additional delay to the last device can be realized based on the finally obtained total latency information. , It may be provided in a device other than the sink device (for example, a source device or a repeater) on the voice transmission path. Further, the additional delay does not necessarily have to be added to the last device on the voice transmission path, and may be provided to any device on the voice transmission path. Further, the additional delay may be distributed and added to a plurality of devices.
Hereinafter, some specific embodiments of the present invention based on the above concept will be described.
(Embodiment 1) 1. System configuration FIG. 2 is a diagram showing a configuration of a video / audio reproduction system to which the concept of lip-sync correction of the present invention is applied.
The video / audio playback system includes a source device 100 that generates a video / audio signal, repeaters 110 and 130 that amplify the audio signal, a repeater 120 that amplifies the video signal, and a sink device 140 that displays the video based on the video signal. And a sink device 150 that outputs audio based on the audio signal.
1.1 Source device The source device is a device that is located at the uppermost stream of the video / audio signal transmission path and is an output source of the video signal and the audio signal. In the present embodiment, the source device 100 is a DVD player that is a playback device of the DVD media 101, and generates and outputs a video signal 200 and an audio signal 300.
Figure 3 shows the hardware configuration. The source device 100 includes a pickup 11 that reads information from the DVD media 101 and converts it into an electric signal, a front-end processor 13 that receives an output signal from the pickup 11 and generates a video signal and an audio signal, and an operation of the entire source device 100. A system controller 15 that controls the above, a RAM 17 that operates as a work area, and a ROM 19 that stores predetermined information are provided. Further, the source device 100 includes an HDMI interface unit 21 for exchanging video signals and the like with an external device, and an IEC interface unit 22 for exchanging audio signals and the like. In the source device 100, the system controller 15 executes a predetermined program to realize the functions and processing units described below. As shown in FIG. 2, the source device 100 has a voice variable delay unit 102, whereby the voice signal 300 is delayed by a predetermined time (20 ms in this example) and output.
The source device is not limited to a DVD player, and may be another type of media playback device such as a hard disk player.
1.2 Sink device A sink device is a device that is located at the most downstream of a video / audio signal transmission path and outputs video or audio.
In the present embodiment, the sink device 140 is a video display device and has the hardware configuration shown in FIG. The sink device 140 includes an HDMI interface unit 31 that receives a digital video signal in accordance with the HDMI interface, a video decoder 33 that decodes the received digital video signal, and a liquid crystal display device (hereinafter referred to as "LCD") that displays video. ) 143, a driver 37 that generates a video signal from the decoded video signal to drive the LCD 143, a controller 39 that controls the operation of the entire sink device 140, a RAM 41 that functions as a work area of the controller 39, and a predetermined number. It is equipped with a ROM 43 that stores information. In the sink device 140, the controller 39 executes a predetermined program to realize the functions and processing units described below. The video decoder 33 and the driver 37 in FIG. 4 constitute the video signal processing unit 142 shown in FIG. Latent time of the video signal of the video signal processing unit 142 (video) latency) Let Lv be 80ms. The latency time is the delay time from the input of a signal to the output of the signal in the device. The information of the latency time Lv is stored in ROM43 of the sink device 140.
The sink device 150 is an audio output device and has the hardware configuration shown in FIG. As shown in FIG. 5, the sink device 150 has an IEC interface unit 51 that receives a digital audio signal in accordance with IEEE60958, an audio decoder 152 that decodes the received digital audio signal, and converts the decoded signal into an analog signal. A D / A converter 55 for conversion, an amplifier 57 for amplifying an analog signal, a speaker 153 for outputting audio according to the output of the amplifier 57, a controller 61 for controlling the operation of the entire sink device 150, and a work of the controller 61. It includes a RAM 63 that functions as an area and a ROM 65 that stores predetermined information. With reference to FIG. 2, the sink device 150 has a voice variable delay unit 151 that delays the voice signal by a predetermined time (10 ms in this example). The latency (audio latency) La of the audio signal of the decoder 152 is 10 ms. The information of the latency time La is stored in ROM 65 of the sink device 150.
The number and types of sink devices are not limited to those described above.
1.3 Repeater A repeater is a device that intervenes in the middle of a video / audio signal transmission path, for example, an amplifier. Figure 6 shows the hardware configuration. The repeater 110 processes the received digital signal with the HDMI interface unit 71 that receives the video / audio digital signal in compliance with HDMI, the IEC interface unit 72 that receives the video / audio digital signal in compliance with IEEE1394, and the received digital signal. It includes a controller 73, a RAM 75 that functions as a work area of the controller 73, and a ROM 79 that stores predetermined information such as a program. When the controller 73 executes a predetermined program, the functions and processing units described below are realized.
In FIG. 2, the repeater 110 inputs the video signal 200 and the audio signal 300 from the source device 100, amplifies them, and outputs them as the video signal 201 and the audio signal 301. The repeater 110 has a re-encoding 112, and the re-encoding 112 once decodes the audio signal 300, reads the information, and re-encodes the read information. The latency time La of this re-encoder 112 is 10 ms. Information on the latency time La is stored in ROM 79 of the repeater 110.
The repeater 120 has a video signal processing unit 122 that amplifies the video signal, inputs the video signal 201, amplifies it, and then outputs it as the video signal 202. The latency time Lv of the video signal processing unit 122 is set to 20 ms. The repeater 130 includes a signal processing unit 132 that performs predetermined processing on the voice signal 301, and outputs the processed signal as the voice signal 302. The latency time La of the signal processing unit 132 is set to 50 ms. The repeaters 120 and 130 have the same configuration as that shown in FIG. 6, and their respective latency times are stored in their respective ROMs.
The number and types of repeaters are not limited to those described above.
1.4 Interface In this embodiment, HDMI (High-Definition Multimedia Interface) is used as a video signal transmission interface. HDMI is a digital video / audio input / output interface standard established in December 2002, mainly for home appliances and AV equipment. Video, audio, and control signals can be transmitted and received together with a single cable, and control signals can be optionally transmitted in both directions.
In the present invention, it is possible to transmit a high-speed digital video signal from the source device 100 to the sink device 140 and sequentially transmit the profile information of the sink device 140 to the repeater 120, the repeater 110, and the source device 100 by HDMI. , Realize the function of the uplink. Hereinafter, the function of this uplink is referred to as "EDID (Extended Display Identification Data) line".
Figure 7 shows the HDMI configuration diagram. As shown in the figure, HDMI has three data channels and one clock channel, through which video data, audio data and other data are transmitted. In addition, HDMI has a display data channel (hereinafter referred to as "DDC") that exchanges configurations and status between devices. In addition, HDMI has an optional CEC line that allows control signals to be transmitted bidirectionally between various AV devices. In this embodiment, DDC is used to transmit information on the latency time of video and audio of the device. Information on the latency time may be transmitted via a CEC line instead of the DDC. The same function can be realized by using IEE1394 instead of HDMI.
Details of HDMI and EDID are disclosed in the following documents, for example. "High-Definition Multimedia Interface Specification Version 1.1", Hitachi, Ltd., 6 other companies, May 20, 2004, Internet <http: // www .hdmi.org/download/HDMI_Specification_1.1.pdf>
In the present embodiment, the IEEE60958 interface, which is a unidirectional interface, is used for the transmission of the voice signal. In the following embodiments, HDMI, which is a bidirectional interface, may be used for audio signal transmission.
2. Operation In the above system configuration, the video signal 200 is transmitted from the source device 100 to the sink device 140 via the repeater 110 and the repeater 120 in accordance with the HDMI interface. The audio signal 300 is transmitted from the source device 100 to the sink device 150 via the repeater 110 and the repeater 130 in accordance with the IEC60958 interface.
In the video signal transmission path, the total latency time TLv of the video signal from the source device 100 to the sink device 140 is the latency time Lv (20 ms) of the video signal processing unit 122 of the repeater 120 and the video signal processing unit of the sink device 140. It is the sum of 142 latency time Lv (80ms) and 100ms.
On the other hand, in the audio signal transmission path, the total latency time La of the audio signal from the source device 100 to the sink device 150 is the latency time La (10 ms, 10 ms, respectively) of the repeater 110, the repeater 130, and the sink device 150. It is the sum of 50ms and 10ms), which is 70ms.
That is, while the total latency time TLv of the video signal is 100 ms, the total latency time La of the audio signal is 70 ms (note that the delay time of the audio variable delay unit 151 whose time can be adjusted is included in this). Not.). Therefore, if nothing is done, the audio signal will be reproduced faster by 30 ms. In this system, the process for resolving the time difference between the video signal and the audio signal during reproduction will be described below.
2.1 Overall flow FIG. 8 is a flowchart showing the overall operation of this system. First, the most upstream source device acquires the total latency time (total latency time) TLv of the video signals of each device on the video transmission path (S11). The details of this process will be described later.
Next, in the audio transmission path, the total latency time TLv of the video signal from the source device 100 to the sink device 150, and the total delay time and latency time of the audio signals of each device (hereinafter referred to as "cumulative delay time") sDa. Send (S12). At that time, the cumulative delay time sDa of the voice signal is sequentially sent to each device on the voice transmission path, and is sent downstream in order while the value of the latency time of each device is added.
Finally, in the sink device located at the most downstream, the voice is delayed and output based on the value of the cumulative delay time sDa of the voice signal (S13). By adjusting the audio output time in this way, the difference in video and audio output time can be eliminated.
2.2 Acquisition of total latency time TLv of video signal The details of the acquisition operation of the total latency time TLv of the video signal (step S11 in FIG. 8) will be described.
2.2.1 Acquisition of TLv by source device 100 With reference to FIG. 9, acquisition of the total latency time TLv of the video signal by the source device 100 will be described. The source device 100 transmits a TLv transmission command to devices (repeaters, sink devices) downstream of the video transmission path (S41). When a TLv transmission command is transmitted from the source device, the command is sequentially transmitted to downstream devices, and the information sLv to which the latency Lv of each device is sequentially added is transmitted from the downstream side to the source device 100. Finally, the integrated value (hereinafter referred to as "cumulative latency time") sLv of the latency time Lv of each device (repeater, sink device) is transmitted to the source device 100. The details of this process will be described later.
When the source device 100 receives the cumulative latency sLv from the downstream device (S42), it reads its own latency Lv from ROM 19 (S43), and adds the read latency Lv to the cumulative latency sLv received from the downstream device. (S44). As a result, the total latency time TLv of the video signal can be obtained.
2.2.2 Sending sLv from repeater The operation of the repeater 120 when the TLv transmission command is received will be described with reference to FIG. When the repeaters 110 and 120 receive the TLv transmission command (S31), they transfer the TLv transmission command to the downstream device (repeater, sink device) (S32) and wait for the transmission of the cumulative latency sLv from the downstream device (S33). .. Upon receiving the cumulative latency sLv from the downstream device, the repeaters 110 and 120 read their own latency Lv from ROM 79 (S34). The read latency Lv is added to the cumulative latency sLv from the downstream device (S35), and the newly obtained cumulative latency sLv is transmitted to the upstream device (repeater, source device) (S36).
2.2.3 Sending sLv from the sink The operation of the sink device 140 when the TLv transmission command is received will be described with reference to FIG. When the sink device 140 receives the TLv transmission command (S21), it reads its own latency Lv from the ROM 43 (S22). The read latency Lv is transmitted to the upstream equipment (repeater, source device) as the cumulative latency sLv (S23).
2.3 Transmission of total latency time TLv of video signal and cumulative delay time sDa of audio signal The details of the transmission operation of the total latency time TLv of the video signal and the cumulative delay time sDa of the audio signal (step S12 in FIG. 8) will be described.
In the audio transmission path, the total latency time TLv of the video signal and the cumulative delay time sDa of the audio signal are sequentially transmitted from the source device 100 to the sink device 150. At this time, the latent time of each device on the voice transmission path is sequentially added to the cumulative delay time sDa of the voice signal, and finally the voice transmission is performed except for the most downstream device (sink device 150). The total value of the latency La of all the devices on the route is transmitted to the most downstream device (sink device 150).
2.3.1 Transmission of TLv and sDa by source device With reference to FIG. 12, the transmission operation of the total latency time TLv of the video signal and the cumulative delay time sDa of the audio signal from the source device 100 will be described.
After receiving the total latency time TLv of the video signal (S51), the source device 100 reads the latency time La of its own audio signal from ROM 19 (S52), and sets the read latency time La as the cumulative delay time sDa of the audio signal. , Transmits to downstream equipment in the audio signal transmission path (S53).
2.3.2 Repeater transmission of TLv and sDa With reference to FIG. 13, transmission of the total latency time TLv of the video signal and the cumulative delay time sDa of the audio signal will be described.
When the repeaters 110 and 130 receive the total latency time TLv of the video signal and the cumulative delay time sDa of the audio signal from the upstream device (S61), the repeaters 110 and 130 read the latency time La of their own audio signal from the ROM 79 (S62). The repeaters 110 and 130 add the read latent time La to the received cumulative delay time sDa and transmit it to downstream devices (repeater, sink device) (S63).
2.4 Audio output time adjustment The details of the audio output time adjustment operation (step S13 in FIG. 8) of the sink device 150 will be described with reference to FIG.
When the sink device 150 receives the total latency time TLv of the video signal and the cumulative delay time sDa of the audio signal from the upstream device (S71), the sink device 150 obtains the delay time (AD) for adjusting the output time (S72). The delay time (AD) is obtained by subtracting the cumulative delay time sDa of the audio signal from the total latency time TLv of the video signal. Then, the audio is output with a delay of the delay time (AD) (S73).
2.5 Specific example of acquisition of total latency time TLv The individual processes have been described above, but the flow of the entire system with the configuration shown in Fig. 2 will be described.
The source device 100 issues a TLv transmission command in order to acquire the total latency time TLv of the video signal from the source device 100 to the sink device 140 using the EDID line. By the TLv transmission command, the sink device 140 transmits "sLv = 80ms", which is a parameter indicating its own latency Lv (80ms), to the repeater 120 via the EDID line. The repeater 120 transmits "sLv = 100ms" indicating the value obtained by adding its own latency time (20ms) to the value indicated by the received sLv = 80ms to the upstream repeater 110. Since the repeater 110 does not have the latency time of the video signal, "sLv = 100ms" is transferred to the source device 100 as it is.
The source device 100 sets the received sLv value (= 100 ms) as the total latency time TLv (100 ms) of the video signal. The source device 100 multiplexes and transmits the total latency time TLv (100ms) of this video signal to the audio signal 300 with the fixed parameter TLv = 100ms. At the same time, the source device 100 transmits the cumulative delay time sDa of the audio signal to the sink device 150 as ancillary information of the audio signal 300. Here, in the example of FIG. 2, the delay time of 20 ms is set for the voice variable delay unit 102 in the source device 100. Therefore, the source device 100 multiplexes and transmits the parameter sDa = 20ms indicating the delay time 20ms of the voice variable delay unit 102 to the voice signal 300.
The repeater 110 receives the parameter "sDa = 20ms" from the source device 100, accumulates its own latency time La (10ms) on the value indicated by the parameter, and multiplexes and transmits "sDa = 30ms" to the audio signal 301. Similarly, the repeater 130 adds its own latent time La (50 ms) to the cumulative delay time sDa (30 ms), and multiplexes sDa = 80 ms into the voice signal 302 and outputs it.
The sink device 150 reads the cumulative delay time sDa (80 ms), adds its own latent time La (10 ms) to that value, and obtains 90 ms as the total delay time of the audio signal before correction. At the same time, the sink device 150 reads the fixed parameter "TLv = 100ms" indicating that the total latency time TLv of the video signal is 100ms. In this way, the sink device 150 can know the total latency time (100 ms) of the video signal and the total delay time (90 ms) of the audio signal before correction, and the difference of 10 ms is used as the correction value (AD). Controls the voice variable delay unit 151. As a result, the total delay time of the audio signal is corrected to 100 ms.
If the audio and video of the DVD media 101 are synchronized with the source device 100, the video signal reproduced by the sink device 140 and the audio signal reproduced by the sink device 150 are both delayed by 100 ms, so that the final result is achieved. It enables audio-video synchronization, that is, playback with lip-sync.
3. Summary As described above, in the present embodiment, each device on the video transmission path sequentially accumulates the delay time (latent time) of its own video signal and transmits it to the video transmission path in accordance with the request from the source device. The source device acquires the total delay time of the video signal in the transmission path from the source device to the sink device, and transmits the parameter of the total delay time of the video signal to the audio signal path together with the audio signal and the audio signal delay time. Each device on the voice transmission path sequentially accumulates the delay time of its own voice signal and transmits it to the voice transmission path. As a result, in the device at the final stage of the audio transmission path, the total delay time of the video signal and the total delay time of the audio signal can be known, and the lip sync shift is performed by performing an additional delay of the audio signal so as to reduce the time difference. Can be corrected. These may have different video and audio paths in the end, and can remove topological restrictions. Further, since the constant parameter information of the latency time and the delay time of each device in each route can be used instead of the time stamp, the data transmission frequency can be reduced.
The functions described above operate as a whole system constituting the network by adding necessary functions to each of the individual devices and the interface connecting them. Therefore, if each device has a required function, the device can be replaced in the network, and the same effect can be obtained even if the topology of the network changes. In the following embodiments, application examples to various network connection forms will be described.
(Embodiment 2) In this embodiment, an application example of the present invention in a system configuration in which a source device and a sink device are directly connected will be described. The source device is a DVD player, and the sink device is a digital television receiver (hereinafter referred to as "digital TV").
Figures 15 (a), (b), and (c) show the system configuration in this embodiment. A video signal and an audio signal are transmitted from a DVD player 100, which is a source device, to a digital TV 140, which is a sink device, and both a video signal and an audio signal are reproduced on the digital TV 140.
The source device 100 is a playback means for the DVD media 101 and outputs both the video signal 200 and the audio signal 300 in accordance with HDMI. The sink device 140 incorporates a video signal processing unit 142 and an audio decoder 152, and outputs video and audio through the LCD 143 and the speaker 153. Speaker 153 includes an amplifier for audio signals. Here, the latency time Lv of the video signal of the video signal processing unit 142 is 80 ms, and the latency time La of the audio signal of the decoder 152 is 10 ms.
FIG. 15 (a) shows an example in the case where the source device 100 is a conventional device. Therefore, the source device 100 cannot acquire the total latency time TLv of the video signal from the sink device 140, nor can add the information to the audio signal. As a matter of course, the information of the cumulative delay time sDa of the voice signal cannot be added to the voice signal. On the other hand, the sink device 140 recognizes that the audio signal of the source device 100 does not have these additional information. Since there is no additional information, lip-sync correction is performed by the sink device 140 itself. Specifically, based on the latency time Lv (80ms) of the video signal in the sink device 140 and the latency time La (10ms) of the decoder 152, the audio delay unit 151 corrects the time difference of 70ms (= 80-10). Causes an additional delay. With this additional delay, the total value of the delay can be set to 80 ms for both the video signal and the audio signal, and lip-sync correction can be performed.
15 (b) and 15 (c) show an example of the case where the source device 100 is a device to which the idea of the present invention is applied. The source device 100 can acquire the cumulative delay time of the audio signal in addition to the total latency time TLv of the video signal through the EDID line, and can further add this information to the audio signal. FIG. 15B shows an example in which the source device 100 has the function of the controller 420 in FIG. 1 and an additional delay is added to the source device 100. FIG. 15 (c) shows an example in which the functions of the controller 420 in FIG. 1 are distributed to the source device 100 and the sink device 140, and the additional delay is also distributed to the source device 100 and the sink device 140.
In FIG. 15 (b), the source device 100 can acquire the latency time La (10 ms) of the audio signal from the sink device 140 through the EDID line in addition to the total latency time TLv (80) of the video signal. Therefore, the source device 100 can calculate the additional delay (70 ms) from these acquired information, and outputs the audio signal with a delay of the additional delay (70 ms). Due to this additional delay, the source device 100 transmits the audio signal as ancillary information with the cumulative delay time sDa of the audio signal as 70 ms. The sink device 140 can recognize that the total latency time of the video signal is 80 ms and the cumulative delay time of the audio signal is 70 ms. Further, since the sink device 140 can recognize that the latency time of the decoder 152 is 10 ms, it can recognize that the total delay time of the audio signal is 80 ms, and it can recognize that no further delay processing is required. .. In this way, lip-sync correction becomes possible as a system.
Further, FIG. 15 (c) is the same as that of FIG. 15 (b). However, the additional delay in the source device 100 is limited to 40 ms. The source device 100 delays the audio signal by an additional delay of 40 ms and outputs the signal. Due to this additional delay, the source device 100 adds the cumulative delay time sDa of the audio signal to 40 ms and transmits the audio signal 300. Since the sink device 140 can recognize the total latency time TLv (80 ms) of the video signal, the cumulative delay time sDa (40 ms) of the audio signal, and the latency time 10 ms of the decoder 152, the difference between the delay times of the video signal and the audio signal. Can be recognized as 30ms. Therefore, the sink device 140 performs an additional delay process of 30 ms at the audio delay unit 151. As a result, the total delay time of both the video signal and the audio signal can be corrected to 80 ms. In this way, lip-sync correction becomes possible as a system.
(Embodiment 3) The present embodiment differs from the second embodiment in that a repeater is inserted in the voice transmission path. 16 (a), (b), and (c) show the system configuration in this embodiment. A sink device 140, which is a digital TV, is connected to a source device 100, which is a DVD player, via a repeater 110, which is a multi-channel amplifier.
The source device 100 outputs the video signal 200 and the audio signal 300 reproduced from the DVD media 101 in accordance with HDMI.
The repeater 110 has an audio signal processing function, and is used for outputting audio with higher sound quality than the amplifier and speaker built in the sink device 140, or for multi-channelization. The repeater 110 has a re-encoder 112 and transmits its output to the sink device 140. In the repeater 110, the video signal is passed through without delay, but the audio signal has a latent time La of 10 ms in the re-encoder 112.
The sink device 140 is a digital TV, which incorporates a video signal processing unit 142 and an audio decoder 152, and presents video and audio through an LCD 143 and a speaker 153. In the sink device 140, the latent time Lv of the video signal processing unit 142 is 80 ms, and the latent time La of the decoder 152 is 10 ms.
FIG. 16A shows an example in the case where the source device 100 is a conventional device. Therefore, the source device 100 cannot recognize the total latency time TL of the video signal and cannot add the information to the audio signal. Further, the cumulative delay time sDa of the voice signal cannot be added to the voice signal.
Since the repeater 110 does not receive the information of the cumulative delay time sDa of the audio signal added to the audio signal 300 via the EDID line, it recognizes that there is no additional information. Therefore, the repeater 110 operates so as to perform the same function as the source device of the present invention. That is, the repeater 110 receives the latent time Lv (80 ms) of the video signal from the sink device 140 via the EDID line, and recognizes the total latent time TLv of the video signal as 80 ms. The repeater 110 transmits the information of the total latency time TLv (80 ms) to the sink device 140 as additional information of the audio signal. Further, the repeater 110 also multiplexes the latent time of the audio signal as the cumulative delay time sDa (10 ms) and transmits it to the sink device 140.
Further, the output of the repeater 112 in the repeater 110 is supplied to the speaker 113 via the audio variable delay unit 111 inside the repeater 110, a multi-channel amplifier (not shown), and the like. At that time, the repeater 110 controls the delay time of the audio variable delay unit 111 based on the difference (70 ms) between the total latency time TLv (80 ms) of the video signal and the cumulative delay time sDa (10 ms) of the audio signal. As a result, the lip-sync deviation between the audio output from the speaker 113 and the video reproduced by the LCD 143 is corrected. In this way, lip-sync correction is possible for the audio signal from the repeater 110 added for the purpose of high sound quality.
Furthermore, lip-sync correction is possible for the sound from the speaker of the sink device 140 that can be used as an auxiliary as follows. Based on the information from the repeater 110, the sink device 140 has a total latency time (TLv) of 80 ms, a cumulative audio delay time (sDa) of 10 ms, and a latency time of 10 ms of its own decoder 152. Calculate 60ms as. The sink device 140 controls the voice variable delay unit 151 based on this 60 ms information, and corrects the total delay time of each to 80 ms. In this way, lip-sync correction is possible for the sound of the sink device 140 as well.
16 (b) and 16 (c) show an example of the case where the source device 100 is a device to which the idea of the present invention is applied. The source device 100 can acquire the cumulative delay time sDa of the audio signal in addition to the total latency time TLv of the video signal through the EDID line, and can add this information to the audio signal.
In FIG. 16B, the source device 100 acquires the total latency time of each of the video signal and the audio signal from the repeater 110 and the sink device 140 through the EDID line. The total latency time TLv of the video signal can be obtained by issuing a TLv transmission command to a downstream device as described in the first embodiment. The source device 100 acquires the total latency time of the video signal of 80 ms.
Further, the source device 100 issues a command to the downstream device to acquire the total latency time of the voice signal. When each downstream device receives this command, it transmits it to the upstream device while adding its own latency time La to the cumulative latency time of the audio signal transmitted from the downstream, as in the case of the TLv transmission command. In this way, the source device 100 can recognize that the total latency time of the voice signal is 20 ms (= 10 ms + 10 ms).
Since the total latency time of the video signal is 80 ms and the total latency time of the audio signal is 20 ms, the source device 100 sets an additional delay of 60 ms (= 80 ms-20 ms) and adds an additional delay to the audio signal. The source device 100 transmits information on the total latency time TLv (80 ms) of the video signal and information on the cumulative delay time sDa (60 ms) of the audio signal as additional information of the audio signal.
Next, the repeater 110 receives this information and controls the audio variable delay unit 111 to correct the difference of 20 ms between the total latency time TLv (80 ms) of the video signal and the cumulative delay time sDa (60 ms) of the audio signal. .. As a result, the deviation of the lip sync between the sound output from the speaker 113 and the image reproduced by the LCD 143 is corrected. Lip-sync correction is possible mainly for the audio signal from the repeater 110 added for the purpose of high sound quality.
Further, as follows, lip-sync correction is also possible for the sound from the speaker of the sink device 140 used as an auxiliary. Based on the information from the repeater 110, the sink device 140 recognizes that the total latency time TLv of the video signal is 80 ms and the cumulative delay time sDa of the audio signal is 70 ms, and the delay time difference between the video signal and the audio signal is 10 ms. Is calculated. Since this time difference of 10 ms is equal to the latency time (10 ms) of the sink device 140, the sink device 140 outputs an audio signal without adding an additional delay. In this way, lip-sync correction can be performed on the sound of the sink device 140 as well.
FIG. 16 (c) is the same as that of FIG. 16 (b). However, the additional delay in the source device 100 is limited to 40 ms. Therefore, the source device 100 adds an additional delay of 40 ms to the audio signal and outputs the signal. Due to this additional delay, the source device 100 transmits information on the cumulative delay time sDa (40 ms) of the audio signal in addition to the latent time TLv (80 ms) of the video signal as ancillary information of the audio signal 300. The repeater 110 and the sink device 140 receive this attached information so that the repeater 110 has an additional delay of the voice variable delay unit 111 of 30 ms, and the sink device 140 has an additional delay of the voice variable delay unit 151 of 20 ms. Control each. In this way, the total delay time of the audio can be corrected to 80 ms for both the audio output of both the main and sub, and the lip-sync correction of the entire system becomes possible.
(Embodiment 4) In the present embodiment, an application example of the present invention will be described in a system configuration in which an audio signal path is branched and an amplifier is connected to the branched path in a configuration in which a source device and a sink device are connected. Connect the source device (DVD player) 100, sink device (multi-channel amplifier) 150, and sink device (digital TV) 140 via HDMI.
Figures 17 (a), (b), and (c) show the system configuration in this embodiment. The sink device 150 is a multi-channel amplifier and includes a decoder 152, an audio variable delay unit 151, a plurality of amplifiers, and a speaker. The source device 100 can acquire the total latency time TLv of the video signal and the cumulative delay time sDa of the audio signal and add them to the audio signal.
FIG. 17A shows an example in which the additional delay of the audio signal is not performed in the source device 100. The source device 100 transmits the information of the total latency time TLv (80 ms) of the video signal to the sink device 140 as additional information of the audio signal 300 via the EDID line. At the same time, the source device 100 adds the cumulative delay time sDa (0 ms) of the audio signal to the audio signal 300 and transmits it to the sink device 140 and the sink device 150.
The sink device 140 calculates a delay time difference of 70 ms from the total latency time TLv (80 ms) of the received video, the cumulative delay time sDa (0 ms) of the audio signal, and the latency time La (10 ms) of the audio signal of the sink device 140. Then, the voice variable delay unit 251 is controlled to add an additional delay of 70 ms.
Similarly, the sink device 150 calculates a time difference of 60 ms from the latency time La (20 ms) of the voice signal of the sink device 150, and controls the voice variable delay unit 151 to add an additional delay of 60 ms. In this way, lip-sync correction of the LCD 143, the speaker 153, and the speaker 253 becomes possible.
17 (b) and 17 (c) are cases where an additional delay of the audio signal is added in the source device 100.
In FIG. 17B, the source device 100 can know in advance the latency time La of the audio signals of the sink device 140 and the sink device 150 via the EDID line. Based on this, the source device 100 sets an additional delay of 60 ms in the voice variable delay unit 102 with reference to the path to the sink device 150 that gives the maximum latency time La. Information on the total latency time TLv (80 ms) and the cumulative delay time sDa (60 ms) of the voice signal is multiplexed on the voice signal 300 and transmitted to the sink device 140 and the sink device 150. In this case, the sink device 150 does not require an additional delay. The sync device 140 is based on the total latency time TLv (80 ms) of the received video, the cumulative delay time sDa (60 ms) of the audio signal, and the latency time La (10 ms) of the audio signal of the sync device 140. The output time difference of 10 ms is calculated, and the voice variable delay unit 251 is controlled to execute an additional delay of 10 ms. In this way, lip-sync correction of the LCD 143, the speaker 153, and the speaker 253 becomes possible.
Next, in FIG. 17C, the source device 100 can know in advance the latency time La of the audio signals of the sink device 140 and the sink device 150 via the EDID line. Based on this, the source device 100 sets an additional delay of 30 ms in the voice variable delay unit 102 with reference to the voice transmission path that gives the maximum latency time La. The source device 100 multiplexes the total latency time TLv (80 ms) of the video and the cumulative delay time sDa (30 ms) of the audio signal on the audio signal 300, and transmits the audio signal to the sink device 140 and the sink device 150. In this case, no additional delay is required in the sink device 140. The sink device 150 uses the audio signal latency time La (20 ms) of the sink device 150 itself based on the total latency time TLv (80 ms) of the received video and the cumulative delay time sDa (30 ms) of the audio signal, and the time difference is 30 ms. Is calculated, and the voice variable delay unit 151 is controlled to execute an additional delay of 30 ms. This enables lip-sync correction of the LCD 143, the speaker 153, and the speaker 253.
(Embodiment 5) In the present embodiment, an application example of the present invention will be described in a system configuration in which an audio signal path is branched and an amplifier is connected to the branched path in a configuration in which a source device and a sink device are connected.
18 (a), (b), and (c) show the system configuration in this embodiment. In this embodiment, the source device (DVD player) 100 and the sink device (digital TV) 140 are connected by HDMI. The source device (DVD player) 100 and the sink device (multi-channel amplifier) 150 are connected by an S / PDIF interface. The S / PDIF interface is an interface that transmits digital audio, etc. using the coaxial or optical cable specified in IEC 60958, and since it is a unidirectional transmission from the source device to the sink device, it is transmitted from the sink device to the source device. It does not have an uplink.
FIG. 18A shows an example in which the source device 100 does not perform the additional delay of the audio signal. The source device 100 transmits the information of the total latent time TLv (80 ms) of the video to the sink device 140 as additional information of the audio signal 300 via the EDID line, and transmits the cumulative delay time sDa (0 ms) of the audio signal to the audio signal. It is multiplexed into 300 and transmitted to the sink device 140 and the sink device 150.
The sync device 140 delays video and audio from the total latency time TLv (80 ms) of the received video, the cumulative delay time sDa (0 ms) of the audio signal, and the audio signal latency La (10 ms) of the sync device 140. The time difference of 70 ms is calculated, and the voice variable delay unit 251 is controlled to execute an additional delay of 70 ms.
Similarly, the sink device 150 has a delay time difference of 60 ms from the total latency time TLv (80 ms) of the received video, the cumulative delay time sDa (0 ms) of the audio signal, and the audio signal latency time La (20 ms) of the sink device 150. Is calculated, and the voice variable delay unit 151 is controlled to execute an additional delay of 60 ms. In this way, lip-sync correction of the LCD 143, the speaker 153, and the speaker 253 becomes possible.
FIGS. 18 (b) and 18 (c) show an example in which the additional delay of the audio signal is performed in the source device 100.
In FIG. 18B, since the sink device 150 is connected to the source device 100 by the S / PDIF interface which is unidirectional transmission, the source device 100 knows the latency time (La) of the audio signal of the sink device 150 in advance. Can't. However, in the source device 100, an additional delay of 70 ms is preset in the voice variable delay unit 102. The source device 100 multiplexes the total latency time TLv (80 ms) and sDa = 70 ms as the cumulative delay time of the audio signal into the audio signal 300 and transmits it to the sink device 140 and the sink device 150.
In this case, the sink device 140 does not require an additional delay. In the sink device 150, the delay time difference is -10 ms from the total latency time TLv (80 ms) of the received video, the cumulative delay time sDa (70 ms) of the audio signal, and the audio signal latency time La (20 ms) of the sink device 150 itself. Is calculated. However, in order to correct the delay time difference of -10 ms, it is necessary to delay the video signal, but this negative delay time difference is ignored because it is impossible in this embodiment. As a result, the audio signal is delayed by 10 ms with respect to the video signal. In the natural world, it is a phenomenon that can often occur due to the difference between the speed of light and the speed of sound, and it is interpreted that humans do not feel uncomfortable if the time difference is up to 100 ms. There is no problem.
In FIG. 18C, the source device 100 can know in advance the latency time La of the audio signal of the sink device 140 via the EDID line. Based on this, the source device 100 sets an additional delay of 30 ms in the voice variable delay unit 102. The total latency time TLv (80 ms) of the video and the cumulative delay time sDa (30 ms) of the audio signal are multiplexed on the audio signal 300 and transmitted to the sink device 140 and the sink device 150. In this case, the sink device 140 does not require an additional delay. In the sink device 150, a delay time difference of 30 ms is calculated from the total latency time TLv (80 ms) of the received video, the cumulative delay time sDa (30 ms) of the audio signal, and the latency time La (20 ms) of the audio signal of the sink device 150. Then, the voice variable delay unit 151 is controlled to execute an additional delay of 30 ms. In this way, lip-sync correction of the LCD 143, the speaker 153, and the speaker 253 becomes possible.
(Embodiment 6) In the present embodiment, an application example of the present invention in a system configuration in which a source device and a sink device are connected and a voice transmission path is connected from the sink device to another sink device will be described.
19 (a) and 19 (b) show the system configuration in this embodiment. The source device (DVD player) 100 is connected to the sink device (digital TV) 140, and the audio signal is connected from the sink device (digital TV) 140 to the sink device (multi-channel amplifier) 150 via HDMI.
In FIG. 19 (a), the source device 100 is a conventional device. Therefore, the source device 100 cannot recognize the total latency time TLv of the video signal or add it to the audio signal, and cannot add the cumulative delay time of the audio signal.
On the other hand, the sink device 140 recognizes that the audio signal output from the source device 100 does not have these additional information. Since there is no additional information, the sink device 140 itself performs lip-sync correction. Specifically, the sink device 140 acquires the latency time La (20 ms) of the audio signal of the sink device 150 in advance on the EDID line, adds this to its own latency time La (10 ms), and adds the latency time La (10 ms). 30ms is calculated as. The sink device 140 causes an additional delay of 50 ms in order to correct the time difference in the audio variable delay unit 251 from the latent time (Lv) of 80 ms of the image and the calculated latent time La (30 ms). As a result, the sync device 140 can set the total delay time of both the video signal and the audio signal to 80 ms, and can realize lip-sync correction.
FIG. 19B shows an example in the case where the source device 100 is a device to which the idea of the present invention is applied. Since the source device 100 can recognize the latency times of the sink devices 140 and 150, delay correction can be performed more flexibly. By performing an additional delay of 50 ms in the source device 100, the total delay time of both the video signal and the audio signal can be finally set to 80 ms in the sink device 150, and lip-sync correction can be realized.
(Embodiment 7) In the present embodiment, an application example of the present invention will be described in a system configuration in which a repeater (video processor) 120 inserted between the source device 100 and the sink device 140 is further added to the system configuration of FIG. Both devices are connected via an HDMI interface. In this example, the repeater 120 is a video processor that performs a predetermined process for improving the image quality of a video signal.
Figures 20 (a), (b) and (c) show the system configuration in this embodiment. The repeater 120 has a latent time Lv of a video signal of 50 ms. The source device 100 can recognize the total latency time TLv (130 ms) of the video signal via the repeater 120, can add this information to the audio signal, and further add the cumulative delay time of the audio signal to the audio signal. You can also do it.
FIG. 20A shows a case where the source device 100 does not perform the additional delay of the audio signal. The source device 100 transmits information of the total latency time TLv (130 ms) to the audio signal 300, and also transmits the cumulative delay time sDa (0 ms) of the latency time of the audio signal by multiplexing the audio signal. The repeater 120 adds an additional delay of 40 ms to the voice signal based on the received information.
The sink device 140 has a delay time difference of 80 ms from the total latency time TLv (130 ms) of the received video, the cumulative delay time sDa (40 ms) of the audio signal, and the audio signal latency time La (10 ms) of the sink device 140 itself. It calculates and controls the voice variable delay unit 251 to execute an additional delay of 80 ms.
Similarly, the sink device 150 has a delay time difference from the total latency time TLv (130 ms) of the received video, the cumulative delay time sDa (40 ms) of the audio signal, and the audio signal latency time La (20 ms) of the sink device 150 itself. It calculates 110ms and controls the voice variable delay unit 151 to execute an additional delay of 110ms. In this way, lip-sync correction of the LCD 143, the speaker 153, and the speaker 253 becomes possible.
20 (b) and 20 (c) show the case where the source device 100 performs an additional delay of the audio signal. In FIG. 20B, the source device 100 sets an additional delay of 110 ms in the voice variable delay unit 102. In this case, the sink device 150 does not require an additional delay.
In the sink device 140, a delay time difference of 10 ms is obtained from the total latency time TLv (130 ms) of the received video, the cumulative delay time sDa (110 ms) of the audio signal, and the audio signal latency time La (10 ms) of the sink device 140 itself. It calculates and controls the voice variable delay unit 251 to execute an additional delay of 10 ms. In this way, lip-sync correction of the LCD 143, the speaker 153, and the speaker 253 becomes possible.
In FIG. 20 (c), the source device 100 sets an additional delay of 80 ms in the voice variable delay unit 102. The source device 100 multiplexes and transmits the total latency time TLv (130 ms) and the cumulative delay time sDa (80 ms) of the audio signal to the audio signal 300. In this case, the sink device 140 does not require an additional delay.
On the other hand, in the sink device 150, the delay time difference is 30 ms from the total latency time TLv (130 ms) of the received video signal, the cumulative delay time sDa (80 ms) of the audio signal, and the audio signal latency time La (20 ms) of the sink device 150 itself. Is calculated, and the voice variable delay unit 151 is controlled to execute an additional delay of 30 ms. This enables lip-sync correction between the LCD 143 and the speaker 153 and the speaker 253.
(Embodiment 8) In the present embodiment, the repeater (video processor) 110 is inserted between the source device (DVD player) 100 and the sink device (digital TV) 140, and the sink device (multi-channel amplifier) 150 is further inserted via the repeater 110. An application example of the present invention in a connected system configuration will be described. The repeater 110 and the sink device 150 are connected by an S / PDIF interface. Others are connected via HDMI.
21 (a), (b), and (c) show the system configuration in this embodiment. In FIG. 21 (a), the source device 100 is a conventional device and cannot recognize the total latency time (TLv) of the video signal.
The repeater 110 can recognize the total latency time TLv (80ms) of the video signal via the EDID line. The repeater 110 outputs the total latency time TLv (80 ms) of the video signal and the cumulative delay time sDa (20 ms) of the audio signal to the downstream device.
The sink device 140 calculates a delay time difference of 50 ms from the total latency time TLv (80 ms) of the video signal received from the repeater 110, the cumulative delay time sDa (20 ms) of the audio signal, and its own latency time La (10 ms). .. The voice variable delay unit 251 is controlled based on the calculated time difference of 50 ms.
The sink device 150 receives the total latency time TLv (80ms) of the video signal and the cumulative delay time sDa (20ms) of the audio signal via S / PDIF, and from these values and its own latency time La (40ms). A delay time difference of 40 ms is calculated, and the voice variable delay unit 151 is controlled based on the calculated value. In this way all lip sync corrections are possible.
21 (b) and 21 (c) show an example in which the additional delay of the audio signal is performed in the source device 100. In FIG. 21 (b), since the repeater 110 and the sink device 150 are connected by the S / PDIF interface, the source device 100 cannot know the latency time (La) of the audio signal of the sink device 150 in advance. The source device 100 sets an additional delay of 50 ms in the voice variable delay unit 102 in advance.
The source 100 multiplexes the total latency time TLv (80 ms) of the video signal and the cumulative delay time sDa (50 ms) of the audio signal into the audio signal and transmits the audio signal to the repeater 110. In this case, the sink device 140 does not require an additional delay.
The sink device 150 calculates a delay time difference of -10 ms from the total latency time TLv (80 ms) of the received video, the cumulative delay time sDa (50 ms) of the audio signal, and the audio signal latency time La (20 ms) of the sink device 150. .. In order to correct the delay time difference of -10ms, it is necessary to delay the video signal instead of the audio signal, but this is not possible in this embodiment and is ignored. As already explained, there is no problem in ignoring such a negative delay time difference.
In FIG. 21 (c), the source device 100 can know in advance the latency time La (50 ms) of the audio signal of the sink device 140 via the EDID line. Based on this value, the source device 100 sets an additional delay of 10 ms in the voice variable delay unit 102. This eliminates the need for additional delay in the sink device 140.
On the other hand, in the sink device 150, a delay time difference of 30 ms is calculated from the total latency time TLv (80 ms) of the received video, the cumulative delay time sDa (30 ms) of the audio signal, and the audio signal latency time La (20 ms) of the sink device 150. Then, the voice variable delay unit 151 is controlled to execute an additional delay of 30 ms. In this way, lip-sync correction of all LCDs and speakers is possible.
(Embodiment 9) A method of superimposing additional information on the voice signal will be described with reference to FIG. 22. FIG. 22 is a diagram illustrating an example of transmitting additional information using a user bit (hereinafter referred to as Ubit) specified in IEC 60958.
Ubit is included in the sample data of the audio signal, and 1 bit is added for each L and R PCM data. This Ubit is a bit string and is defined by a block of 8 bits starting from the start bit. Ubit has decided to use the General user data format. Therefore, additional information such as latency time information is additionally specified in accordance with the General user data format. The general user data format is described in detail in IEC 60958-3 Section 6.2.4.1.
FIG. 23 is a diagram showing the format of additional information such as latency time information, which is additionally specified in accordance with the General user data format. The first IU1 indicates that it is latent time information, and IU2 specifies the information word length. IU3 is a copy of the category code. Eight IUs from IU4 to IU11 are used to construct 6-byte information including the following information. Audio Latency Valid: The validity bit of the information of the cumulative delay time sDa of the audio signal. Audio Unit type: Indicates the unit of information of the cumulative delay time sDa of the audio signal. Audio Latency: Information on the cumulative delay time sDa of the audio signal, which has a 16-bit word length. Video Latency Valid: The validity bit of the total latency (TLv) information of the video signal. Video Unit type: Indicates the unit of information of the total latency time (TLv) of the video signal. Video Latency: Information on the total latency (TLv) of the video signal, which has a 16-bit word length.
The reason for transmitting using Ubit in this way is to maintain wide generality regardless of the category. It is also possible to transmit SMPTE standard time code.
In this way, by using the digital audio interface conforming to the IEC standard, it is possible to transmit the latency information required for the present invention. That is, a digital audio interface conforming to the IEC standard can be applied to the S / PDIF interface in the above embodiment.
For HDMI, the audio transmission format is compliant with IEC 60958. In other words, the IEC 60958 format is wrapped in the HDMI format for transmission. That is, by using Ubit in exactly the same way as IEC 60958, it is possible to transmit all the latency information used in the present invention in the form of IEC 60958 even with HDMI. Furthermore, the latency information of the present invention can be transmitted by the 1394 interface in exactly the same manner by using the format called IEC conformant specified by the AMP protocol of IEC61883 (1394).
Although the present invention has been described for a particular embodiment, many other modifications, modifications, and other uses will be apparent to those skilled in the art. Therefore, the present invention is not limited to the specific disclosures herein, but may be limited only by the appended claims. This application is related to the Japanese patent application, Japanese Patent Application No. 2005-131963 (submitted on April 28, 2005), and their contents are incorporated in the text by reference.
The present invention is useful for a wide range of devices, AV devices and interfaces such as AV devices that handle video signals and audio signals and networks that are configured by using interfaces that connect them.
<figref num="1">Diagram for explaining the basic concept of the present invention</figref><figref num="2">A system block diagram showing a system configuration according to the first embodiment of the present invention.</figref><figref num="3">Hardware configuration diagram of the source device (DVD player)</figref><figref num="4">Hardware configuration diagram of sink device (video display device)</figref><figref num="5">Hardware configuration diagram of sink device (audio output device)</figref><figref num="6">Repeater hardware configuration diagram</figref><figref num="7">HDMI configuration diagram</figref><figref num="8">A flowchart showing the overall operation of the system according to the first embodiment.</figref><figref num="9">A flowchart showing the acquisition process of the total latency time TLv of the video signal by the source device in the first embodiment.</figref><figref num="10">Flowchart showing repeater processing when a TLv transmission command is received</figref><figref num="11">A flowchart showing the processing of the sink device when the TLv transmission command is received.</figref><figref num="12">A flowchart showing the transmission processing of the total latency time TLv of the video signal and the cumulative delay time sDa of the audio signal from the source device.</figref><figref num="13">Flowchart showing transmission processing of total latency time TLv of video signal and cumulative delay time sDa of audio signal</figref><figref num="14">Flow chart showing the adjustment operation process of the audio output time in the sink device</figref><figref num="15">A system block diagram showing a system configuration according to a second embodiment of the present invention.</figref><figref num="16">A system block diagram showing a system configuration according to a third embodiment of the present invention.</figref><figref num="17">A system block diagram showing a system configuration according to a fourth embodiment of the present invention.</figref><figref num="18">A system block diagram showing a system configuration according to a fifth embodiment of the present invention.</figref><figref num="19">A system block diagram showing a system configuration according to a sixth embodiment of the present invention.</figref><figref num="20">A system block diagram showing a system configuration according to a seventh embodiment of the present invention.</figref><figref num="21">A system block diagram showing a system configuration according to an eighth embodiment of the present invention.</figref><figref num="22">Schematic diagram for explaining a transmission method of additional information added to a voice signal</figref><figref num="23">A diagram showing the format of additional information such as latency time information, which is additionally specified in accordance with the General user data format.</figref><figref num="24">System block diagram showing the connection configuration of the conventional lip-sync correction device</figref>
Code description
100 sources 101 DVD media 102 Audio variable delay section 110 repeater 112 re-encoder 120 repeater 122 Video signal processing unit 130 repeater 132 Signal processing unit 140 sink 142 Video signal processing unit 143 LCD (Liquid Crystal Display) 150 sink 151 Audio variable delay section 152 Decoder 153 speaker
24 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2002290932A | Cites | Japan | Search report |
| JP2003101958A | Cites | Japan | Search report |
| JP2004282667A | Cites | Japan | Search report |
| JPH0282887A | Cites | Japan | Examiner |
| JP02082887A | Cites | Japan | – |
| JP2004282667A | Cites | Japan | – |
| JP2003101958A | Cites | Japan | – |
| JP2002290932A | Cites | Japan | – |
13 members in 4 offices
Members13
| Document | Office | Kind | |
|---|---|---|---|
| WO2006118106A1 | World Intellectual Property Organization (WIPO) | A1 | |
| CN101171838A | China | A | |
| JPWO2006118106A1 | Japan | A1 | |
| US2009073316A1 | United States of America | A1 | |
| CN101171838B | China | B | |
| JP2011223639A | Japan | A | |
| JP4912296B2This record | Japan | B2 | |
| JP5186028B2 | Japan | B2 | |
| US8451375B2 | United States of America | B2 | |
| US2013250174A1 | United States of America | A1 | |
| US8687118B2 | United States of America | B2 | |
| US2014160351A1 | United States of America | A1 | |
| US8891013B2 | United States of America | B2 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4912296
- Application
- 2007514741
Titles2
- Japanese
- リップシンク補正システム、リップシンク補正装置及びリップシンク補正方法
- English
- Lip sync correction system, lip sync correction device and lip sync correction method
Classification
- CPC, 6
- H04N5/04
- H04N21/2368
- H04N21/4341
- H04N21/43632
- H04N21/44227
- H04N21/43072
- IPC, 3
- H04N7 173
- H04N5 60
- H04N7 52
