Audio-video synchronizing
31 claims: 3 independent, 28 dependent
- 1(57)【特許請求の範囲】 【請求項1】 オーディオのプレゼンテーションタイムスタンプの値を、特定のオーディオデータパケット及びビデオデータパケットはオーディオデータとビデオデータを関連付けて再生するための望ましい再生時刻を表すプレゼンテーションタイムスタンプの値と、該特定のパケット間の再生されるオーディオ出力フレームの数を表すオーディオフレーム数とを含むオーディオデータパケットとして入力されるオーディオ入力データをデマルチプレクス及びデコードして得られる連続したオーディオフレームの一部であり、ビデオデータパケットとして入力されるビデオ入力データをデマルチプレクス及びデコードして得られる連続したビデオフレームの一部であるビデオ出力フレームとともに再生されるオーディオ出力フレームに関連付けるオーディオとビデオの同期方法において、 オーディオ及びビデオのデマルチプレクス処理の間にオーディオ及びビデオのプレゼンテーションタイムスタンプの値を特定のオーディオ及びビデオデータパケットに記憶するステップと、 オーディオのデマルチプレクス処理の間に上記記憶されたオーディオプレゼンテーションタイムスタンプの値のうちの1つに対応している各オーディオフレームカウンタにオーディオフレームの数を記憶するステップと、 オーディオ及びビデオデータを特定のオーディオデータパケット及びビデオデータパケットとして連続的にデコードし、オーディオ及びビデオのフレームをそれぞれ生成するステップと、 オーディオのフレームとビデオのフレームを利用者に対して同期して再生させるステップと、 オーディオのフレームの再生に応答してオーディオフレームカウンタの値を選択的に減少させるステップと、 値が0であるオーディオフレームカウンタの1つを検出するステップと、 上記オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値を読み出すステップと、 オーディオ及びビデオのフレームの再生を選択的に修正して、オーディオとビデオを利用者に対して同期して出力させるステップとを有するオーディオとビデオの同期方法。
- 2【請求項2】オーディオとビデオの同期方法において、オーディオのプレゼンテーションタイムスタンプの値を読み出すステップの後に、 システムタイムカウンタの値が、オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値とシステムタイムカウンタの現在の値との差分と略等しくなるようにオーディオクロックの拡張を行うステップと、 システムタイムカウンタがオーディオのフレームと同期して再生するように、上記オーディオクロックの拡張によってシステムタイムカウンタの現在の値を調整するステップとを更に有する請求項1に記載のオーディオとビデオの同期方法。
- 3【請求項3】ビデオデコード処理の開始時に、ビデオデコード処理の持続時間と略等しい持続時間を持った、各ビデオフレームカウンタのビデオフレームの枚数を、特定のビデオデータパケットとして記憶するステップと、 ビデオのフレームの再生に応答して全てのビデオフレームカウンタの値を減少させるステップと、 値が0であるビデオフレームカウンタの1つを検出するステップと、 上記ビデオフレームカウンタの1つに対応したビデオのプレゼンテーションタイムスタンプの値を読み出すステップと、 ビデオフレームカウンタの1つに対応したビデオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値にオーディオクロックの拡張分を加算した値と略等しいと判定し、オーディオフレームとビデオフレームが同期すると判定するステップとを更に有する請求項2に記載のオーディオとビデオの同期方法。
- 4【請求項4】値が0であるビデオフレームカウンタの1つに応答して、ビデオのフレームが、現在利用者に出力されているオーディオのフレームと同期せずに速く出力されていると判定するステップと、 ビデオのフレームがオーディオのフレームと同期して出力されるようにビデオのフレームを反復するステップとを更に有する請求項3に記載のオーディオとビデオの同期方法。
- 5【請求項5】ビデオフレームカウンタの1つに対応したビデオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値にオーディオクロックの拡張分を加算した値より大きいと判定し、ビデオのフレームがオーディオのフレームよりも速く出力されていると判定するステップを更に有する請求項4に記載のオーディオとビデオの同期方法。
- 6【請求項6】値が0であるビデオフレームカウンタの1つに応答して、ビデオのフレームが、現在利用者に出力されているオーディオのフレームと同期せずに遅く出力されていると判定するステップと、 ビデオのフレームがオーディオのフレームと同期して出力されるようにビデオのフレームをスキップするステップとを更に有する請求項5に記載のオーディオとビデオの同期方法。
- 7【請求項7】ビデオフレームカウンタの1つに対応したビデオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値にオーディオクロックの拡張分を加算した値より小さいと判定し、ビデオのフレームがオーディオのフレームよりも遅く出力されていると判定するステップを更に有する請求項6に記載のオーディオとビデオの同期方法。
- 8【請求項8】ビデオフレームカウンタの1つに対応したビデオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値にオーディオクロックの拡張分を足してオーディオデータのデコード処理における遅延時間とオーディオ及びビデオの同期において許容される程度のエラーとを加算した値に略等しい制限値を減算したものより小さいと判定するステップを更に有する請求項7に記載のオーディオとビデオの同期方法。
- 9【請求項9】ビデオの再生フレームとオーディオの再生フレームの同期を決定するために、オーディオマスター同期が選択されているかを検出するステップを有する請求項8に記載のオーディオとビデオの同期方法。
- 10【請求項10】ビデオデコード処理の開始時に、ビデオデコード処理の持続時間と略等しい持続時間を持った、各ビデオフレームカウンタのビデオフレームの枚数を、特定のビデオデータパケットとして記憶するステップと、 ビデオのフレームの再生に応答して全てのビデオフレームカウンタの値を減少させるステップと、 値が0であるビデオフレームカウンタの1つを検出するステップと、 上記ビデオフレームカウンタの1つに対応したビデオのプレゼンテーションタイムスタンプの値を読み出すステップと、 システムタイムカウンタの値が、ビデオフレームカウンタの1つに対応したビデオのプレゼンテーションタイムスタンプの値とシステムタイムカウンタの現在の値との差分と略等しくなるようにビデオクロックの拡張を行うステップと、 システムタイムカウンタがビデオのフレームと同期して再生するように、上記ビデオクロックの拡張によってシステムタイムカウンタの現在の値を調整するステップとを更に有する請求項1に記載のオーディオとビデオの同期方法。
- 11【請求項11】値が0であるオーディオフレームカウンタの1つに応答して、オーディオのフレームが、現在利用者に出力されているビデオのフレームと同期せずに速く出力されていると判定するステップと、 オーディオのフレームがビデオのフレームと同期して出力されるようにオーディオのフレームを反復するステップとを更に有する請求項10に記載のオーディオとビデオの同期方法。
- 12【請求項12】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値にビデオクロックの拡張分を加算した値より大きいと判定し、オーディオのフレームがビデオのフレームよりも速く出力されていると判定するステップを更に有する請求項11に記載のオーディオとビデオの同期方法。
- 13【請求項13】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値にビデオクロックの拡張分と最短のビデオフレームの持続時間の略半分と略等しい制限値とを加算した値より大きいと判定するステップを更に有する請求項12に記載のオーディオとビデオの同期方法。
- 14【請求項14】値が0であるオーディオフレームカウンタの1つに応答して、オーディオのフレームが、現在利用者に出力されているビデオのフレームと同期せずに遅く出力されていると判定するステップと、 オーディオのフレームがビデオのフレームと同期して出力されるようにオーディオのフレームをスキップするステップとを更に有する請求項13に記載のオーディオとビデオの同期方法。
- 15【請求項15】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値にビデオクロックの拡張分を加算した値より小さいと判定し、オーディオのフレームがビデオのフレームよりも遅く出力されていると判定するステップを更に有する請求項14に記載のオーディオとビデオの同期方法。
- 16【請求項16】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値にビデオクロックの拡張分を加算して最短のビデオフレームの持続時間の略半分と略等しい制限値を減算した値より小さいと判定するステップを更に有する請求項15に記載のオーディオとビデオの同期方法。
- 17【請求項17】オーディオの再生フレームとビデオの再生フレームの同期を決定するために、ビデオマスター同期が選択されているかを検出するステップを有する請求項16に記載のオーディオとビデオの同期方法。
- 18【請求項18】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値より大きいと判定し、オーディオのフレームがビデオのフレームよりも速く出力されていると判定するステップを更に有する請求項11に記載のオーディオとビデオの同期方法。
- 19【請求項19】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値に最短のビデオフレームの持続時間の略半分と略等しい制限値を加算した値より大きいと判定するステップを更に有する請求項18に記載のオーディオとビデオの同期方法。
- 20【請求項20】値が0であるオーディオフレームカウンタの1つに応答して、オーディオのフレームが、現在利用者に出力されているビデオのフレームと同期せずに遅く出力されていると判定するステップと、 オーディオのフレームがビデオのフレームと同期して出力されるようにオーディオのフレームをスキップするステップとを更に有する請求項19に記載のオーディオとビデオの同期方法。
- 21【請求項21】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値より小さいと判定し、オーディオのフレームがビデオのフレームよりも遅く出力されていると判定するステップを更に有する請求項20に記載のオーディオとビデオの同期方法。
- 22【請求項22】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値に最短のビデオフレームの持続時間の略半分と略等しい制限値を減算した値より小さいと判定するステップを更に有する請求項21に記載のオーディオとビデオの同期方法。
- 23【請求項23】値が0であるビデオフレームカウンタの1つに応答して、ビデオのフレームが、現在利用者に出力されているオーディオのフレームと同期せずに速く出力されていると判定するステップと、 ビデオのフレームがオーディオのフレームと同期して出力されるようにビデオのフレームを反復するステップとを更に有する請求項22に記載のオーディオとビデオの同期方法。
- 24【請求項24】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値より大きいと判定し、オーディオのフレームがビデオのフレームよりも速く出力されていると判定するステップを更に有する請求項23に記載のオーディオとビデオの同期方法。
- 25【請求項25】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値に最短のビデオフレームの持続時間の略半分と略等しい制限値を加算した値より大きいと判定するステップを更に有する請求項24に記載のオーディオとビデオの同期方法。
- 26【請求項26】値が0であるオーディオフレームカウンタの1つに応答して、オーディオのフレームが、現在利用者に出力されているビデオのフレームと同期せずに遅く出力されていると判定するステップと、 オーディオのフレームがビデオのフレームと同期して出力されるようにオーディオのフレームをスキップするステップとを更に有する請求項25に記載のオーディオとビデオの同期方法。
- 27【請求項27】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値より小さいと判定し、オーディオのフレームがビデオのフレームよりも遅く出力されていると判定するステップを更に有する請求項26に記載のオーディオとビデオの同期方法。
- 28【請求項28】オーディオフレームカウンタの1つに対応したオーディオのプレゼンテーションタイムスタンプの値が、システムタイムカウンタの現在の値に最短のビデオフレームの持続時間の略半分と略等しい制限値を減算した値より小さいと判定するステップを更に有する請求項27に記載のオーディオとビデオの同期方法。
- 29【請求項29】クロックの拡張が決定及び使用されず、ビデオとオーディオが独立してシステムタイムカウンタと同期するマスター同期が選択されていないかを検出するステップを更に有する請求項27に記載のオーディオとビデオの同期方法。
- 30【請求項30】特定のオーディオデータパケット及びビデオデータパケットはプレゼンテーションタイムスタンプの値と、該特定のパケット間の再生されるオーディオ出力フレームの数を表すオーディオフレーム数とを含むオーディオの出力データのフレームとビデオの出力データのフレームを同期して出力させるオーディオとビデオの同期方法において、 プレゼンテーションタイムスタンプの値を含むオーディオデータパケットをデマルチプレクスするオーディオデマルチプレクス処理の間に各オーディオFIFOのメモリ位置にオーディオデータを連続的に記憶するステップと、 第1のデマルチプレクス処理の間に、第1のメモリ位置のプレゼンテーションタイムスタンプの値をオーディオプレゼンテーションタイムスタンプテーブルに連続的に記憶するステップと、 第1のデマルチプレクス処理の間に、プレゼンテーションタイムスタンプの値を含む特定のビデオデータパケットとしてビデオFIFOメモリ位置に書き込まれたデータに対応した書込ポインタの位置を、オーディオプレゼンテーションタイムスタンプテーブルの第2のメモリ位置に記憶するステップと、 第1のデマルチプレクス処理の間に、オーディオフレームの数をプレゼンテーションタイムスタンプテーブルのカウンタ位置に記憶するステップと、 CPUへの開始オーディオフレーム割込と同時にオーディオデコード処理を開始するステップと、 開始オーディオフレーム割込に応答してオーディオFIFOの読出ポインタの位置を獲得するステップと、 読出ポインタの位置と略等しい、オーディオプレゼンテーションタイムスタンプテーブルの書込ポインタの位置を検出するステップと、 オーディオFIFOから読み出されているオーディオデータをデコードするステップと、 連続したオーディオ出力フレームを利用者に提供するステップと、 オーディオプレゼンテーションタイムスタンプテーブルの全てのカウンタメモリ位置から1ずつ減算するステップと、 システムタイムカウンタの現在の値を、値が0であるカウンタに対応したオーディオプレゼンテーションタイムスタンプテーブルのプレゼンテーションタイムスタンプの値に合わせるステップとを有するオーディオとビデオの同期方法。
- 31【請求項31】特定のオーディオデータパケット及びビデオデータパケットはプレゼンテーションタイムスタンプの値を含むオーディオデータパケット及びビデオデータパケットを含むオーディオ及びビデオデータをデマルチプレクス及びデコードして得られるオーディオとビデオデータの出力フレームを同期して再生させるデジタルビデオプロセッサにおいて、 オーディオデータパケット及びビデオデータパケットのオーディオ及びビデオデータを連続的に記憶するオーディオ及びビデオFIFOメモリと、 プレゼンテーションタイムスタンプの値を特定のオーディオデータパケットに連続的に記憶する第1のメモリ位置と、プレゼンテーションタイムスタンプの値を含む特定のオーディオデータパケットのオーディオFIFOメモリ位置に書き込まれたデータに対応した書込ポインタ位置を表す値を記憶する第2のメモリ位置と、次のプレゼンテーションタイムスタンプの値が出るまで再生されるオーディオフレームの数を記憶するカウンタ位置として機能する第3のメモリ位置とを有するオーディオプレゼンテーションタイムスタンプメモリテーブルと、 プレゼンテーションタイムスタンプの値を特定のビデオデータパケットに連続的に記憶する第1のメモリ位置と、プレゼンテーションタイムスタンプの値を含む特定のビデオデータパケットのビデオFIFOメモリ位置に書き込まれたデータに対応した書込ポインタ位置を表す値を記憶する第2のメモリ位置と、利用者に対して再生される前にビデオデータのフレームをデコードするのに必要な遅延時間と略等しいビデオのフレームの数を表す値を記憶するカウンタ位置として機能する第3のメモリ位置とを有するビデオプレゼンテーションタイムスタンプメモリテーブルとを備えるデジタルビデオプロセッサ。
Independent claims31
143 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
The present invention relates to a method for synchronizing audio and video and a digital video processor, for example, regarding digital processing of an image displayed on the screen of a television receiver, and particularly to a video in which an audio signal and a video signal are digitally synchronized. Regarding the process of outputting to the display device.
【0002】
[Conventional technology]
Most of the television receivers manufactured today are, for example, video tape recorders (hereinafter referred to as VTRs), digital video discs (hereinafter referred to as DVDs) players, cables, digital satellite systems (hereinafter referred to as DSS), etc. It is possible to connect to the signal source devices of various programs of, and from the signal source devices of these programs, an audio signal for generating sound and a video signal for displaying an image corresponding to the sound Is output. Among these signal source devices, digital audio signals and video signals conforming to the audio / video digital compression standard of Moving Picture Expert Group-2 (hereinafter referred to as MPEG-2: Moving Picture Expert Group-2) are output. There is also something to do. Therefore, it is desirable that current television receivers and / or DVD players have the ability to process compressed digital input signals and output digital output signals to display the desired image. In many cases, these digital signals are converted into analog signals and displayed on the display device of a conventional analog television receiver.
【0003】
In order to realize the display of an image based on the signal from the audio / video signal source device of the program and the digital signal processing for outputting the corresponding sound, it is not found in the conventional analog audio / video signal processing. Various designs have been tried. For example, in digital signal processing, the audio signal is separate from the video signal, that is, the audio and the image are processed independently. The audio and the image must be reproduced in synchronization so that the desired audio signal and the video signal supplied from the signal source device of the program correspond to each other and are reproduced as one.
【0004】
The program signal source device outputs audio data and video data as data packets, for example, in the MPEG-2 format. Each packet of audio data and video data is received as a continuous data stream from the program signal source device. Each packet of video data has a header block followed by a data block. The data block can contain multiple frames of video data, such as 1-20, and these frames are the complete fields of the video data, or unique header blocks for identifying the type and display order of the pictures. Can include a coded group of pictures with. The header block of the video data packet contains control information such as, for example, the format identifier of the video data, the type of compression, the size of the image if necessary, the display order, and all other parameters. Packets of audio data also have a header block, which is also used to identify the format of the audio data, including instructions for decoding the audio data and the desired enhancement processing method applied. Further, an audio data packet has a header block followed by a data block, and the audio data block includes a plurality of audio data blocks or frames such as 1 to about 20.
【0005】
One specific header block for audio and video data packets contains a presentation time stamp (PTS: Presentation Time Stamp) value, which is time management information for playback output, which is the data. A time stamp that can be applied to the packet. The value of this PTS is a time based on the system time clock (hereinafter referred to as STC: System Time Clock) operating during the generation or recording of audio data and video data. Also, when the audio data and video data are played back, the same STC is operating, and if the audio data and video data are played back at the time indicated by their respective PTSs, the audio data and video data are ideal. It is output to the user in synchronization. That is, the PTS is used to display or reproduce audio data and video data in synchronization.
【0006】
When decoding audio data, the audio data must be stretched, reconstructed and enhanced to match the capabilities of the program's signal source and audio player. In some applications, a packet of audio data may contain up to 6 channels of raw audio data. The audio player can play, for example, 2 to 6 channels of audio data, depending on the number of channels, selectively using the channels of raw audio data to provide multiple audio channels. , The audio data of these channels is stored in the first-in-first-out memory for audio (hereinafter referred to as FIFO).
【0007】
Decoding video data usually requires decompression, conversion of partial frames to all frames, and recognition of all frames. At the same time as the decoding process, frames of audio data and video data are output, that is, played back to the user, and this playback must be synchronized so that the frames of the audio data and the video data correspond to each other and become one. Must be.
【0008】
As is clear from the prior art, demultiplexing audio and video data packets decomposes the data packets so that the decoding and playback of the data are synchronized, and the instructions required for decoding as well as the content data itself. It is a complicated process of memorizing. According to one well-known technique, audio and video content data, i.e. raw data, is stored in each audio and video FIFO memory. These FIFOs include write and read pointers controlled by a memory controller, which is under the general control of the CPU. The write pointer operates as a function required for demultiplexing processing, and the write pointer continuously distributes data to each FIFO. The read pointer is independent of the decoding process and operates as a parallel function, and data is continuously read from the FIFO. In demultiplexing, in addition to loading the raw data into the FIFO, the corresponding PTS value, if any, is written to the storage in each audio and video PTS table. In order to make the PTS value correspond to the data in the FIFO, in addition to the PTS value, the position in each FIFO where the first byte of the data received following the PTS was written is also respectively. Written to the PTS table for audio and video.
【0009】
The audio and video data are written to the FIFO memory by demultiplexing, while the audio and video data are simultaneously and in parallel from their respective FIFOs during the decoding and playback of the audio and video data. Read out. While both of these processes are being performed, a monitoring process is required to monitor whether the audio data and the video data read by the decoding process of the audio data and the video data are synchronized in time. In the well-known technique described above, this monitoring process is performed by associating the read pointer of the FIFO operated by the decoding process with the storage area stored in the PTS table. When the read pointer is close enough to the storage corresponding to the value of the PTS, it can be determined that the PTS indicates the current time for the decoding process. By identifying the value of the PTS in this way, the PTS can be used for comparison to determine if one decoding process precedes or follows the other decoding process.
【0010】
However, this method has drawbacks due to the following reasons. During the audio and video data decoding process, the read pointer to each FIFO is operated automatically and continuously by the decoding process that interacts directly with the memory controller independent of any instructions from the CPU. There is. This is unavoidable because the entire process of demultiplexing, decoding, and outputting audio and video data must be performed synchronously and continuously.
【0011】
The above-mentioned techniques for synchronizing audio and video decoding and playback have serious problems such as an inherent delay between the monitoring process and the various decoding processes. For example, when decoding audio data, the audio decoder applies a start audio frame interrupt to the CPU executing the monitoring process each time it starts decoding one audio frame. When decoding of one audio frame is started, the monitoring process starts with the data currently being read from the audio FIFO and the corresponding PTS, that is, the audio when the currently read data is written to the audio FIFO. It must correspond to the PTS that has the value written in the PTS table. Theoretically, if the position of the read pointer at the start of the audio frame is known, compare the position of this read pointer with the position of the write pointer stored in the PTS table during the demultiplexing process. Can be done. When it is found that the position of the current read pointer and the position of the stored write pointer match, the PTS corresponding to the stored write pointer is the audio data identified by the current read pointer. It will match the PTS of. If the value of PTS corresponding to the data to be read can be accurately determined, the decoding and playback process skips or repeats the frames and outputs the audio and video frames in synchronization by a well-known method. Will be done.
【0012】
In the above-mentioned processing, there are two conditions for the audio data and the video data to be reproduced in a state where they cannot be synchronized. The first condition occurs because the CPU executing the monitoring process must monitor the audio data and video data decoding processing and the demultiplexing processing in a time-division manner. Therefore, the CPU must respond to each monitoring process using priority interrupts based on the communication method. In addition, the CPU communicates with a memory controller and other circuits via a time-division multiplex communication bus or channel. Therefore, even if the start audio frame is interrupted to the CPU, the start audio frame interrupt process is not immediately executed when the CPU is processing another interrupt of the same or higher priority. Further, even if the CPU immediately executes the start audio frame interrupt process, the CPU must communicate with the memory controller via the time division multiplexing bus. Access to the bus is arbitrated, and the CPU does not always have the highest priority. During the first delay due to the start audio frame interrupt process and the second delay due to communication with the memory controller via the time division multiplexing bus, the audio data from the FIFO is subjected to the decoding process for the audio data FIFO. Reading continues. Therefore, when the start audio frame interrupt is accepted by the CPU and the CPU is able to communicate with the memory controller, the position of the read pointer of the audio data FIFO that the CPU obtains is usually the start audio frame interrupt first. It is different from the position of the read pointer when it is accepted by the CPU. That is, when the CPU responds to the interrupt and gets the current read pointer of the audio FIFO from the memory controller, the read pointer is no longer the value at which the interrupt occurred. Therefore, if there is a delay, the CPU will get the incorrect read pointer position.
【0013】
The second condition is that the audio packet to be read is small and processed at high speed. The PTS table has two PTS entries with FIFO position values, which are very close together. Therefore, if the incorrect read pointer position is compared to the write pointer position value in the PTS table, the wrong PTS entry will be associated with the start of the audio frame being decoded. As a result, the audio and video frames are played out of sync. The user does not notice such a small deviation of synchronization once.
【0014】
[Problems to be Solved by the Invention]
By the way, when the processing time for interrupting the start audio frame is long and the time required for reading the audio data packet is very short, or when such a plurality of audio data packets are continuously generated, The out-of-sync will be greater. Further, the synchronization deviation in the audio processing is accumulated with the synchronization deviation in the other decoding processing. Therefore, if the synchronization deviations are accumulated, the user may notice or be bothered by the synchronization deviations.
【0015】
That is, in the above-mentioned device, it is necessary to make the PTS value stored in the PTS table correspond to the audio data and the video data read from the respective FIFO memories during the decoding process.
【0016】
An object of the present invention is to provide an audio-video synchronization method and a digital video processor that more accurately synchronize and reproduce audio frames and video frames from a program signal source device.
【0017】
[Means for solving problems]
The method for synchronizing audio and video according to the present invention represents a value of an audio presentation time stamp, and a specific audio data packet and a video data packet represent a desirable playback time for playing back the audio data and the video data in association with each other. Continuous audio obtained by demultiplexing and decoding audio input data input as an audio data packet, including the stamp value and the number of audio frames representing the number of audio output frames played between the particular packets. With audio associated with an audio output frame that is part of a frame and is played with a video output frame that is part of a continuous video frame obtained by demultiplexing and decoding video input data that is input as a video data packet. In a video synchronization method, the step of storing the value of the audio and video presentation time stamp in a specific audio and video data packet during the audio and video demultiplexing process and the storage during the audio demultiplexing process. A step to store the number of audio frames in each audio frame counter corresponding to one of the values of the audio presentation time stamp given, and the audio and video data being continuous as specific audio data packets and video data packets. The step of decoding to generate audio and video frames, respectively, the step of playing back the audio frame and the video frame synchronously to the user, and the step of playing the audio frame in response to the playback of the audio frame counter. A step that selectively reduces the value, a step that detects one of the audio frame counters that has a value of 0, a step that reads the value of the audio presentation time stamp that corresponds to one of the audio frame counters, and audio and Selectively modify the playback of video frames to provide audio and video to the userIt has a step of synchronizing and outputting.
【0018】
Further, in the method of synchronizing audio and video according to the present invention, a specific audio data packet and a video data packet have a value of a presentation time stamp and a number of audio frames representing the number of audio output frames reproduced between the specific packets. Audio and video output data frames including and video output data frames are synchronized and output. The audio and video synchronization method is audio demultiplexing processing that demultiplexes audio data packets containing presentation time stamp values. During the step of continuously storing audio data in the memory location of each audio FIFO during and during the first demultiplexing process, the value of the presentation timestamp of the first memory location is stored in the audio presentation timestamp table. Between the step of continuously storing and the first demultiplexing process, the position of the write pointer corresponding to the data written to the video FIFO memory location as a specific video data packet containing the value of the presentation timestamp. , A step to store the number of audio frames in the counter position of the presentation time stamp table during the first demultiplexing process and a step to store it in the second memory location of the audio presentation time stamp table, and to the CPU. An audio presentation time stamp that is approximately equal to the position of the read pointer, the step of starting the audio decoding process at the same time as the start audio frame interrupt, and the step of acquiring the position of the read pointer of the audio FIFO in response to the start audio frame interrupt. All of the audio presentation timestamp table, including the step of detecting the position of the write pointer in the table, the step of decoding the audio data read from the audio FIFO, the step of providing the user with continuous audio output frames, and the step The step of subtracting 1 from the counter memory position of and the current value of the system time counterIt has a step to match the value of the presentation timestamp in the audio presentation timestamp table corresponding to the counter whose value is 0.
【0019】
Further, the digital video processor according to the present invention obtains a specific audio data packet and video data packet by demultiplexing and decoding audio data packet including an audio data packet including a presentation time stamp value and audio and video data including a video data packet. In a digital video processor that synchronizes and reproduces output frames of audio and video data, the audio and video FIFO memory that continuously stores the audio and video data of audio data packets and video data packets, and the value of the presentation time stamp are used. A value that represents the first memory location that is continuously stored in a particular audio data packet and the write pointer position that corresponds to the data written to the audio FIFO memory location of the particular audio data packet, including the value of the presentation time stamp. With an audio presentation time stamp memory table that has a second memory position to store data and a third memory position that acts as a counter position to store the number of audio frames played until the next presentation time stamp value is reached. Corresponds to the first memory location that continuously stores the presentation time stamp value in a specific video data packet and the data written to the video FIFO memory location of the specific video data packet that contains the presentation time stamp value. Represents a second memory location that stores a value that represents the write pointer position and the number of video frames that is approximately equal to the delay time required to decode the frame of the video data before it is played back to the user. It has a third memory position that functions as a counter position to store the value.
【0020】
BEST MODE FOR CARRYING OUT THE INVENTION
Hereinafter, the audio-video synchronization method and the digital video processor according to the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing a specific configuration of a television receiver to which the present invention is applied.
【0021】
As shown in FIG. 1, the television receiver 18 includes a digital audio / video processor 20, which produces a desired image via the input terminal 22. A series of data packets containing audio and video information necessary for the processor is continuously supplied. Such information, that is, audio data and video data, is, for example, in a video tape recorder (hereinafter referred to as VTR), a DVD player, a cable, a receiver of a digital satellite system (hereinafter referred to as DSS), and a television receiver 18. It is sourced from one of the devices selected by the displayed menu command, and the video information is sourced in one of several different formats. Generally, audio data and video data packets are supplied in a format conforming to the standards of Moving Picture Expert Group-2 (hereinafter referred to as MPEG-2: Moving Picture Expert Group-2).
【0022】
The supplied audio and video data packets are demultiplexed into parallel data streams that are independent of each other. Also, the decoding and playback processing of the output frames of the audio data and the video data is continuously performed in a parallel data stream independent of the demultiplexing processing. Further, the demultiplexing process is a process that changes remarkably in real time depending on the properties of the supplied audio data and video data. Furthermore, the number and display order of the displayed video frames cannot be determined from the raw video data supplied. The creation of the video frame and the determination of its display order are performed during the decoding process, which is first determined by the control data in the header portion of the video data packet. Similarly, the raw audio data delivered as a data packet is very similar to the output and displayed audio data, and the displayed audio data frames are generated during the audio data decoding process. ..
【0023】
By the way, the length of the output audio frame can be any length in real time, multiple audio frames can correspond to one video frame, and conversely, one audio frame can be multiple. It can also be output while the image generated from the video frame in is displayed. It is necessary that the audio data frame and the video data frame are reproduced in synchronization, correspond to each other, and provided to the user as a unit. One specific header block for audio data and video data packets is a presentation time stamp (PTS: Presentation Time) for synchronizing audio data frames and video data frames for output. It's called Stamp. ) Is included, and this PTS is the time based on the system time counter operating during the generation or recording of audio data and video data. A similar system counter is operating in the CPU 26 during decoding and playback of audio and video data frames, and PTS tables for audio and video are generated during demultiplexing. When audio and video frames are output in perfect synchronization, their stored PTS value is equal to the current value of the system time counter. By the way, since audio data and video data are processed differently as independent parallel bit streams, such accurate time control is not easy. Therefore, for this reason, frames of audio data and video data may be reproduced out of synchronization with respect to the system time counter. As a result, the audio data and the video data are reproduced out of synchronization with each other for the user.
【0024】
The detailed operation of each circuit constituting the digital audio / video processor 20 will be described with reference to a flowchart. FIG. 2 is a flowchart showing the normal operation of the demultiplexer 19 shown in FIG.
【0025】
In step S202, the demultiplexer 19 continuously receives data including audio data packets, video data packets, and subpicture data packets in a random order from a DVD player, for example, via the input terminal 22, as an input bit stream. ..
【0026】
In step S204, the packet header block is extracted. In step S206, it is determined whether the packet is a video data packet based on the header block, and if applicable, the process proceeds to step S208. In step S208, a video demultiplex interrupt is applied to the CPU 26 shown in FIG. In step S210, the video data is continuously stored in the first-in-first-out (hereinafter referred to as FIFO) buffer 320 for the video data shown in FIG. 4 provided in the memory 24 shown in FIG.
【0027】
Similarly, in step S212, if the packet is determined to be an audio data packet based on the header block, the process proceeds to step S214. In step S214, an audio demultiplex interrupt is applied to the CPU 26. In step S216, the audio data is continuously stored in the FIFO buffer 300 for the audio data shown in FIG. 4 provided in the memory 24.
【0028】
Similarly, in step S218, if the packet is determined to be a subpicture data packet based on the header block, the process proceeds to step S262. In step S220, the sub-picture demultiplex interrupt is applied to the CPU 26. In step S222, the sub-picture data is continuously stored in the sub-picture data FIFO buffer (not shown) provided in the memory 24. The demultiplexer 19 performs other processing, but since these processings are not related to audio and video synchronization, further description of the operation of the demultiplexer 19 will be omitted.
【0029】
FIG. 3 is a flowchart showing the demultiplexing process subsequently executed in the CPU 26, and FIG. 4 shows how various parts of the audio data and the video data are distributed in the memory 24. It is a schematic diagram which shows. In the memory 24, in addition to the FIFO buffer 300 for audio data and the FIFO buffer 320 for video data shown in FIG. 4, a PTS table 302 for audio and a PTS table 324 for video are provided.
【0030】
As shown in FIG. 3, in step S250, the CPU 26 responds to an interrupt from the demultiplexer 19. In step S252, the CPU 26 determines whether the interrupt is an audio PTS interrupt, and if applicable, proceeds to step S254. In step S254, the PTS value of the header block of the audio data packet is loaded or stored in, for example, position 304 of the PTS table 302 for audio. Further, the position of the write pointer 306 of the FIFO buffer 300 corresponding to the position of the first byte of the audio data loaded in the FIFO buffer 300 for audio data is set to the position 308 of the PTS table 302 for audio, for example. Be remembered.
【0031】
As mentioned above, the PTS is contained only in certain audio data packets, but the header blocks of these audio data packets also include the current PTS of the current audio data packet and the next audio data packet. A field with an audio frame number indicating the number of audio data frames to be output is provided between the PTS and the PTS. Further, the PTS table 302 for audio is provided with a frame counter, for example, in the storage area 309. When demultiplexing an audio data packet containing a PTS value, this PTS value is loaded into an appropriate location, for example position 304. Further, the number of audio frames up to the next PTS value is added to the value of the corresponding frame counter, that is, the frame counter 310, and the total value is added to the frame counter corresponding to the next PTS, that is, the frame counter 316. Written. For example, when a PTS with 7 frames until the next audio data packet is loaded at position 304, that number 7 is added to the current value 10 of the frame counter 310, and this total value 17 is It is written to the frame counter 316 corresponding to the next audio data packet having a PTS value.
【0032】
While the audio demultiplexing process is being executed, the decoding and playback processing of the output audio frame is also being executed in parallel with the demultiplexing processing. In the audio reproduction process, when each audio frame is output, the values of all the frame counters in the PTS table 302 for audio are decremented. Therefore, when the next PTS count value is loaded into the frame counter 316, the value loaded into the frame counter 310 is less than the original value. As a result, the value of the frame counter in storage 309 of the PTS table 302 for audio changes each time an audio data packet is demultiplexed and written to the FIFO buffer 300, and is being executed in parallel at the same time. The audio data is read from the FIFO buffer 300, decoded, and the output audio frame is output to the user.
【0033】
In step S256 shown in FIG. 3, the CPU 26 determines whether the interrupt is a video PTS interrupt, and if applicable, proceeds to step S258. In step S258, the value of the PTS (hereinafter, simply referred to as video PTS) of the header block of the video data packet is stored in the PTS table 324 for video shown in FIG. 4, for example, at position 322. Also, the position of the write pointer 326 of the FIFO buffer 320 for video data when the first byte of the video data of the packet is stored is written to, for example, position 328 of the PTS table 324 for video. The PTS table 324 for video is also provided with a frame counter, and during demultiplexing, the CPU 26 uses, for example, a frame counter 330 corresponding to the PTS at position 322, such as FF in hexadecimal notation. Set to a meaningless value.
【0034】
In step S260 shown in FIG. 3, the CPU 26 determines whether the interrupt is an interrupt of the sub-picture PTS, and if applicable, proceeds to step S262. In step S262, the CPU 26 stores the PTS of the subpicture in an internal register (not shown). The PTS value is included in all sub-picture data packets, and when the system time clock (STC) becomes the same value as the stored PTS, the corresponding sub-picture is output. Since sub-picture synchronization is not as important as audio and video synchronization, further details regarding the processing of sub-picture data will be omitted.
【0035】
After the interrupt in the demultiplex is applied, the demultiplex process described with reference to FIGS. 1 to 4 is repeated in the same manner. The raw data of the next audio data packet is continuously loaded into the FIFO buffer 300 for audio data. If the next audio data packet does not have a PTS in its header block, no PTS entry will be created for the PTS table 302 for audio. By the way, when the next audio data packet has a PTS in its header block, the PTS is written to, for example, position 312 of the PTS table 302 for audio. The position of the write pointer, where the write pointer 306 loads the first audio data of the current audio data packet into the FIFO buffer 300, is loaded at position 314. Also, the number of frames of audio data between the current PTS and the next PTS is added to the value of the non-rem counter 316 and loaded into the frame counter 317. Further, when the raw video data of each contiguous video data packet is continuously loaded into the FIFO buffer 320 for video data and the data is appropriate, the PTS value, along with each write pointer position, Loaded into PTS table 324 for video. Demultiplexing of packets of audio data, video data and subpicture data is done in a similar contiguous manner, with the appropriate data loaded into the FIFO buffer and PTS table.
【0036】
As is clear from the above description, during the demultiplexing process, the data is written to the FIFO buffers 300 and 320 for audio data and video data as a function required for the demultiplexing process. Also, during the demultiplexing process, the PTS value does not correspond to each audio and video data and is stored in each PTS table 302, 324. In the simultaneous independent parallel processing, the audio data and the video data are read from the FIFO buffers 300 and 320, decoded by the audio decoder 25 and the video decoder 23 shown in FIG. 1, and output to the user. To. During the decoding process, the read pointers 318 and 332 of the FIFO buffers 300 and 320 for audio data and video data are automatically and continuously moved by the controller of the memory 24, and these read pointers are usually moved. , Not controlled by special instructions from CPU26. PTS table 302 for audio and video to read audio and video data streams from each FIFO buffer 300, 320 to play back audio frames and video data frames synchronously during the decoding process. , The appropriate PTS value stored in 324 must be matched again.
【0037】
When the audio decoder 25 shown in FIG. 1 detects that the header block has been extracted in step S204 shown in FIG. 2, it initializes the audio decoding process. Then, the audio decoder 25 interrupts the CPU 26 for audio decoding and starts the decoding process of the next audio data packet. When the audio decoder 25 interrupts the CPU 26 for audio decoding, the read pointer 318 shown in FIG. 4 continuously moves the current position in the FIFO buffer 300 in which the audio decoder 25 is reading data. Shown. By the way, the CPU 26 has to process a large number of interrupts with different priorities and real-time processing requirements. Therefore, when the CPU 26 is processing an interrupt having a higher priority than the audio decoder interrupt, there is a delay before the audio decoder interrupt is processed. Further, the CPU 26 communicates with the memory controller in the memory 24 and other circuits via the time division multiplexing bus or channel 29. During the first delay due to the processing of the start audio frame interrupt and the second delay due to communication with the memory controller via time division multiplexing bus 29, the audio decoder 25 is removed from the FIFO buffer 300 for audio data. The audio data is continuously read, and the read pointer 318 is also continuously moving.
【0038】
In step S400 shown in FIG. 5, the CPU 26 responds to the audio decoding interrupt. In step S402, the CPU 26 reads the value at the current position of the read pointer 318. By the way, the current position of the read pointer 318 is different from the position of the read pointer 318 when the audio decode interrupt is first supplied to the CPU 26. In step S404, the CPU 26 scans the positions of all write pointers stored in the PTS table 302 for audio and shows the value of the write pointer closest to the value of the current read pointer, eg, FIG. Detects the value of the write pointer at position 350. In step S408, the value of the closest write pointer is subtracted from the value of the current read pointer to obtain the difference. In step S410, the difference is compared with the maximum allowable difference value (hereinafter referred to as the maximum allowable difference value). This maximum allowable difference value is represented by a function that predicts the movement of the write pointer 318 within the maximum delay time that occurs when the CPU 26 responds to the audio decoding interrupt in step S400.
【0039】
In step S410, it is determined whether the difference between the value of the closest write pointer at the position 350 and the value of the current read pointer is equal to or greater than the maximum allowable difference value, and if applicable, the process ends. That is, there is no PTS value corresponding to the audio data being read by the current read pointer 318, including the PTS value at position 352 of the PTS table 302 for audio. If the difference is less than the maximum allowable difference value in step S410, the process proceeds to step S412. That is, it is determined that the PTS value at position 352 corresponding to the detected position 350 of the write pointer corresponds to the audio data currently being read by the read pointer 318.
【0040】
In step S412, the value of the frame counter 354 corresponding to the value of the closest write pointer at position 350 is evaluated, which is the duration or time of the audio decoding process, i.e. the length, as determined in terms of the number of audio frames. Is determined to be approximately equal to. The real time required to decode a frame of audio data can be predicted fairly accurately. Further, the decoding processing time can be calibrated or measured according to the number of audio frames of the data to be decoded. Since the frame counter is decremented as each frame of audio data is output, the value of the frame counter is expected to be equal to the number of audio frames in the decoder delay. Therefore, at the end of the decoding process, if the value of the frame counter 354 in the PTS table 302 for audio is approximately equal to the expected delay in decoding the current audio data, then the audio data being played is for audio. This means that the value of the frame counter in the PTS table 302 of is 0, which corresponds to the value of PTS.
【0041】
Since a bit error may occur during the decoding process, it may be determined in step S412 that the value of the counter is not substantially equal to the number of audio frames within the time of the audio decoding process. In this case, in step S414, the value of the frame counter 354 is changed to set the number of audio frames within the decoder delay time. Therefore, every time an error occurs in this process, the error value in one of the frame counters is detected and corrected. When the value of one frame counter is corrected, the values of the other frame counters under the frame counter 354 of the PTS table 302 for audio will also change with the same changes made to the frame counter 354. Be changed. That is, during the multiplex processing, the PTS value did not correspond to the audio data in the FIFO buffer 300, but during the decoding processing, the CPU 26 was for the audio data read from the FIFO buffer 300 and for audio. Responds to interrupts by associating with the PTS value in the PTS table 302 of.
【0042】
As shown in FIG. 6, the video decoding process, like the video demultiplexing process, is executed continuously and in parallel independently of the audio decoding process. The audio decoder 25 shown in FIG. 1 initializes the audio decoding process, while the CPU 26 initializes the video decoding process. In step S500 shown in FIG. 6, the CPU 26 detects the video header extracted from the video data packet in step S204 shown in FIG. The subsequent video decoding process is functionally the same as the audio decoding process. That is, in step S502, the CPU 26 reads the current position of the write pointer 332 shown in FIG. 4 of the FIFO buffer 320 for video data. As described above, this operation causes a delay in the CPU 26 accessing the memory controller in the memory 24 via the time division multiplexing communication bus 29. Therefore, the read pointer 332 corresponding to the FIFO buffer 320 for video data moves from the position at the time when the CPU 26 issues the instruction to read the pointer position and the position at the time when the CPU 26 recognizes the current read pointer. It ends up. In step S506, the CPU 26 sets the value of the write pointer position closest to the value of the current read pointer position in the PTS table 324 for video, for example, the value of the write pointer position at position 360 that matches the value of the read pointer position. To detect. In step S508, the difference between the value of the current read pointer position and the value of the nearest write pointer position is obtained. In step S510, this difference is compared to the maximum value. This maximum value is expressed as a function of the delay time that occurs when the CPU 26 determines the current position of the write pointer 332, as in the audio decoding process.
【0043】
When this difference is greater than or equal to the maximum value, it means that the video data corresponding to the current value of the read pointer does not have the corresponding PTS value stored in the PTS table 324 for video. On the other hand, when the difference between the current read pointer value and the nearest write pointer value is less than the maximum value, the PTS value at position 362 corresponding to the nearest write pointer value at position 360 is It is determined that it corresponds to the video data being read by the current read pointer 332. Further, in step S512, the value of the counter at position 264 of the PTS table 324 for video is set to the same value equal to the delay time of the video decoding process determined from the viewpoint of the number of frames to be decoded. As with audio data, the real time required to decode video data can be predicted fairly accurately. Also, the time required to process the video data corresponding to the current position of the read pointer can be calibrated or measured according to the number of frames of the video data. Therefore, the number of frames of video data that the video decoder 23 is about to process before the video data corresponding to the current position of the read pointer is played back from the video decoder 23 is the frame counter in the PTS table 324 for video. Loaded on 364. Then, in step S514, the CPU 26 supplies the video decoder 23 with a command to start the video decoding process.
【0044】
As described above, the audio data and the video data are continuously read from the audio and video FIFO buffers 300 and 320 shown in FIG. 4, and the read data is the audio and video decoder 25. , 23 are processed continuously, independently and in parallel. During the audio decoding process, the audio decoder 25 generates an audio frame that is ready for output to the user. At that time, the audio decoder 25 interrupts the CPU 26 with an audio frame, and immediately after that, performs a process of starting playback of the audio frame packet. As shown in FIG. 7, in step S600, the CPU 26 decrements the values of all the frame counters 309 of the PTS table 302 for audio shown in FIG. 4 by 1 in response to the audio frame interrupt. In step S604, the CPU 26 detects if there is a frame counter whose value is 0. For example, when the CPU 26 detects that the value of the frame counter 370 is 0, the CPU 26 proceeds to step S606. In step S606, the CPU 26 reads, for example, the value of the audio PTS corresponding to the frame counter whose value is 0 from position 372.
【0045】
When it detects that the state of the frame counter is 0, in step S608, the CPU 26 determines whether the audio has been selected as the master, and if applicable, proceeds to step S610. There are several different ways to synchronize audio and video data. As mentioned above, the system time counter provided in the CPU 26 is operating while playing audio and video frames. When perfectly synchronized, when audio and video frames are output, their stored PTS values will be equal to the current value of the system time counter. By the way, either the audio stream or the video stream of the output frame may not be synchronized with the system time counter or may not be synchronized with each other. The first way to synchronize is to select the audio as the master and sync the video to the audio. The second method is to select the video as the master and sync the audio to the video. A third method is to synchronize both audio and video with the system time counter. Each of these methods has advantages and disadvantages that depend on the signal source, the format of the signal source, the capabilities of the audio / video processor, and the like. The choice of mastering audio, mastering video, or not mastering is usually a system parameter chosen by the manufacturer of the television receiver. Alternatively, the user can select this. Another application is to dynamically and adaptively select the master by the audio / video processor.
【0046】
When the CPU 26 determines that the audio is the master in step S608 of FIG. 7, in step S610, the CPU 26 corresponds to the current value of the system time counter and the frame counter where the value such as position 372 is 0. The system time counter is extended and the value is offset so that the difference from the PTS value is equal to the system time counter. This extension of the system time counter is used to synchronize the video, as described below. When the CPU26 determines in step S608 that the audio is not the master, in step S612 the CPU26 changes the PTS value corresponding to the counter whose PTS table value is 0 to the current value of the system time counter. Determines whether the value is within a certain range with respect to the value obtained by adding the extension of the system time counter. Similar to the case where the audio in step S608 described above is the master, when the video master mode is used, the extended value of the system time counter is generated and updated in the video playback process. By the way, when neither audio nor video is selected as the master, the extension value of the system time counter is 0. In step S612, the value of the selected PTS is compared with the value obtained by adding the system time counter, the extension of the system time counter, and a certain limit value. In audio, the magnitude of this limit is the delay due to the CPU communicating with the memory controller, the delay until the CPU responds to the PTS interrupt, and the allowable error in audio and video synchronization. Is approximately equal to the value obtained by adding.
【0047】
In step S612, the value of PTS read from the PTS table 302 is greater than the current value of the system time counter plus its extension and limit, that is, the audio frame is what the system itself asks for the audio frame. If the playback time is earlier than or precedes the playback time, the process proceeds to step S614. That is, the audio frame being played and the video frame are clearly out of sync. In step S614, the CPU 26 supplies the audio decoder 25 with instructions to iterate over the audio frames in order to correct the out-of-sync. By repeating the audio frame while the video frame is naturally progressing, the audio frame is brought back in time and resynchronized with the video frame.
【0048】
If the value of PTS read from PTS table 302 in step S612 is less than the sum of the current value of the system time counter and its extension and limit, in step S616, the CPU 26 determines the PTS. It is determined whether the value of PTS in table 302 is smaller than the current value of the system time counter and the value obtained by subtracting the limit value from the added value of the extension, and if applicable, the process proceeds to step S618. When the PTS value is small, it means that the audio frame is being played later or later than the playback time required by the system itself for the audio frame. In step S618, the CPU 26 supplies the audio decoder 25 with instructions to skip audio frames in order to correct the out-of-sync. By skipping the audio frame while the video frame is naturally progressing, the audio frame is in a time-advanced state and is resynchronized with the video frame.
【0049】
As shown in FIG. 8, the output processing of the video frame is similar to the output processing of the audio frame shown in FIG. 7 except for the first part. In step S700, the CPU 26 responds to the interrupt generated by the raster scan process that controls the playback of the video frame. After responding to the interrupt in step S700, in step S702, the CPU 26 initializes various internal registers for playing the video. After that, in step S704, the CPU 26 decrements the values of all the frame counters such as the frame counters 336 and 364 by one. The CPU does not decrease the video frame counter having a value of FF. In step S708, the CPU determines if there is a frame counter with a value of 0. When the CPU 26 detects that the value of the frame counter 336 is 0, for example, the CPU 26 reads the value of the video PTS.
【0050】
In step S712 shown in FIG. 8, the CPU 26 determines whether the video is the master, and if applicable, proceeds to step S714. In step S714, the CPU 26 sets the system time so that the system time counter is equal to the difference between the current value of the system time counter and the PTS value corresponding to the frame counter where the value of, for example, position 372 is 0. Extend the counter or give an offset to its value. When the CPU26 determines in step S712 that the audio is not the master, in step S716 the CPU26 determines that the PTS value corresponding to the counter whose PTS table counter value is 0 is relative to the current value of the system counter. Determine if it is within a certain range. In step S716, the CPU 26 determines if the value of the video PTS at position 338 in the PTS table is greater than the current value of the system time counter plus its extension and limit. For video, this limit is approximately equal to half the duration of the shortest video frame. In step S716, if the value of the video PTS at position 338 in the PTS table is greater than the current value of the system time counter plus its extensions and limits, then the video frame is a video frame that the system itself becomes a video frame. It means that the playback time is earlier than or preceded by the desired playback time. In step S718, the CPU 26 supplies the video decoder 23 with instructions to iterate over the video frames in order to correct the out-of-sync. By repeating the video frame while the audio frame is naturally progressing, the video frame is brought back in time and resynchronized with the audio frame.
【0051】
In step S716, if the value of the video PTS at position 338 in the PTS table is less than the current value of the system time counter plus its extensions and limits, the video frame is calculated by the system itself for the video frame. It means that the playback time is later than or after the playback time. In step S722, the CPU 26 supplies the video decoder 23 with instructions to skip video frames in order to correct the out-of-sync. By skipping the video frame while the audio frame is naturally in progress, the video frame is in a time-advanced state and is resynchronized with the audio frame.
【0052】
That is, the audio and video output processing is performed in parallel, continuously, and independently so that the audio and video are synchronized with the user. During this process, the PTS of the currently displayed video frame is compared with the PTS of the currently output audio frame, and if out of sync is detected, the video frame is repeated or skipped. Try to get back in sync.
【0053】
The television receiver to which the present invention is applied has important advantages over conventional devices. In traditional devices, the combination of delayed response to audio PTS interrupts and the need to process packets of audio data very quickly causes the wrong audio PTS value to be selected from the PTS table for audio. There was a risk of In this case, the wrong audio PTS value will correspond to the audio data being read from the audio data FIFO. In a conventional device, as described in FIG. 6, the counter of the PTS table 302 for audio operates in the same manner as the video counter. Therefore, the counter provided in the storage 309 of the PTS table for audio is loaded during the audio decoding process, along with the number of audio frames that match the time of the audio decoder. Therefore, in a conventional device, if the wrong PTS value is selected from the PTS table, the wrong frame counter value is stored in the PTS table 302 after the delay due to decoding. Further, this wrong PTS value is used for the purpose of synchronizing the audio and the video for playback, and as a result, the audio and the video are synchronized and output to the user.
【0054】
In the present invention, the number of audio frames included in the header information of the audio data packet is one of the demultiplexing processes when generating the frame counter of the PTS table 302 for audio and the corresponding PTS value. Used as a part. Also, the value of the frame counter is decremented during the decoding process to output each audio frame, so the number of audio frames decoded and output based on the PTS table 302 for audio is the number of each audio data packet. It is precisely controlled during the decoding process. This method is much more accurate than the traditional method of comparing memory pointers to identify PTS values in associating the PTS value with the output decoded data. In the apparatus according to the present invention, the memory pointer comparison is used only to confirm whether the PTS table counter value for audio is substantially equal to the delay value of the audio decoder in order to correct the error that has occurred. There is.
【0055】
In the present invention, the number of audio frames from the header block is used to put the PTS value in the decoding process into the audio frame counter in the CPU, so that the audio PTS value and the audio PTS value are demultiplexed during the playback process. It can continue to match the audio data exactly. The header portion of the video data packet does not contain similar information about the number of video frames between the expected PTS values. However, in the present invention, the reproduction of the audio frame is very accurate, so by adjusting the audio and video to synchronize according to the processing shown in FIGS. 7 and 8, the audio can be reproduced in a state very close to the ideal reproduction. And video can be synchronized and output with high accuracy.
【0056】
The present invention is not limited to the above-described embodiment. For example, the header block of an audio data packet contains a PTS value and bytes of audio data between consecutive audio PTS values. It contains the number, that is, the number of bytes of the audio data of the current data packet plus the number of bytes of audio data that has the next PTS value but up to the data packet. This data may be stored in the PTS tables 302 and 324 for audio and video instead of the position of the write pointer.
【0057】
In FIG. 4, the value of the frame counter is obtained by adding the value of the previous counter to the number of audio frames from the header block of the audio data packet, and this added value is the frame corresponding to the current PTS value. I am trying to write to the counter. Then, the values of all the frame counters are decremented by 1 each time each audio frame is output. Instead of such processing, for example, the number of audio frames from the header block of the audio data packet may be written to the frame counter. In this case, the value of only the currently operating counter is decremented with the output of the audio frame. When the value of the counter reaches 0, the value of the next counter is decremented.
【0058】
In the above-described embodiment, for example, a system time counter of an apparatus is used to synchronize audio and video. This system time counter is a counter that indicates the state of the count that is directly related to the value of PTS when playing back audio data and video data. Instead of such a system time counter, audio and video may be synchronized using, for example, another counter or clock that can be adjusted by the value of PTS.
【0059】
[Effect of the invention]
In the present invention, the number of audio frames from the header block is used to put the PTS value in the decoding process into the audio frame counter in the CPU, so that the audio PTS value and the audio PTS value are demultiplexed during the playback process. It can continue to match the audio data exactly. The header portion of the video data packet does not contain similar information about the number of video frames between the expected PTS values. However, in the present invention, the reproduction of the audio frame is very accurate, so by adjusting the audio and the video to be synchronized, the audio and the video are synchronized and output with high accuracy in a state very close to the ideal reproduction. can do.
[Simple explanation of drawings]
[Figure 1]
It is a block diagram which shows the specific structure of the television receiver using the digital audio / video processor to which this invention is applied.
[Figure 2]
It is a flowchart for demonstrating the operation of the demultiplexer which constitutes a digital audio / video processor.
[Fig. 3]
It is a flowchart for demonstrating demultiplex processing executed by CPU.
[Fig. 4]
It is a schematic diagram which shows the arrangement of the data in the memory which constitutes a digital audio / video processor.
[Fig. 5]
It is a flowchart for demonstrating the audio decoding process executed by a CPU.
[Fig. 6]
It is a flowchart for demonstrating the video decoding process executed by a CPU.
[Fig. 7]
It is a flowchart for demonstrating the reproduction process of a frame of audio data executed by a CPU.
[Fig. 8]
It is a flowchart for demonstrating the reproduction process of a frame of video data executed by a CPU.
[Explanation of symbols]
18 television receiver, 19 demultiplexer, 20 digital audio / video processor, 23 video decoder, 24 memory, 25 audio decoder, 26 CPU, 27 audio I / O, 28 video mixer, 29 time division multiplexing bus, 38 encoder, 40 D / A, 42 video monitor, 48 host processor
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP10271457A | Cites | Japan |
| JP10262246A | Cites | Japan |
| JP10136355A | Cites | Japan |
| JP9128895A | Cites | Japan |
5 members in 3 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 08901090 | United States of America | – | |
| 90109097 | United States of America | A |
Members5
| Document | Office | Kind | |
|---|---|---|---|
| EP0895427A2 | European Patent Office (EPO) | A2 | |
| JPH11191286A | Japan | A | |
| US5959684A | United States of America | A | |
| JP3215087B2This record | Japan | B2 | |
| EP0895427A3 | European Patent Office (EPO) | A3 |
8 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 |
Numbers
- Publication
- 3215087
- Application
- 10227619
Titles2
- Japanese
- オーディオとビデオの同期方法及びデジタルビデオプロセッサ
- English
- INDUSTRIAL APPLICABILITY: Audio and video synchronization method and digital video processor
Classification
- CPC, 4
- H04N21/4341
- H04N19/44
- H04N19/42
- H04N21/43072
- IPC, 11
- H04N5 60
- G09G5 00
- G09G5 12
- G11B27 10
- H04N5 92
- H04N5 93
- H04N7 04
- H04N7 045
- H04N7 26
- H04N21 2368
- H04N21 434
