Synchronization of audio and video data in wireless communication system
Abstract
Problem to be solved.To provide encoding of an audio video stream that is transmitted over a network, for example, a wireless or IP network such that an entire frame of audio and an entire frame of video are transmitted simultaneously within a period required to render the audio video stream frames by an application in a receiver.
Solution.This encoding includes receiving audio and video RTP streams and assigning an entire frame of RTP video data to communication channel packets that occupy the same period, or less, as the video frame rate. An entire frame of RTP audio data is assigned to communication channel packets that occupy the same period, or less, as the audio frame rate. The video and audio communication channel packets are transmitted simultaneously. Receiving and assigning RTP streams can be performed in a remote station, or a base station.
Copyright (C)2011,JPO&INPIT
Term
Projected expiry 29 November 2030.
- Priority
- Filed
- Published
- Today
- Projected expiry
29 claims: 18 independent, 11 dependent
- 1A first decoder configured to receive a first encoded data stream and output a decoded first data stream, and here the first encoded data stream is during an information interval. Has a first bit rate;第1符号化されたデータストリームを受け取りそしてデコードされた第1データストリームを出力するように構成された第1デコーダと、なおここでは、前記第1符号化されたデータストリームは情報間隔の間、第1ビットレートを有する;A second decoder configured to receive a second encoded data stream and output a decoded second data stream, and here the second encoded data stream is during the information interval. , Has a second bit rate;第2符号化されたデータストリームを受け取りそしてデコードされた第2データストリームを出力するように構成された第2デコーダと、なおここでは、前記第2符号化されたデータストリームは前記情報間隔の間、第2ビットレートを有する;With a first buffer configured to store the decoded first data stream for at least one information interval and output frames of the decoded first data stream every interval period;少なくとも1情報間隔の間、前記デコードされた第1データストリームを蓄積し、そして間隔期間毎に前記デコードされた第1データストリームのフレームを出力するように構成された第1バッファと;With a second buffer configured to store the decoded second data stream for at least one information interval and output frames of the decoded second data stream every interval period;and decoded A combiner configured to receive the frame of the first data stream and the frame of the decoded second data stream and output the synchronized frame of the first and second decoded data streams. When;少なくとも1情報間隔の間、前記デコードされた第2データストリームを蓄積し、そして間隔期間毎に前記デコードされた第2データストリームのフレームを出力するように構成された第2バッファと;そして デコードされた第1データストリームの前記フレームおよびデコードされた第2データストリームの前記フレームを受け取りそして第1および第2のデコードされたデータストリームの同期をとられたフレームを出力するように構成された結合器と;A data stream synchronizer equipped with. を備えるデータストリーム同期装置。
- 5With a video decoder configured to receive encoded video data and output the decoded video data;符号化されたビデオデータを受け取りそしてデコードされたビデオデータを出力するように構成されたビデオデコーダと;With an audio decoder configured to receive encoded audio data and output the decoded audio data;符号化されたオーディオデータを受け取りそしてデコードされたオーディオデータを出力するように構成されたオーディオデコーダと;With a video buffer configured to store the decoded video data for at least one frame period and output frames of the video data every frame period;少なくとも1フレーム期間、デコードされたビデオデータを蓄積し、そしてフレーム期間毎にビデオデータのフレームを出力するように構成されたビデオバッファと;With an audio buffer configured to store decoded audio data for multiple frame periods and output frames of audio data for each frame period;and to receive said frames of video data and said frames of audio data and With a coupler configured to output synchronized frames of audio-video data;複数のフレーム期間、デコードされたオーディオデータを蓄積し、そしてフレーム期間毎にオーディオデータのフレームを出力するように構成されたオーディオバッファと;そして ビデオデータの前記フレームおよびオーディオデータの前記フレームを受け取りそしてオーディオビデオデータの同期をとられたフレームを出力するように構成された結合器と;Remote station device equipped with. を備える遠隔局装置。
- 9With a video communication channel interface configured to receive a video RTP stream and allocate all frames of RTP video data to communication channel packets that occupy the same or less time than the video frame rate;ビデオRTPストリームを受け取りそしてRTPビデオデータの全フレームをビデオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるように構成されたビデオ通信チャネルインタフェースと;With an audio communication channel interface configured to receive an audio RTP stream and allocate all frames of RTP audio data to communication channel packets that occupy the same or less time than the audio frame rate;and the video and audio communication channel packets. With transmitters configured to receive and send;オーディオRTPストリームを受け取りそしてRTPオーディオデータの全フレームをオーディオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるように構成されたオーディオ通信チャネルインタフェースと;そして 前記ビデオおよびオーディオの通信チャネルパケットを受け取りそして送信するように構成された送信機と;Remote station device equipped with. を備える遠隔局装置。
- 11With a video decoder configured to receive encoded video data and output the decoded video data;符号化されたビデオデータを受け取りそしてデコードされたビデオデータを出力するように構成されたビデオデコーダと;With an audio decoder configured to receive encoded audio data and output the decoded audio data;符号化されたオーディオデータを受け取りそしてデコードされたオーディオデータを出力するように構成されたオーディオデコーダと;With a video buffer configured to store the decoded video data during the video frame period and output frames of the video data at each frame period;ビデオフレーム期間、デコードされたビデオデータを蓄積し、そしてフレーム期間毎にビデオデータのフレームを出力するように構成されたビデオバッファと;With an audio buffer configured to store the decoded audio data during the audio frame period and output a frame of audio data at each frame period;and receive and audio the frame of the video data and the frame of the audio data. With a coupler configured to output synchronized frames of video data;オーディオフレーム期間、デコードされたオーディオデータを蓄積し、そしてフレーム期間毎にオーディオデータのフレームを出力するように構成されたオーディオバッファと;そして ビデオデータの前記フレームおよびオーディオデータの前記フレームを受け取りそしてオーディオビデオデータの同期をとられたフレームを出力するように構成された結合器と;Base station equipment equipped with. を備える基地局装置。
- 15With a video communication channel interface configured to receive a video RTP stream and allocate all frames of RTP video data to communication channel packets that occupy the same or less time than the video frame rate;ビデオRTPストリームを受け取りそしてRTPビデオデータの全フレームをビデオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるように構成されたビデオ通信チャネルインタフェースと;With an audio communication channel interface configured to receive an audio RTP stream and allocate all frames of RTP audio data to communication channel packets that occupy the same or less time than the audio frame rate;and the video and audio communication channel packets. With transmitters configured to receive and send;オーディオRTPストリームを受け取りそしてRTPオーディオデータの全フレームをオーディオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるように構成されたオーディオ通信チャネルインタフェースと;そして 前記ビデオおよびオーディオの通信チャネルパケットを受け取りそして送信するように構成された送信機と;Base station equipment equipped with. を備える基地局装置。
- 17A video communication channel interface configured to receive a video RTP stream and allocate all frames of RTP video data to communication channel packets that occupy the same or less time than the video frame rate, and receive an audio RTP stream and of RTP audio data. An audio communication channel interface configured to allocate all frames to communication channel packets that occupy the same or less time than the audio frame rate, and a transmitter configured to receive and transmit said video and audio communication channel packets. And with the base station equipment that has;ビデオRTPストリームを受け取りそしてRTPビデオデータの全フレームをビデオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるように構成されたビデオ通信チャネルインタフェースと、 オーディオRTPストリームを受け取りそしてRTPオーディオデータの全フレームをオーディオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるように構成されたオーディオ通信チャネルインタフェースと、 前記ビデオおよびオーディオの通信チャネルパケットを受け取りそして送信するように構成された送信機と、 を有する基地局装置と;A video decoder configured to receive video communication channel packets and output decoded video data, an audio decoder configured to receive audio communication channel packets and output decoded audio data, and a video frame. A video buffer configured to store period, decoded video data, and output frames of video data every frame period, and audio frame period, store decoded audio data, and every frame period. A combination configured to output a frame of audio data and an audio buffer configured to receive the frame of video data and the frame of audio data and output a synchronized frame of audio-video data. With a vessel and a remote station device with;ビデオ通信チャネルパケットを受け取りそしてデコードされたビデオデータを出力するように構成されたビデオデコーダと、 オーディオ通信チャネルパケットを受け取りそしてデコードされたオーディオデータを出力するように構成されたオーディオデコーダと、 ビデオフレーム期間、デコードされたビデオデータを蓄積し、そしてフレーム期間毎にビデオデータのフレームを出力するように構成されたビデオバッファと、 オーディオフレーム期間、デコードされたオーディオデータを蓄積し、そしてフレーム期間毎にオーディオデータのフレームを出力するように構成されたオーディオバッファと、そして ビデオデータの前記フレームおよびオーディオデータの前記フレームを受け取りそしてオーディオビデオデータの同期をとられたフレームを出力するように構成された結合器と、 を有する遠隔局装置と;A wireless communication system including. を備える無線通信システム。
- 18A video communication channel interface configured to receive a video RTP stream and allocate all frames of RTP video data to communication channel packets that occupy the same or less time than the video frame rate, and receive an audio RTP stream and of RTP audio data. An audio communication channel interface configured to allocate all frames to communication channel packets that occupy the same or less time than the audio frame rate, and a transmitter configured to receive and transmit said video and audio communication channel packets. And with remote station equipment with;ビデオRTPストリームを受け取りそしてRTPビデオデータの全フレームをビデオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるように構成されたビデオ通信チャネルインタフェースと、 オーディオRTPストリームを受け取りそしてRTPオーディオデータの全フレームをオーディオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるように構成されたオーディオ通信チャネルインタフェースと、 前記ビデオおよびオーディオの通信チャネルパケットを受け取りそして送信するように構成された送信機と、 を有する遠隔局装置と;A video decoder configured to receive video communication channel packets and output decoded video data, an audio decoder configured to receive audio communication channel packets and output decoded audio data, and a video frame. A video buffer configured to store period, decoded video data, and output frames of video data every frame period, and audio frame period, store decoded audio data, and every frame period. A combination configured to output a frame of audio data and an audio buffer configured to receive the frame of video data and the frame of audio data and output a synchronized frame of audio-video data. With a vessel and a base station device with;ビデオ通信チャネルパケットを受け取りそしてデコードされたビデオデータを出力するように構成されたビデオデコーダと、 オーディオ通信チャネルパケットを受け取りそしてデコードされたオーディオデータを出力するように構成されたオーディオデコーダと、 ビデオフレーム期間、デコードされたビデオデータを蓄積し、そしてフレーム期間毎にビデオデータのフレームを出力するように構成されたビデオバッファと、 オーディオフレーム期間、デコードされたオーディオデータを蓄積し、そしてフレーム期間毎にオーディオデータのフレームを出力するように構成されたオーディオバッファと、そして ビデオデータの前記フレームおよびオーディオデータの前記フレームを受け取りそしてオーディオビデオデータの同期をとられたフレームを出力するように構成された結合器と、 を有する基地局装置と;A wireless communication system including. を備える無線通信システム。
- 19Receiving the first encoded data stream, decoding it, and outputting the decoded first data stream, where the first encoded data stream is the first bit rate during the information interval. Have;第1符号化されたデータストリームを受け取り、デコードしそしてデコードされた第1データストリームを出力することと、なおここでは、前記第1符号化されたデータストリームは情報間隔の間、第1ビットレートを有する;Receiving the second encoded data stream, decoding and outputting the decoded second data stream, where the second encoded data stream is the second bit during the information interval. Have a rate;第2符号化されたデータストリームを受け取り、デコードしそしてデコードされた第2データストリームを出力することと、なおここでは、前記第2符号化されたデータストリームは前記情報間隔の間、第2ビットレートを有する;Accumulating the decoded first data stream for at least one information interval and outputting a frame of the decoded first data stream every interval period;少なくとも1情報間隔の間、前記デコードされた第1データストリームを蓄積し、そして間隔期間毎に前記デコードされた第1データストリームのフレームを出力することと;Accumulating the decoded second data stream for at least one information interval and outputting a frame of the decoded second data stream every interval period;and said to the decoded first data stream. Combining a frame with said frame of a decoded second data stream and outputting a synchronized frame of the first and second decoded data streams;少なくとも1情報間隔の間、前記デコードされた第2データストリームを蓄積し、そして間隔期間毎に前記デコードされた第2データストリームのフレームを出力することと;そして デコードされた第1データストリームの前記フレームとデコードされた第2データストリームの前記フレームとを結合し、そして第1および第2のデコードされたデータストリームの同期をとられたフレームを出力することと;How to decode and synchronize a data stream, including. を含む、データストリームをデコードし同期させる方法。
- 20A method of decoding and synchronizing audio and video data, receiving encoded video data and outputting the decoded video data;オーディオおよびビデオデータをデコードし同期させる方法であって、 符号化されたビデオデータを受け取りそしてデコードされたビデオデータを出力することと;To receive encoded audio data and output decoded audio data;符号化されたオーディオデータを受け取りそしてデコードされたオーディオデータを出力することと;With the video frame period, accumulating the decoded video data, and outputting the frame of the video data for each frame period;ビデオフレーム期間、デコードされたビデオデータを蓄積し、そしてフレーム期間毎にビデオデータのフレームを出力することと;For the audio frame period, accumulating the decoded audio data, and outputting the frame of the audio data for each frame period;and combining the frame of the video data and the frame of the audio data, and for each video frame period. To output synchronized frames of audio-video data to;オーディオフレーム期間、デコードされたオーディオデータを蓄積し、そしてフレーム期間毎にオーディオデータのフレームを出力することと;そして ビデオデータの前記フレームとオーディオデータの前記フレームとを結合し、そしてビデオフレーム期間毎にオーディオビデオデータの同期をとられたフレームを出力することと;How to include. を含む方法。
- 21A method of encoding audio and video data, receiving a video RTP stream and allocating all frames of the RTP video data to communication channel packets that occupy the same or less time than the video frame rate;and assigning the audio RTP stream. Receiving and allocating all frames of RTP audio data to communication channel packets that occupy the same or less time than the audio frame rate;オーディオおよびビデオデータを符号化する方法であって、 ビデオRTPストリームを受け取りそしてRTPビデオデータの全フレームをビデオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てることと;そして オーディオRTPストリームを受け取りそしてRTPオーディオデータの全フレームをオーディオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てることと;Including methods. を含む、方法。
- 22A computer-readable medium that embodies a method of decoding and synchronizing a data stream, the method of receiving a first encoded data stream, decoding and outputting the decoded first data stream, and still Here, the first encoded data stream has a first bit rate during the information interval;データストリームをデコードし同期させる方法を具現化するコンピュータ可読媒体であって、前記方法は 第1符号化されたデータストリームを受け取り、デコードしそしてデコードされた第1データストリームを出力することと、なおここでは、前記第1符号化されたデータストリームは情報間隔の間、第1ビットレートを有する;Receiving the second encoded data stream, decoding and outputting the decoded second data stream, where the second encoded data stream is the second bit during the information interval. Have a rate;第2符号化されたデータストリームを受け取り、デコードしそしてデコードされた第2データストリームを出力することと、なおここでは、前記第2符号化されたデータストリームは前記情報間隔の間、第2ビットレートを有する;Accumulating the decoded first data stream for at least one information interval and outputting a frame of the decoded first data stream every interval period;少なくとも1情報間隔の間、前記デコードされた第1データストリームを蓄積し、そして間隔期間毎に前記デコードされた第1データストリームのフレームを出力することと;Accumulating the decoded second data stream for at least one information interval and outputting a frame of the decoded second data stream every interval period;and said to the decoded first data stream. Combining a frame with said frame of a decoded second data stream and outputting a synchronized frame of the first and second decoded data streams;少なくとも1情報間隔の間、前記デコードされた第2データストリームを蓄積し、そして間隔期間毎に前記デコードされた第2データストリームのフレームを出力することと;そして デコードされた第1データストリームの前記フレームとデコードされた第2データストリームの前記フレームとを結合し、そして第1および第2のデコードされたデータストリームの同期をとられたフレームを出力することと;Computer-readable media, including. を含む、コンピュータ可読媒体。
- 23A computer-readable medium that embodies a method of decoding and synchronizing audio and video data, the method of receiving encoded video data and outputting the decoded video data;オーディオおよびビデオデータをデコードし同期させる方法を具現化するコンピュータ可読媒体であって、前記方法は 符号化されたビデオデータを受け取りそしてデコードされたビデオデータを出力することと;To receive encoded audio data and output decoded audio data;符号化されたオーディオデータを受け取りそしてデコードされたオーディオデータを出力することと;With the video frame period, accumulating the decoded video data, and outputting the frame of the video data for each frame period;ビデオフレーム期間、デコードされたビデオデータを蓄積し、そしてフレーム期間毎にビデオデータのフレームを出力することと;For the audio frame period, accumulating the decoded audio data, and outputting the frame of the audio data for each frame period;and combining the frame of the video data and the frame of the audio data, and for each video frame period. To output synchronized frames of audio-video data to;オーディオフレーム期間、デコードされたオーディオデータを蓄積し、そしてフレーム期間毎にオーディオデータのフレームを出力することと;そして ビデオデータの前記フレームとオーディオデータの前記フレームとを結合し、そしてビデオフレーム期間毎にオーディオビデオデータの同期をとられたフレームを出力することと;Computer-readable media, including. を含む、コンピュータ可読媒体。
- 24A computer-readable medium that embodies a method of encoding audio and video data, said method is a communication channel that receives a video RTP stream and occupies all frames of the RTP video data at the same or less time than the video frame rate. Assigning to packets;and receiving audio RTP streams and allocating all frames of RTP audio data to communication channel packets that occupy the same or less time than the audio frame rate;オーディオおよびビデオデータを符号化する方法を具現化するコンピュータ可読媒体であって、前記方法は、 ビデオRTPストリームを受け取りそしてRTPビデオデータの全フレームをビデオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てることと;そして オーディオRTPストリームを受け取りそしてRTPオーディオデータの全フレームをオーディオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てることと;Computer-readable media, including. を含む、コンピュータ可読媒体。
- 25A means for receiving the first encoded data stream and outputting the decoded first data stream, and here the first encoded data stream has a first bit rate during the information interval. Have;第1符号化されたデータストリームを受け取り、そしてデコードされた第1データストリームを出力するための手段と、なおここでは、前記第1符号化されたデータストリームは情報間隔の間、第1ビットレートを有する;A means for receiving a second encoded data stream and outputting a decoded second data stream, and here the second encoded data stream is a second bit during the information interval. Have a rate;第2符号化されたデータストリームを受け取り、そしてデコードされた第2データストリームを出力するための手段と、なおここでは、前記第2符号化されたデータストリームは前記情報間隔の間、第2ビットレートを有する;As a means for accumulating the decoded first data stream for at least one information interval and outputting a frame of the decoded first data stream for each interval period;少なくとも1情報間隔の間、前記デコードされた第1データストリームを蓄積し、そして間隔期間毎に前記デコードされた第1データストリームのフレームを出力するための手段と;A means for accumulating the decoded second data stream for at least one information interval and outputting a frame of the decoded second data stream for each interval period;and a decoded first data stream. With means for combining said frames of the first and second decoded data streams with said frames of the decoded second data stream and outputting synchronized frames of the first and second decoded data streams;少なくとも1情報間隔の間、前記デコードされた第2データストリームを蓄積し、そして間隔期間毎に前記デコードされた第2データストリームのフレームを出力するための手段と;そして デコードされた第1データストリームの前記フレームとデコードされた第2データストリームの前記フレームとを結合し、そして第1および第2のデコードされたデータストリームの同期をとられたフレームを出力するための手段と;A data stream synchronizer. を備える、データストリーム同期装置。
- 26With means for receiving encoded video data and outputting decoded video data;符号化されたビデオデータを受け取りそしてデコードされたビデオデータを出力するための手段と;With means for receiving encoded audio data and outputting decoded audio data;符号化されたオーディオデータを受け取りそしてデコードされたオーディオデータを出力するための手段と;With a means for accumulating decoded video data during a video frame period and outputting a frame of video data for each frame period;ビデオフレーム期間、デコードされたビデオデータを蓄積し、そしてフレーム期間毎にビデオデータのフレームを出力するための手段と;A means for accumulating decoded audio data during an audio frame period and outputting a frame of audio data at each frame period;and combining said frames of video data with said frames of audio data, and audio video. As a means to output synchronized frames of data;オーディオフレーム期間、デコードされたオーディオデータを蓄積し、そしてフレーム期間毎にオーディオデータのフレームを出力するための手段と;そして ビデオデータの前記フレームとオーディオデータの前記フレームとを結合し、そしてオーディオビデオデータの同期をとられたフレームを出力するための手段と;Remote station device equipped with. を備える遠隔局装置。
- 27A means for receiving a video RTP stream and allocating all frames of RTP video data to communication channel packets that occupy the same or less time than the video frame rate;and receiving an audio RTP stream and allocating all frames of RTP audio data to audio frames. Means for allocating communication channel packets that occupy the same or less time than the rate;ビデオRTPストリームを受け取りそしてRTPビデオデータの全フレームをビデオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるための手段と;そして オーディオRTPストリームを受け取りそしてRTPオーディオデータの全フレームをオーディオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるための手段と;Remote station device equipped with. を備える遠隔局装置。
- 28With means for receiving encoded video data and outputting decoded video data;符号化されたビデオデータを受け取りそしてデコードされたビデオデータを出力するための手段と;With means for receiving encoded audio data and outputting decoded audio data;符号化されたオーディオデータを受け取りそしてデコードされたオーディオデータを出力するための手段と;With a means for accumulating decoded video data during a video frame period and outputting a frame of video data for each frame period;ビデオフレーム期間、デコードされたビデオデータを蓄積し、そしてフレーム期間毎にビデオデータのフレームを出力するための手段と;A means for accumulating decoded audio data during an audio frame period and outputting a frame of audio data at each frame period;and combining said frames of video data with said frames of audio data and audio video data. With a means to output synchronized frames;オーディオフレーム期間、デコードされたオーディオデータを蓄積し、そしてフレーム期間毎にオーディオデータのフレームを出力するための手段と;そして ビデオデータの前記フレームとオーディオデータの前記フレームとを結合しそしてオーディオビデオデータの同期をとられたフレームを出力するための手段と;Base station equipment equipped with. を備える基地局装置。
- 29A means for receiving a video RTP stream and allocating all frames of RTP video data to communication channel packets that occupy the same or less time than the video frame rate;and receiving an audio RTP stream and allocating all frames of RTP audio data to audio frames. Means for allocating communication channel packets that occupy the same or less time than the rate;ビデオRTPストリームを受け取りそしてRTPビデオデータの全フレームをビデオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるための手段と;そして オーディオRTPストリームを受け取りそしてRTPオーディオデータの全フレームをオーディオフレームレートと同じあるいはより少ない期間を占める通信チャネルパケットに割り当てるための手段と;Base station equipment equipped with. を備える基地局装置。
Independent claims18
76 paragraphs, as filed
Priority claim
(Claim of priority under 35 USC 119) This patent application is assigned to the assignee and expressly incorporated herein by reference, "Multimedia Packets Carried by" filed on May 13, 2004. Claims the priority of US Patent Provisional Application No. 60 / 571,673 entitled "CDMA Physical Layer Products)".
(Refer to the pending related patent application) This patent application relates to the following co-pending US patent application: "Delivery Of Information Over A Communication Channel" of agent reference number 030166UI, which is filed at the same time, transferred to the assignee, and explicitly incorporated herein by reference in its entirety; At the same time, the method and apparatus for assigning information to the channels of the communication system of agent reference number 030166U2, which is filed at the same time, transferred to the assignee, and explicitly incorporated herein by reference in its entirety. For Allocation Of Information To Channels Of A Communication System) "; "Header Compression Of Multimedia Data Transmitted Over A Wireless Communication System) ".
background
(Field) The present invention generally relates to the distribution of information on wireless communication systems, and more specifically to the synchronization of audio and video data transmitted on wireless communication systems.
(background) Various techniques have been developed for transmitting multimedia or real-time data such as audio or video data over various communication networks. One such technology is the real-time transport protocol (RTP). RTP provides end-to-end network transport functions suitable for applications that send real-time data over multicast or unicast network services. RTP does not support resource reservation and does not guarantee quality-of-service for real-time services. Allows monitoring of data delivery in a way that is extensible to large multicast networks and minimal control and identification Data transmission is enhanced by the Control Protocol (RTCP), which provides functionality. RTP and RTCP are designed to be layer-independent of the underlying transport and network. The protocol supports the use of RTP-level translators and mixers. Further details on RTP can be incorporated here in its entirety, "RTP: A Transport Protocol for Real-Time Applications (RTP)" (H. Schulzrinne [University of Columbia]. ], S. Casner [Packet Design], R. Frederick [Blue Court Systems], V. Jacobson [Packet Design], RFC-3550 Draft Standard Standard, Internet Engineering Steering Group, July 2003) be able to.
An example that illustrates the aspects of RTP is the conference call, where RTP remains at the top of the Internet Protocol (IP) services for voice communications. Through the allocation mechanism, the source of the conference gets the multicast group address and the paired ports. One port is used for audio data and the other is used for control (RTCP) packets. This address and port information will be distributed to the target participants. The conference call application used by each conference participant sends audio data in a small partition, for example a partition with a duration of 20ms. Each partition of audio data is preceded by an RTP header, and the combined RTP header and data are encapsulated in UDP packets. The RTP header contains information about the data, for example, what type of audio coding is included in each packet, such as PCM, ADPCM or LPC, the time the RTP packet is due to be given Time Indicates a Stamp (TS), Sequence Number (SN) of the packet that can be used to detect lost / duplicate packets, and so on. This changes the type of encoding that the sender uses during the conference, for example, to accommodate new participants connected through low bandwidth links, or to respond to network congestion indications. Allows you to.
According to the RTP standard, if both audio and video media are used in an RTP conference, they will be sent as separate RTP sessions. That is, separate RTP and RTCP packets are sent to each media using two different UDP port pairs and / or multicast addresses. Directly at the RTP level between audio and video sessions, except that users participating in both sessions must use the same name in both RTCP packets so that the sessions can be associated. There is no combination of.
The motivation for sending audio and video as separate RTP sessions is to allow some participants in the conference to receive only one medium, if they choose. Despite the isolation, synchronized playback of the source audio and video can be achieved using the timing information carried in the RTP / RTCP packets of both sessions.
Packet networks such as the Internet may sometimes lose or reorganize packets. In addition, individual packets may experience indefinite delays at their respective transmission times. To address these failures, the RTP header contains timing information and sequence numbers that allow the receiver to reconstruct the timing generated by the source. This timing reconstruction is done separately for each source of the session's RTP packets.
Even if the RTP header contains timing information and sequence numbers, the audio and video are delivered in separate RTP streams, so there is either lip-synch or AV-synch between the streams. There is a potential time slip, called. Applications on the receiving side will have to resynchronize these streams before giving them audio and video. In addition, in applications where RTP streams such as audio and video are transmitted over wireless networks, the likelihood of packet loss increases, which makes stream resynchronization more difficult.
Therefore, there is a technical need to improve the synchronization of audio and video RTP streams transmitted over the network.
The embodiments disclosed herein are described above by encoding a data stream, such as an audio-video stream, transmitted over a network, eg, a wireless or IP network, so that the data streams are synchronized. Corresponds to sex. For example, an entire frame of audio and an entire frame of video are within the frame period required to provide audio and video frames by the application at the receiving end. Will be sent. For example, a data stream synchronizer should receive a first encoded data stream and output a decoded first data stream. It may include a configured first decoder, where the first encoded data stream is during an information interval (during an information). It has an interval) and a first bit rate. The synchronized data may also include a second decoder configured to receive the second encoded data stream and output the decoded second data stream, where here the second code. The converted data stream has a second bit rate during the information interval. The first buffer is configured to store the first decoded data stream for at least one information interval and output frames of the first decoded data stream for each interval period. Has been done. The second buffer is configured to store the decoded second data stream for at least one information interval and output a frame of the decoded second data stream every interval period. Then a combiner (a) configured to receive the frames of the first decoded data stream and the frames of the second decoded data stream. combiner) outputs the synchronized frame of the first and second decoded data streams. The first encoded data stream can be video data, and the second encoded data stream can be audio data.
One aspect of this technology is that it receives audio and video RTP streams, and that occupy the same period, or less, as the video frame, that occupy the same period, or less, as the video frame. rate) Includes assigning to communication channel packets. Also, all frames of RTP audio data are allocated to communication channel packets that occupy the same or less time period as the audio frame rate. Video and audio communication channel packets are transmitted at the same time. Receiving and allocating RTP streams is a remote station. Or it can be executed at the base station.
Another aspect is receiving communication channel packets containing audio and video data. The audio and video data is decoded and the data is stored for a period equal to the frame period of the audio and video data. At the end of the frame period, the video frame and the audio frame are combined. The audio and video frames are synchronized because the audio and video frames are transmitted simultaneously and each transmission occurs within the frame period. Decoding and accumulating can be performed at a remote station or base station.
<figref num="1">FIG. 1 is an explanatory diagram of a part of a communication system configured according to the present invention.</figref><figref num="2">FIG. 2 is a block diagram showing an exemplary packet data network and various air interface options for delivering packet data over the wireless network in the system of FIG.</figref><figref num="3">FIG. 3 is a chart showing synchronization problems in conventional techniques for the transmission of separate RTP streams over wireless communication channels.</figref><figref num="4">FIG. 4 is a chart showing a technique for transmitting separate RTP streams over a wireless communication channel according to the present invention.</figref><figref num="5">FIG. 5 is a block diagram of a portion of a wireless audio / video receiver configured to receive communication channel packets.</figref><figref num="6">FIG. 6 is a block diagram of a portion of a wireless audio / video transmitter configured to transmit communication channel packets.</figref><figref num="7">FIG. 7 is a flow chart of transmission of an independent RTP stream over a wireless communication link.</figref><figref num="8">FIG. 8 is a flowchart of received audio and video data on the wireless communication channel.</figref><figref num="9">FIG. 9 is a block diagram of a wireless communication device or mobile station (MS) configured according to an exemplary embodiment of the present invention.</figref>
Detailed explanation
The term "exemplary" as used herein means "an example, an instance, or an illustration." Any embodiment described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments.
The term "streaming" as used herein refers to essentially continuous multimedia data on a dedicated or shared channel in an interactive unicast or broadcast application, such as audio, speech or video information. Means real time delivery. As used herein, the phrase "multimedia frame" for video means a video frame that can be displayed / drawn on a display device after decoration. Video frames can be further divided into independently decodable units. In video jargon, these are called "slices." In the case of audio and speech, the term "multimedia frame" used here is a time window (a time) in which the speech or audio is compressed for decoding at the transmission and receiver. It means the information in window). The phrase "information unit interval" used herein refers to the time duration of the multimedia frame described above. For example, in the case of video, the information unit interval is 100 ms for 10 frames / second video. Furthermore, as an example, in the case of speech, the information unit interval is typically 20 milliseconds in cdma2000, GSM® and WCDMA. From this description, it is clear that audio / speech frames are usually not further divided into independently decodable units, and video frames are usually divided into independently decodable slices. .. The phrases "multimedia frame", "information unit spacing", etc. are clear from the context when referring to multimedia data for video, audio and speech.
Techniques for synchronizing RTP streams transmitted over a set of constant bit rate communication channels are described. The technology includes partitioning the information units transmitted in an RTP stream into data packets, where the size of the data packet matches the physical layer data packet size of the communication channel ( match) is selected. For example, audio and video data that are synchronized with each other can be encoded. The encoder can be constrained so that it encodes the data to a size that fits the available physical layer packet size of the communication channel. RTP streams are transmitted simultaneously or continuously, but the data packet size must be one or more available physical layer packet sizes because audio and video packets are required to be delivered synchronously within the timeframe. Supports sending multiple synchronized RTM streams by constraining them to fit. For example, if audio and video RTP streams are transmitted and the data packets are suppressed so that their size matches the available physical layer packets, then the audio and video data will be in display time. Sent and synchronized. Different physical layers when the amount of data required to represent an RTP stream changes, as described in the simultaneously pending applications listed above (see Pending Related Patent Applications). The communication channel capacity changes depending on the packet size selection.
Examples of information units such as RTP streams include variable bit rate data streams, multimedia data, video data, and audio data. Information units can occur at constant repetition rates. For example, the information unit may be a frame of audio / video data.
Different national and international standards have been established to support a variety of air interfaces, such as Advanced Mobile Phone Service (AMPS), Pan-European Digital Mobile Phone System (Pan-European Digital Mobile Phone System). Global System for Mobile (GSM), General Packet Radio Service (GPRS), Enhanced Data GSM Environment (EDGE), Provisional Standard 95 (IS-95) and its derivatives, IS-95A, IS-95B, ANSI J-STD-008 (often collectively referred to here as IS-95), and new high data rate systems such as cdma 2000, Universal Mobile Telecommunications. Service) (UMTS), wideband CDMA (wideband) CDMA), WCDMA, etc. are included. These standards are disseminated by the American Telecommunications Industry Association (TIA), the Third Generation Partnership Project (3GPP), the European Telecommunications Standards Institute (ETSI) and other well-known standards bodies.
FIG. 1 shows a communication system 100 configured according to the present invention. Communication system 100 includes infrastructure 101, multiple wireless communication devices (WCD) 104 and 105, and landline communication devices 122 and 124. WCD will also be referred to as mobile stations (MS) or mobiles. In general, WCDs are mobile or fixed. Terrestrial communication devices 122 and 124 can include serving nodes, or content servers, for services that provide various types of multimedia data, such as streaming data. In addition, the MS can transmit streaming data such as multimedia data.
Infrastructure 101 also includes other components such as base stations 102, base station controllers 106, mobile switching centers 108, a switching network 120 and the like. It may contain things. In one embodiment, the base station 102 is integrated with the base station controller 106, and, in other embodiments, the base station 102 and base station controller 106 separate Component is cement. Different types of switching networks 120, such as IP networks, or public switched telephone networks (PSTNs), may be used to send signals in communication system 100.
The term "forward link" or "downlink" refers to the signal path from Infrastructure 101 to MS, and the term "reverse link" or "uplink". "link)" refers to the signal path from the MS to the infrastructure. As shown in FIG. 1, MS104 and 105 receive signals 132 and 136 on the forward link and transmit signals 134 and 138 on the reverse link. Generally, the signals transmitted from the MS 104 and 105 are intended to be received by another communication device, such as another remote unit, or terrestrial communication line communication devices 122 and 124, and pass through an IP network or a switching network 120. Will be sent. For example, if the signal 134 transmitted from the starting WCD (initiating WCD) 104 is the destination MS (destination). If intended to be received by MS) 105, the signal is sent through infrastructure 101 and signal 136 is sent over the forward link to the destination MS 105. Similarly, signals initiated in infrastructure 101 may be broadcast to MS105. For example, the content provider may send multimedia data, such as streaming multimedia data, to the MS105. Typically, a communication device such as an MS or terrestrial line communication device can be both an initiator and a destination of the signal.
Examples of MS104 include mobile phones, wirelessly communicable personal computers, and personal digital assistants (PDAs), and other wireless devices. Communication system 100 may be designed to support one or more wireless standards. For example, the standards are Pan-European Digital Mobile Phone System (GSM), General Line Radio Service (GPRS), Extended Data GSM Environment (EDGE), TIA / EIA-95-B (IS-95), TIA / EIA. -98-C (IS-98), IS2000, HRPD, cdma2000, standards called Broadband CDMA (WCDMA), and others may be included.
FIG. 2 is a block diagram illustrating an exemplary packet data network and various air interface options for delivering packet data over a wireless network. The techniques described can be implemented in packet-switched data networks such as those illustrated in Figure 2. As shown in the example of FIG. 2, the packet-switched data network system includes a radio channel 202, multiple recipient nodes or MS204, a sending node or content server 206, a serving node 208, and a controller 210. May include. The transmitting node 206 can be coupled to the serving node 208 via a network 212 such as the Internet.
The serving node 208 may include, for example, a packet data serving node (PDSN) or a serving GPRS support node (SGSN) or a gateway GPRS support node (GGSN). The serving node 208 can receive packet data from the transmitting node 206 and supply a packet of information to the controller 210. The controller 210 may include, for example, a base station controller / packet control function (BSC / PCF) or a radio network controller (RNC). In one embodiment, the controller 210 communicates with the serving node 208 on a Radio Access Network (RAN). The controller 210 communicates with the serving node 208 and sends a packet of information over the radio channel 202 to at least one receiving node 204.
In one embodiment, the serving node 208 and / or transmitting node 206 may also include an encoder for encoding the data stream, a decoder for decoding the data stream, or both. For example, an encoder could encode a video stream, thereby generating variable-sized frames of data, and a decoder could generate variable-sized frames of data. You will be able to receive and decode them. Since the frames are of various sizes but the video frame rate is constant, a variable bit rate stream of data is generated. Similarly, the MS may include an encoder for encoding the data stream, a decoder for decoding the received data stream, or both. The term "codec" is used to describe a combination of encoder and decoder.
In the example illustrated in FIG. 2, data such as multimedia data is delivered from a transmitting node 206 connected to a network or Internet 212 to a receiving node or MS204, serving node, or packet data serving node (PDSN). ) 206, and can be sent via the controller, or base station controller / packet control function (BSC / PCF) 208. The radio channel 202 interface between the MS204 and the BSC / PCF210 is an air interface and typically uses many channels for signaling and bearers, or payloads, data. be able to.
The air interface 202 can be operated according to any of many wireless standards. For example, the standard is a standard based on TDMA, such as Pan-European Digital Mobile Phone System (GSM), General Line Radio Service (GPRS), Extended Data GSM Environment (EDGE), or TIA / EIA-95- CDMA-based standards such as B (IS-95), TIA / EIA-98-C (IS-98), IS2000, HRPD, cdma2000, wideband CDMA (WCDMA), and others can be included.
FIG. 3 is a chart showing synchronization problems in conventional techniques for the transmission of separate RTP streams over wireless communication channels. In the example shown in FIG. 3, video frames and audio data are encoded into RTP streams and then assigned to communication channel packets. FIG. 3 shows a stream of video frame 302. Typically, video frames occur at a constant rate. For example, video frames may occur at a 10 Hz rate, i.e. new frames occur every 100 milliseconds.
As shown in Figure 3, individual video frames may contain different amounts of data, as indicated by the height of the bars that represent each frame. For example, if the video data is encoded as Motion Picture Expert Group (MPEG) data, then the video stream will have intra frames (I frames) and predictive frames (predictive frames) ( P frame). An I-frame is self-contained, that is, it contains all the information needed to draw or display a single complete video frame. P-frames are not self-contained and typically have differential information relative to the previous frame, such as motion vector and differential texture. information) etc. are included. Typically, I-frames may be up to 8-10 times larger than P-frames, depending on the content and encoder settings. Video frames may have different amounts of data, but they still occur at a constant rate. Frames I and P can also be partitioned into multiple video slices. Video slices represent smaller areas of the display screen and can be individually decoded by the decoder.
In FIG. 3, video frames N and N + 4 could represent I frames, and video frames N + 1, N + 2, N + 3, and N + 5 could represent P frames. As shown, the I frame is indicated by the height of the bars that represent the frame and contains a larger amount of data than the P frame. The video frame is then packetized into packets in the RTP stream 304. As shown in FIG. 3, the RTP packets N and N + 4 corresponding to the video I frames N and N + 4 are the video P frames N + 1, N + 2, and N as shown by their width. Greater than RTP packets N + l, N + 2, and N + 3 corresponding to +3.
The video RTP packet is assigned to communication channel packet 306. In traditional communication channels such as CDMA or GSM, the communication channel data packet 306 is transmitted at a constant size and at a constant rate. For example, the communication channel data packet 306 can be transmitted at a rate of 50 Hz, i.e. new data packets are transmitted every 20 milliseconds. Since communication channel packets are of constant size, sending larger RTP packets takes more communication channel packets. Thus, sending an RTP packet corresponding to +4 of I-video frames N and N is because it sends a smaller RTP packet corresponding to P-video frames N + 1, N + 2, and N + 3. Takes more communication channel packets 306 than the communication channel packets required for. In the example shown in FIG. 3, video frame N occupies block 308 of nine communication channel packets 306. Video frames N + 1, N + 2, and N + 3 occupy blocks 310, 312, and 314, respectively, with four communication channel packets 306. Video frame N + 4 occupies block 316 of nine communication channel packets 306.
There is corresponding audio data for each frame of video data. FIG. 2 shows a stream of audio frame 320. Each audio frame N, N + 1, N + 2, N + 3, N + 4, and N + 5 correspond to their respective video frames and occur at a 10Hz rate, ie a new audio frame is 100ms. It starts every time. In general, audio data is less complex than associated video data, so it can be represented by fewer bits, and the size at which RTP packet 322 can be transmitted over a communication channel within a frame period. Is typically encoded. In addition, typical audio frames are generated once every 20 milliseconds in CDMA, GSM, WDCMA, etc. Multiple audio frames are bundled in such cases so that audio and video packets represent the same time period for RTP packetization. For example, RTP packets N, N + 1, N + 2, N + 3, N + 4, and N + 5 allow each RTP packet to be transmitted over a communication channel within a 100 ms frame period. Each RTP packet is the size that can be assigned to the communication channel packet 324.
As shown in FIG. 3, audio frame packets N, N + l, N2, N + 3, N + 4, and N + 5, respectively, occupy blocks 326,328, 330,332, 334, and 336, respectively, respectively. It includes five communication channel packets 324.
A comparison of the allocation of video frames to their respective communication channel packets of audio frames shows the loss of synchronization between audio and video frames. In the example shown in FIG. 3, block 308 of nine communication channel packets 306 is required to transmit video frame N. The audio frame N associated with the video frame N was transmitted in block 326 of the five communication channel packets 324. Since the video and audio in the communication channel packet are transmitted simultaneously, during the transmission of video frame N, audio frame N as well as four of the five communication channel packets in block 328 of audio frame N + 1. Will be sent.
For example, in FIG. 3, for example, if the video and associated audio, frame rate is 10 Hz, and the communication channel packet rate is 50 Hz, then all during the 100 ms period of frame N. Audio data is transmitted, but only part of the video data is transmitted. In this example, all video data in frame N requires 180 ms for transmission compared to 100 ms for full transmission of audio frame N, with another four communication channel packets 306 transmitted. It will not be transmitted until it reaches the complete video frame N. Since the audio and video RTP streams are independent, a portion of the audio frame N + 1 data is transmitted while the video frame N data is transmitted. This loss of synchronization between the video and audio streams can cause a "slip" between the video and audio at the receiver of the communication channel.
Video encoders such as H.263, AVC / H.264, MPEG-4, etc. are essentially due to predictive coding and also due to the use of variable length coding (VLC) for many parameters. Because of the variable rate in effect, real-time delivery of variable rate streams on circuit-switched and packet-switched networks is generally achieved by buffered traffic shaping at the source and receiver. Will be done. Traffic shaping buffers result in additional delay, which is typically undesirable. For example, further delays can be annoying during teleconferencing, when there is a delay between when one person speaks and when another person hears the story.
For example, a delay in the communication channel can cause a pause during playback, since the video on the receiving side of the communication channel is played at the same rate as the original video frame rate. In FIG. 3, video frame N cannot be played until all frames of data have been received. Since the entire frame data is not received during the frame period, playback must be paused until all of the video data for frame N is received. In addition, all data from audio frame N needs to be preserved until all of the video data for frame N has been received so that the audio and video playback can be synchronized. It also means that the audio data from frame N + l received while the video data from frame N is still being received must be stored until all of the video data from frame N + 1 has been received. It will also be noticed. Due to the variable size of the video frame, a large traffic shaping buffer is required to achieve synchronization.
FIG. 4 is a chart showing a technique for transmitting separate RTP streams over a wireless communication channel according to the present invention. FIG. 4 shows a stream of variable-sized video frames 302 and a stream of audio frames 320, respectively, encoded into independent RTP streams 304 and 322, as in FIG. Video and audio frames occur at a constant rate, such as a 10Hz rate.
In FIG. 4, as in FIG. 3, video frames N and N + 4 could represent I frames, and video frames N + 1, N + 2, N + 3, and N. +5 could represent a P-frame. Video frames are packetized into packets in RTP stream 304. As shown in FIG. 4, the RTP packets N and N + 4 corresponding to the video I frames N and N + 4 are the video P frames N + 1, N + 2, and N as shown by their width. Larger than RTP packets N + 1, N + 2, and N + 3 corresponding to +3.
The video RTP packet is assigned to communication channel packet 406. The capacity of the communication channel is variable using techniques such as those described in the co-pending applications listed above (see Pending Related Patent Applications). Due to the variable capacity of communication channel packet 406, video frame N can be transmitted in block 408 containing 5 communication channel packets 406.
Traditional communication channels such as CDMA based standards, such as TIA / EIA-95-B (IS-95), TIA / EIA-98-C (IS-98), IS2000, HRPD, cdma2000, and broadband CDMA. In (WCDMA) and the like, the communication channel data packet 406 can be transmitted at a rate of 50 Hz, i.e., new data packets are transmitted every 20 milliseconds. Since the capacity of the communication channel packet 406 can be changed, the coding of the video frame N can be constrained so that the entire video frame N can be transmitted during the frame period. As shown in FIG. 4, when transmitting the RTP packet N corresponding to the video frame N, the capacity of the communication channel packet 406 is increased so that the entire packet can be transmitted during the frame period. The techniques described can also be applied to communication channels based on GSM, GPRS, or EDGE.
As shown in FIG. 4, video frames N, N + 1, N + 2, N + 3, N + 4, and N + 5 are encoded in RTP packets, and 408,410, in the communication channel block, Assigned to 412, 414, 416, and 418, respectively. It is also noted that by changing the communication channel capacity, the entire video frame is transmitted within the frame period. For example, if the video frame rate is 10 Hz, then all frames of the video data will be transmitted within a frame period of 100 milliseconds.
For each frame of video data 302, there is a corresponding audio frame 320. Each audio frame N, N + 1, N + 2, N + 3, N + 4, and N + 5 correspond to their respective video frames and occur at a rate of 10Hz, ie a new audio frame is 100ms. Starts every time. As described in connection with FIG. 3, audio data can be represented by fewer bits than the associated video data, as it is generally less complex, and typically RTP packet 322 is of a frame. Encoded to be sized to be transmitted over a communication channel within a 100 millisecond period. That is, audio RTP packets N, N + 1, N + 2, N + 3, N + 4, and N + 5 are such that each RTP packet is assigned to blocks 326,328, 330,332, 334, and 336 of the communication channel packet, respectively. It is a size that can be used. Thus, if the video frame rate is 10 Hz, then each video frame can be transmitted over the communication channel within a 100 ms frame period. Like video, if the audio packet size is large, the communication channel capacity can also be modified to support the transmission of the entire audio frame during the frame period.
In FIG. 4, a comparison of the allocation of video and audio frames to their respective communication channel packets shows that the video and audio frames remain in sync. In other words, every frame period, the entire video and audio frame are transmitted. The entire video and audio frame is transmitted during each frame period, so no further buffering is needed. The received video and audio data only needs to be accumulated during the frame period, after which it can be played out. Video and audio frames remain in sync because there is no delay brought in by the communication channel.
As shown in Figure 3, video frames N + 1, N + 2, and N + 3 only requested four video communication channel packets 306 to transmit the entire frame of video data. It is noted that. As shown in FIG. 4, the video communication channel packet 406 can be reduced in size or a blank packet can be transmitted so that the video data fits into 5 packets. .. Similarly, if there is excess capacity available in the audio communication channel, blank packets can be sent. In this way, video and audio data is a communication channel packet in which all frames of audio and video data occupy the same or less frame rate (that occupy the same period, or less, or the respective frame rate). It is encoded so that it can be assigned to.
Depending on the aspects of the communication network, different techniques can be used to synchronize the RTP streams, as described below. For example, the communication network may be over provisioned, i.e. it may have excessive capacity, or the communication network may be guaranteed quality of service. May have Service). In addition, RTP streams may be modified to stay in sync when transmitted over a communication network. Each of these techniques is described below.
Overprepared communication network When the communication link between PDSN 208 and caller 206 is over-prepared, i.e. there is excess capacity available for transmitting data over the wireline internet, in a scenario where it is congested. There is no delay due. Transmission can be accommodated by the communication link because there is excess capacity in the communication link and there is no need to delay the transmission. There is no "time slip" between audio and video packets when they arrive at the infrastructure, such as PDSN, without transmission delays. That is, as described in the present invention, the audio and video data remain synchronized with each other up to the PDSN, and synchronization is maintained between the PDSN and MS.
Audio-visual synchronization is easily achieved in over-prepared scenarios. For example, video data may have a frame rate of 10 frames per second (fps) based on 100 ms frames, and related audio may have 50 fps frames based on 20 ms speech frames. May have a rate. In this example, five frames of received audio data will be buffered, so it will be synchronized to the video frame rate. That is, five frames of audio data will be buffered corresponding to 100 ms of audio data, so it will be synchronized with 100 ms video frames.
Communication network with guaranteed QoS for maximum latency It is possible to adapt to lower frame rate video frames by buffering an appropriate number of higher frame rate speech frames. In general, if a video packet is distributed with a quality of service (QoS) delay guarantee: QoS_delay = nTms However, n is the delay in frames, and T = 1000 / frames_per_second Second, a buffer sized to store nT / w speech frames is required to store enough speech frames to ensure that the speech and video are in sync, where w Is the duration of the speech frame in milliseconds. In cdma2000UMTS, the speech frame duration, w, is 20 ms, and in other communication channels, the speech frame duration is different or can vary.
Another technique for synchronizing audio and video data involves buffering both data streams. For example, if the communication system is D<sub>Q</sub>If you have a guaranteed maximum delay of milliseconds, D<sub>Q</sub>Means the maximum delay that can be experienced during the transmission of audio and video streams, where a properly sized buffer can be used to maintain synchronization.
For example, D<sub>Q</sub>With the guaranteed maximum delay of, then D<sub>Q /</sub>T video frame (T is the duration of the video frame in milliseconds) and D<sub>Q /</sub>Buffering w speech frames (where w is the duration of the speech frame in milliseconds) will guarantee audio-video synchronization (AV-synchronization). These additional buffer spaces are commonly referred to as a de-jitter buffer.
The technique has described synchronization of audio and video data streams. The technology can be used with any data stream that needs to be synchronized. If there are two data streams, a first higher bit rate data stream and a second lower bit rate data stream that have the same information interval and need to be synchronized, then Buffering higher bitrate data allows it to be synchronized with lower bitrate data. The size of the buffer can be determined depending on the QoS as described above. Similarly, higher bitrate data streams and lower bitrate data streams can be buffered and synchronized as described above.
The technique described includes a first decoder configured to receive a first encoded data stream and output a decoded first data stream. It can be performed by a data stream synchronizer, where the first encoded data stream is during the information interval (during an information). interval), has the first bit rate. The second decoder is then configured to receive the second encoded data stream and output the decoded second data stream, where the second encoded data stream is second during the information interval. Has a bit rate. The synchronized data stream is also a first buffer configured to accumulate the decoded first data stream for at least one information interval and output a frame of the decoded first data stream every interval period. And then, for at least one information interval, it includes a second buffer configured to store the decoded second data stream and output frames of the decoded second data stream every interval period. In addition, the combiner should receive the frames of the first decoded data stream and the frames of the second decoded data stream and output the synchronized frames of the first and second decoded data streams. It is composed of. In one example, the first encoded data stream may be video data and the second encoded data stream may be audio data, but the first bit rate is higher than the second bit rate. It's like.
Single RTP stream with multiplexed audio and video Another embodiment transports audio and video within a single RTP stream. As noted, transmitting audio and video as a single RTP stream is not common in IP networks. RTP was designed to allow participants with different resources, such as terminals capable of both video and audio, and terminals capable of audio only, to communicate in the same multimedia conference. ..
The restrictions of sending audio and video as separate RTP streams may not be applicable in wireless networks for video services. In this case, the new RTP profile can be designed to carry specific speech and video codec payloads. Coupling audio and video into a common RTP stream eliminates any time slip between audio and video data without the need for an overprepared communication network. Thus, audio-video synchronization can be achieved using the techniques described in relation to over-prepared networks, as described above.
FIG. 5 is a block diagram of a portion of the wireless audio / video receiver 500 configured to receive communication channel packets. As shown in FIG. 5, the audio / video receiver 500 includes a communication channel interface 502 configured to receive communication channel packets. The communication channel interface 502 outputs the video communication channel packet to the video decoder 504 and the audio communication channel packet to the audio decoder 506. The video decoder 504 decodes the video communication channel packet and outputs the video data to the video buffer 508. The audio decoder 506 decodes the audio communication channel packet and outputs the audio data to the audio buffer 510. The video buffer 508 and the audio buffer store video and audio data, respectively, during the frame period. The video buffer 508 and the audio buffer 510 output video frames and audio frames to the combiner 512, respectively. The combiner 512 is configured to combine video and autoframes and output synchronized audio and video signals. The operation of video buffer 508, audio buffer 510, and combiner 512 may be controlled by controller 514.
FIG. 6 is a block diagram of a portion of a wireless audio / video transmitter 600 configured to transmit communication channel packets. As shown in FIG. 6, the audio / video transmitter 600 includes a video communication channel interface 602 configured to receive a video data RTP stream. The video communication channel interface allocates RTP packets to communication channel packets. As noted, the capacity of the communication channel packet is an entire frames worth of RTP video for the communication channel packet that occupies the same period as the video frame. can change to assign data). The audio / video transmitter 600 also includes an audio communication channel interface 604 configured to receive an audio data RTP stream. Audio communication channel interface 604 allocates RTP packets to communication channel packets. As noted, the capacity of a communication channel packet will generally be sufficient to allocate an entire frame of RTP audio data to a communication channel packet that occupies the same period as the audio frame. If the channel capacity is not sufficient, then it is similar to the video communication channel packet so that the communication channel packet occupying the same period as the audio frame has enough capacity to allocate the entire frame of RTP audio data. Can be changed.
Video and audio communication channel packets are output by video and audio communication channel interfaces 602 and 604, respectively, and propagated to combiner 606. The combiner 606 is configured to accept and combine video and audio communication channel packets and output a composite signal. The output of combiner 606 is transmitted to transmitter 608, which transmits its composite signal to the radio channel. The operation of the video communication channel interface 602, the audio communication channel interface 604, and the combiner 606 may be controlled by the controller 614.
FIG. 7 is a flow chart of transmission of an independent RTP stream over a wireless communication link. The flow starts at block block 702, where video and audio RTP data streams are received. The flow then proceeds to block 704, where the video RTP stream is assigned to the communication channel packet. At block 706, the audio RTP stream is assigned to the communication channel packet. At block 708, video and audio communication channel packets are combined and transmitted over the radio channel.
FIG. 8 is a flowchart of received audio and video data on the wireless communication channel. The flow begins at block 802, where video and audio data are received over the wireless communication channel. The flow proceeds to block 804, where video and audio data are decoded. At block 806, the decoded video and audio data is assembled into the respective video and audio frames. At block 810, video and audio data are combined into synchronized video / audio frames. Block 810 outputs synchronized video / audio frames.
FIG. 9 is a block diagram of a wireless communication device or mobile station (MS) configured according to an exemplary embodiment of the present invention. The communication device 902 includes a network interface 906, a codec 908, a host processor 910, a memory device 912, a program product 914, and a user interface 916.
Signals from the infrastructure are received by network interface 906 and sent to host processor 910. The host processor 910 receives the signal and responds with an appropriate action according to the content of the signal. For example, the host processor 910 may decode the received signal itself, or it may route the received signal to codec 908 for decoding. In another embodiment, the received signal is sent directly from network interface 906 to codec 908.
In one embodiment, network interface 906 may be transceivers and antennas that interface to the infrastructure on wireless channels. In another embodiment, the network interface 906 may be a network interface card used to interface the infrastructure over a terrestrial line. Codec 908 may be implemented as a digital signal processor (DSP), or a general purpose processor such as a central processing unit (CPU).
Both host processor 910 and codec 908 are connected to memory device 912. The memory device 812 can be used not only to store the program code that will be executed by the host processor 910 or DSP9, but also to store the data while the WCD is in operation. For example, the host processor, codec, or both can operate under the control of programming instructions that are temporarily stored in memory device 912. The host processor 910 and codec 908 can also include their own program storage memory. When programming instructions are performed, the host processor 910 and / or codec 908 have their functions, such as audio / video. It encodes or decodes multimedia streams such as data, and assembles audio and video frames. In this way, the programming step implements the functionality of the respective host processor 910 and codec 908, so that the host processor and codec, respectively, have the ability to decode or encode the content stream and assemble the frames as desired. Can be made to run. Programming steps can be received from program product 914. Program product 914 can store programming steps and transfer them to memory 912 for execution by the host processor, codec, or both.
The program product 914 may be a semiconductor memory chip such as a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, in addition to a storage device such as a hard disk, a removable disk, or a CD-ROM. Alternatively, it may be any form of storage medium known in the art capable of storing computer-readable instructions. In addition, program product 914 may be a source file containing program steps that are received from the network, stored in memory, and then executed. As described above, the processing steps required for the operation according to the present invention can be embodied on the program product 914. In FIG. 9, an exemplary storage medium is shown coupled to the host processor 910 so that the host processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the host processor 910.
User interface 916 is connected to both host processor 910 and codec 908. For example, the user interface 916 can include a display and speakers used to output multimedia data to the user.
Those skilled in the art will recognize that the steps of the method described in connection with embodiments can be interchanged without departing from the scope of the invention.
Those skilled in the art will also appreciate that information and signals can be represented using any of a variety of different techniques and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be voltage, current, electromagnetic waves, magnetic fields or or magnetic particles, light fields or optical particles, or theirs. It can be represented by any combination.
Those skilled in the art will further implement the various descriptive logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein as electronic hardware, computer software, or a combination of both. You will recognize that it can be done. To articulate this compatibility of hardware and software, various exemplary components, blocks, modules, circuits, and steps are generally described above in terms of their functionality. Whether their functionality is implemented as hardware or software depends on the specific application and design constraints imposed on the entire system. Skilled skilled craftsmen may implement functionality described in different ways for each particular application, but decisions on such implementation deviate from the scope of the invention. Should not be interpreted.
The various explanatory logic blocks, modules, and circuits described in connection with the embodiments disclosed herein include general purpose processors, digital signal processors (DSPs), application specific ICs (ASICs), and field programmable gates. It can be implemented or performed using an array (FPGA) or other programmable logic device, discrete gate or transistor logic, a discrete hardware component, or any combination of these designed to perform the functions described herein. .. The general purpose processor may be a microprocessor, but otherwise the processor may be any conventional processor, controller, microcontroller, or state machine. Processors are also implemented as a combination of arithmetic units, such as a combination of DSP and microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. May be done.
The steps of the method or algorithm described in connection with the embodiments disclosed herein can be embodied in software modules executed by a processor directly in hardware or in combination of the two. Software modules are present in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or any other form of technically known storage medium. sell. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write the information to the storage medium. Alternatively, the storage medium may be built into the processor. The processor and storage medium may reside in the ASIC. The ASIC may exist in the user terminal. Alternatively, the processor and storage medium may be present as discrete components within the user terminal.
The above description of the disclosed embodiments is provided to allow anyone of ordinary skill in the art to make or use the present invention. Various variations of these embodiments will be readily apparent to those of skill in the art, and the comprehensive principles defined herein apply to other embodiments without departing from the spirit or scope of the invention. Can be done. Therefore, the present invention is not intended to be limited to the embodiments presented herein, but should be provided with the broadest scope consistent with the principles and novel features disclosed herein.
104 members in 14 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 57167304 | United States of America | P | |
| 57167304 | United States of America | P | |
| 60571673 | United States of America | – | |
| 2004571673 | – | – | – |
| US20040571673P | – | – | – |
Members104
| Document | Office | Kind | |
|---|---|---|---|
| US2005259613A1 | United States of America | A1 | |
| US2005259623A1 | United States of America | A1 | |
| US2005259690A1 | United States of America | A1 | |
| US2005259694A1 | United States of America | A1 | |
| CA2565977A1 | Canada | A1 | |
| CA2566124A1 | Canada | A1 | |
| CA2566125A1 | Canada | A1 | |
| CA2566126A1 | Canada | A1 | |
| CA2771943A1 | Canada | A1 | |
| CA2811040A1 | Canada | A1 | |
| WO2005114919A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2005114943A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2005114950A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2005115009A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2005114943A3 | World Intellectual Property Organization (WIPO) | A3 | |
| TW200618544A | Taiwan Province of China | A | |
| TW200618564A | Taiwan Province of China | A | |
| TW200623737A | Taiwan Province of China | A | |
| KR20070013330A | Republic of Korea | A | |
| KR20070014200A | Republic of Korea | A | |
| KR20070014201A | Republic of Korea | A | |
| EP1751955A1 | European Patent Office (EPO) | A1 | |
| EP1751956A2 | European Patent Office (EPO) | A2 | |
| EP1751987A1 | European Patent Office (EPO) | A1 | |
| MXPA06013186A | Mexico | A | |
| MXPA06013193A | Mexico | A | |
| EP1757027A1 | European Patent Office (EPO) | A1 | |
| KR20070023731A | Republic of Korea | A | |
| MXPA06013210A | Mexico | A | |
| MXPA06013211A | Mexico | A | |
| CN1969562A | China | A | |
| CN1973515A | China | A | |
| CN1977516A | China | A | |
| CN1985477A | China | A | |
| BRPI0510952A | Brazil | A | |
| BRPI0510953A | Brazil | A | |
| BRPI0510961A | Brazil | A | |
| BRPI0510962A | Brazil | A | |
| JP2007537681A | Japan | A | |
| JP2007537682A | Japan | A | |
| JP2007537683A | Japan | A | |
| JP2007537684A | Japan | A | |
| KR20080084866A | Republic of Korea | A | |
| KR100870215B1 | Republic of Korea | B1 | |
| KR100871305B1 | Republic of Korea | B1 | |
| EP1757027B1 | European Patent Office (EPO) | B1 | |
| ATE417436T1 | Austria | T1 | |
| DE602005011611D1 | Germany | D1 | |
| EP1751955B1 | European Patent Office (EPO) | B1 | |
| ATE426988T1 | Austria | T1 | |
| KR20090039809A | Republic of Korea | A | |
| ES2318495T3 | Spain | T3 | |
| DE602005013517D1 | Germany | D1 | |
| ES2323011T3 | Spain | T3 | |
| KR100906586B1 | Republic of Korea | B1 | |
| KR100918596B1 | Republic of Korea | B1 | |
| MY139431A | Malaysia | A | |
| JP4361585B2 | Japan | B2 | |
| JP4448171B2 | Japan | B2 | |
| MY141497A | Malaysia | A | |
| EP2182734A1 | European Patent Office (EPO) | A1 | |
| EP2214412A2 | European Patent Office (EPO) | A2 | |
| JP4554680B2 | Japan | B2 | |
| EP1751987B1 | European Patent Office (EPO) | B1 | |
| ATE484157T1 | Austria | T1 | |
| MY142161A | Malaysia | A | |
| DE602005023983D1 | Germany | D1 | |
| CN1977516B | China | B | |
| EP2262304A1 | European Patent Office (EPO) | A1 | |
| ES2354079T3 | Spain | T3 | |
| EP1751956B1 | European Patent Office (EPO) | B1 | |
| ATE508567T1 | Austria | T1 | |
| DE602005027837D1 | Germany | D1 | |
| KR101049701B1 | Republic of Korea | B1 | |
| JP2011142616AThis record | Japan | A | |
| CN1969562B | China | B | |
| KR101068055B1 | Republic of Korea | B1 | |
| ES2366192T3 | Spain | T3 | |
| TWI353759B | Taiwan Province of China | B | |
| TW201145943A | Taiwan Province of China | A | |
| US8089948B2 | United States of America | B2 | |
| CA2566125C | Canada | C | |
| EP2262304B1 | European Patent Office (EPO) | B1 | |
| CN1985477B | China | B | |
| EP2214412A3 | European Patent Office (EPO) | A3 | |
| TWI381681B | Taiwan Province of China | B | |
| CN1973515B | China | B | |
| CN102984133A | China | A | |
| TWI394407B | Taiwan Province of China | B | |
| EP2592836A1 | European Patent Office (EPO) | A1 | |
| CA2565977C | Canada | C | |
| JP5356360B2 | Japan | B2 | |
| EP2182734B1 | European Patent Office (EPO) | B1 | |
| CA2566124C | Canada | C | |
| US8855059B2 | United States of America | B2 | |
| US2014362740A1 | United States of America | A1 | |
| US2015016427A1 | United States of America | A1 | |
| CA2771943C | Canada | C | |
| CN102984133B | China | B | |
| US9674732B2 | United States of America | B2 |
13 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 |
Numbers
- Publication
- 2011142616
- Publication, DOCDB
- 2011142616
- Publication, EPODOC
- JP2011142616
- Application
- 265208
- Application, DOCDB
- 2010265208
- Application, EPODOC
- JP20100265208
Titles2
- Japanese
- 無線通信システムにおけるオーディオおよびビデオデータの同期
- English
- Synchronization of audio and video data in wireless communication systems
Classification
- CPC, 36
- H04L69/04
- H04W28/06
- H04N21/2381
- H04N21/41407
- H04N21/44004
- H04N21/4788
- H04N21/6131
- H04N21/6181
- H04N21/6437
- H04N21/64707
- H04W28/065
- H04W72/1263
- H04W80/00
- H04W84/04
- H04W88/181
- H04L65/80
- H04L69/166
- H04L69/22
- H04L69/161
- H04N19/102
- H04N19/115
- H04N19/61
- H04N19/124
- H04N19/152
- H04N19/164
- H04N19/174
- H04L69/321
- H04L47/36
- H04W4/06
- H04L65/764
- H04L65/00
- H04L9/40
- H04L65/75
- H04L65/1101
- H04W72/044
- H04W88/02
- IPC, 13
- H04N7 173
- H04B1 7073
- H04B7 00
- H04B7 216
- H04L12 28
- H04L12 56
- H04L12 66
- H04L47 36
- H04N7 26
- H04W28 06
- H04W72 12
- H04W84 04
- H04W88 18