Synchronization of audio and video data in a wirel
Abstract
The present invention discloses a technique for encoding an audio video stream transmitted over a network (e.g., a wireless network or an IP network) such that a complete audio frame and a complete video message are simultaneously transmitted in a desired cycle. a frame for reproducing the audio video stream frames by an application in a receiver. Aspects of such techniques include receiving audio and video RTP streams, and assigning a complete frame of RTP video data to a communication channel packet occupying the same period as the video frame rate or less than the video frame rate. . A complete frame of RTP audio data is also assigned to a communication channel packet occupying the same period as the audio frame rate or less than the audio frame rate. These video and audio communication channel packets are simultaneously transmitted. Receiving and assigning RTP streams can be performed at a remote station or a base station.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
20 claims: 12 independent, 8 dependent
- 1A data stream synchronizer, comprising:a first decoder configured to receive a first encoded data stream and output a first decoded data stream, wherein the first encoded data stream is an information The interval has a first bit rate;a second decoder configured to receive a second coded data stream and output a second decoded data stream, wherein the second coded data stream is in the information The interval period has a second bit rate, and the second coded data stream is audio data;a first buffer is configured to accumulate the first decoded data stream of at least one information interval, and the Output a frame of the first decoded data stream during each interval;a second buffer configured to accumulate the second decoded data stream of at least one information interval, and output the second decoded data stream during each interval A frame of the second decoded data stream;and a combiner configured to receive the first decoded data stream frame and the second decoded data stream frame, and output the first and second decodes One of the data streams is a synchronized frame. 一種資料串流同步器,其包含:一第一解碼器,其經組態以接收一第一編碼資料串流並輸出一第一解碼資料串流,其中該第一編碼資料串流在一資訊間隔期間具有一第一位元速率;一第二解碼器,其經組態以接收一第二編碼資料串流並輸出一第二解碼資料串流,其中該第二編碼資料串流在該資訊間隔期間具有一第二位元速率,且其中該第二編碼資料串流為音頻資料;一第一緩衝器,其經組態以累積至少一資訊間隔之該第一解碼資料串流,並在每一間隔期間輸出該第一解碼資料串流之一訊框;一第二緩衝器,其經組態以累積至少一資訊間隔之該第二解碼資料串流,並在每一間隔期間輸出該第二解碼資料串流之一訊框;及一組合器,其經組態以接收該第一解碼資料串流訊框及該第二解碼資料串流訊框,並輸出第一及第二解碼資料串流之一同步訊框。
- 4A remote site equipment, which includes:a video decoder, which is configured to receive encoded video data and output Decode video data;an audio decoder, which is configured to receive encoded audio data and output decoded audio data;a video buffer, which is configured to accumulate decoded video data for at least one frame period, and in each signal Frame period output a video data frame;an audio buffer configured to accumulate decoded audio data of multiple frame periods, and output an audio data frame in each frame period;and a combiner, which It is configured to receive the video data frame and the audio data frame, and output a synchronous frame of audio and video data. 一種遠端站台設備,其包含:一視頻解碼器,其經組態以接收編碼視頻資料並輸出 解碼視頻資料;一音頻解碼器,其經組態以接收編碼音頻資料並輸出解碼音頻資料;一視頻緩衝器,其經組態以累積至少一訊框週期之解碼視頻資料,並在每一訊框週期輸出一視頻資料訊框;一音頻緩衝器,其經組態以累積多個訊框週期之解碼音頻資料,並在每一訊框週期輸出一音頻資料訊框;及一組合器,其經組態以接收該視頻資料訊框及該音頻資料訊框,並輸出音頻視頻資料之一同步訊框。
- 8A base station equipment comprising:a video decoder configured to receive encoded video data and output decoded video data;an audio decoder configured to receive encoded audio data and output decoded audio data;a video Buffer, which is configured to accumulate decoded video data of a video frame period, and output a video data frame in each frame period;an audio buffer, which is configured to accumulate a video frame period of untie Code audio data, and output an audio data frame in each frame period;and a combiner configured to receive the video data frame and the audio data frame, and output a synchronization signal of the audio and video data frame. 一種基地台設備,其包含:一視頻解碼器,其經組態以接收編碼視頻資料並輸出解碼視頻資料;一音頻解碼器,其經組態以接收編碼音頻資料並輸出解碼音頻資料;一視頻緩衝器,其經組態以累積一視頻訊框週期之解碼視頻資料,並在每一訊框週期輸出一視頻資料訊框;一音頻緩衝器,其經組態以累積一音頻訊框週期之解 碼音頻資料,並在每一訊框週期輸出一音頻資料訊框;及一組合器,其經組態以接收該視頻資料訊框及該音頻資料訊框,並輸出音頻視頻資料之一同步訊框。
- 12A wireless communication system, which includes:a base station device, which includes: a video communication channel interface, which is configured to receive a video real-time transmission protocol (RTP) stream and assign a complete frame of RTP video data Pre-occupied communication channel packets of the same or less period than the video frame rate;an audio communication channel interface, which is configured to receive an audio RTP stream, and one of the RTP audio data The complete frame is assigned to communication channel packets occupying the same or less period than the audio frame rate;a transmitter configured to receive and transmit the video and audio communication channel packets;A remote site equipment, which includes: a video decoder configured to receive video communication channel packets And output decoded video data;an audio decoder, which is configured to receive audio communication channel packets and output decoded audio data;a video buffer, which is configured to accumulate decoded video data for a video frame period, and Output a video data frame for each frame period;an audio buffer configured to accumulate decoded audio data for an audio frame period, and output an audio data frame for each frame period;and a combination The device is configured to receive the video data frame and the audio data frame, and output a synchronous frame of audio and video data. 一種無線通信系統,其包含:一基地台設備,其包含:一視頻通信頻道介面,其經組態以接收一視頻即時傳送協定(RTP)串流,並將RTP視頻資料之一完整訊框指派予佔用與該視頻訊框速率相同或比該視頻訊框速率更少之週期的通信頻道封包;一音頻通信頻道介面,其經組態以接收一音頻RTP串流,並將RTP音頻資料之一完整訊框指派予佔用與該音頻訊框速率相同或比該音頻訊框速率更少之週期的通信頻道封包;一傳輸器,其經組態以接收並傳輸該等視頻及音頻通信頻道封包;一遠端站台設備,其包含:一視頻解碼器,其經組態以接收視頻通信頻道封包 並輸出解碼視頻資料;一音頻解碼器,其經組態以接收音頻通信頻道封包並輸出解碼音頻資料;一視頻緩衝器,其經組態以累積一視頻訊框週期之解碼視頻資料,並在每一訊框週期輸出一視頻資料訊框;一音頻緩衝器,其經組態以累積一音頻訊框週期之解碼音頻資料,並在每一訊框週期輸出一音頻資料訊框;及一組合器,其經組態以接收該視頻資料訊框及該音頻資料訊框,並輸出一音頻視頻資料之同步訊框。
- 13A wireless communication system includes:a remote station device, including: a video communication channel interface, which is configured to receive a video RTP stream, and assign a complete frame of RTP video data to the A communication channel packet with the same video frame rate or a period less than the video frame rate;an audio communication channel interface, which is configured to receive an audio RTP stream and assign a complete frame of RTP audio data Pre-occupied communication channel packets of the same or less period than the audio frame rate;a transmitter configured to receive and transmit the video and audio communication channel packets;a base station equipment , Which contains: A video decoder, which is configured to receive video communication channel packets and output decoded video data;an audio decoder, which is configured to receive audio communication channel packets and output decoded audio data;a video buffer, which is configured to receive audio communication channel packets and output decoded audio data;Mode to accumulate the decoded video data of a video frame period, and output a video data frame in each frame period;an audio buffer, which is configured to accumulate the decoded audio data of an audio frame period, and Each frame period outputs an audio data frame;and a combiner configured to receive the video data frame and the audio data frame, and output a synchronization frame of audio and video data. 一種無線通信系統,其包含:一遠端站台設備,其包含:一視頻通信頻道介面,其經組態以接收一視頻RTP串流,並將RTP視頻資料之一完整訊框指派予佔用與該視頻訊框速率相同或比該視頻訊框速率更少之週期的通信頻道封包;一音頻通信頻道介面,其經組態以接收一音頻RTP串流,並將RTP音頻資料之一完整訊框指派予佔用與該音頻訊框速率相同或比該音頻訊框速率更少之週期的通信頻道封包;一傳輸器,其經組態以接收並傳輸該等視頻及音頻通信頻道封包;一基地台設備,其包含: 一視頻解碼器,其經組態以接收視頻通信頻道封包並輸出解碼視頻資料;一音頻解碼器,其經組態以接收音頻通信頻道封包並輸出解碼音頻資料;一視頻緩衝器,其經組態以累積一視頻訊框週期之解碼視頻資料,並在每一訊框週期輸出一視頻資料訊框;一音頻緩衝器,其經組態以累積一音頻訊框週期之解碼音頻資料,並在每一訊框週期輸出一音頻資料訊框;及一組合器,其經組態以接收該視頻資料訊框及該音頻資料訊框,並輸出一音頻視頻資料之同步訊框。
- 14A method for decoding a synchronous data stream, comprising:receiving a first coded data stream, decoding the first coded data and outputting a first decoded data stream, wherein the first coded data stream is an information The interval has a first bit rate;receiving a second coded data stream, decoding the second coded data and outputting a second decoded data stream, wherein the second coded data stream has a A second bit rate, wherein the second encoded data stream is audio data;the first decoded data stream of at least one information interval is accumulated, and one of the first decoded data streams is output in each interval period Frame;accumulate the second decoded data stream of at least one information interval, and output a frame of the second decoded data stream in each interval period;and Combine the first decoded data stream frame and the second decoded data stream frame, and output a synchronization frame of one of the first and second decoded data streams. 一種用於解碼同步資料串流之方法,其包含:接收一第一編碼資料串流、解碼該第一編碼資料並輸出一第一解碼資料串流,其中該第一編碼資料串流在一資訊間隔期間具有一第一位元速率;接收一第二編碼資料串流、解碼該第二編碼資料並輸出一第二解碼資料串流,其中該第二編碼資料串流在該資訊間隔期間具有一第二位元速率,且其中該第二編碼資料串流為音頻資料;累積至少一資訊間隔之該第一解碼資料串流,並在每一間隔週期輸出該第一解碼資料串流之一訊框;累積至少一資訊間隔之該第二解碼資料串流,並在每一間隔週期輸出該第二解碼資料串流之一訊框;及 組合該第一解碼資料串流訊框及該第二解碼資料串流訊框,並輸出第一及第二解碼資料串流之一同步訊框。
- 15A method for decoding and synchronizing audio and video data. The method includes:receiving encoded video data and outputting decoded video data;receiving encoded audio data and outputting decoded audio data;accumulating decoded video data for at least one video frame period, and Output a video data frame in each frame period;accumulate decoded audio data for at least one audio frame period, and output an audio data frame in each frame period;and combine the video data frame and the audio data Frame, and output a synchronous frame of audio and video data in each video frame period. 一種用於解碼並同步音頻及視頻資料之方法,該方法包含:接收編碼視頻資料並輸出解碼視頻資料;接收編碼音頻資料並輸出解碼音頻資料;累積至少一視頻訊框週期之解碼視頻資料,並在每一訊框週期輸出一視頻資料訊框;累積至少一音頻訊框週期之解碼音頻資料,並在每一訊框週期輸出一音頻資料訊框;及組合該視頻資料訊框及該音頻資料訊框,並在每一視頻訊框週期輸出音頻視頻資料之一同步訊框。
- 16A computer-readable medium implementing a method for decoding and synchronizing data streams, the method comprising:receiving a first encoded data stream, decoding the first encoded data, and outputting a decoded first data stream, wherein The first encoded data stream has a first bit rate during an information interval;receiving a second encoded data stream, decoding the second encoded data and outputting a decoded second data stream, wherein the second encoding The data stream has a second bit rate during the information interval, and the second encoded data stream is audio data;the first decoded data stream for at least one information interval is accumulated, and output in each interval period A frame of the first decoded data stream;the second decoded data stream of at least one information interval is accumulated, and each Output a frame of the second decoded data stream at an interval period;and combine the first decoded data stream frame and the second decoded data stream frame, and output the first and second decoded data streams A sync frame. 一種實施一用於解碼及同步資料串流之方法的電腦可讀取媒體,該方法包含:接收一第一編碼資料串流、解碼該第一編碼資料並輸出一解碼第一資料串流,其中該第一編碼資料串流在一資訊間隔期間具有一第一位元速率;接收一第二編碼資料串流、解碼該第二編碼資料並輸出一解碼第二資料串流,其中該第二編碼資料串流在該資訊間隔期間具有一第二位元速率,且其中該第二編碼資料串流為音頻資料;累積至少一資訊間隔之該第一解碼資料串流,並在每一間隔週期輸出該第一解碼資料串流之一訊框;累積至少一資訊間隔之該第二解碼資料串流,並在每 一間隔週期輸出該第二解碼資料串流之一訊框;及組合該第一解碼資料串流訊框及該第二解碼資料串流訊框,並輸出第一及第二解碼資料串流之一同步訊框。
- 17A computer-readable medium that implements a method for decoding and synchronizing audio and video data. The method includes:receiving encoded video data and outputting decoded video data;receiving encoded audio data and outputting decoded audio data;accumulating a video frame Periodically decode video data, and output a video data frame in each frame period;accumulate decoded audio data for an audio frame period, and output an audio data frame in each frame period;and combine the video data The frame and the audio data frame, and output a synchronous frame of audio and video data. 一種實施一用於解碼及同步音頻及視頻資料之方法的電腦可讀取媒體,該方法包含:接收編碼視頻資料並輸出解碼視頻資料;接收編碼音頻資料並輸出解碼音頻資料;累積一視頻訊框週期之解碼視頻資料,並在每一訊框週期輸出一視頻資料訊框;累積一音頻訊框週期之解碼音頻資料,並在每一訊框週期輸出一音頻資料訊框;及組合該視頻資料訊框及該音頻資料訊框,並輸出音頻視頻資料之一同步訊框。
- 18A data stream synchronizer, comprising:a component for decoding a first encoded data stream and outputting a first decoded data stream, wherein the first encoded data stream has a first during an information interval Bit rate;a component used to decode a second coded data stream and output a second decoded data stream, wherein the second coded data stream has a second bit rate during the information interval, and wherein The second encoded data stream is audio data;a component for accumulating the first decoded data stream of at least one information interval, and outputting a frame of the first decoded data stream in each interval period;A member for accumulating the second decoded data stream of at least one information interval, and outputting a frame of the second decoded data stream in each interval period;and for combining the first decoded data stream frame And the second decoded data stream frame, and output a synchronization frame of the first and second decoded data streams. 一種資料串流同步器,其包含:用於解碼一第一編碼資料串流,並輸出一第一解碼資料串流之構件,其中該第一編碼資料串流在一資訊間隔期間具有一第一位元速率;用於解碼一第二編碼資料串流,並輸出一第二解碼資料串流之構件,其中該第二編碼資料串流在該資訊間隔期間具有一第二位元速率,且其中該第二編碼資料串流為音頻資料;用於累積至少一資訊間隔之該第一解碼資料串流,並在每一間隔週期輸出該第一解碼資料串流之一訊框的構件; 用於累積至少一資訊間隔之該第二解碼資料串流,並在每一間隔週期輸出該第二解碼資料串流之一訊框的構件;及用於組合該第一解碼資料串流訊框及該第二解碼資料串流訊框,並輸出一第一及第二解碼資料串流之同步訊框的構件。
- 19A remote station equipment comprising:a component for receiving encoded video data and outputting decoded video data;a component for receiving encoded audio data and outputting decoded audio data;a component for accumulating decoded video data for a video frame period, And a component for outputting a video data frame in each frame period;a component for accumulating decoded audio data for an audio frame period and outputting an audio data frame for each frame period;and for combining the The video data frame and the audio data frame, and a component that outputs a synchronous frame of audio and video data. 一種遠端站台設備,其包含:用於接收編碼視頻資料並輸出解碼視頻資料之構件;用於接收編碼音頻資料並輸出解碼音頻資料之構件;用於累積一視頻訊框週期之解碼視頻資料,並在每一訊框週期輸出一視頻資料訊框的構件;用於累積一音頻訊框週期之解碼音頻資料,並在每一訊框週期輸出一音頻資料訊框的構件;及用於組合該視頻資料訊框及該音頻資料訊框,並輸出一音頻視頻資料之同步訊框的構件。
- 20A base station equipment comprising:a component for receiving encoded video data and outputting decoded video data;a component for receiving encoded audio data and outputting decoded audio data;a component for accumulating decoded video data for a video frame period, and A component for outputting a video data frame in each frame period;a component for accumulating decoded audio data for an audio frame period and outputting an audio data frame for each frame period;and for combining the video The data frame and the audio data frame, and a component that outputs a synchronous frame of audio and video data. 一種基地台設備,其包含:用於接收編碼視頻資料並輸出解碼視頻資料之構件;用於接收編碼音頻資料並輸出解碼音頻資料之構件;用於累積一視頻訊框週期之解碼視頻資料,並在每一訊框週期輸出一視頻資料訊框的構件;用於累積一音頻訊框週期之解碼音頻資料,並在每一訊框週期輸出一音頻資料訊框的構件;及用於組合該視頻資料訊框及該音頻資料訊框,並輸出一音頻視頻資料之同步訊框的構件。
Independent claims12
84 paragraphs, as filed
Synchronization of audio and video data in wireless communication systems
The present invention generally relates to the transmission of information in a wireless communication system, and more specifically relates to the synchronization of audio and video data transmitted in a wireless communication system.
Various technologies have been developed for transmitting multimedia or real-time data such as audio or video data on various communication networks. One of this technology is Real Time Transfer Protocol (RTP). RTP provides an end-to-end network transmission function that can be used to transmit real-time data in multi-point play or single-point play network services. RTP does not deal with resource reservations and does not guarantee the service quality of real-time services. A control protocol (RTCP) is used to enhance data transmission to allow monitoring of data transmission to be upgraded to a large multipoint broadcast network in one way, and provide minimal control and identification functionality. RTP and RTCP are designed to be independent of the underlying transport and network layers. The protocol supports the use of RTP horizontal translators and mixers. July 2003 RFC-3550 draft standard, H. Schulzrinne [Columbia University], S. Casner [Packet Design], R. Frederick [Blue Coat Systems Inc.], V. Jacobson [Packet Design] in Internet Engineering Steering Group, etc. Human "RTP: A Transport Protocol for Real-Time More details about RTP can be found in "Applications", the full text of the case is incorporated into this article by reference.
An example illustrating the state of RTP is audio conferencing, where RTP is used for voice communication on the upper layer of the Internet Protocol (IP) service of the Internet. With a distribution organization, a conference originating station obtains a multicast group address and a pair of ports. One port is used for audio data, and the other port is used for control (RTCP) packets. Assign this address and port information to the desired participants. The audio conference application used by each conference participant sends audio data in small segments, such as a segment with a duration of 20 ms. An RTP header is added before each division of audio data; and the combined RTP header and data are encapsulated in a UDP packet. The RTP header includes information about the data, for example, it indicates the type of audio coding contained in each packet (such as PCM, ADPCM or LPC), and the time stamp (TS) of the time to reproduce the RTP packet, which can be used to detect loss/copy The sequence number (SN) of multiple consecutive packets of a packet, etc. For example, this allows the sender to change the encoding type used during the conference to accommodate a new participant connected via a low-bandwidth link, or to respond to indications of network congestion.
According to the RTP standard, if both audio and video media are used in an RTP conference, they will be transmitted as an independent RTP session. That is, two different UDP port pairs and/or multicast addresses are used for each media to transmit independent RTP and RTCP packets. There is no direct coupling between audio and video sessions at the RTP level, except that users participating in the two sessions use the same name for the two in the RTCP packet to make the sessions related.
The motivation for transmitting audio and video as independent RTP sessions is to allow certain participants in the conference to receive only one type of media (if they choose). Despite the separation, the timing information carried in the RTP/RTCP packets of the two sessions can still be used to achieve synchronized replay of the source audio and video.
Packet networks like the Internet may occasionally lose or reschedule packets. In addition, individual packets can experience a variable amount of delay in their respective transmission time. To deal with this damage, the RTP header contains timing information and a sequence number that allows the receiver to reconstruct the timing generated by the source. Perform this timing reconstruction on each source of the RTP packet in a session.
Although the RTP header includes timing information and a sequence number, because audio and video are transmitted in independent RTP streams, there is a potential time slip between these streams, also known as lip sync or AV sync. The application at the receiver must resynchronize these streams before reproducing audio and video. In addition, in applications that transmit RTP streams such as audio and video over wireless networks, the possibility of packet loss increases, thereby making it more difficult to resynchronize the streams.
Therefore, it is necessary to improve the synchronization of audio and video RTP streams transmitted over the network in this technology.
The embodiments disclosed herein deal with the aforementioned needs by encoding data streams such as audio and video streams transmitted on a network (such as a wireless network or an IP network) to synchronize the data streams. For example, a complete audio frame and a complete video frame are transmitted within a required frame period to reproduce the audio and video frames by the application program in the receiver. For example, the data stream synchronizer may include a first decoder configured to receive a first encoded data stream and output and decode the first data stream, wherein the first encoded data stream has a first decoder during the information interval One bit rate. The data synchronizer may also include a second decoder configured to receive a second encoded data stream and output a decoded second data stream, wherein the second encoded data stream has a second bit during the information interval rate. The first buffer is configured to accumulate the first decoded data stream of at least one information interval and output a first decoded data stream frame in each interval period. The second buffer is configured to accumulate the second decoded data stream of at least one information interval and output a second decoded data stream frame in each interval period. Then, a combiner configured to receive the first decoded data stream frame and the second decoded data stream frame outputs synchronization frames of the first and second decoded data streams. The first encoded data stream may be video data, and the second encoded data stream may be audio data.
One aspect of this technology includes receiving an audio and video RTP stream and assigning a complete frame of RTP video data to communication channel packets. The communication channel packets occupy the same rate as the video frame or higher than the video frame rate. Few cycles. A complete frame of RTP audio data is also assigned to communication channel packets, which occupy a period that is the same as the audio frame rate or less than the audio frame rate. Simultaneously transmit these video and audio communication channel packets. Receiving and assigning RTP streams can be performed in the remote site or base station.
Another aspect of the present invention is to receive communication channel packets including audio and video data. Decode audio and video data and accumulate data with a period equal to the frame period of the audio and video data. At the end of the frame period, a video frame and an audio frame are combined. Because the audio frame and the video frame are transmitted at the same time, and each transmission occurs within a frame period, the audio and video frames are synchronized. Decoding and accumulation can be performed in the remote station or base station.
The word "exemplary" as used herein means "serving as an example, illustration, or illustration." Any embodiment described herein as "exemplary" need not be construed as better or superior to other embodiments.
The term "streaming" as used in this article refers to the real-time delivery of essentially continuous multimedia data such as audio, voice or video information on dedicated and shared channels in conversational, single-point play, and broadcast applications. For video, the phrase "multimedia frame" used in this article means a video frame that can be displayed/reproduced on a display device after decoding. A video frame can be further divided into independent decodable units. In video terms, these units are called "segments". In the case of audio and speech, the term "multimedia frame" as used herein means information in a time window on which speech or audio is compressed for transmission and decoding at the receiver. The phrase "information unit interval" used in this article refers to the duration of the above-mentioned multimedia frame. For example, in the case of video, in the case of a video with 10 frames per second, the information unit interval is 100 milliseconds. In addition, as an example, in the case of voice, the information unit interval is usually 20 milliseconds in cdma2000, GSM and WCDMA. From this description, it should be understood that the audio/voice frame usually does not need to be further divided into independently decodable segments and the video frame is usually further divided into independently decodable segments. Obviously, when the phrases "multimedia frame", "information unit interval", etc. refer to multimedia data such as video, audio and voice, they should be understood to form this text.
Describes the technology used to synchronize the transmitted RTP stream on a set of constant bit rate communication channels. These technologies include dividing the information unit transmitted in the RTP stream into data packets, where the size of the data packet is selected to match the physical layer data packet size of the communication channel. For example, audio and video data that are synchronized with each other can be encoded. The encoder can be constrained so that the encoder encodes the data to a size that matches the size of the available physical layer packet of the communication channel. Restrict the data packet size to match one or more available physical layer packet sizes. Support the transmission of multiple simultaneous RTP streams, because these RTP streams are transmitted simultaneously or continuously, but the audio and video packets need to be reproduced simultaneously in the time frame . For example, if the audio and video RTP streams are transmitted, and the data packet is constrained so that its size matches the available physical layer packet, the audio and video data are transmitted and synchronized within the display time. As described in the applications listed in REFERENCE TO CO-PENDING APPLICATIONS FOR PATENTS above, when the amount of data required for the RTP stream changes, the communication channel capacity can be changed by selecting different physical layer packet sizes .
Examples of information units such as RTP streams include variable bit rate data streams, multimedia data, video data, and audio data. The information units can appear at a constant repetition rate. For example, the information units can be audio/video data frames.
Different domestic and international standards have been established to support various radio interfaces, including (for example) Advanced Mobile Phone Service (AMPS), Global Mobile Communication System (GSM), General Packet Radio Service (GPRS), Data Incremental GSM Environment (EDGE) ), Transitional Standard 95 (IS-95) and its derivative standards (IS-95A, IS-95B, ANSI J-STD-008 (referred to as IS-95 in this article)), and such as cdma2000, Global Mobile Telecommunications Service (UMTS) ), broadband CDMA, WCDMA and other emerging high data rate systems. These standards are issued by the Telecommunications Industry Association (TIA), the Third Generation Mobile Communications Partnership Project (3GPP), the European Telecommunications Standards Institute (ETSI) and other well-known standardization bodies.
Fig. 1 shows a communication system 100 constructed in accordance with the present invention. The communication system 100 includes an infrastructure 101, a plurality of wireless communication devices (WCD) 104 and 105, and land-based communication devices 122 and 124. WCD will also be called Action Station (MS) or Action Body. Generally, WCD can be mobile or fixed. The land-based communication devices 122 and 124 may include, for example, service nodes or content servers, which provide various types of multimedia data such as streaming multimedia data. In addition, the MS can transmit streaming data such as multimedia data.
The infrastructure 101 may also include other components, such as the base station 102, the base station controller 106, the mobile switching center 108, the switching network 120, and the like. In one embodiment, the base station 102 and the base station controller 106 are integrated, and in other embodiments, the base station 102 and the base station controller 106 are independent components. Different types of switching networks 120 such as an IP network or a public switched telephone network (PSTN) can be used in the communication system 100 to project signals.
The term "forward link" or "downlink" refers to the signal path from the infrastructure 101 to the MS, and the term "reverse link" or "uplink" refers to the signal path from the MS to the infrastructure . As shown in FIG. 1, MS 104 and 105 receive signals 132 and 136 on the forward link, and transmit signals 134 and 138 on the reverse link. Generally, the signals transmitted from MS 104 and 105 are intended to be received at another communication device such as another remote unit or land communication device 122 and 124, and transmitted via the switching network 120. For example, if the signal 134 transmitted from the initial WCD 104 is intended to be received by the destination MS 105, the signal is delivered via the infrastructure 101, and the signal 136 is transmitted to the destination MS 105 on the forward link. Similarly, in the infrastructure 101, the initial signal can be broadcast to the MS 105. For example, a content provider can send multimedia data such as streaming multimedia data to the MS 105. Generally, a communication device such as an MS or a land-based communication device can be the start end of the signal and the destination end of the signal.
Examples of MS 104 include mobile phones, wireless communication-enabled personal computers and personal digital assistants (PDAs), and other wireless devices. The communication system 100 can be designed to support one or more wireless standards. For example, these standards may include what are known as Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Data Incremental GSM Environment (EDGE), TIA/EIA-95-B (IS-95) , TIA/EIA-98-C (IS-98), IS2000, HRPD, cdma2000, wideband CDMA (WCDMA) and other standards.
Figure 2 is a block diagram illustrating an exemplary packet data network and various radio interface options for transmitting packet data over a wireless network. The technique can be implemented in a packet-switched data network 200 such as that illustrated in FIG. 2 . As shown in the example of FIG. 2, the packet switching data network system may include a wireless channel 202, a plurality of receiving nodes or MS 204, a sending node or content server 206, a service node 208, and a controller 210. The sending node 206 may be coupled to the service node 208 via a network 212 such as the Internet.
The serving node 208 may include, for example, a packet data serving node (PDSN) or a serving GPRS support node (SGSN) or a gateway GPRS support node (GGSN). The service node 208 can receive the packet data from the sending node 206 and provide the information packet to the controller 210. The controller 210 may include, for example, a base station controller/packet control function (BSC/PCF) or a radio network controller (RNC). In an embodiment, the controller 210 communicates with the service node 208 on a radio access network (RAN). The controller 210 communicates with the service node 208 and transmits a packet of information to at least one of the receiving nodes 204 via the wireless channel 202, such as an MS.
In one embodiment, the serving node 208 or the sending node 206 or both may also include an encoder for encoding the data stream, or a decoder for decoding the data stream, or both. For example, the encoder can encode audio/video streams and generate data frames therefrom, and the decoder can receive the data frames and decode them. Similarly, the MS may include an encoder for encoding a data stream, or a decoder for decoding a received data stream, or both. The term "codec" is used to describe the combination of encoder and decoder.
In an example illustrated in FIG. 2, data such as multimedia data from the sending node 206 can pass through the service node or packet data service node (PDSN) 208 and a controller or base station controller/packet control function (BSC/PCF). ) 210 is sent to a receiving node or MS 204, and the sending node 206 is connected to the network or the Internet 212. The wireless channel interface 202 between the MS 204 and the BSC/PCF 210 is a radio interface, and usually many channels can be used for signal transmission and carrying or payload data.
The radio interface 202 can operate according to any of a large number of wireless standards. For example, these standards may include standards based on TDMA such as Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Data Incremental GSM Environment (EDGE), or based on standards such as TIA/EIA-95- B (IS-95), TIA/EIA-98-C (IS-98), IS2000, HRPD, cdma2000, broadband CDMA (WCDMA) CDMA standards and other standards.
Fig. 3 is a diagram illustrating synchronization difficulties in the conventional technology of transmitting independent RTP streams on a wireless communication channel. In the example illustrated in FIG. 3, video and audio data frames are encoded into RTP streams and then assigned to communication channel packets. FIG. 3 illustrates the streaming of the video frame 302. Usually, the video frame appears at a constant rate. For example, a video frame can appear at a rate of 10 Hz, that is, a new frame appears every 100 milliseconds.
As shown in Figure 3, a single video frame can contain different amounts of data, as indicated by the height of the bars representing each frame. For example, if the video data is encoded as moving image compression standard (MPEG) data, the video stream consists of an internal frame (I frame) and a predictive frame (P frame). The I frame is self-contained, that is, it includes all the information needed to reproduce or display a complete video frame. The P frame is non-self-contained and will usually contain difference information, such as motion vector and difference texture information, relative to the previous frame. Generally, the I frame can be up to 8 to 10 times larger than the P frame, depending on the content and encoder settings. Although video frames can have different amounts of data, they still appear at a constant rate. The I and P frames can be further divided into multiple video clips. A video segment represents a small area in the display screen and can be decoded by the decoder alone.
In FIG. 3, video frames N and N+4 can represent I frames, and video frames N+1, N+2, N+3, and N+5 can represent P frames. As shown in the figure, the I frame includes a larger amount of data than the P frame, which is indicated by the height of the bars representing the frame. The video frame is then packetized into packets in the RTP stream 304. As shown in FIG. 3, RTP packets N and N+4 corresponding to video I frames N and N+4 are larger than RTP packets N+1, N+2, and N+3 corresponding to video P frames N+1, N+2, and N+3, as indicated by their widths.
The video RTP packet is allocated to the communication channel packet 306. In a conventional communication channel such as CDMA or GSM, the communication channel data packet 306 has a constant size and is transmitted at a constant rate. For example, the communication channel data packets 306 can be transmitted at a rate of 50 Hz, that is, a new data packet is transmitted every 20 milliseconds. Because the communication channel packets are of constant size, more communication channel packets are needed to transmit larger RTP packets. Therefore, compared to the communication channel packet required to transmit the smaller RTP packets corresponding to the P video frame N+1, N+2, and N+3, the communication channel packet 306 required to transmit the RTP packet corresponding to the I video frame N and N+4 is more many. In the example illustrated in FIG. 3, the video frame N occupies the block 308 of the nine communication channel packet 306. Video frames N+1, N+2, and N+3 occupy blocks 310, 312, and 314, respectively, and each block has four communication channel packets 306. Video frame N+4 occupies block 316 of packet 306 of nine communication channels.
For each frame of video data, there is a corresponding audio data. FIG. 2 illustrates the streaming of the audio frame 320. Each audio frame N, N+1, N+2, N+3, N+4, and N+5 corresponds to the respective video frame and appears at a rate of 10 Hz, that is, a new audio frame starts every 100 milliseconds. Generally, audio data is less complex, so that the audio data can be represented by fewer bits than the associated video data. The audio data is usually encoded so that the RTP packet 322 can be transmitted via the communication channel within a frame period.ofsize. In addition, a typical audio frame is generated every 20 milliseconds in CDMA, GSM, WCDMA, etc. In these cases, multiple audio frames are bundled together so that the audio and video packets represent the same duration of the RTP packet. For example, RTP packets N, N+1, N+2, N+3, N+4, and N+5 have a size, and each RTP packet can be assigned to the communication channel packet 324 so that each RTP packet is transmitted via the communication channel within a frame period of 100 milliseconds .
As shown in FIG. 3, audio frame packets N, N+1, N+2, N+3, N+4, and N+5 occupy blocks 326, 328, 330, 332, 334, and 336, respectively, and each block has five communication channel packets 324.
The comparison between assigning a video frame and an audio frame to their respective communication channel packets indicates that the synchronization between the audio frame and the video frame is lost. In the example illustrated in FIG. 3, a block 308 of nine communication channel packets 306 is required to transmit the video frame N. The audio frame N associated with the video frame N is transmitted in the block 326 of the five communication channel packet 324. Because the video and audio in the communication channel packet are transmitted at the same time, during the transmission of the video frame N, the audio frame N and four of the five communication channel packets in the block 328 of the audio frame N+1 are transmitted.
For example, in Figure 3, if the frame rate of video and associated audio is 10 Hz and the communication channel packet rate is 50 Hz, then all audio data is transmitted during the 100 millisecond period of frame N, but only Transmit part of the video data. In this example, all the video data of frame N is not completely transmitted until the other four communication channel packets 306 are transmitted, which results in the completion of the transmission of video frame N compared to 100 milliseconds for completing the transmission of audio frame N It takes 180 milliseconds. Since the audio and video RTP streams are independent, part of the data of the audio frame N+1 is transmitted during the transmission of the data of the video frame N. This loss of synchronization between the video and audio streams can cause a "slip" between the video and audio at the receiver of the communication channel.
Because the use of many parameters due to predictive coding and also due to a number of variable length coding (VLC), such as H.263, AVC/H.264, MPEG-4, etc., video encoders actually have inherent variable Therefore, the sender and receiver usually use buffers to perform string traffic shaping to realize the real-time transmission of variable-rate bit streams via circuit-switched networks and packet-switched networks. The string traffic shaping buffer introduces additional delays that are generally undesirable. For example, the extra delay between when one person speaks and when another person hears the voice during a conference call may be annoying.
For example, because the video is replayed at the receiver of the communication channel at the same rate as the original video frame rate, a delay in the communication channel can cause a pause in playback. In Figure 3, the video frame N cannot be replayed until the complete frame of the received data. Because the complete frame data is not received during the frame period, the playback must be paused until all the video data of frame N is received. In addition, it is necessary to store all the data from the audio frame N until all the video data of the frame N is received, so that the audio and video playback can be synchronized. It should also be noted that although the video data from frame N is still being received, the audio data from frame N+1 must be stored until all the video data from frame N+1 is received. Due to the variable size of the video frame, a large string traffic shaping buffer is required to achieve synchronization.
FIG. 4 is a diagram illustrating a technique of transmitting independent RTP streams via a wireless communication channel according to the present invention. Similar to FIG. 3, FIG. 4 illustrates that the stream of the video frame 302 and the stream of the audio frame 320 of varying sizes are encoded into independent RTP streams 304 and 322, respectively. The video and audio frames appear at a constant rate, such as a rate of 10 Hz.
Similar to FIG. 3, in FIG. 4, video frames N and N+4 can represent I frames, and video frames N+1, N+2, N+3, and N+5 can represent P frames. These video frames are packetized into packets in the RTP stream 304. As shown in FIG. 4, RTP packets N and N+4 corresponding to video I frames N and N+4 are larger than RTP packets N+1, N+2, and N+3 corresponding to video P frames N+1, N+2, and N+3, as indicated by their widths.
The video RTP packet is allocated to the communication channel packet 406. The capacity of the communication channel can be changed by using the technology described in the application case in the same application as listed in REFERENCE TO CO-PENDING APPLICATIONS FOR PATENT above. Since the capacity of the communication channel packet 406 is variable, the video frame N can be transmitted in a block 408 containing five communication channel packets 406.
For example, conventional communications based on CDMA standards (such as TIA/EIA-95-B (IS-95), TIA/EIA-98-C (IS-98), IS2000, HRPD, cdma2000, and broadband CDMA (WCDMA)) In the channel, the communication channel data packet 406 can be transmitted at a rate of 50 Hz, that is, a new data packet is transmitted every 20 milliseconds. Because the capacity of the communication channel packet 406 is variable, the encoding of the video frame N can be constrained so that the complete video frame N can be transmitted during one frame period. As shown in FIG. 4, when the RTP packet N corresponding to the video frame N is transmitted, the capacity of the communication channel packet 406 increases, so that a complete packet can be transmitted during the frame period. The described technologies can also be applied to communication channels based on GSM, GPRS or EDGE.
As illustrated in FIG. 4, video frames N, N+1, N+2, N+3, N+4, and N+5 are encoded into RTP packets and assigned to communication channel blocks 408, 410, 412, 414, 416, and 418, respectively. It should also be noted that by changing the communication channel capacity, a complete video frame can be transmitted within a frame period. For example, if the video frame rate is 10 Hz, the complete frame of the video data is transmitted during the frame period of 100 milliseconds.
For each frame of the video data 302, there is a corresponding audio frame 320. Each audio frame N, N+1, N+2, N+3, N+4, and N+5 corresponds to a respective video frame and appears at a rate of 10 Hz, that is, a new audio frame starts every 100 milliseconds. Compared with the discussion in FIG. 3, the audio data is usually less complex, so that the audio data can be represented by fewer bits than the associated video data, and the audio data is usually encoded so that the RTP packet 322 has a frame that can be The size of the transmission via the communication channel in a 100 millisecond period. That is, the audio RTP packets N, N+1, N+2, N+3, N+4, and N+5 have a size to assign each RTP packet to the blocks 326, 328, 330, 332, 334, and 336 of the communication channel packet, respectively. Therefore, if the video frame rate is 10 Hz, each video frame can be transmitted via the communication channel within a frame period of 100 milliseconds. Similar to video, if the audio packet size is larger, the communication channel capacity can also be changed to support the transmission of a complete audio frame during a frame period.
In Figure 4, the comparison between assigning video frames and audio frames to their respective communication channel packets shows that the video and audio frames are still synchronized. In other words, each frame period transmits a complete video frame and a complete audio frame. Because each frame cycle transmits a complete frame of video and audio, no additional buffering is required. During a frame period, only the received video and audio data need to be accumulated, and then they can be played out. Because there is no delay introduced by the communication channel, the video and audio frames remain synchronized.
As illustrated in Figure 3, it should be noted that the video frames N+1, N+2, and N+3 only require four video communication channel packets 306 to transmit the complete frame of the video data. As shown in FIG. 4, the size of the video communication channel packet 406 can be reduced, so that the size of the video data can fit five packets, or a blank packet can be transmitted. Similarly, if there is excess available capacity in the audio communication channel, blank packets can be transmitted. Therefore, video and audio data are encoded so that complete frames of audio and video data are assigned to communication channel packets occupying the same or less cycles or their respective frame rates.
As described below, depending on the state of the communication network, different technologies can be used to synchronize RTP streams. For example, the communication network may be over-provisioned, that is, it has excess capacity, or the communication network may have a quality of service guarantee. In addition, the RTP stream can be modified to facilitate synchronization during transmission over the communication network. Each of these technologies will be discussed below.
<b>Oversupply communication network</b>
In the case that the communication link between the PDSN 208 and the transmitter 206 is over-provisioned (that is, there is an excess capacity available for data transmission on the wired Internet), there is no delay due to congestion. Because there is excess capacity in the communication link, there is no need to delay the transmission so that the communication link can accommodate the transmission. When there is no delay in transmission, when the voice and video packets arrive at the infrastructure (such as the PDSN), there is no "time slip" between the voice and video packets. In other words, the audio and video data are kept synchronized with each other up to the PDSN, and as described in the present invention, between the PDSN and the MS.
In the case of oversupply, it is easy to achieve audio-visual synchronization. For example, the video data may have a frame rate of 10 frames per second (fps) based on a 100 millisecond frame, and the associated audio may have a frame rate of 50 fps based on a 20 millisecond voice frame. In this example, five frames of the received audio data are buffered to synchronize with the video frame rate. That is, five frames corresponding to 100 milliseconds of audio data will be buffered to synchronize with the 100 milliseconds of video frame.
Communication network with guaranteed service quality for maximum delay
By buffering an appropriate number of higher frame rate voice frames, a lower frame rate video frame can be matched. Generally, if the following quality of service (QoS) delay guarantees are used to deliver video packets: QoS_delay=nT ms Equation 1 where n is the delay in the frame; and T=1000/frames_per_second, it needs to have the size to store nT/w voice frames Buffer to store enough voice frames to ensure that the voice and video can be synchronized, where w is the duration of the voice frame in milliseconds. In cdma2000 UMTS, the duration w of the voice frame is 20 milliseconds, and the duration of the voice frame in other communication channels can be different or changeable.
Another technique for synchronizing audio and video data includes buffering the data streams of both. For example, if a communication system has D<sub>Q</sub>Guaranteed maximum delay in milliseconds, meaning D<sub>Q</sub>For the maximum delay that can be experienced during the transmission of audio and video streams, a buffer of appropriate size can be used to maintain synchronization.
For example, with D<sub>Q</sub>The maximum delay guarantee, then buffer D<sub>Q</sub>/T video frames (T is the duration of the video frame in milliseconds) and D<sub>Q</sub>/w voice frames (w is the duration of the audio frame in milliseconds) will ensure audio and video synchronization (AV synchronization). This extra buffer space is usually called a de-jitter buffer.
These technologies describe the synchronization of audio and video data streams. These technologies can be used with any data stream that needs to be synchronized. If there are two data streams with the same information interval and need to be synchronized with a first higher bit rate data stream and a second lower bit rate data stream, buffering the higher bit rate data allows it to be synchronized with Data synchronization at a lower bit rate. It can be determined that the size of the buffer depends on the above-mentioned QoS. Likewise, as described above, both higher and lower bit rate data streams can be buffered and synchronized.
The described technique can be performed by a data stream synchronizer, which includes: a first decoder configured to receive a first encoded data stream and output a decoded first data stream , Wherein the first encoded data stream has a first bit rate during an information interval; and a second decoder configured to receive a second encoded data stream and output a decoded second data stream, The second coded data stream has a second bit rate during the information interval. The data stream synchronizer also includes a first buffer configured to accumulate a first decoded data stream of at least one information interval and output a first decoded data stream frame in each interval period, and a configured A second buffer that accumulates the second decoded data stream of at least one information interval and outputs a second decoded data stream frame in each interval period. Then a combiner is configured to receive the first decoded data stream frame and the second decoded data stream frame and output the synchronization frames of the first and second decoded data streams. In an example, the first encoded data stream may be video data and the second encoded data stream is audio data, such that the first bit rate is higher than the second bit rate.
<b>Single RTP stream with multiplexed audio and video</b>
Another embodiment is to carry audio and video in a single RTP stream. It should be noted that in common practice, audio and video are not transmitted as a single RTP stream in an IP network. RTP is designed to enable participants with different resources (for example, terminals with video and audio functions and terminals with only audio functions) to communicate in the same multimedia conference.
The restriction of transmitting audio and video as independent RTP streams cannot be used in wireless networks for video services. In this case, a new RTP profile can be designed to carry a specific voice and video codec payload. Combining audio and video into a common RTP stream eliminates any time slip between audio and video data without the need to oversupply the communication network. Therefore, audio and video synchronization can be achieved by using the technology described in conjunction with the oversupply network described above.
Figure 5 is a block diagram of a portion of a wireless audio/video receiver 500 configured to receive communication channel packets. As shown in FIG. 5, the audio/video receiver 500 includes a communication channel interface 502 configured to receive communication channel packets. The communication channel interface 502 outputs the video communication channel packet to a video decoder 504 and outputs the audio communication channel packet to an audio decoder 506. The video decoder 504 decodes the video communication channel packet and outputs video data to a video buffer 508. The audio decoder 506 decodes the audio communication channel packet and outputs audio data to an audio buffer 510. The video buffer 508 and the audio buffer respectively accumulate video and audio data for one frame period. The video buffer 508 and the audio buffer 510 respectively output a video frame and an audio frame to a combiner 512. The combiner 512 is configured to combine video and audio frames and output a synchronized audio and video signal. The operation of the video buffer 508, the audio buffer 510, and the combiner 512 can be controlled by the controller 514.
Figure 6 is a block diagram of a portion of a wireless audio/video transmitter 600 configured to transmit communication channel packets. As shown in FIG. 6, the audio/video transmitter 600 includes a video communication channel interface 602 configured to receive an RTP stream of video data. The video communication channel interface assigns RTP packets to communication channel packets. It should be noted that the capacity of the communication channel packet is variable in order to assign the complete frame value (worth) of the RTP video data to the communication channel packet occupying the same period as the video frame. The audio/video transmitter 600 also includes an audio communication channel interface 604 configured to receive an audio data RTP stream. The audio communication channel interface 604 assigns the RTP packet to the communication channel packet. It should be noted that usually the capacity of the communication channel packet is sufficient to assign a complete frame of RTP audio data to the communication channel packet occupying the same period as the audio frame. If the channel capacity is insufficient, it is similar to a video communication channel packet. The audio communication channel capacity can be changed so that there is enough capacity to assign a complete frame of RTP audio data to a communication channel packet occupying the same period as the audio frame.
The video and audio communication channel interfaces 602 and 604 respectively output video and audio communication channel packets and are connected to the combiner 606. The combiner 606 is configured to accept video and audio communication channel packets and combine them to output a composite signal. The output of the combiner 606 is connected to a transmitter 608 that transmits the composite signal to the wireless channel. The operation of the video communication channel interface 602, the audio communication channel interface 604, and the combiner 606 can be controlled by the controller 614.
Fig. 7 is a serial flow chart of transmitting independent RTP streams via a wireless communication link. The streaming process starts at block 702, where video and audio RTP data streams are received. The streaming flow then continues to block 704, where the video RTP stream is assigned to the communication channel packet. In block 706, the audio RTP stream is assigned to the communication channel packet. In block 708, the video and audio communication channel packets are combined and transmitted via the wireless channel.
Fig. 8 is a sequence flow chart of receiving audio and video data via a wireless communication channel. The string flow starts at block 802, where video and audio data are received via a wireless communication channel. The string flow continues to block 804, where video and audio data are decoded. In block 806, the decoded video and audio data sets are translated into respective video and audio frames. In block 808, the video and audio data are combined into a synchronized video/audio frame. In block 810, the synchronized video/audio frame is output.
FIG. 9 is a block diagram of a wireless communication device or a mobile station (MS) constructed according to an exemplary embodiment of the present invention. The communication device 902 includes a network interface 906, a codec 908, a host processor 910, a memory device 912, a program product 914, and a user interface 916.
The network interface 906 receives signals from the infrastructure and sends them to the host processor 910. The host processor 910 receives the signal and responds with appropriate actions depending on the content of the signal. For example, the host processor 910 can decode the received signal itself, or it can deliver the received signal to the codec 908 for decoding. In another embodiment, the received signal is directly sent from the network interface 906 to the codec 908.
In one embodiment, the network interface 906 can be a transceiver and an antenna to establish an interface with the infrastructure via wireless channels. In another embodiment, the network interface 906 may be a network interface card, which is used to establish an interface with the infrastructure via a land communication line. The codec 908 may be implemented as a digital signal processor (DSP), or a general-purpose processor such as a central processing unit (CPU).
Both the host processor 910 and the codec 908 are connected to a memory device 912. During the WCD operation, the memory device 912 can be used to store data and store program codes to be executed by the host processor 910 or the DSP 908. For example, the host processor, the codec, or both can operate under the control of program instructions, which are temporarily stored in the memory device 912. The host processor 910 and the codec 908 may also include their own program storage memory. When the program instructions are executed, the host processor 910 or the codec 908 or both perform their functions, such as decoding or encoding multimedia streams (such as audio/video data) and assembling such audio and video frames. Therefore, the program steps respectively implement the functionality of the host processor 910 and the codec 908, so that the host processor and the codec can respectively perform the functions of decoding or encoding content streams and assembling frames as needed. These program steps can be received from a program product 914. The program product 914 can store the program steps and transfer them to the memory 912 for execution by the host processor, the codec, or both.
The program product 914 can be a semiconductor memory chip, such as RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, register, and other storage devices, such as hard disks, removable disks, CD-ROM or any other form of storage medium known in the art that can store instructions that can be read by a computer. In addition, the program product 914 may be a source file including program steps, the source file is received from the network and stored in the memory and then executed. In this way, the processing steps required for the operation of the present invention can be embodied on the program product 914. In FIG. 9, the exemplary storage medium shown is coupled to the host processor 910 so that the host processor can read information from the storage medium and write information into the storage medium. Alternatively, the storage medium can be integrated into the host processor 910.
The user interface 916 is connected to the host processor 910 and the codec 908. For example, the user interface 916 may include a display and speakers for outputting multimedia data to the user.
Those familiar with the art will understand that the steps of the method described in the embodiments can be interchanged without departing from the scope of the present invention.
Those familiar with this technology will also understand that any of a variety of different technologies and processes can be used to represent information and signals. For example, the data, commands, commands, information, signals, bits, symbols referred to throughout the above description can be represented by voltage, electric stream, electromagnetic waves, magnetic fields or particles, light fields or particles, or any combination thereof And chips.
Those familiar with the technology will further understand that the various illustrative logic blocks, modules, circuits, and algorithm steps described can be implemented as electronic hardware, computer software, or a combination of the two in combination with the embodiments disclosed herein. In order to clearly illustrate the interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have generally been described above based on their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Skilled technicians can implement the described functionality in different ways for each specific application, but this implementation method should not be construed as causing a departure from the scope of the present invention.
It can be used by general-purpose processors, digital signal processors (DSP), special application integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic, Discrete hardware components or any combination thereof designed to perform the functions described herein implement or execute the various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein. The general-purpose processor may be a microprocessor, but or the processor may be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors connected to the same DSP core, or any other such configuration.
The steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be directly embodied in hardware, in a software module executed by a processor, or a combination of the two. The software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, scratchpad, hard disk, removable disc, CD-ROM or known in this technology In any other form of storage media. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information into the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and the storage medium can be located in an ASIC. The ASIC can reside in the user terminal. Alternatively, the processor and the storage medium may reside as discrete components in the user terminal.
The previous description of the disclosed embodiments is provided to enable anyone familiar with the art to make or use the present invention. Those familiar with the art will easily understand various modifications to these embodiments, and the general principles defined herein can be applied to other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown in this document, but should conform to the broadest scope consistent with the principles and novel features disclosed in this document.
<p>100. . . Communication Systems</p><p>101. . . basicly construct</p><p>102. . . Base station</p><p>104,105. . . Wireless communication device</p><p>106. . . Base station controller</p><p>108. . . Mobile switching center</p><p>120. . . Switched network</p><p>122,124. . . Land communication device</p><p>132,134,136,138. . . Signal</p><p>200. . . Packet-switched data network</p><p>202. . . Radio interface</p><p>204. . . Receiving node/MS</p><p>206. . . Sending node</p><p>208. . . Service Node/PDSN</p><p>210. . . Controller/(BSC/PCF)</p><p>212. . . Internet/Internet</p><p>302. . . Video frame</p><p>304. . . RTP streaming/RTP packet</p><p>306. . . Communication channel packet</p><p>308,310,312,314,316. . . Block</p><p>320. . . Audio frame</p><p>322. . . RTP streaming/RTP packet</p><p>324. . . Communication channel packet</p><p>326,328,330,332,334,336. . . Block</p><p>406. . . Communication channel packet</p><p>408,410,412,414,416,418. . . Communication channel block</p><p>500. . . Audio/video receiver</p><p>502. . . Communication channel interface</p><p>504. . . Video decoder</p><p>506. . . Audio decoder</p><p>508. . . Video buffer</p><p>510. . . Audio buffer</p><p>512. . . Combiner</p><p>514. . . Controller</p><p>602. . . Video communication channel interface</p><p>604. . . Audio communication channel interface</p><p>608. . . Transmitter</p><p>610. . . Controller</p><p>902. . . Communication device</p><p>906. . . Network interface</p><p>908. . . Codec</p><p>910. . . Host processor</p><p>912. . . Memory device/memory</p><p>914. . . Program product</p><p>916. . . user interface</p>
Fig. 1 is an illustration of a part of a communication system constructed in accordance with the present invention.
FIG. 2 is a block diagram illustrating an exemplary packet data network and various radio interface options for transmitting packet data on the wireless network in the system of FIG. 1. FIG.
Fig. 3 is a diagram illustrating synchronization difficulties in the conventional technology of transmitting independent RTP streams on a wireless communication channel.
FIG. 4 is a diagram illustrating a technique of transmitting independent RTP streams on a wireless communication channel according to the present invention.
Figure 5 is a block diagram of a portion of a wireless audio/video receiver configured to receive communication channel packets.
Figure 6 is a block diagram of a portion of a wireless audio/video transmitter configured to transmit communication channel packets.
Fig. 7 is a flow chart of the transmission of an independent RTP stream on a wireless communication link.
Figure 8 is a sequence flow chart of receiving audio and video data on a wireless communication channel.
FIG. 9 is a block diagram of a wireless communication device or mobile station (MS) constructed according to an exemplary embodiment of the present invention.
10 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| TWI487360B | Cited by | Taiwan Province of China | Examiner |
| US9160546B2 | Cited by | United States of America | Applicant |
104 members in 14 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 57167304 | United States of America | P | |
| 57167304 | United States of America | P | |
| 60571673 | United States of America | – | |
| 60571673 | – | – | – |
| US20040571673P | – | – | – |
Members104
| Document | Office | Kind | |
|---|---|---|---|
| US2005259613A1 | United States of America | A1 | |
| US2005259623A1 | United States of America | A1 | |
| US2005259690A1 | United States of America | A1 | |
| US2005259694A1 | United States of America | A1 | |
| CA2565977A1 | Canada | A1 | |
| CA2566124A1 | Canada | A1 | |
| CA2566125A1 | Canada | A1 | |
| CA2566126A1 | Canada | A1 | |
| CA2771943A1 | Canada | A1 | |
| CA2811040A1 | Canada | A1 | |
| WO2005114919A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2005114943A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2005114950A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2005115009A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2005114943A3 | World Intellectual Property Organization (WIPO) | A3 | |
| TW200618544A | Taiwan Province of China | A | |
| TW200618564A | Taiwan Province of China | A | |
| TW200623737A | Taiwan Province of China | A | |
| KR20070013330A | Republic of Korea | A | |
| KR20070014200A | Republic of Korea | A | |
| KR20070014201A | Republic of Korea | A | |
| EP1751955A1 | European Patent Office (EPO) | A1 | |
| EP1751956A2 | European Patent Office (EPO) | A2 | |
| EP1751987A1 | European Patent Office (EPO) | A1 | |
| MXPA06013186A | Mexico | A | |
| MXPA06013193A | Mexico | A | |
| EP1757027A1 | European Patent Office (EPO) | A1 | |
| KR20070023731A | Republic of Korea | A | |
| MXPA06013210A | Mexico | A | |
| MXPA06013211A | Mexico | A | |
| CN1969562A | China | A | |
| CN1973515A | China | A | |
| CN1977516A | China | A | |
| CN1985477A | China | A | |
| BRPI0510952A | Brazil | A | |
| BRPI0510953A | Brazil | A | |
| BRPI0510961A | Brazil | A | |
| BRPI0510962A | Brazil | A | |
| JP2007537681A | Japan | A | |
| JP2007537682A | Japan | A | |
| JP2007537683A | Japan | A | |
| JP2007537684A | Japan | A | |
| KR20080084866A | Republic of Korea | A | |
| KR100870215B1 | Republic of Korea | B1 | |
| KR100871305B1 | Republic of Korea | B1 | |
| EP1757027B1 | European Patent Office (EPO) | B1 | |
| ATE417436T1 | Austria | T1 | |
| DE602005011611D1 | Germany | D1 | |
| EP1751955B1 | European Patent Office (EPO) | B1 | |
| ATE426988T1 | Austria | T1 | |
| KR20090039809A | Republic of Korea | A | |
| ES2318495T3 | Spain | T3 | |
| DE602005013517D1 | Germany | D1 | |
| ES2323011T3 | Spain | T3 | |
| KR100906586B1 | Republic of Korea | B1 | |
| KR100918596B1 | Republic of Korea | B1 | |
| MY139431A | Malaysia | A | |
| JP4361585B2 | Japan | B2 | |
| JP4448171B2 | Japan | B2 | |
| MY141497A | Malaysia | A | |
| EP2182734A1 | European Patent Office (EPO) | A1 | |
| EP2214412A2 | European Patent Office (EPO) | A2 | |
| JP4554680B2 | Japan | B2 | |
| EP1751987B1 | European Patent Office (EPO) | B1 | |
| ATE484157T1 | Austria | T1 | |
| MY142161A | Malaysia | A | |
| DE602005023983D1 | Germany | D1 | |
| CN1977516B | China | B | |
| EP2262304A1 | European Patent Office (EPO) | A1 | |
| ES2354079T3 | Spain | T3 | |
| EP1751956B1 | European Patent Office (EPO) | B1 | |
| ATE508567T1 | Austria | T1 | |
| DE602005027837D1 | Germany | D1 | |
| KR101049701B1 | Republic of Korea | B1 | |
| JP2011142616A | Japan | A | |
| CN1969562B | China | B | |
| KR101068055B1 | Republic of Korea | B1 | |
| ES2366192T3 | Spain | T3 | |
| TWI353759BThis record | Taiwan Province of China | B | |
| TW201145943A | Taiwan Province of China | A | |
| US8089948B2 | United States of America | B2 | |
| CA2566125C | Canada | C | |
| EP2262304B1 | European Patent Office (EPO) | B1 | |
| CN1985477B | China | B | |
| EP2214412A3 | European Patent Office (EPO) | A3 | |
| TWI381681B | Taiwan Province of China | B | |
| CN1973515B | China | B | |
| CN102984133A | China | A | |
| TWI394407B | Taiwan Province of China | B | |
| EP2592836A1 | European Patent Office (EPO) | A1 | |
| CA2565977C | Canada | C | |
| JP5356360B2 | Japan | B2 | |
| EP2182734B1 | European Patent Office (EPO) | B1 | |
| CA2566124C | Canada | C | |
| US8855059B2 | United States of America | B2 | |
| US2014362740A1 | United States of America | A1 | |
| US2015016427A1 | United States of America | A1 | |
| CA2771943C | Canada | C | |
| CN102984133B | China | B | |
| US9674732B2 | United States of America | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Annulment or lapse of patent due to non-payment of feesLapsedMM4A | MM4A |
Numbers
- Publication
- I353759
- Publication, DOCDB
- I353759
- Publication, EPODOC
- TWI353759B
- Application
- 94115718
- Application, DOCDB
- 94115718
- Application, EPODOC
- TW20050115718
Titles5
- Chinese
- 在無線通信系統中音頻及視頻資料的同步
- English
- SYNCHRONIZATION OF AUDIO AND VIDEO DATA IN A WIRELESS COMMUNICATION SYSTEM
- English
- Synchronization of audio and video data in wireless communication systems
- Unlabeled
- 在無線通信系統中音頻及視頻資料的同步
- Unlabeled
- Synchronization of audio and video data in wireless communication systems
Classification
- CPC, 36
- H04L69/04
- H04W28/06
- H04N21/2381
- H04N21/41407
- H04N21/44004
- H04N21/4788
- H04N21/6131
- H04N21/6181
- H04N21/6437
- H04N21/64707
- H04W28/065
- H04W72/1263
- H04W80/00
- H04W84/04
- H04W88/181
- H04L65/80
- H04L69/166
- H04L69/22
- H04L69/161
- H04N19/102
- H04N19/115
- H04N19/61
- H04N19/124
- H04N19/152
- H04N19/164
- H04N19/174
- H04L69/321
- H04L47/36
- H04W4/06
- H04L65/764
- H04L65/00
- H04L9/40
- H04L65/75
- H04L65/1101
- H04W72/044
- H04W88/02
- IPC, 12
- H04L29 06
- H04N7 26
- H04B7 00
- H04B7 216
- H04L12 28
- H04L12 56
- H04L12 66
- H04L47 36
- H04W28 06
- H04W72 12
- H04W84 04
- H04W88 18