Delivery of information over a communication channel
Abstract
Methods and devices for transmitting information units over multiple constant bit rate communication channels are described. This technique involves encoding a unit of information, thereby creating multiple data packets. The data packet size is encoded so that it matches the physical layer packet size of the communication channel. The information unit may include a variable bit rate data stream, multimedia data, and audio data. Communication channels include CDMA channels, WCDMA, GSM channels and EDGE channels.

Term
Term ended
Projected expiry passed 13 May 2025, 1.4 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
71 claims: 8 independent, 63 dependent
- 1無線通信システムにおいて情報を送信する方法において、 複数の利用可能な固定ビットレート通信チャネルの可能な物理層パケットサイズを決定することと、 分割が、前記複数の利用可能な固定ビットレート通信チャネルにより供給される前記利用可能な物理層パケットサイズの少なくとも1つの物理層パケットサイズを超えないような大きさであるように情報単位を分割するための制約を確立することと、を備えた方法。
- 2情報を分割することは、変化するサイズの分割を発生することができるレート制御モジュールを備えたソースエンコーダーを備えた、請求項1に記載の方法。
- 3前記情報単位は可変ビットレートデータストリームを備えた、請求項1に記載の方法。
- 4前記情報単位はマルチメディアデータを備えた、請求項1に記載の方法。
- 5前記情報単位はビデオデータを備えた、請求項1に記載の方法。
- 6前記情報単位はオーディオデータを備えた、請求項1に記載の方法。
- 7前記固定ビットレート通信チャネルはCDMAチャネルである、請求項1に記載の方法。
- 8前記固定ビットレート通信チャネルは補足チャネルを含む、請求項7に記載の方法。
- 9前記固定ビットレート通信チャネルは専用制御チャネルを含む、請求項7に記載の方法。
- 10前記固定ビットレート通信チャネルはパケットデータチャネルを含む、請求項7に記載の方法。
- 11前記固定ビットレート通信チャネルはGSMチャネルである、請求項1に記載の方法。
- 12前記固定ビットレート通信チャネルはEDGEチャネルである、請求項1に記載の方法。
- 13前記固定ビットレート通信チャネルはGPRSチャネルである、請求項1に記載の方法。
- 14前記制約は前記情報単位のエンコーディングの期間に使用される、請求項1に記載の方法。
- 15前記情報単位は固定間隔で生じる、請求項1に記載の方法。
- 16無線通信システムにおいて情報を送信する方法において、 複数の利用可能な固定ビットレート通信チャネルの利用可能な物理層パケットサイズを決定することと、 情報単位をデータパケットにエンコードすることであって、個々のデータパケットサイズは、前記利用可能な固定ビットレート通信チャネルの前記物理層パケットサイズの1つを超えないように選択されることと、を備えた方法。
- 17エンコーディング情報は、変化するサイズの分割を発生することができるレート制御モジュールを備えたソースエンコーダーを備えた、請求項16に記載の方法。
- 18前記情報単位はマルチメディアストリームを備えた、請求項16に記載の方法。
- 19前記情報単位はビデオデータを備えた、請求項16に記載の方法。
- 20前記情報単位はオーディオデータを備えた、請求項16に記載の方法。
- 21前記固定ビットレート通信チャネルはCDMAチャネルである、請求項16に記載の方法。
- 22前記固定ビットレート通信チャネルは補足チャネルを含む、請求項21に記載の方法。
- 23前記固定ビットレート通信チャネルは専用制御チャネルを含む、請求項21に記載の方法。
- 24前記固定ビットレート通信チャネルはパケットデータチャネルを含む、請求項21に記載の方法。
- 25前記固定ビットレート通信チャネルはGSMチャネルである、請求項16に記載の方法。
- 26前記固定ビットレート通信チャネルはEDGEチャネルである、請求項16に記載の方法。
- 27前記固定ビットレート通信チャネルはGPRSチャネルである、請求項16に記載の方法。
- 28前記情報単位は固定間隔で生じる、請求項16に記載の方法。
- 29複数の固定ビットレート通信チャネルに対応するように構成された受信機と、 前記受信した複数の固定ビットレート通信チャネルに対応するように構成され、前記固定ビットレートチャネルをデコードするように構成されたデコーダーであって、前記デコードされた固定ビットレートチャネルは累算されてデータの可変ビットレートストリームを生成するデコーダーと、を備えた無線通信装置。
- 30前記デコーダーは、前記通信チャネルから受信されたデータパケットのサイズを推定する、請求項29に記載の無線通信装置。
- 31前記通信チャネルから受信されたデータパケットのサイズは、付加信号伝達において示される、請求項29に記載の無線通信装置。
- 32前記可変ビットレートストリームはマルチメディアストリームである、請求項29に記載の無線通信装置。
- 33前記可変ビットレートストリームはビデオデータを備える、請求項29に記載の無線通信装置。
- 34前記可変ビットレートストリームはオーディオデータを備える、請求項29に記載の無線通信装置。
- 35前記複数の固定ビットレートチャネルはCDMAチャネルである、請求項29に記載の無線通信装置。
- 36前記複数の固定ビットレートチャネルはGSMチャネルである、請求項29に記載の無線通信装置。
- 37前記複数のコテイビットレートチャネルはGPRSチャネルである、請求項29に記載の無線通信装置。
- 38前記複数の固定ビットレートチャネルはEDGEチャネルである、請求項29に記載の無線通信装置。
- 39複数の利用可能な固定ビットレート通信チャネルから物理層パケットサイズのセットを決定するように構成されたコントローラーと、 情報単位をデータパケットに分割するように構成されたエンコーダーであって、個々のデータパケットサイズは前記利用可能な固定ビットレート通信チャネルの前記物理層パケットのサイズの少なくとも1つを超えないように選択される、エンコーダーと、を備えた無線通信装置。
- 40前記エンコーダーは、変化するサイズの分割を発生することができるレート制御モジュールをさらに備えた、請求項39に記載の無線通信装置。
- 41前記物理層パケットを送信するように構成された送信器をさらに備えた、請求項39に記載の無線通信装置。
- 42前記情報単位は、可変ビットレートストリームを備えた、請求項39に記載の無線通信装置。
- 43前記情報単位はマルチメディアデータを備えた、請求項39に記載の無線通信装置。
- 44前記情報単位はビデオデータを備えた、請求項39に記載の無線通信装置。
- 45前記複数の固定ビットレートチャネルはCDMAチャネルである、請求項39に記載の無線通信装置。
- 46前記複数の固定ビットレートチャネルはGSMチャネルである、請求項39に記載の無線通信装置。
- 47前記複数の固定ビットレートチャネルはGPRSチャネルである、請求項39に記載の無線通信装置。
- 48前記複数の固定ビットレートチャネルはEDGEチャネルである、請求項39に記載の無線通信装置。
- 49無線通信システムにおけるエンコーダーにおいて、前記エンコーダーは、情報単位に対応するように構成され、前記情報単位をデータパケットに分割するように構成され、前記データパケットは、利用可能な固定ビットレート通信チャネルの少なくとも1つの物理層パケットサイズを超えないような大きさであるエンコーダー。
- 50前記情報単位は固定レートで生じる、請求項49に記載のエンコーダー。
- 51前記情報単位は可変レートデータストリームを備えた、請求項49に記載の無線通信装置。
- 52前記情報単位はマルチメディアデータを備えた、請求項49に記載のエンコーダー。
- 53前記情報単位はビデオデータを備える、請求項49に記載のエンコーダー。
- 54前記情報単位はオーディオデータを備える、請求項49に記載のエンコーダー。
- 55前記固定ビットレート通信チャネルはCDMAチャネルである、請求項49に記載のエンコーダー。
- 56前記固定ビットレート通信チャネルはGSMチャネルである、請求項49に記載のエンコーダー。
- 57前記固定ビットレート通信チャネルはGPRSチャネルである、請求項49に記載のエンコーダー。
- 58前記固定ビットレート通信チャネルはEDGEチャネルである、請求項49に記載のエンコーダー。
- 59前記パケットの集合は、あらかじめ選択されたビットの最大数に限定されるように前記エンコーダーが制約される、請求項46に記載のエンコーダー。
- 60無線通信システムにおけるデコーダーにおいて、前記デコーダーは、複数の固定ビットレート通信チャネルからのデータストリームに対応するように構成され、前記データストリームをデコードし、前記デコードされた複数のデータストリームを可変ビットレートデータストリームに累算するように構成されたデコーダー。
- 61前記通信チャネルから受信されたデータパケットのサイズが推定される、請求項60に記載のデコーダー。
- 62前記通信チャネルから受信されたデータパケットのサイズは、付加信号伝達で示される、請求項60に記載のデコーダー。
- 63前記可変ビットレートストリームはマルチメディアストリームである、請求項60に記載のデコーダー。
- 64前記可変ビットレートストリームはビデオストリームである、請求項60に記載のデコーダー。
- 65前記可変ビットレートストリームはオーディオストリームである、請求項60に記載のデコーダー。
- 66前記固定ビットレート通信チャネルはCDMAチャネルである、請求項60に記載のデコーダー。
- 67前記固定ビットレート通信チャネルはGSMチャネルである、請求項60に記載のデコーダー。
- 68前記固定ビットレート通信チャネルはGPRSチャネルである、請求項60に記載のデコーダー。
- 69前記固定ビットレート通信チャネルはEDGEチャネルである、請求項60に記載のデコーダー。
- 70データをエンコードする方法を具現化するコンピューター読み取り可能媒体において、 前記方法は、情報単位を分割し、それにより複数のデータパケットを作成し、各データパケットは、利用可能な固定ビットレート通信チャネルに対応する物理層パケットサイズのセットから少なくとも1つの物理層パケットサイズのサイズを超えないような大きさであるコンピューター読み取り可能媒体。
- 71ブロードキャストコンテンツをデコードする方法を具現化するコンピューター読み取り可能媒体において、前記方法は、 複数の固定ビットレート通信チャネルからのデータストリームを受け入れることと、 前記デーストリームをデコードして、前記デコードされた複数のデータストリームを可変ビットレートデータストリームに累算することと、を備えたコンピューター読み取り可能媒体。
Independent claims71
156 paragraphs, as filed
Description of related application
Priority Claim under 35 U.SC § 119 This patent application is expressly incorporated herein by reference to a co-pending patent application, filed May 13, 2004, "CDMA". Claims priority over US Provisional Application No. 60 / 571,673 entitled "Multimedia Packets Carried by CDMA Physical Layer Product".
This patent application is related to the following co-pending US patent applications: "Communication" with agent reference number 030166U2, filed at the same time as this specification, transferred to the assignee of this specification and expressly incorporated herein by reference in its entirety. Method And MFP For Allocation Of Information To Channels Of A Communication System.
And with agent reference number 030166U3, filed at the same time as this specification, transferred to the transferee of this specification and expressly incorporated herein by reference in its entirety. Header Compression Of Multimedia Data Transmitted Over A Wireless Communication System.
And with agent reference number 030166U4, filed at the same time as this specification, transferred to the transferee of this specification and expressly incorporated herein by reference in its entirety. Synchronization Of Audio And Video Data In A Wireless Communication System.
The present invention generally relates to the distribution of information via a communication system, and particularly to the division of information units in order to match physical layer packets of a constant bit rate communication link.
The demand for distribution of various communication networks is increasing. For example, consumers want to deliver video over various communication channels such as the Internet, wireline networks and wireless networks. Multimedia data can be in different formats and data rates. Various communication networks use different mechanisms for the transmission of real-time data over their respective communication channels.
One type of communication network that has become common is mobile wireless networks for wireless communication. Wireless communication systems have many applications, including, for example, mobile phones, paging, wireless local loops, personal digital assistance (PDAs), Internet telephones, and satellite communication systems. A particularly important application is the form telephone system for mobile subscribers. As used herein, the term "cellular" includes both cellular and personal communications services (PCS) frequencies. Various wireless interfaces have been developed for form telephone systems including frequency division multiple access (FDMA), time division multiple access (TDMA) and code division multiple access (CDMA).
Different national and international standards have been established, such as Advanced Modular Telephone Services (AMPS), Global Systems for Mobile (GSM), General Packet Radio Service (GPRS), Enhanced Data GSM Environment (EDGE). , Interim Standard 95 (IS-95), and its derivatives, IS-95A, IS-95B, ANSI J-STD-008 (often collectively referred to here collectively as IS-95), and various wireless interfaces. To support. And high data rate systems such as cdma2000, Universal Mobile Telecommunications Services (UMTS) and Broadband CDMA (WCDMA) have emerged. These standards are recommended by the Telecommunications Machinery Manufacturers Association (TIA), the Third Generation Partnership Project (3GPP), and the European Telecommunications Standards Association.
Users or customers of mobile wireless networks, such as mobile phone networks, want to receive streaming media such as video, multimedia, and Internet Protocol (IP) over wireless communication links. For example, a customer wants to be able to receive video such as video conferencing or television broadcasts on a mobile phone or other portable wireless communication device. Other examples of the types of data that customers want to receive using wireless communication devices include multimedia multicast / broadcast and Internet access.
There are different types of multimedia sources and different types of communication channels for which streaming data is desired to be transmitted. For example, multimedia data sources can generate data at constant bit rate (CBR) or variable bit rate (VBR). In addition, the communication channel can transmit data in CBR or VBR. Table 1 below lists the various combinations of data sources and communication channels.<tables num="1"><img file="JP2007537682A_D0001.tif" /></tables>
Communication channels typically transmit data in chunks called physical layer packets or physical layer frames. The data generated by the multimedia source may be a stream of contiguous bytes, such as a mu-law or A-law encoded audio signal. More complicatedly, the data generated by the multimedia source resides in a group of bytes called data packets. For example, an MPEG-4 video encoder compresses visual information as a sequence of information units, referred to here as video frames. Visual information typically must be encoded by the encoder at a fixed video frame rate of 25Hz or 30Hz and rendered by the decoder at the same rate. The video frame period is the time between two video frames and can be calculated as the reciprocal of the video frame rate. For example, a 40ms video frame period corresponds to a 25Hz video frame rate. Each video frame is encoded into a variable number of data packets, and all data packets are sent to the decoder. If a portion of the data packet is lost, the packet becomes unusable by the decoder. The decoder, on the other hand, may reconstruct the video frame even if some of the data packets are lost, but at the expense of the quality degradation in the resulting sequence. Therefore, each data packet contains a portion of the description of the video frame, and therefore the number of packets is variable from one video frame to another.
If the source produces data at a fixed bit rate and the communication channel sends data at a fixed rate, the communication system resource assumes that the communication channel data rate is at least the rate of the source data rate, or so. If not, if the two data rates match, it will be used efficiently. In other words, if the fixed data rate of the source is the same as the fixed data rate of the channel, the resources of the channel are fully utilized and the source data can be transmitted without delay. Similarly, if the source produces data at a variable rate and the channel sends at a variable rate, then the two data rates can match, as long as the channel data rate can support the source data rate, in this case. As before, the channel source is fully utilized and all source data can be transmitted without delay.
If the source produces data at a fixed data rate and the channel is a variable data rate channel, channel resources may not be used as efficiently as possible. For example, in this mismatched case, the statistical multiplexing gain (SMG) is less than the SMG compared to the CBR source on the matched CBR channel. Statistical multiplexing gain occurs when the same communication channel can be used or multiplexed among multiple users. For example, when transmitting voice using a communication channel, the speaker usually does not speak continuously. That is, there will be a "talk" eruption from the speaker, followed by silence (listening). If the ratio of time for "talk" eruption to silence is, for example 1: 1, it is possible to multiplex the same communication channel on average, or to support two users. However, if the data source has a fixed data rate and is delivered via a variable rate channel, there is no SMG. This is because there is no time for the communication channel to be used by other users. That is, there is no interruption during the "silence" period for the CBR source.
In the last case shown in Table 1 above, the source of the multimedia data is a variable bit rate stream such as a multimedia stream such as video, and a constant bit rate such as a wireless communication channel with a fixed bit rate allocation. It is a situation where it is transmitted via a communication channel having. In this case, the delay is typically introduced between the source and the communication channel, creating a "spout" of data so that the communication channel can be used efficiently. In other words, the variable rate data stream is stored in a buffer and is delayed for a long enough time. Therefore, the buffer output can be emptied at a fixed data rate to match the channel fixed data rate. The buffer, the data must be well stored or delayed, so a fixed output can be maintained without buffering "empty". Therefore, the CBR communication channel is fully utilized and the resources of the communication channel are not wasted.
The encoder periodically generates video frames according to the video frame period. A video frame is composed of data packets, and the total amount of data in the video frame is variable. The video decoder must provide video frames at the same video frame rate used by the encoder to ensure an acceptable result for the viewer.
Transmission of video frames with variable amounts of data at a fixed video frame rate over a fixed rate communication channel can result in inefficiencies. For example, the total amount is much too large for the data in the video frame, if not be transmitted within the video frame period at the bit rate of the channel, the decoder, all in time to render the entire frame according to the video frame rate Body May not receive frames. In practice, traffic shaping buffers are used to smooth out such large changes for delivery over fixed rate channels. If the fixed video frame rate is to be maintained by the decoder, this introduces a delay in serving the video.
Another problem is that if data from multiple video frames is contained in a single physical layer packet, the loss of the single physical layer packet will result in degradation of the multiple video frames. Loss of one physical layer packet can result in degradation of multiple video frames, even in situations where the data packet is close to the physical layer packet size.
Therefore, there is a need for techniques and devices that can improve the transmission of variable data rate multimedia data over fixed data rate channels.
Outline of the invention
The embodiments disclosed herein address the above-mentioned need by providing methods and devices for transmitting information units over a constant bit rate communication channel. This technique involves splitting an information unit into data packets. In this case, the size of the data packet is chosen to match the physical layer data packet size of the communication channel. For example, the number of bytes contained in each information unit may vary over time and with respect to the number of bytes that each physical layer data packet that the communication channel can hold may vary independently. Good. Described that this technique divides information units and thereby creates multiple data packets. For example, the encoder may be constrained to encode a unit of information into a data packet of a size that does not exceed or "matches" the physical layer packet size of the communication channel. The data packet is then assigned to the physical layer data packet of the communication channel.
The phrase "multimedia frame" for video is used here to mean a video frame that can be displayed / drawn on a display device after being decoded. The video frame can be further divided into independently decodable units. In video terms, these are called "slices." In the case of audio and speech, the term "multimedia frame" is used here to mean information within a time window in which the speech or audio is compressed for transport and for decoding at the receiver. Will be done. The phrase "information unit interval" is used herein to describe the time duration of the multimedia frames described above. For example, in the case of video, the information unit interval is 100 milliseconds for video at 10 frames per second. Further, as an example, in the case of speech, the information unit interval is typically 20 milliseconds in cdma2000, GSM, and WCDMA. From this description, it must be clear that typically audio / speech frames are not further divided into independently decodable units, and typically video frames are further divided into independently decodable slices. Must be. It must be clear from the context when phrases such as "multimedia frame", "information unit spacing" refer to multimedia data in video, audio and speech.
This technology can be a Global System for Mobile Communications (GSM), General Line Radio Service (GPRS), Enhanced Data GSM Environment (EDGE), or TIA / EIA-95-B (IS-95), TIA / EIA. It can be used with a variety of wireless interfaces such as -98-C (IS-98), IS-2000, HRPD, Broadband CDMA (WCDMA), and other CDMA-based standards.
Perspectives include determining the possible physical layer packet size of at least one available constant bit rate communication channel. The information unit is divided, thereby creating multiple data packets so that the size of each data packet does not exceed or match at least one of the physical layer packets of the constant bit rate communication channel. The data packet is then encoded and assigned to the physical layer packet on the matched constant bit rate communication channel. Encoding the information can include a source encoder with a rate control module that can generate variable size divisions.
Using the techniques described, the information unit is encoded into a stream of data packets transmitted over one or more fixed bit rate channels. Since the units of information vary in size, they may be encoded into data packets of different sizes. Data packets may be transmitted using different combinations of constant bit rate channels with different available physical layer packet sizes. For example, the information unit may contain video data contained in video frames of different sizes, so different combinations of constant bit rate communication channel physical layer packets may be selected and adapted for transmission of video frames of different sizes. May be good.
Other aspects include determining the physical layer packet size and available data rates for multiple constant bit rate communication channels. The information unit is then assigned to the data packet. In this case, the individual data packet size is chosen to fit one physical layer packet on each constant bit rate communication channel. The combination of individual constant bit rate channels may be selected so that the physical layer packet size matches the variable bit rate stream packet size. For example, one or more different combinations of constant bit rate channels may be selected depending on the variable bit rate data stream.
Another aspect is an encoder configured to accept information units. The information unit is then divided into data packets. In this case, the size of the individual data packets does not exceed or match the size of one physical layer packet for the available constant bit rate communication channels.
Another aspect is a decoder configured to accept data streams from multiple constant bit rate communication channels. The data stream is decoded and the decoded data stream is accumulated in the variable bit rate data stream.
Examples of constant bit rate communication channels are GSM, GPRS, EDGE, or TIA / EIA-95-B (IS-95), TIA / EIA-98-C (IS-98), IS-2000, HRPD, and broadband. Includes CDMA-based standards such as CDMA (WCDMA).
Other features and advantages of the invention will become apparent from the following description of the exemplary embodiments that illustrate the viewpoints of the invention, as an example.
The term "exemplary" is used herein to mean act as an example, instance, or exemplification. Any embodiment described herein as "exemplary" is not necessarily construed as suitable or advantageous over other embodiments.
The term "streaming" here refers to the real-time delivery of virtually continuous multimedia data such as audio, speech, or video information over dedicated and shared channels in interactive unicast and broadcast applications. Used to mean. The phrase "multimedia frame" for video is used here to mean a video frame that can be displayed / drawn on a display device after decoding. Video frames can be further divided into independently decodable units. In video terms, these are called "slices." In the case of audio and speech, the term "multimedia frame" is used here to mean information within a time window in which the speech or audio is compressed for transport and for decoding at the receiver. Will be done. The phrase "information unit interval" is used herein to describe the time duration of the multimedia frames described above. For example, in the case of video, the information unit interval is 100 milliseconds for video at 10 frames per second. Further, as an example, in the case of speech, the information unit interval is typically 20 milliseconds in cdma2000, GSM and WCDMA. From this description, it should be clear that typically audio / speech frames are not further divided into independently decodable units, and typically video frames are further divided into independently decodable slices. Must be. It must be clear from the context when phrases such as "multimedia frame", "information unit spacing" refer to multimedia data in video, audio and speech.
Techniques for transmitting information units over multiple constant bit rate communication channels are described. This technique involves splitting an information unit into data packets. The size of the data packet is chosen to match the physical layer data packet of the communication channel. For example, the information unit may occur at a fixed rate. The communication channel may then transmit physical layer packets at different rates. It is described that this technique divides an information unit and thereby generates a plurality of data packets. For example, the encoder may be constrained to encode the unit of information to a size that matches the physical layer packet size of the communication channel. The encoded data packet is then assigned to the physical layer data packet of the communication channel. The information unit may include a variable bit rate data stream, multimedia data, video data and audio data. Communication channels are GSM, GPRS, EDGE, or TIA / EIA-95-B (IS-95), TIA / EIA-) 98-C (IS-98), IS2000, HRPD, cdma2000, Broadband CDMA (WCDMA) and others. Includes CDMA based standards such as.
Perspectives include determining the possible physical layer packet size of at least one available constant bit rate communication channel. The information unit is divided, thereby creating multiple data packets so that the size of each data packet matches one of at least one physical layer packet on a constant bit rate communication channel. The data packet is then encoded and assigned to the physical layer packet of the matching constant bit rate communication channel. In this way, the unit of information is encoded into a stream of data packets transmitted over one or more constant bit rate channels. Since the units of information vary, they may be encoded into data packets of different sizes. Data packets may then be transmitted using different combinations of constant bit rate channels with different available physical layer packet sizes. For example, the information unit may include video data contained in frames of different sizes. Then, different combinations of constant bit rate communication channel physical layer packets may be selected to accommodate the transmission of video frames of different sizes.
Other aspects include determining the physical layer packet size and available data rates for multiple constant bit rate communication channels. The information unit is then assigned to the data packet. In this case, the individual data packet size is chosen to fit one physical layer packet on each constant bit rate communication channel. The combination of individual constant bit rate channels may be selected so that the physical layer packet size matches the variable bit rate data stream packet size. For example, one or more different combinations of constant bit rate channels may be selected depending on the variable bit rate data stream.
Another aspect is an encoder configured to accept information units. The information unit is then divided into data packets. In this case, the size of the individual data packets matches the physical layer packet size of one of the available constant bit rate communication channels.
Another aspect is a decoder configured to accept data streams from multiple constant bit rate communication channels. The data stream is decoded and the decoded data stream is accumulated in the variable bit rate data stream.
Examples of information units include variable bit rate data streams, multimedia data, video data, and audio data. Information units may occur at a fixed iteration rate.
For example, the information unit may be a frame of video data. Examples of constant bit rate communication channels include CDMA channels, GSM channels GPRS channels and EDGE channels.
Examples of protocols and formats for transmitting information units such as variable bit rate data, multimedia data, video data, speech data or audio data from sources on content servers or wired networks to mobile are also provided. The techniques described are applicable to any type of multimedia application such as unicast streaming, interactive and broadcast streaming applications. For example, this technique can be used to transmit multimedia data. For example, video data (from a content server on wired streaming to wireless mobile), as well as other multimedia applications, such as broadcast / multicast services, or audio and interactive services such as video calls between two mobiles.
FIG. 1 shows a communication system 100 constructed according to the present invention. Communication system 100 includes infrastructure 101, multiple wireless communication devices (WCD) 104 and 105, and terrestrial communication line communication devices 122 and 124. WCDs are also called mobile stations (MS) or mobile. In general, WCDs may be mobile or fixed. Terrestrial communication devices 122 and 124 can include serving nodes or content servers that supply various types of multimedia data, such as streaming data. In addition, MSs can send streaming data such as multimedia data.
Infrastructure 101 may also include other components such as base station 102, base station controller 106, mobile switching center 108, switching network, and the like. In one embodiment, the base station 102 is integrated with the base station controller 106, and in other embodiments, the base station 102 and the base station controller 106 are separate components. Different types of switching networks 120 may be used to send signals to communication system 100, such as IP networks or public switched telephone networks (PSTNs).
The term "forward link" or "downlink" refers to the signal path from infrastructure 101 to MS, and the term "reverse link" or "uplink" refers to the signal path from MS to infrastructure. As shown in FIG. 1, MSs 104 and 105 receive signals 132 and 136 on the forward link and transmit signals 134 and 138 over the reverse link. Generally, the signals transmitted from the MS 104 and 105 are intended to be received by other remote devices or other communication devices such as ground warfare communication devices 122, 124 and sent over an IP network or switching network. There is. For example, if the signal 134 transmitted from the initiating WCD104 is intended to be received by the destination MS105, the signal is sent over infrastructure 101 and the signal 136 is sent over the forward link to the destination MS105. Will be done. Similarly, signals initiated in Infrastructure 101 may be broadcast to MS105. For example, the content provider may send multimedia data, such as streaming multimedia data, to the MS105. Typically, a communication device such as an MS or terrestrial communication device may be both a signal initiator and a signal destination.
Examples of MS104 include mobile phones, wirelessly communicable personal computers, personal digital assistants (PDAs) and other wireless devices. Communication system 100 may be designed to support one or more radio standards. For example, the standards are Global System for Mobile Communications (GSM), General Line Radio Service (GPRS), Enhanced Data GSM Environment (EDGE), TIA / EIA-95-B (IS-95), TIA / EIA. It may include standards called -98-C (IS-98), IS2000, HRPD, cdma2000, Broadband CDMA (WCDMA), and others.
FIG. 2 is a block diagram illustrating various air interface options for delivering packet data over an exemplary packet data network and a wireless network. The techniques described may be implemented in packet-switched data network 200 as illustrated in FIG. As shown in the example of FIG. 2, the packet-switched data network system may include a radio channel 202, a plurality of receiving nodes or MS204s, a transmitting node or a content tuning server 206, a serving node 208 and a controller 210. The transmitting node 206 may be connected to the serving node 208 via a network 212 such as the Internet.
The serving node 208 may include, for example, a packet data serving node (PDSN) or a serving GPRS support node (SGSN) and a gateway GPRS support node (GGSN). The serving node 200 may receive packet data from the transmitting node 206, or may supply a packet of information to the controller 210. The controller 210 may include, for example, a base station controller / packet control function (BSC / PCF) or a wireless network controller (RNC). In one embodiment, the controller 210 communicates with the serving node 208 via a radio access network (RAN). The controller 210 communicates with the serving node 208 and sends a packet of information over the radio channel 202 to at least one of the receiving nodes, such as the MS.
In one embodiment, the serving node 200 and / or transmitting node 206 may include an encoder for encoding the data stream and / or a decoder for decoding the data stream. For example, an encoder may encode a video stream to generate frames of variable size data, and a decoder may receive and decode frames of variable size data. Frames vary in size, but the video frame rate is constant, so a variable bitrate stream of data is generated. Similarly, the MS may include an encoder for encoding the data stream and / or a decoder for receiving the data stream. The term "codec" is used to describe a combination of encoder and decoder.
In one example illustrated in Figure 2, data such as multimedia data from a transmitting node 206 connected to a network or the Internet, serving node, packet data serving node (PDSN) 206, controller, or base station controller / packet control function. It can be sent to the receiving node or MS204 via (BSC / PCF) 208. The radio channel 202 interface between the MS204 and the BSC / PCF210 is an air interface and can typically use many channels for signaling, bearers, or payloads, data.
Air Interface Air Interface 202 may operate according to any of a number of wireless standards. For example, the standard is a TDMA-based standard such as Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Enhanced Data GSM Environment (EDGE), or TIA / EIA-95-B. CDMA-based standards such as (IS-95), TIA / EIA-98-C (IS-98), IS2000, HRPD, cdma2000, Broadband CDMA (WCDMA), and others may be included.
In systems based on cdma2000, the data is multichannel, such as basic channel (FCH), dedicated control channel (DCCH), supplemental channel (SCH), and packet data commonly used to transmit voice. It can be transmitted over a channel (PDCH) as well as other channels.
The FCH provides communication channels for the transmission of speech at multiple fixed rates, such as full rate, half rate, 1/4 rate and 1/8 rate. FCH provides these rates and if the user's speech behavior requires less than the full rate to achieve the target voice quality, the system will be used by other users in the system with one of the lower data rates. To reduce interference with. The advantage of lowering the source to increase system capacity is well known in CDMA networks.
DCCH is similar to FCH, but provides only full-rate traffic at one of two fixed rates: 9.6kbps (RC3) in wireless configuration 3 and 14.4kbps (RC5) in wireless configuration 5. This is called the Ix traffic rate.
SCH can be configured to provide traffic rates on cdma2000 at 1x, 2x, 4x, 8x and 16x. Both DCCH and SCH can abort transmission when there is no data to transmit. That is, it does not transmit any data and, also called dtx, can be guaranteed to reduce interference with other users in the system or stay within the transmitted power of the base station transmitter. PDCH can be configured to send data packets n * 45, but n = {1, 2, 4, 8}.
FCH and DCCH channels provide fixed delay and low data packet loss for data communication, for example to enable interactive services. The SCH and PDCH channels provide constant bit rate channels that provide higher bandwidth than FCH and DCCH, eg, 300 kbps to 3 Mbps. Also, since SCH and PDCH are shared among many users, they have variable delay. In the case of SCH, multiple users are multiplexed in time. This captures different delays depending on the system load. For PDCH, bandwidth and delay depend on, for example, radio conditions, agreed quality of service (QoS), and other scheduling considerations. Similar channels are available in systems based on TIA / EIA-95-B (IS-95), TIA / EIA-98-C (IS-98), IS2000, HRPD, UMTS, and Wideband CDMA (WCDMA). ..
It is noteworthy that the FCH provides multiple fixed bit data rates (full rate, half rate, 1/4 rate and 1/8 rate) to store the power requested by the voice user. Typically, voice encoders or vocoders will use lower data rates when the time-frequency structure of the transmitted signal allows for higher compression without excessive degradation. This technique is commonly referred to as source-controlled variable bit rate bocoding. Therefore, it can be used to transmit data in systems based on TIA / EIA-95-B (IS-95), TIA / EIA-98-C (IS-98), IS2000, HRPD, UMTS, or cdma2000. There are multiple fixed bit rate channels.
In CDMA-based systems such as cdma2000, communication channels are divided into continuous streams of "slots". For example, the communication channel may be divided into 20ms segments or time slots. This is also called the "transmission time interval" (TTI). The data transmitted during these time slots is assembled into packets. In this case, the size of the data packet depends on the available data rate or bandwidth of the channel. Therefore, it is possible that there are individual data packets being transmitted over their respective communication channels during any individual time slot period. For example, during a single time slot period, data packets may be sent over the DCCH channel, or different data packets may be sent over the SCH channel at the same time.
Similarly, in systems based on GSM, or GPRS or EDGE, data can be transmitted between BSC208 and MS204 using multiple time slots within a frame. FIG. 3 is a block diagram illustrating two radio frames 302 and 304 in a GSM air interface. As shown in FIG. 3, the GSM air interface radio frames 302 and 304 are each divided into eight time slots. Individual time slots are assigned to special users in the system. In addition, GSM transmission and reception use two different frequencies. Also, forward and reverse links are offset by three time slots. For example, in FIG. 3, the downlink radio frame 302 will start at time t0 and be transmitted on one frequency. The uplink radio frame 304 will then be transmitted at different frequencies. The downlink radio frame 302 is offset by three time slots TS0-TS2 from the uplink radio frame. Having an offset between the downlink and uplink radio frames allows the radio communication device, or terminal, to operate without the need to be able to transmit and receive at the same time.
Advances in GSM radios or terminals have resulted in GSMs capable of receiving multiple time slots during the same radio frame period. These are referred to as "multi-slot classes" and can be found in Appendix B of 3GPP TS 45.002, which are incorporated herein by reference in their entirety. Therefore, in systems based on GSM, GPRS, or EDGE, there are multiple fixed time slots available to transmit data.
VBR Multimedia Characteristics Variable bit rate (VBR) multimedia data, such as video, usually contains common characteristics. For example, video data is typically captured in a fixed frame by a sensor such as a camera. Multimedia transmitters generally require a finite processing time with an upper limit to encode a video stream. Multimedia receivers generally require a finite amount of processing time to decode a video stream.
It is generally desirable to reconstruct multimedia frames at the same frame rate at which multimedia frames are generated. For example, in the case of video, it is desirable to display the reconstructed video frame at the same rate as the video was captured by the sensor or camera. Having the same reconstruction and capture rate makes it easier to synchronize with other multimedia elements. For example, synchronizing a video stream with attached audio, speech, and streams is simplified.
For video, it is usually desirable to maintain a consistent level of quality from the perspective of human perception. In general, it is more cumbersome and painful for a person to process a continuous multimedia stream of varying quality than to process a multimedia stream of consistent quality. For example, processing a video stream containing quality artifacts such as still images or blockiness is usually a hassle.
Delay considerations Transporting multimedia content, such as audio / video, typically results in delays. Some of these delays are due to code settings, and some are due, among other things, to network settings such as Radio Link Protocol (RLP) transmissions that allow retransmissions and sorting transmitted via air interfaces and the like. An objective methodology for assessing multimedia transmission delays is to observe the encoded stream. For example, a transmission cannot be decoded until a completely independently decodeable packet is received. Therefore, the delay may be affected by the size of the packet and the rate of transmission.
For example, if the packet size is 64 Kbytes and is sent over a channel of 64 Kbytes per second, the packet cannot be decoded and must be delayed by 1 second before the entire packet is received. All packets received will need to be delayed sufficiently to fit the largest packet. This allows packets to be decoded at a fixed rate. For example, if a video packet or resizing is sent, the receiver must delay or buffer all received packets in an amount equal to the delay required to fit the maximum packet size. There will be. The delay will allow the decoded video to be drawn or displayed at a fixed rate. If the maximum packet size is not known in advance, the maximum packet size and correlation delay estimation can be made based on the parameters used during the packet encoding period.
The techniques described now can be used to assert delays for any video codec (H.263, AVC / H.264, MPEG-4, etc.). In addition, if only video decoders are specified as standard by the Motion Picture Experts Group (MPEG) and the International Telecommunication Union (ITU), they will be introduced by implementing different encoders for mobile in typical wireless arrangements. It is useful to have an objective measure that can be used to estimate the delay to be made.
Video streams will generally have more latency than other types of data in multimedia services. It will have more delay than, for example, voice, audio, timed text, etc. Other multimedia data that needs to be synchronized with the video data due to the longer delay typically experienced by the video stream usually needs to be intentionally delayed to stay in sync with the video. There will be.
Encoder / decoder delay In some multimedia encoding techniques, multimedia data frames are encoded or decoded using information from previous reference multimedia data frames. For example, a video codec that implements the MPEG-4 standard will encode and decode different types of video frames. In MPEG-4, video is typically encoded in "I" and "P" frames. The I-frame is built-in. That is, an I-frame contains all the information needed to draw or display a single complete video frame. P-frames are not built-in and typically contain differential information related to previous frames, such as movement vectors and different texture information. Typically, I-frames are about 8 to 10 times larger than P-frames, depending on the content and encoder settings. The encoding and decoding of multimedia data introduces delays that may depend on available processing resources. A typical implementation of that type of scheme utilizes a ping-pong buffer to allow processing resources to capture or display one frame at a time and process another.
Video encoders such as H.263, AVC / H.264, MPEG-4, etc. are essentially variable rates. This is due to the use of predictive coding and variable length coding (VLC) of many parameters. Real-time delivery of variable-rate bitstreams over circuit-switched and packet-switched networks is typically accomplished by traffic shaping with buffers at the transmitter and receiver. Traffic shaping buffers typically introduce additional delays that are not desirable. For example, if there is a delay between when one person speaks and when another person hears a speech, further delays during a video conference can be frustrating.
Encoder and decoder delays can affect the amount of time an encoder and decoder must process multimedia data. For example, the upper limit of the time that allows the encoder and decoder to process the data and maintain the desired frame rate is given by:<maths num="1"><img file="JP2007537682A_D0002.tif" /></maths>
However, Δe and Δd represent the encoder delay and the decoder delay, respectively. And f is the desired frame rate in frames per second (fps) for a given service.
For example, video data typically has a desired frame rate of 15 fps, 10 fps, or 7.5 fps. The upper limits of the time allowed for the encoder and decoder to process the data and maintain the desired frame rate result in 66.6ms, 100ms, and 133ms for frame rates of 15fps, 10fps, and 7.5fps, respectively.
Rate Control Buffer Delay In general, different numbers of bits are required for different frames to maintain consistent perceptual quality for multimedia services. For example, a video codec may need to use a different number of bytes to encode an I-frame rather than a P-frame to maintain consistent quality. Therefore, maintaining consistent quality and fixed frame rates results in a video stream becoming a variable bit rate stream. Consistent quality in the encoder can be achieved by setting the encoder "quantization parameter" (Qp) to a fixed value or approximately below the variable of the target value Qp.
FIG. 4 is a chart illustrating an example of changes in frame size for a typical video sequence entitled "Car Phone". Car phones are standard video sequences well known to those of skill in the art and are used to supply "common" video sequences used to evaluate various techniques such as video compression, error correction, and transmission. Will be done. Figure 4 shows the frame size in bytes for the number of sample frames of mobile phone data encoded using the MPEG-4 and AVC / H.264 encoding techniques shown by criteria 402 and 404, respectively. An example of change is shown. The desired quality of encoding can be achieved by setting the encoder parameter "Qp" to the desired value. In FIG. 4, the car phone is encoded using an MPEG encoder with Qp = 33 and AVC / H.264 with Qp = 33. When the encoded data stream illustrated in Figure 4 is transmitted over a fixed bit rate (CBR) like a typical radio channel, frame size changes have a fixed or negotiated QoS bit rate. It will need to be "flattened" to maintain. Typically, the "smoothing" of time variation at this frame size results in the introduction of an additional delay, commonly referred to as the buffering delay Δ &.
FIG. 5 is a block diagram showing how a buffering delay can be used to support the transmission of variable size frames transmitted over a CBR channel. As shown in FIG. 5, changing size 502 data frames are input to buffer 504. Buffer 504 will store a sufficient number of frames of data so that it can output data frames of fixed size from buffer 506 for transmission over CBR channel 508. This type of buffer is commonly referred to as a "leakable bucket" buffer. A "leakable bucket" buffer outputs data at a fixed rate, much like a bucket with a hole in the bottom. If the rate at which water enters the bucket changes, the bucket must hold a sufficient amount of water in the bucket and must prevent the bucket from drying out when the rate of water entering the bucket falls below the leaking rate. is there. Similarly, the bucket must be large enough to prevent the bucket from overflowing when the rate of water entering the bucket exceeds the rate of leakage. The buffer 504 operates in a manner similar to a bucket, and the amount of data the buffer needs to store to prevent underflow results in a delay corresponding to the length of time the data stays in the buffer.
FIG. 6 is a graph illustrating the buffering delay introduced by streaming a variable bit rate (VBR) multimedia stream through the CBR channel in the system of FIG. As illustrated in Figure 6, the video signal is encoded using the VBR encoding scheme, MPEG4, to produce a VBR stream. The number of bytes in the VBR stream is illustrated in FIG. 6 by line 602, which represents the cumulative or total number of bytes required to transmit a given number of video frames. In this example, the MPEG-4 stream is encoded at an average bit rate of 64 kbps and transmitted over the 64 kbps CBR channel. The number of bytes transmitted by the CBR channel is represented by a fixed slope line 604 corresponding to a fixed transmission rate of 64 kbps.
The display or playback 606 in the decoder needs to be delayed to avoid underflowing the buffer in the decoder due to insufficient data received in the decoder that allows decoding of full video frames. In this example, the delay is 10 frames or 1 second for the desired display rate of 10 fps. In this example, a fixed rate of 64 kbps is used for the channel, but if an MPEG-4 stream with an average data rate of 64 kbps is transmitted over the 32 kbps CBR channel, the buffering delay increases with the length of the sequence. Will do. For example, for the 50-frame sequence illustrated in Figure 6, the buffering delay increases to 2 seconds.
In general, the buffering delay Δb due to the buffer flow constraint can be calculated as follows.<maths num="2"><img file="JP2007537682A_D0003.tif" /></maths>
However, B (i) is the buffer occupancy period (video frame #i) in the encoder in bytes at time i.
R (i) is the encoder output in bytes at time i C (i) can be transmitted in one frame period f is the desired number of frames per second BW (i) is available at time i Bandwidth CBR transmission special case<maths num="3"><img file="JP2007537682A_D0004.tif" /></maths>
It should be noted that
To avoid underflowing the decoder buffer, playback must be delayed by the time required to transmit the maximum buffer occupancy in the encoder during the entire display period. Therefore, the buffering delay can be expressed as follows.<maths num="4"><img file="JP2007537682A_D0005.tif" /></maths>
The denominator in Equation 5 represents the average data rate for the entire session in period I. For CBR channel allocation, the denominator is C. Using the above analysis, the nominal encoder buffer size required to avoid overflow in the encoder can be estimated by calculating max {Be (i)} for all i in the set of model sequences. ..
Examples of MPEG-4 and AVC / H.264 buffer delays Figure 7 shows milliseconds for various 50-frame sequenced video clips encoded using a nominal rate of 64 kbps and constant Qp for AVC / H.264 and MPEG-4. It is a bar graph which illustrates the buffer delay Δb of a second. As shown in FIG. 7, the MPEG-4 frame sequence of FIG. 6 is represented by a bar 702 indicating a buffer delay of 1000 ms. The same video sequence encoded using AVC / H.264 is represented by bar 704, which indicates a buffer delay of 400 ms. A further example of a 50-frame sequence of video clips is shown in Figure 7. In this case, the buffer delay associated with each encoded sequence using both MPEG-4 and AVC / H.264 is shown.
FIG. 8 is a bar graph illustrating video quality represented by the peak signal-to-noise ratio (PSNR) of the sequence illustrated in FIG. As shown in FIG. 8, a car phone sequence encoded using MPEG-4 with Qp = 15 is represented by a bar 802 showing a PSNR of about 28 dB. The same sequence encoded using AVC / H.264 with QP = 33 is indicated by bar 804, which shows a PSNR of about 35 dB.
Transmission channel delay Transmission delay Δt depends on the number of retransmissions used and a time constant for a given network. It can be assumed that Δt has a nominal value when retransmissions are not used. For example, it may be assumed that Δt has a nominal value of 40 ms when retransmission is not used. If retransmissions are used, the frame erasure rate (FER) will decrease but the delay will increase. The delay will depend, at least in part, on the number of retransmissions and the associated overhead delay.
Error Elasticity Consideration When transmitting RTP flows over a wireless link or channel, there will generally be some residual packet loss. This is because RTP streams are delay sensitive and it is impractical to guarantee 100% reliable transmission by means of retransmission protocols such as RLP or RLC. To aid in understanding the effects of channel errors, descriptions of various protocols such as the RTP / UDP / IP protocol are provided below. FIG. 9 illustrates the various levels of encapsulation that exist when transmitting multimedia data such as video data over wireless links using the RTP / UDP / IP protocol.
As shown in FIG. 9, the video codec generates a payload 902 containing information describing the video frame. Payload 902 may consist of several video packets (not drawn). Payload 902 contains Slice_Header (SH) 904. Therefore, the application layer data packet 905 is composed of the video data 902 and the associated Slice_Header 904. Since the payload travels through networks such as the Internet, additional header information may be added. For example, a real-time protocol (RTP) header 906, a user datagram protocol (UDP) header 908 and an internet protocol (IP) header 910 may be added. These headers provide the information used to send from the source to the destination.
Upon entering the wireless network, a point-to-point protocol (PPP) header 912 is added to provide framing information for ordering packets into a stream of contiguous bits. The wireless link protocol, eg, RLP in cdma2000 or RL in W-CDMA, packs a stream of bits into RLP packet 914. Among other things, the wireless link protocol allows the retransmission and sorting of packets transmitted over the air interface. Finally, the air interface MAC layer takes one or more RLP packets 914, packs them into MUX layer packets 916, and adds a multiplexing header (MUX) 918. The physical layer channel coder then adds a checksum (CRC) 920 to detect decoding errors and tail 922 forming the physical layer packet 925.
The continuous, uncoordinated encapsulation illustrated in FIG. 9 has some consequences in the transmission of multimedia data. One such result is that there may be a discrepancy between the application layer data packet 905 and the physical layer packet 925. As a result of this discrepancy, every time a physical layer packet 925 containing a portion of one or more application layer packets 925 is lost, the corresponding entire application layer 905 is lost. Losing one physical layer packet 925 can result in loss of the entire application layer packet 905, as part of a single application layer data packet 905 may be contained in two or more physical layer data packets 925. There is sex. This is because the entire application layer data packet 905 needs to be properly decoded. Another result is that if a portion of two or more application layer data packets 905 is contained in a physical layer data packet, then a loss of a single physical layer data packet 925 can result in a loss of two or more application layer data packets 905. There is sex.
FIG. 10 illustrates an example of generally allocating an application data packet 905, such as a multimedia data packet, to a physical layer data packet 925. Shown in Figure 10 are two application data packets 1002 and 1004. The application data packet can be a multimedia data packet. For example, each data packet 1002 and 1004 can represent a video frame. The disjointed encapsulation illustrated in FIG. 10 can result from a single application data packet or a physical layer packet with data from two or more application data packets. As shown in FIG. 10, the first physical layer data packet 1006 can contain data from a single application layer packet 1002, and the second physical layer data packet 1008 is two or more application data packets 1002. And data from 1004 can be included. In this example, if the first physical layer data packet 1006 is "lost" or is corrupted during transmission, then a single application layer data packet 1002 is lost. On the other hand, if the second physical layer packet 1008 is lost, then the two application data packets 1002 and 1004 are also lost.
For example, if the application layer data packet is two consecutive video frames, the loss of the first physical layer data packet 1006 results in the loss of a single video frame. However, the loss of the second physical layer data packet results in the loss of both video frames. This is because some of both video frames are lost and neither video frame can be properly decoded or recovered by the decoder.
Explicit Bit Rate (EBR) Control The use of a technique called Explicit Bit Rate Control (EBR), CBR or VBR, can improve the transmission of VBR sources over the CBR channel. In EBR, the information unit is divided into data packets so that the size of the data packet matches the size of the available physical layer packet. For example, a VBR stream of data, such as video data, may be split into data packets so that the physical layer data packets and application layer data packets of the communication channel on which the data is transported match. For example, in EBR, any such as GSM, GPRS, EDGE, TIA / EIA-95-B (IS-95), TIA / EIA-98-C (IS-98), cdma2000, wideband CDMA (WCDMA) and others. The encoder is constrained to output bytes at time i (previously displayed R (i)) that matches the "capacity" of the physical channel used to deliver the data stream in the wireless standard of It may be configured. In addition, encoded packets may be constrained. Thereby, the encoder generates a data packet having the same number of bytes as or less than the size of the physical layer data packet of the communication channel. In addition, the encoder can be constrained so that each application layer data packet output by the encoder can be decoded independently. EBR technology simulations for AVC / H.264 reference encoders are a perceptible loss in quality when the encoder is constrained according to EBR technology, provided that an appropriate number of explicit rates are used to constrain the VBR encoding. Indicates that there is no. Examples of constraints for some channels are given below as examples.
Multimedia encoding and decoding As mentioned above, multimedia encoders, such as video encoders, may generate variable sized multimedia frames. For example, in some compression techniques, each new multimedia frame may contain all the information needed to fully render the frame content, while other frames may be from previously fully rendered content. May include information about changes to the content of. For example, as mentioned above, in systems based on MPEG-4 compression technology, video frames may typically be of two types. That is, an I frame or a P frame. Each I-frame is built-in, similar to a JPEG file, in that it contains all the information needed to draw or display one complete frame. In contrast, the P-frame typically contains information related to the previous frame, such as difference information related to the previous frame and motion vector. Therefore, since P-frames depend on previous frames, P-frames are not built-in and cannot draw or display a complete frame without a dependency on the previous frame. In other words, P-frames cannot be self-decoded. Here, the term "decoded" is used to mean completely reconstructed to display a frame. Typically, I-frames are larger than P-frames. For example, it is about 8 to 10 times larger depending on the content and encoder settings.
Generally, each frame of data can be divided into parts or "slices". For example, each slice can be independently decoded as described further below. In one case, the frame of data may be contained in a single slice, and in the other case, the frame of data is divided into multiple slices. For example, if the frame of data is video information, the video frame may be contained within an independently decodable slice. Alternatively, the frame may be divided into two or more independently decodeable slices. In one embodiment, each encoded slice is configured such that the size of the slice matches the available size of the communication channel physical layer data packet. If the encoder encodes the video information, each slice is configured so that each of the video slices matches the available size of the physical layer packet. In other words, the frame slice size matches the physical layer packet size.
The advantage of slicing a size that matches the available communication channel physical layer data size is that there is a one-to-one correspondence between application packets and physical layer data packets. This helps alleviate some of the problems associated with disjointed encapsulation illustrated in Figure 10. Therefore, if the physical layer data packet is corrupted or lost during transmission, only the corresponding slice is lost. Also, if each slice of the frame can be decoded independently, the loss of the frame slice will not prevent decoding of the other slices of the frame. For example, if a video frame is divided into 5 slices so that each slice can be decoded independently and matches the physical layer data packet, then one deterioration or loss of the physical layer data packet is only the corresponding slice. Will cause a loss. Then, the physically layer packet transmitted successfully can be successfully decoded. Therefore, the entire video frame may not be decoded, but some may be decoded. In this example, four of the five video slices are successfully decoded, thereby allowing the video frame to be drawn or displayed despite the reduced performance.
For example, in a cdma2000 based system using DCCH and SCH channels, if the video slices are communicated from the transmitting node to the MS, the video slices will be sized to match these available channels. .. As mentioned above, DCCH channels can be configured to support multiple fixed data rates. For example, a system based on cdma2000 can support data transmission rates of 9.60 kbps or 14.4 kbps depending on the selected rates (RS), RS1 and RS2, respectively. The SCH channel can also be configured to support multiple fixed data rates, depending on the SCH radio configuration (RC). SCH supports multiples of 9.6kbps when configured as RC3 and multiples of 14.4kbps when configured as RC5. The SCH data rates are:<maths num="5"><img file="JP2007537682A_D0006.tif" /></maths>
However, n is 1, 2, 4, 8, or 16 depending on the channel configuration.
Table 2 below illustrates the possible physical layer data packet sizes for DCCH and SCH channels in a communication system based on cdma2000. The first column identifies the case or possible configuration. The second and third columns are DCCH rate sets and SCH radio configurations, respectively. The fourth column has four entries. The first is the dtx case where no data is transmitted over the DCCH or SCH. The second is the physical layer data packet size of the 20ms time slot for the DCCH channel. The third entry is the physical layer data packet size of the 20ms time slot for the SCH channel. The fourth entry is the physical layer data packet size of the 20ms time slot for the combination of DCCH and SCH channels.<tables num="2"><img file="JP2007537682A_D0007.tif" /></tables>
Keep in mind that there are trade-offs to consider when application layer data packets are too large to fit DCCH physical layer data packets or SCH physical layer data packets and instead combined DCCH plus SCH packets are used. There is a need to. The trade-off in deciding to encode the application layer data packet to fit the combined DCCH plus SCH data packet size, as opposed to the application layer data packet size making two packets, is Larger application layer packets, or slices, generally produce better compression efficiency, and smaller slices generally produce better error tolerance. For example, larger slices generally require less overhead. Referring to FIG. 9, each slice 902 has its own slice header 904. Therefore, if two slices are used instead of one, there are two slice headers attached to the payload, which requires more data to encode the packet, thereby reducing compression efficiency. If, on the other hand, two slices are used, one is transmitted over DCCH and the other is transmitted over SCH, then only one of the DCCH data packets or SCH data packets is degraded or lost in the other data packets. Will allow recovery. Therefore, it improves error tolerance.
To help you understand Table 2, the derivations of Case 1 and Case 9 will be explained in detail. In Case 1, the DCCH is configured as an RSI corresponding to a data rate of 9.6kbps. Since the channel is divided into 20ms time slots, the amount of data or physical layer packet size that can be sent over the DCCH configured as RSI within each time slot is 9600 bits / sec * 20 msec = 192 bits. = 24 bytes ... (Equation 7).
Only 20 bytes are available for application layer data packets, including slices and slice headers, due to additional overhead added to physical layer packets, such as RLP for error correction. Therefore, in case 1, the first entry in the fourth column of Table 2 is 20.
The SCH for Case 1 is configured as 2x in RC #. RC3 corresponds to a base data rate of 9.6kbps, and 2X means that the channel data rate is twice the base data rate. Therefore, within each time slot, the amount of data or physical layer packet size that can be transmitted on the SCH configured as 2x in RC3 is 2 * 9600 bits / sec * 20 msec = 384 bits = 48 bytes ... (Equation 8). Here, only 40 bytes are available for the application layer data packet, including slices and slice headers, due to the additional overhead added to the physical layer packet. Therefore, in case 1, the second entry in the fourth column of Table 2 is 40. The third entry in the fourth column of Table 2 for Case 1 is the sum of the first and second entries, or 60.
Case 9 is similar to Case 1. In both cases, the DCCH is configured as an RSI corresponding to a physical layer packet size of 20 bytes. The SCH channel in Case 9 is configured as 2X on RC5. RC5 corresponds to a base data rate of 14.4kbps, and 2X means that the channel data rate is twice the base data rate. Therefore, within each time slot, the amount of data or physical layer packet size that can be transmitted on the SCH configured as 2X in RC5 is as follows.
2 * 14 400 bits / sec * 20 msec = 576 bits = 72 bytes ... (Equation 9) Here, for the application layer data packet including slices and slice headers, due to the additional overhead added to the physical layer packet. Only 64 bytes are available. Therefore, in case 9, the second input in the fourth column of Table 2 is 64. In case 9, the third entry in the fourth column of Table 2 is the sum of the first and second entries, or 84.
Other entries in Table 2 are determined in a similar manner. In this case, RS2 corresponds to a DCCH with a data rate of 14.4 kbps, which corresponds to 36 bytes in a 20 msec time slot, 31 of which are available to the application layer. It should be noted that there is a dtx operation available for all cases, which has a zero payload size and no data is sent to any channel. When user data can be transmitted below the available physical layer slots (20 ms each), dtx is used in the next slot to reduce interference with other users in the system.
As illustrated in Table 2 above, by configuring the available fixed data rate channels, such as DCCH and SCH, a set of CBR channels can behave similarly to VBR channels. .. That is, by configuring a plurality of fixed rate channels, it is possible to create a CBR channel that operates as a pseudo VBR channel. Techniques that utilize pseudo-VBR channels determine the possible physical layer data packet size corresponding to the bit rate of the CBR channel from multiple available fixed bit rate communication channels and encode the fixed bit rate stream of data. It involves creating multiple data packets such that each size of the data packet matches one size of the physical layer data packet size.
In one embodiment, the configuration of the communication channel is established at the beginning of the session and then remains unchanged or rarely changes throughout the session. For example, the SCH discussed in the example above is typically set to a configuration and stays in that configuration throughout the session. That is, the described SCH is a fixed rate SCH. In other embodiments, the channel configuration can be changed dynamically during the session. For example, the variable rate SCH (V-SCH) can change its configuration for each time slot. That is, during one time slot, the V-SCH can be configured in one configuration, such as 2xRC3, and in the next time slot, the V-SCH can be configured in a different configuration, such as 16xRC3, or any other. Can be configured in the V-SCH configuration of. V-SCH provides additional flexibility and can improve system performance in EBR technology.
If the communication channel configuration is fixed for the entire session, application layer packets or slices are selected so that they fit into one of the available physical layer data packets. For example, if DCCH and SCH are configured as RSI and 2xRC3, as illustrated in Case 1 in Table 2, the application layer slices should fit 0-byte, 20-byte, 40-byte, or 60-byte packets. Will be selected. Similarly, if the channels are configured as RS1 and 16xRC3, as illustrated in Case 4 of Table 2, the application layer slices are selected to fit 0-byte, 20-byte, 320-byte or 340-byte packets. Will be. If V-SCH channels are used, it is possible to change between two different configurations on a slice-by-slice basis. If DCCH is configured as RSI and V-SCH is configured as RC3, change between V-SCH configurations 2xRC3, 4xRC3, 8xRC3 or 16xRC3, corresponding to Cases 1-4 in Table 2. It is possible. As illustrated in Case 1-4 of Table 2, the choices between these various configurations are 0 bytes, 20 bytes, 40 bytes, 60 bytes, 80 bytes, 100 bytes, 160 bytes, 180 bytes, 320 bytes. , Or provides a 340-byte physical layer data packet. Therefore, in this example, using the V-SCH channel allows the application layer slice to be selected to fit any of the 10 different physical layer data packet sizes listed in Cases 1-4 of Table 2. To. In the case of cdma2000, the size of the delivered data is estimated by MS and this process is called "blind detection".
Similar techniques can be used in wideband CDMA (WCDMA) using data channels (DCH). DCH, similar to V-SCH, supports different physical layer packet sizes. For example, DCH can support rates from 0 to nx in multiples of 40 octets. In this case, "nx" corresponds to the maximum allocated rate of the DCH channel. Typical values for nx include 64kbps, 128kbps and 256kbps. In a technique called "explicit display", the size of the data delivered can be indicated using additional signaling, thereby eliminating the need for blind detection. For example, in the case of WCDMA, the size of the delivered data packet may be displayed using the "Transport Format Combination Indicator" (TFCI), so the MS does not need to perform blind detection. This reduces the computational load on the MS when variable size packets are used, as in EBR. The described EBR concept is applicable to both blind detection and explicit display of packet size.
By selecting the application layer data packet so that the application layer data packet fits the physical layer data packet, the combination of fixed bit rate communication channels with total data rate is similar to and in some cases VBR communication channels. It can transmit VBR data streams with better performance than VBR communication channels. In one embodiment, the variable bit rate data stream is encoded into a stream of data packets that are sized to match the physical layer data packet size of the available communication channels. In other embodiments, the bit rate of the variable bit rate data stream varies, so that the bit rate of the variable bit rate data stream may be encoded into data packets of different sizes. Data packets may then be transmitted using different combinations of constant bit rate channels.
For example, different frames of video data may be of different sizes, so different combinations of constant bit rate communication channels may be selected to fit the transmission of video frames of different sizes. In other words, variable bitrate data is a constant bit rate data by assigning a data packet to at least one of the constant bit rate communication channels so that the total bit rate of the constant bit rate communication channel matches the bit rate of the variable bit rate stream. It can be transmitted efficiently via the rate channel.
Another aspect is that the encoder can be constrained to limit the total number of bits used to represent a variable bit rate data stream to the maximum number of preselected bits. That is, if the variable bit rate data stream is a frame of multimedia data such as video, the frame may be split into slices. In this case, each slice can be decoded independently, and the slice is selected so that the number of bits in the slice is limited to the number of bits selected in advance. For example, if the DCCH and SCH channels are configured as RSI and 2xRC3, respectively (in Case 1 of Table 2), the slice can be constrained to not be larger than 20 bytes, 40 bytes or 60 bytes, respectively.
In other embodiments, using EBR to transmit multimedia data can use cdma2000 packet data channel (PDCH). PDCH can be configured to send data packets that are n * 45 bytes. However, n = {1, 2, 4, 8}. Again, using PDCH for multimedia data, such as video data, can be divided into "slices" that match the available physical layer packet size. In cdma2000, PDCH has different data rates available for forward PDCH (F-PDCH) and reverse PDCH (R-PDCH). In cdma2000, F = PDCH has slightly less available bandwidth than R-PDCH. Although this difference in bandwidth is available, in some cases it is advantageous to limit the R-PDCH to the same bandwidth as the F-PDCH. For example, if the first MS sends the video stream to the second MS, the video stream is sent on the R-PDCH by the first MS and received on the F-PDCH by the second MS. If the first MS used the entire bandwidth of the R-PDCH, some of the data streams would have had to be removed to match the bandwidth of the F-PDCH transmission to the second MS. Let's go. To alleviate the difficulties associated with reformating the transmission from the first MS so that the transmission can be transmitted to the second MS on the channel with the smaller bandwidth, R- The bandwidth of PDCH can be limited to be the same as F-PDCH. One way to limit the F-PDCH bandwidth is to limit the application data packet size sent over the R-PDCH to the packet size supported by the F-PDCH and the rest in the R-PDCH physical layer packet. Is to add a "stuffing bit" to the bit of. In other words, the stuffing bit is an F-PDCH data packet.
Using the techniques just described, Table 3 shows the possible physical layer data packet sizes for F-PDCH and R-PDCH for four possible data rate cases, one for each value of n. And list the number of "stuffing bits" that will be added to the R-PDCH.<tables num="3"><img file="JP2007537682A_D0008.tif" /></tables>
When a multimedia stream such as a video stream is split into slices, such as EBR with DCCH plus SCH, smaller slice sizes generally improve error tolerance, but may jeopardize compression efficiency. Absent. Similarly, if larger slices are used, there will generally be an increase in compression efficiency. However, system performance may be degraded by lost packets. This is because the loss of individual packets results in the loss of more data.
Similarly, techniques for matching multimedia data, such as video slices, to the available size of physical layer packets can be performed in systems based on other radio standards. For example, in a system based on GSM, or GPRS, or EDGE, multimedia frames such as video slices can be sized to match the available time slots. As mentioned above, many GSM, GPRS and EDGE devices can receive multiple time slots. Therefore, depending on the number of time slots available, the frame-encoded stream can be constrained, and thus the video slice can be matched to a physical packet. In other words, like the GSM time slot, multimedia data can be encoded so that the packet size matches the available size of the physical layer packet, and the total data rate of the physical layer packet used is multimedia data. Supports data rates.
EBR Performance Consideration As mentioned above, when an encoder for a multimedia data stream operates in EBR mode, the encoder produces multimedia slices that match the physical layer. Therefore, there is no loss of compression efficiency compared to true VBR mode. For example, a video codec that operates according to EBR technology produces video slices that match the particular physical layer to which the video is transmitted. In addition, it has the advantages of error tolerance, lower latency, and lower transmission overhead. Details of these benefits are further described below.
Performance at Channel Error As mentioned with reference to Figure 10, it can be seen that in general encapsulation, the loss of physical layer packets may result in the loss of two or more application layers. In EBR technology, each physical packet loss within a wireless link results in exactly one application layer packet loss.
Figure 11 illustrates an example of encoding an application layer packet according to EBR technology. As mentioned above, application layer packets may be of various sizes. As mentioned in Tables 2 and 3, physical layer packets may be of various sizes. For example, the physical layer may consist of channels that use different sized physical layer data packets. In the example of FIG. 11, four application packets 1102, 1104, 1106, and 1108 and four physical layer packets 1110, 1112, 1114, 1116 are illustrated. Three different examples of matching application tier packets to physical tier packets are illustrated. First, a single application tier packet can be encoded so that a single application tier packet is sent within multiple physical tier packets. In the example shown in FIG. 11, a single physical layer packet 1102 is encoded into two physical layer packets 1110 and 1112. For example, if DCCH and SCH are configured as RSI and 2xRC3 respectively (Case 1 in Table 2) and the application data packet is 60 bytes, then the application data packet is via two physical layer packets that correspond to the DCCH and SCH packet combination. It is also possible to send it. It is envisioned that a single application layer packet can be encoded into any number of physical layer packets corresponding to the available communication channels. The second example shown in FIG. 11 is that a single application layer packet 1104 is encoded into a single physical layer packet 1104. For example, if the application layer data packet is 40 bytes, the application layer data packet can be transmitted using only the SCH object ideal data in Case 1 of Table 2. In both of these examples, the loss of a single physical layer packet results in the loss of only a single application layer packet.
The third example shown in FIG. 11 is when multiple application layer packets can be encoded into a single physical layer packet 1116. In the example shown in FIG. 11, the two application layers 1106 and 1108 are encoded and transmitted in a single physical layer packet. It is envisioned that three or more application layer packets may be encoded to fit within a single physical layer packet. The drawback of this example is that the loss of a single physical layer packet 1116 will result in the loss of multiple application layer packets 1106 and 1108. However, there may be trade-offs such as full use of the physical layer that will ensure that multiple application layers transmitted within a single physical layer packet are encoded.
FIG. 12 is a block diagram showing an embodiment of a codec that transmits a VBR data stream via IP / UDP / RTP such as the Internet. As shown in FIG. 12, the codec generates a payload, or application layer data packet 1202 containing slice 1204 and slice header 1206. Application layer 1202 passes through the network. In this case, the IP / UDP / RTP header information 1208 is added to the application layer data packet 1202. The packet then passes through the wireless network. In this case, the RLP header 1210 and the MUX header 1212 are added to the packet. Since the sizes of IP / UDP / RTP headers 1208, RLP headers 1210, and MUX headers 1214 are well known, the codec chooses the size for slice 1203. The slice and all associated headers are thereby fitted to the physical layer data packet or payload 1216.
Figure 13 shows various examples of encoded video sequences using a true VBR transmit channel and using EBR transmit utilizing DCCH plus SCH and PDCH when the channel packet loss is 1%. It is a bar graph which shows the relative decrease in the peak signal-to-noise ratio (PSNR). The video sequence shown in Figure 13 is a standard video sequence well known to those of skill in the art and is a "common" video used to evaluate various techniques such as video compression, error correction, and transmission. Used to provide a sequence. As shown in FIG. 13, the true VBR1302 sequence has the maximum PSNR reduction, followed by EBR with PDCH1306, followed by EBR with DCCH plus SCH1304. For example, in a car phone sequence, the true VBR1302 sequence received a reduction of about 1.5 dB in PSNR, and EBR1306 with PDCH and EBR1304 with DCCH and SCH had a reduction of about 0.8 dB and 0.4 dB in PSNR, respectively. I received it. FIG. 13 shows that when the transmit channel experiences 1% packet loss, the distortion measured by the PSNR for the VBR sequence is more severe than for the EBR sequence.
Similar to FIG. 13, FIG. 14 shows a channel loss of 5% for various examples of standard encoded video sequences with true VBR1402, EBR1404 with DCCH plus SCH, and EBR1406 with PDCH. When is, it is a bar graph showing the relative decrease in the peak signal to noise ratio (PSNR). As shown in FIG. 14, the true VBR1402 sequence has the maximum PSNR reduction, followed by EBR1406 with PDCH, followed by EBR1404 with DCCH plus SCH. For example, in a car phone sequence, the true VBR1402 sequence received a reduction in PSNR of about 2.5 dB, and EBR1406 with PDCH and EBR1404 with DCCH plus SCH received a reduction in PSNR of about 1.4 and 0.8 dB, respectively. .. Comparing FIGS. 14 and 13 shows that when transmit channel packet loss increases, the distortion measured by PSNR for VBR sequences is more severe than for EBR sequences.
Figure 15 shows the drawbacks received for the encoded video sequence of Figure 13 using the true VBR1502, EBR1504 with DCCH and SCH, and EBR1506 with PDCH when the channel packet loss is 1%. It is a bar graph showing the percentage of a macroblock with. Figure 16 shows defects received for the encoded video sequence of Figure 14 using the true VBR1602, EBR with DCCH and SCH1604, and EBR1606 with PDCH when the channel packet loss is 5%. It is a bar graph showing the percentage of a macroblock with. A comparison of these graphs shows that in both cases the percentage of defective macroblocks is greater in the VBR sequence than in the EBR sequence. It should be noted that the defective percentage of the slice must be the same as the packet loss rate, as the slice matches the physical layer packet size. However, since slices can contain different numbers of macroblocks, the loss of one data packet corresponding to one slice is different than the loss of different data packets corresponding to different slices containing different numbers of macroblocks. May result in defective macroblocks.
FIG. 17 is a graph showing the rate distortion of one standard encoded video sequence titled "Foreman". As shown in FIG. 17, four different cases showing PSNR vs. bit rate are shown. The first two cases show a video sequence encoded using the VBR1702 and 1704. The next two cases show a video sequence encoded using EBR15. In this case, EBR15 is an EBR with DCCH plus SCH configured as 8x in RS2 and RC5, respectively, as listed in Case 15 of Table 2 above. VBR and EBR data streams are transmitted over "clean" channels 1702 and 1706 and "noisy" channels 1704 and 1708. As mentioned above, there is no packet loss during transmission on clean channels, and noisy channels lose 1% of data packets. As shown in FIG. 17, VBR-encoded sequences transmitted over clean channel 1702 have the highest PSNR for all bit rates. However, EBR15 encoded sequences transmitted over the clean channel 1706 have approximately the same PSNR performance or rate distortion for all bits. Therefore, there is a very small performance degradation between VBR encoding and EBR encoding 15 when the transmit channel is clean. This example shows that in the absence of packet loss during transmission, there can be sufficient accuracy in the EBR encoding configuration to have comparable performance to the true VBR encoding configuration.
When a VBR-encoded sequence is transmitted over the noisy channel 1704, the PSNR drops significantly by more than 3 dB for all bit rates. However, when an encoded sequence of EBRI5 is transmitted over the same noisy channel 1708, its PSNR performance is degraded for all bitrates, but its performance is only degraded by about 1 dB. Therefore, when transmitting over a noisy channel, the PSNR performance of the EBR15 encoded sequence is about 2 dB higher than the VBR encoded sequence transmitted over the same noisy channel. As Figure 17 shows, the rate distortion performance of EBR15 encoding is comparable to VBR encoding in clean channels. And when the channel gets noisy, the rate distortion performance of EBR15 encoding is better than VBR encoding.
FIG. 18 is similar to FIG. 17 and illustrates the rate distortion curve of another encoded video sequence titled "Car Phone". In this case as well, four cases showing the PSNR vs. bit rate are illustrated. The first two cases show video sequences encoded using VBR1802 and 1804. The next two cases show a video sequence encoded using EBR15. In this case, EBR15 is the EBR with DCCH plus SCH configured as 8x in RS2 and RC5, respectively, listed in Case 15 of Table 2 above. VBR and EBR data streams are transmitted over "clean" channels 1802 and 1806 and "noisy" channels 1804 and 1808. In this example, the PSNR performance of the encoded sequence of EBR15 transmitted over clean channel 1806 exceeds the performance of the VBR sequence via clean channel 1802. The PSNR performance of the EBR15 sequence over the noisy channel 1808 exceeds the VBR sequence transmitted over the noisy channel 1804 by about 1.5 dB. In this example, using a car phone sequence on both clean and noisy channels results in the rate distortion performance of the EBR15 encoding, which outperforms the VBR encoding as measured by PSNR.
Latency considerations The use of EBR encoding improves latency performance. For example, using EBR, video slices can be transmitted over radio channels in encoders and decoders without traffic shaping buffers. For real-time services, this is an important advantage as it can enhance the overall user experience.
Due to the variable bit rate (VBR) nature of video encoding to indicate buffering delay, for a typical sequence encoded at an average bit rate of 64 kbps and transmitted over a 64 kbps CBR channel as shown in Figure 6. Consider the transmission plan of. The display represented by curve 608 needs to be delayed to avoid buffer flow in the decoder. In this example, the delay is 10 frames or 1 second for the desired 10 fps display rate.
The delay Δb due to the buffer flow constraint can be calculated as follows.<maths num="6"><img file="JP2007537682A_D0009.tif" /></maths>
However, B (i) is the buffer occupancy in the encoder at the byte in frame i. R (i) is the encoder output in bytes with respect to frame i. C (i) is the byte number f that can be transmitted at the frame interval i. The desired number of frames per second BW (i) is the available bandwidth in bits at frame interval i, in the special case of CBR transmission.<maths num="7"><img file="JP2007537682A_D0010.tif" /></maths>
It should be noted that
Playback must be delayed by the time required to occupy the maximum transmit buffer in the encoder to avoid starvation of the decoder buffer during the entire presentation.<maths num="8"><img file="JP2007537682A_D0011.tif" /></maths>
The denominator above represents the average data rate for the entire session during session period I. For CBR channel allocation, the denominator is C. For the EBR case, if the total channel bandwidth for a given 100-ms period is greater than the frame size, i.e.<maths num="9"><img file="JP2007537682A_D0012.tif" /></maths>
Then there is no buffering delay. Therefore, when the data arrives, the data can be transmitted, resulting in a buffer occupancy of 0 in the encoder.
That is,<maths num="10"><img file="JP2007537682A_D0013.tif" /></maths>
Video frames typically span multiple MAC layer frames K (slots). If it is possible to change C (i) for the K slot so that all of R (i) can be transmitted, then B (i) is 0, so the buffering delay Δb is 0.<maths num="11"><img file="JP2007537682A_D0014.tif" /></maths>
Figure 19 shows an example of a transmission in which a typical EBR stream is encoded at an average rate of 64 kbps. In FIG. 19, frame number vs. cumulative bytes are shown for source 1902, transmission 1904 and multimedia stream display 1906. In the example of FIG. 19, the buffering delay is 0, but there are still delays due to encoding, decoding and transmission. However, these delays are typically much smaller than the VBR buffering delays.
FIG. 20 is a flow chart showing an embodiment of a method of transmitting data. The flow begins at block 2002. The next flow continues to block 2004. In block 2004, the possible physical layer packet size of the available communication channels is determined. For example, if DCCH and SCH channels are used, the configuration of these radio channels will establish the available physical layer packet size, as illustrated in Table 2 above. Then the flow continues to block 2006. In block 2006, a unit of information, eg, one frame of an available bit rate data stream, is received. Examples of variable bit rate data streams include multimedia streams such as video streams. Then the flow continues to block 2008.
In block 2008, the information unit is divided into slices. Splits, or slices, are selected so that their size does not exceed one size of the possible physical layer packet size. For example, the split can be sized so that each of the split sizes is not greater than at least one of the available physical layer packet sizes. The flow then follows block 2010, where the splits are encoded and assigned to physical layer packets. For example, the encoding information can include a source encoder with a rate control module capable of generating variable size splits. Then, in block 2012, it is determined whether all of the frame splits have been encoded and assigned to the physical layer packet. If they do not have a negative result in block 2012, the flow follows block 2010, where the next split is encoded and assigned to the physical layer packet. Returning to block 2012, if all of the frame splits are encoded and assigned to physical layer packets, then block 2012 will have a positive result and the flow will continue to block 2014.
At block 2014, determine if the flow of information ends as it did at the end of the session. If the flow of information is not finished, a negative result will be obtained in block 2014, the flow will follow block 2006, and the next unit of information will be received. Returning to block 2014, if the flow of information ends as in the case of the end of the session, a positive result will be obtained in 2014, the flow will follow block 2016 and the process will stop.
FIG. 21 is a flow diagram illustrating another embodiment of the method of transmitting data. The flow begins at block 2102. The flow then continues to block 2104. At block 2104, the possible physical layer packet size of the available communication channels is determined. For example, if DCCH and SCH channels are used, the configuration of these radio channels will establish the available physical layer packet size, as illustrated in Table 2 above. The flow follows block 2106 and the information unit is received. For example, the information unit may be a multimedia stream, that is, variable bit rate data such as a video stream. The flow then continues to block 2108.
At block 2108, it is determined whether it is desirable to reconfigure the communication channel configuration. If a communication channel that can be reconfigured during the session is used, such as the V-SCH channel, it may be desirable to change the channel configuration during the session. For example, if the frame of data has more data than can be transmitted through the current configuration of the communication channel, change the configuration to a higher bandwidth and the communication channel will support more data. It is desirable to be able to. If it is determined in block 2108 that it is not desirable to reconfigure the communication channel, then block 2108 will have a negative result and the flow will continue to block 2110. In block 2110, the information unit is divided into sizes so that it does not exceed one size of the possible physical layer packet size. Returning to block 2108, if it is determined that it is desirable to reconfigure the communication channel, a positive result will be obtained at block 2108 and the flow will continue to block 2112. At block 2112, the desired physical layer packet size is determined. For example, the received information unit may be parsed and the size of the data packet required to transmit the entire unit may be determined. The flow then continues to block 2114. The desired communication channel configuration is determined at block 2114. For example, different physical layer packet sizes for different configurations of available communication channels can be determined, and configurations with physical layer packets large enough to fit an information unit may be selected. .. The communication channel is then reconfigured accordingly. Next, the flow continues to block 2110. In this case, the information units are divided into sizes so that their size matches one of the possible physical layer packet sizes of the reconfigured communication channel. The flow then continues to block 2116. Block 21 At 16, the split is encoded and assigned to the physical layer data packet. For example, the encoding information can include a source encoder with a rate control module that can generate variable size splits. The flow then continues to block 2118.
At block 2118, it is determined whether all divisions of information units have been encoded and assigned to physical layer packets. If they are not done, a negative result will result in block 2118, the flow will follow block 2110, and the next split will be encoded and assigned to the physical layer packet. Returning to block 2118, if all divisions of the information unit were encoded and assigned to physical layer packets, a positive result would be obtained at block 2118 and the flow would continue to block 2120.
At block 2120, it is determined whether the information flow has ended, as in the case of the end of the session. If the information flow is not finished, a negative result in block 2120, the flow proceeds to block 2106, and the next unit of information is received. If you return to block 2120 and the information flow ends, you will get a positive result in block 2120, the flow will follow block 2122, and the process will stop.
FIG. 22 is a block diagram of a wireless communication device or mobile station (MS) configured according to an exemplary embodiment of the present invention. The communication device 2202 includes a network interface 2206, a codec 2208, a host processor 2210, a memory device 2212, a program product 2214, and a user interface 2216.
Signals from the infrastructure are received by network interface 2206 and sent to the host processor 2210. The host processor 2210 receives the signal according to the content of the signal and responds to the appropriate action. For example, the host processor 2210 may decode the received signal itself or send the received signal to codec 2208 for decoding. In another embodiment, the received signal is transmitted directly from network interface 2206 to the codec.
In one embodiment, the network interface 2206 may be a transceiver and an antenna to interface the infrastructure via a radio channel. In other embodiments, the network interface 2206 may be a network interface card for interfacing the infrastructure over the ground line. Codec 2208 may be a general purpose processor such as a digital signal processor (DSP) or a central processor (CPU).
Both the host processor 2210 and the codec 2208 are connected to the memory device 2212. The memory device 2212 may be used to store data during WCD operation, as well as program code that would be executed by the host processor 2210 or DSP2208. For example, the host processor, codec, or both may operate under the control of programming instructions that are temporarily stored in memory device 2212. Host processor 2210 and codec 2208 can include their own program storage memory. When a programming instruction is executed, the host processor 2210, and / or codec 2208, performs their function, eg, decoding or encoding a multimedia stream. Therefore, since the programming step implements the functionality of each host processor 2210 and codec 2208, each host processor and codec can be made to perform decoding or encoding of the content stream as desired. Programming steps may be received from program product 2214. Program product 2214 may store or transfer programming steps to memory 2212 for execution by the host processor, codec, or both.
Program product 2214 may be a semiconductor memory chip such as RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, as well as hard disk, removable disk, CD-ROM, or computer read. It may be another storage device, such as any other form of technically known storage medium, which may store possible instructions. In addition, program product 2214 may be a source file that contains program steps that are received from the network, stored in memory, and then executed. Thus, the processing steps required for operation in accordance with the present invention may be embodied on program product 2214. In FIG. 22, the exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write the information to the storage medium. Alternatively, the storage medium may be integrated into the processor.
User interface 2216 is connected to both host processor 2210 and codec 2208. For example, user interface 2216 may include a display and speakers used to output multimedia data to the user.
Those skilled in the art will recognize that the steps of the methods described in connection with embodiments may be replaced without departing from the scope of the invention.
Those skilled in the art will appreciate that information and signals may be represented using any of a wide variety of different techniques and techniques. For example, data, instructions, commands, information, signals, bits, symbols and chips that may be referenced throughout the description are represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or particles, or any combination thereof. May be done.
Those skilled in the art will further realize the various exemplary logical blocks, modules, circuits and algorithm steps described in connection with the embodiments disclosed herein as electronic hardware, computer software, or a combination of both. You will understand what can be done. To articulate this compatibility of hardware and software, a variety of exemplary components, blocks, modules, circuits and steps are generally described above in terms of their functionality. Whether such functionality is realized as hardware or software depends on the design constraints imposed on the particular application and the overall system. One of ordinary skill in the art may realize the functionality described in varying ways for each particular application, but decisions of such realization should be construed as causing a deviation from the scope of the invention. is not it.
The various exemplary logic blocks, modules and circuits described in connection with the embodiments disclosed herein include general purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), and field programmable gates. It may be implemented or performed with an array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. .. The general purpose processor may be a microprocessor, but in the alternative the processor may be any conventional processor, controller, microcontroller or state machine. Processors may also be implemented as a combination of computing units, eg, a combination of DSPs and microprocessors, multiple microprocessors, one or more microprocessors coupled with a DSP, or any other combination of such configurations. May be good. The steps of the method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known technically. .. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write the information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and storage medium may reside in the ASIC. The ASIC may reside on the user terminal. Alternatively, the processor and storage medium may reside as separate components within the user terminal.
The above description of the disclosed embodiments is provided to allow one of ordinary skill in the art to make or use the invention. Various changes to these embodiments will be readily apparent to those skilled in the art, and the integrated principles defined herein may be applied to other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not intended to be limited to the embodiments presented herein, and the broadest scope consistent with the principles and novel features disclosed herein should be tolerated.
<figref num="1">FIG. 1 is an explanatory diagram of a part of a communication system 100 constructed according to the present invention.</figref><figref num="2">FIG. 2 is a block diagram illustrating various wireless air interface options for delivering packet data over an exemplary packet data network and the wireless network in the system of FIG.</figref><figref num="3">FIG. 3 is a block diagram illustrating two radio frames 302 and 304 in the system of FIG. 1 utilizing a GSM air interface.</figref><figref num="4">FIG. 4 is a chart illustrating an example of frame size variation for a typical video sequence in the system of FIG.</figref><figref num="5">FIG. 5 is a block diagram illustrating a buffering delay used to support the transmission of frames of various sizes transmitted over the CBR channel in the system of FIG.</figref><figref num="6">FIG. 6 is a graph illustrating the buffering delay introduced by streaming a variable bit rate (VBR) multimedia stream through the CBR channel in the system of FIG.</figref><figref num="7">Figure 7 illustrates the buffer delay Ab in milliseconds for various 50-frame sequenced video clips encoded using a nominal rate of 64 kbps and Qp for AVC / H.264 and MPEG-4 in the system. It is a bar graph.</figref><figref num="8">FIG. 8 is a bar graph illustrating visual quality represented by the well-understood objective metric "Peak Signal to Noise Ratio" of the sequence illustrated in FIG.</figref><figref num="9">FIG. 9 illustrates the various levels of encapsulation that exist when transmitting multimedia data such as video data over wireless links using RTP / UDP / IP in the system.</figref><figref num="10">FIG. 10 illustrates an example of the allocation of application data packets, such as multimedia data packets, to physical layer data packets in the system.</figref><figref num="11">Figure 11 illustrates an example of encoding application layer packets according to EBR technology in the system.</figref><figref num="12">FIG. 12 is a block diagram illustrating an embodiment of a codec that transmits a VBR data stream over an IP / UDP / RTP network such as the Internet.</figref><figref num="13">FIG. 13 is a bar graph illustrating the relative reduction in peak signal-to-noise ratio (PSNR) for various examples of video sequences encoded using different encoding techniques when the channel packet loss is 1%. Is.</figref><figref num="14">FIG. 14 is a bar graph illustrating the relative reduction in peak signal-to-noise ratio (PSNR) when the channel loss is 5% for various examples of encoded video sequences.</figref><figref num="15">FIG. 15 is a bar graph illustrating the percentage of defective data packets received for the encoded video sequence of FIG.</figref><figref num="16">FIG. 16 is a bar graph illustrating the percentage of defective data packets received for the encoded video sequence of FIG.</figref><figref num="17">FIG. 17 is a graph illustrating the PSNR of sample-encoded video sequences vs. bit rates for four different cases.</figref><figref num="18">FIG. 18 is a graph illustrating the PSNR of other encoded video sequences vs. bit rates for four different cases.</figref><figref num="19">FIG. 19 is a graph illustrating a transmission plan for an AVC / H.264 stream with an average rate of 64 kbps.</figref><figref num="20">FIG. 20 is a flow chart illustrating an embodiment of a method of transmitting data.</figref><figref num="21">FIG. 21 is a flow diagram illustrating another embodiment of the method of transmitting data.</figref><figref num="22">FIG. 22 is a block diagram of a wireless communication device or mobile station (MS) constructed according to an exemplary embodiment of the present invention.</figref>
14 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO0021321A1 | Cites | World Intellectual Property Organization (WIPO) | Search report |
| WO0205575A2 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| WO0223745A2 | Cites | World Intellectual Property Organization (WIPO) | Search report |
| WO03103331A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| US2002075831A1 | Cites | United States of America | Search report |
| JP2003051849A | Cites | Japan | Examiner |
| US2003208615A1 | Cites | United States of America | Search report |
| WO2004036816A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| US6647006B1 | Cites | United States of America | Search report |
| JPH07312783A | Cites | Japan | Examiner |
| JPH09312656A | Cites | Japan | Examiner |
104 members in 14 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 57167304 | United States of America | P | |
| 57167304 | United States of America | P | |
| 60571673 | United States of America | – | |
| 2005016837 | United States of America | W | |
| 2005016837 | United States of America | W | |
| 2004571673 | – | – | – |
| 2005016837 | – | – | – |
| US20040571673P | – | – | – |
| WO2005US16837 | – | – | – |
Members104
| Document | Office | Kind | |
|---|---|---|---|
| US2005259613A1 | United States of America | A1 | |
| US2005259623A1 | United States of America | A1 | |
| US2005259690A1 | United States of America | A1 | |
| US2005259694A1 | United States of America | A1 | |
| CA2565977A1 | Canada | A1 | |
| CA2566124A1 | Canada | A1 | |
| CA2566125A1 | Canada | A1 | |
| CA2566126A1 | Canada | A1 | |
| CA2771943A1 | Canada | A1 | |
| CA2811040A1 | Canada | A1 | |
| WO2005114919A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2005114943A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2005114950A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2005115009A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2005114943A3 | World Intellectual Property Organization (WIPO) | A3 | |
| TW200618544A | Taiwan Province of China | A | |
| TW200618564A | Taiwan Province of China | A | |
| TW200623737A | Taiwan Province of China | A | |
| KR20070013330A | Republic of Korea | A | |
| KR20070014200A | Republic of Korea | A | |
| KR20070014201A | Republic of Korea | A | |
| EP1751955A1 | European Patent Office (EPO) | A1 | |
| EP1751956A2 | European Patent Office (EPO) | A2 | |
| EP1751987A1 | European Patent Office (EPO) | A1 | |
| MXPA06013186A | Mexico | A | |
| MXPA06013193A | Mexico | A | |
| EP1757027A1 | European Patent Office (EPO) | A1 | |
| KR20070023731A | Republic of Korea | A | |
| MXPA06013210A | Mexico | A | |
| MXPA06013211A | Mexico | A | |
| CN1969562A | China | A | |
| CN1973515A | China | A | |
| CN1977516A | China | A | |
| CN1985477A | China | A | |
| BRPI0510952A | Brazil | A | |
| BRPI0510953A | Brazil | A | |
| BRPI0510961A | Brazil | A | |
| BRPI0510962A | Brazil | A | |
| JP2007537681A | Japan | A | |
| JP2007537682AThis record | Japan | A | |
| JP2007537683A | Japan | A | |
| JP2007537684A | Japan | A | |
| KR20080084866A | Republic of Korea | A | |
| KR100870215B1 | Republic of Korea | B1 | |
| KR100871305B1 | Republic of Korea | B1 | |
| EP1757027B1 | European Patent Office (EPO) | B1 | |
| ATE417436T1 | Austria | T1 | |
| DE602005011611D1 | Germany | D1 | |
| EP1751955B1 | European Patent Office (EPO) | B1 | |
| ATE426988T1 | Austria | T1 | |
| KR20090039809A | Republic of Korea | A | |
| ES2318495T3 | Spain | T3 | |
| DE602005013517D1 | Germany | D1 | |
| ES2323011T3 | Spain | T3 | |
| KR100906586B1 | Republic of Korea | B1 | |
| KR100918596B1 | Republic of Korea | B1 | |
| MY139431A | Malaysia | A | |
| JP4361585B2 | Japan | B2 | |
| JP4448171B2 | Japan | B2 | |
| MY141497A | Malaysia | A | |
| EP2182734A1 | European Patent Office (EPO) | A1 | |
| EP2214412A2 | European Patent Office (EPO) | A2 | |
| JP4554680B2 | Japan | B2 | |
| EP1751987B1 | European Patent Office (EPO) | B1 | |
| ATE484157T1 | Austria | T1 | |
| MY142161A | Malaysia | A | |
| DE602005023983D1 | Germany | D1 | |
| CN1977516B | China | B | |
| EP2262304A1 | European Patent Office (EPO) | A1 | |
| ES2354079T3 | Spain | T3 | |
| EP1751956B1 | European Patent Office (EPO) | B1 | |
| ATE508567T1 | Austria | T1 | |
| DE602005027837D1 | Germany | D1 | |
| KR101049701B1 | Republic of Korea | B1 | |
| JP2011142616A | Japan | A | |
| CN1969562B | China | B | |
| KR101068055B1 | Republic of Korea | B1 | |
| ES2366192T3 | Spain | T3 | |
| TWI353759B | Taiwan Province of China | B | |
| TW201145943A | Taiwan Province of China | A | |
| US8089948B2 | United States of America | B2 | |
| CA2566125C | Canada | C | |
| EP2262304B1 | European Patent Office (EPO) | B1 | |
| CN1985477B | China | B | |
| EP2214412A3 | European Patent Office (EPO) | A3 | |
| TWI381681B | Taiwan Province of China | B | |
| CN1973515B | China | B | |
| CN102984133A | China | A | |
| TWI394407B | Taiwan Province of China | B | |
| EP2592836A1 | European Patent Office (EPO) | A1 | |
| CA2565977C | Canada | C | |
| JP5356360B2 | Japan | B2 | |
| EP2182734B1 | European Patent Office (EPO) | B1 | |
| CA2566124C | Canada | C | |
| US8855059B2 | United States of America | B2 | |
| US2014362740A1 | United States of America | A1 | |
| US2015016427A1 | United States of America | A1 | |
| CA2771943C | Canada | C | |
| CN102984133B | China | B | |
| US9674732B2 | United States of America | B2 |
19 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 2007537682
- Publication, DOCDB
- 2007537682
- Publication, EPODOC
- JP2007537682
- Application
- 2007513420
- Application, DOCDB
- 2007513420
- Application, EPODOC
- JP20070513420
Titles2
- Japanese
- 通信チャネルを介した情報の配信
- English
- Distribution of information via communication channels
Classification
- CPC, 36
- H04L69/04
- H04W28/06
- H04N21/2381
- H04N21/41407
- H04N21/44004
- H04N21/4788
- H04N21/6131
- H04N21/6181
- H04N21/6437
- H04N21/64707
- H04W28/065
- H04W72/1263
- H04W80/00
- H04W84/04
- H04W88/181
- H04L65/80
- H04L69/166
- H04L69/22
- H04L69/161
- H04N19/102
- H04N19/115
- H04N19/61
- H04N19/124
- H04N19/152
- H04N19/164
- H04N19/174
- H04L69/321
- H04L47/36
- H04W4/06
- H04L65/764
- H04L65/00
- H04L9/40
- H04L65/75
- H04L65/1101
- H04W72/044
- H04W88/02
- IPC, 12
- H04L12 56
- H04Q7 38
- H04B7 00
- H04B7 216
- H04L12 28
- H04L12 66
- H04L47 36
- H04N7 26
- H04W28 06
- H04W72 12
- H04W84 04
- H04W88 18
Designated states4
- Regional, 4
- Zimbabwe
- Turkmenistan
- Türkiye
- Togo