Video generalized reference decoder
Abstract
This record has no abstract on file.
Term
Term ended
Expired 19 September 2022, 4 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
22 claims: 4 independent, 18 dependent
- 1所与 のビデオクリップ についての符号化されたデータのビットストリーム のために提供されるパラメータの複数のセットであって、各々が該 所与 のビデオクリップ についての符号化されたデータのビットストリームための リファレンスデコーダのモデルを特徴付け、 前記所与のビデオクリップについての符号化されたデータのビットストリームための レートパラメータおよびデコーダバッファサイズパラメータを含むパラメータの複数のセットを受信するステップと、 前記パラメータの複数のセット の各々 によって前記 所与 のビデオクリップの符号化されたデータを復号するためのピークレートまたはピークデコーダバッファサイズを含む動作条件が決定されるよう、前記パラメータの複数のセットを処理するステップと を備えたことを特徴とするコンピュータに実装される方法。
- 2前記 所与 のビデオクリップ についての符号化されたデータのビットストリーム のために提供される前記パラメータのセットの数を示す情報を受信するステップをさらに備えたことを特徴とする請求項1に記載の方法。
- 3前記 所与 のビデオクリップ についての符号化されたデータのビットストリーム のために提供される別のパラメータの複数のセットを受信するステップと、 前記別のパラメータの複数のセットによって前記動作条件が決定されるよう、前記別のパラメータの複数のセットを処理するステップとをさらに備えたことを特徴とする請求項1に記載の方法。
- 4前記パラメータの複数のセットは、帯域外で提供されることを特徴とする請求項1に記載の方法。
- 5前記パラメータの複数のセットは、該パラメータの複数のセットが挿入された、前 記ビ デオビットストリームのストリームヘッダから処理を開始されることを特徴とする請求項1に記載の方法。
- 6前記パラメータの複数のセットを処理するステップは、前記動作条件を決定するステップを含むことを特徴とする請求項1に記載の方法。
- 7前記動作条件を決定するステップは、前記パラメータの複数のセットのうちの1つのから1つのパラメータを選択するステップを含むことを特徴とする請求項6に記載の方法。
- 8前記動作条件を決定するステップは、前記パラメータの複数のセットのうちの2つのセットの間を内挿するステップを含むことを特徴とする請求項6に記載の方法。
- 9前記動作条件を決定するステップは、前記パラメータの複数のセットのうちの2つのセットについて外挿するステップを含むことを特徴とする請求項6に記載の方法。
- 10前記ピークレートは最小ピークレートであり、前記動作条件を決定するステップは、前記パラメータの複数のセットの1または2以上のデコーダバッファサイズパラメータ を用いて 前記最小ピークレートを決定するステップを含むことを特徴とする請求項6に記載の方法。
- 11前記動作条件を決定するステップは、前記パラメータの複数のセットの1または2以上の前記レートパラメータ を用いて 前記デコーダバッファサイズを決定するステップを含むことを特徴とする請求項6に記載の方法。
- 12前記パラメータの複数のセットの各々は、初期バッファ満杯パラメータをさらに含むことを特徴とする請求項1ないし11のいずれかに記載の方法。
- 13前記パラメータの複数のセットの各々は、前記 所与 のビデオクリップ についての符号化されたデータのビットストリーム のためのリファレンスデコーダの別のリーキーバケットモデルを特徴付けることを特徴とする請求項1ないし11のいずれかに記載の方法。
- 14前記パラメータの複数のセットの各々は、前記 所与 のビデオクリップ についての符号化されたデータのビットストリーム のためのレート-デコーダバッファサイズ曲線に沿って異なる点を表していることを特徴とする請求項1ないし11のいずれかに記載の方法。
- 15前記デコーダバッファサイズパラメータは、前記パラメータの複数のセットの各々について異なることを特徴とする請求項1ないし11のいずれかに記載の方法。
- 16前記レートパラメータは、前記パラメータの複数のセットの各々について異なることを特徴とする請求項1ないし11のいずれかに記載の方法。
- 17前記ピークレートは、前記符号化されたデータを復号する際のディスクの駆動速度に対応することを特徴とする請求項1ないし11のいずれかに記載の方法。
- 18前記ピークレートは、前記符号化されたデータを復号する際のネットワーク接続の伝送レートに対応することを特徴とする請求項1ないし11のいずれかに記載の方法。
- 19前記符号化されたデータの復号は、前記 所与 のビデオクリップのライブビデオ伝送の間になされることを特徴とする請求項1ないし11のいずれかに記載の方法。
- 20前記符号化されたデータの復号は、前記 所与 のビデオクリップのオンデマンド伝送の間になされることを特徴とする請求項1ないし11のいずれかに記載の方法。
- 21請求項1ないし11のいずれかの方法を実行するようコンピュータシステムに実行させるコンピュータ実行可能な命令を格納したコンピュータ読取可能な媒体。
- 22請求項1ないし11のいずれかの方法を実行するよう調整されたコンピュータシステム。
Independent claims22
1 paragraph, as filed
[0001] [Technical field to which the invention belongs] The present invention relates to decoding video and video signals and other time-varying signals such as audio and audio. [0002] [Conventional technology] In the standard of video coding, the bitstream follows the mathematical model of the decoder connected to the output of the encoder, at least if it can be theoretically decoded. Such model decoders are known as H.263 coding standard hypothetical reference decoders (HRDs) and MPEG coding standard video buffer testers. Generally, an actual decoder device (or terminal) has a decoder buffer, a decoder and a display device. If the actual decoder device is configured according to the mathematical model of the decoder and the bitstream according to the mathematical model is transmitted under certain conditions, the decoder buffer will not overflow or underflow. Decryption is done accurately. [0003] Traditional reference (model) decoders assume that a bitstream is transmitted through a channel at a given constant bit rate and is decoded by a device with a given buffer size (after a given buffer delay). Therefore, such a model becomes very inflexible, broadcasting live video to devices with different buffer sizes, or encoded video over network paths with different peak bit rates. It cannot meet the demands of many of today's most important video applications for streaming, such as on-demand. [0004] In a traditional reference decoder, a video bitstream is received at a given constant bit rate (usually the average rate at bits per second of the stream) and its buffer reaches the desired level of fullness. Is stored in the decoder buffer. For example, data corresponding to at least one initial frame of video information is required before the output frame from the video information can be reproduced by decoding. This desired level is the initial decoder buffer full. Called buffer fullness), at a constant bit rate, it is directly proportional to the transmission or rise (buffer) delay. When this full amount is reached, the decoder instantly (dramatically) removes the bits for the first video frame in the sequence and decodes these bits to display the frame. At subsequent time intervals, the bits for the next frame are also instantly stripped, decrypted, and displayed. [0005] [Problems to be Solved by the Invention] Such a reference decoder operates at a fixed bit rate, buffer size and initial delay. However, in many modern video applications (eg, video streaming over the Internet or ATM networks), the peak bandwidth varies with the network path. For example, the peak bandwidth varies depending on whether the connection to the network is by modem, ISDN, DSL, cable, etc. In addition, network conditions, such as network congestion, the number of connected users, and other known factors, also cause peak bandwidth to fluctuate over time. In addition, handset, personal digital Video bitstreams are sent to a variety of devices with different buffer capacities, including assistants (PDAs), personal computers, pocket-sized computers, television set-top boxes, DVD players, etc., which differ, for example, low-latency streaming, progressive download, etc. Generated for scenarios with delayed requests. [0006] Existing reference decoders do not adjust for such fluctuations. At the same time, encoders generally do not know or know in advance what the changing conditions will be for a given recipient. As a result, resources and / or delays are often unnecessarily wasted and often inappropriate. [0007] [Means for solving problems] Simply put, the present invention provides a generalized reference decoder that has been improved to operate according to any number of rate and buffer parameter sets for a given bitstream. Each set is a leaky bucket It characterizes what is called a bucket) model or its parameter set and contains three values (R, B, F). Here, R is the transmission bit rate, B is the buffer size, and F is the initial decoder buffer full amount. It is understood that F / R is the rising or initial buffer delay. [0008] The encoder can generate a video bitstream containing a desired number of N leaky buckets, or simply calculate N sets of parameters after the bitstream has occurred. The encoder sends the number (at least once) to the decoder along with the corresponding number of (R, B, F) sets in some way, such as in the initial stream header or out of band. [0009] When received by the decoder, when there are at least two sets, the generalized reference decoder chooses one or interpolates between those leaky bucket parameters, and by doing so any desired It can also operate at peak bitrates, buffer sizes or delays. More particularly, given the desired peak transmission rate R'recognized by the decoder, the generalized reference decoder decodes the bitstream unimpeded by buffer overflow or underflow. The minimum buffer size and delay that can be made is possible by choosing one of the ((R, B, F) sets, interpolating between two or more, or extrapolating (R,). Select according to the set of B, F). Alternatively, for a given decoder buffer B', the hypothetical decoder chooses the minimum required peak transmission rate and operates at that rate. [0010] The advantage of a generalized reference decoder is that the content provider once generates a bitstream, and the server sends it to multiple devices with different performances using different channels with different peak transmission rates. It includes being able to do it. Alternatively, the server and terminal can negotiate the best leaky bucket parameters for a given network condition. For example, one that produces the lowest rise (buffer) delay, or one that requires the lowest peak transmission rate for a given buffer size of the device. In practice, the buffer size and delay for the terminal can be reduced by order of magnitude, or peak transmission rate, except for a negligible amount of additional bits for communicating leaky bucket information. Can be reduced by a large factor (eg, 4x), and / or the signal-to-noise ratio (SNR) can be increased, perhaps by a few dB, without increasing the average bit rate. [0011] An object of the present invention is to solve the above-mentioned problems. The invention according to claim 1 uses at least two sets of parameters, including rate data and buffer size data, by 1) selecting the buffer size based on the rate data or 2) selecting the rate based on the buffer size. Performed by a computer characterized in determining operating conditions and, in a time-varying signal decoder, maintaining the data encoded according to the operating conditions in a buffer and decoding the encoded data from the buffer. It is the method to be done. [0012] Further, according to the second aspect of the present invention, in the time-variable signal decoder of the method implemented by the computer according to the first aspect, at least two sets of parameters are received and the operating conditions are determined by the time-variable signal decoder. It is characterized by further providing. [0013] The invention according to claim 3 is characterized in that, in the method implemented by the computer according to claim 2, each set of parameters also includes full amount data received by the time-varying signal decoder. [0014] The invention according to claim 4 is characterized in that at least two sets of parameters are determined by an encoder in the method carried out by the computer according to claim 2. [0015] The invention according to claim 5 is characterized in that, in the method implemented by the computer according to claim 4, at least two sets of parameters are received in a stream header together with information indicating the total number of parameter sets. To do. [0016] Further, in the invention according to claim 6, in the method carried out by the computer according to claim 1, determining the operating conditions using at least two sets of parameters means selecting one of the parameter sets. It is characterized by including. [0017] Further, in the invention according to claim 7, in the method carried out by the computer according to claim 1, determining the operating conditions using at least two sets of parameters is between the data points of at least two sets of parameters. It is characterized by including interpolation by. [0018] Further, in the invention according to claim 8, in the method carried out by the computer according to claim 1, determining the operating conditions using at least two sets of parameters is excluded from the data points of at least two sets of parameters. It is characterized by including insertion. [0019] Further, the invention according to claim 9 includes, in the method implemented by the computer according to claim 1, selecting the buffer size based on the rate data determines the buffer size close to the minimum loading delay. It is characterized by that. [0020] The invention according to claim 10 comprises, in the computer-implemented method of claim 1, selecting a rate based on buffer size includes determining a minimum required peak transmission rate based on buffer size. It is a feature. [0021] [0021] The invention according to claim 11 is characterized in that, in the method implemented by the computer according to claim 1, the operating conditions are changed at least once during the communication of the encoded data to the buffer. .. [0022] The invention according to claim 12 is a method performed by a computer, in which a time-varying signal decoder receives at least two sets of parameters including rate data and buffer size data, and at least two sets of parameters are used. Use 1) select the buffer size based on the rate data or 2) select the rate based on the buffer size to determine the operating conditions, keep the data encoded according to the operating conditions in the buffer, and buffer. It is characterized by decoding the data encoded from. [0023] The invention according to claim 13 is further provided with providing full amount data to the time-varying signal decoder in the method performed by the computer according to claim 12. [0024] Also, the invention of claim 14 defines at least two of the parameters in the computer-implemented method of claim 12, and provides the time-varying signal decoder with two parameter sets. It is characterized by further preparation. [0025] The invention according to claim 15 is characterized in that at least two sets of parameters are set by an encoder in the method implemented by the computer according to claim 14. [0026] The invention according to claim 16 is characterized in that, in the method implemented by the computer according to claim 15, at least two sets of parameters are received in a stream header together with information indicating the total number of parameter sets. To do. [0027] In addition, the invention according to claim 17 specifies one of the parameter sets to determine the operating conditions using at least two sets of parameters in the method performed by the computer according to claim 12. It is characterized by including. [0028] Further, in the invention according to claim 18, in the method carried out by the computer according to claim 12, determining the operating conditions using at least two sets of parameters is between the data points of at least two sets of parameters. It is characterized by including interpolation by. [0029] In addition, the invention according to claim 19 uses at least two sets of parameters to determine operating conditions in the method performed by the computer according to claim 12, in which data points of at least two sets of parameters are used. It is characterized by including extrapolation from. [0030] The invention of claim 20 also includes, in the computer-implemented method of claim 12, selecting a buffer size based on rate data determines a buffer size close to the minimum loading delay. It is characterized by that. [0031] The invention according to claim 21 also includes determining the minimum required peak transmission rate based on the buffer size in selecting the rate based on the buffer size in the method performed by the computer according to claim 12. It is characterized by that. [0032] The invention according to claim 22 is characterized in that, in the method implemented by the computer according to claim 12, the operating conditions are changed at least once during the communication of the encoded data to the buffer. To do. [0033] The invention according to claim 23 includes an encoder that provides a time-variable signal, an encoder buffer and a decoder buffer that are connected to each other by a transmission medium and maintain the time-variable signal in a system for providing a time-variable signal. A decoder that removes the time-varying signal from the decoder, a first mechanism that defines at least two sets of parameters, including rate data and buffer size data, to maintain the decoder buffer so that it does not overflow or underflow, and the rate. It is characterized by including a second mechanism for determining the size of the decoder buffer based on the data or determining the rate at which data is transferred from the encoder buffer to the decoder buffer based on the buffer size data. [0034] The invention according to claim 24 is characterized in that, in the system for providing the time fluctuation signal according to claim 23, a first mechanism for determining at least two sets of parameters is incorporated in the encoder. To do. [0035] The invention according to claim 25 is characterized in that a second mechanism is incorporated in the decoder in the system for providing the time fluctuation signal according to claim 23. [0036] The invention according to claim 26 incorporates a first mechanism for defining at least two sets of parameters into an encoder and a second mechanism for a decoder in the system for providing the time variation signal according to claim 23. Built into, the encoder is characterized by transmitting the parameter set to the decoder. [0037] The invention according to claim 27 is characterized in that, in the system for providing the time variation signal according to claim 26, the encoder transmits a parameter set to the decoder via a stream header. [0038] The invention according to claim 28 is characterized in that, in the system for providing the time variation signal according to claim 26, the encoder identifies the total number of parameter sets. [0039] The invention according to claim 29 is characterized in that, in the system for providing the time variation signal according to claim 23, each set of parameters also includes full amount data received by the decoder. [0040] The invention according to claim 30 is the system for providing the time variation signal according to claim 23, wherein the second mechanism is a decoder based on rate data by selecting one of the parameter sets. It is characterized by determining the size of the buffer or the rate of the transferred data. [0041] The invention of claim 31 is the system for providing the time variation signal of claim 23, wherein the second mechanism is by interpolating between data points of at least two sets of parameters. , The size of the decoder buffer is determined based on the rate data, or the rate of the transfer data is determined. [0042] The invention according to claim 32 is a system for providing the time variation signal according to claim 23, wherein the second mechanism is a rate by extrapolating from data points of at least two sets of parameters. It is characterized in that the size of the decoder buffer is determined or the rate of the transferred data is determined based on the data. [0043] The invention according to claim 33 is the system for providing the time variation signal according to claim 23, wherein the second mechanism determines a buffer size close to the minimum loading delay of the decoder buffer. It is characterized by determining the size. [0044] The invention according to claim 34 is the system for providing the time variation signal according to claim 23, wherein the second mechanism determines a minimum required peak transmission rate corresponding to a predetermined buffer size. By doing so, it is characterized in that rate data is determined. [0045] The invention according to claim 35 is the system for providing the time variation signal according to claim 23, wherein the second mechanism determines a new size of the decoder buffer based on the rate data and the time information. It is characterized by. [0046] The invention according to claim 36 is the system for providing the time variation signal according to claim 23, wherein the second mechanism determines a new rate for transferring data based on the buffer size data and the time information. It is characterized by that. [0047] BEST MODE FOR CARRYING OUT THE INVENTION Other advantages and advantages will become apparent from the detailed description below, along with the drawings below. [0048] [Typical operating environment] FIG. 1 illustrates an example of a suitable operating environment 120 in which the present invention can be practiced, especially for encoding video and / or video data. The operating environment 120 is merely an example of a suitable operating environment and is not intended to imply any limitation with respect to the scope or function of the present invention. Other well-known computer systems, environments and / or configurations suitable for use with the present invention are personal computers, servers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, programmable consumers. Includes, but is not limited to, electronic devices for electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. For example, encoding video and / or video image data is often processed on a computer with even higher processing power than modern handheld personal computers, but encoding cannot be performed on this illustrated device, or There seems to be no reason why it is not decrypted by more powerful devices. [0049] The present invention may be described in the general context of computer-executable instructions, such as program modules executed by one or more computers or other devices. In general, a program module includes routines, programs, objects, components, data structures, etc. that perform a particular task or realize a particular abstract data type. Typically, the functionality of the program module may be combined or distributed as required by the various embodiments. The arithmetic unit 120 typically includes at least some form of computer-readable recording medium. The computer-readable recording medium can be any medium accessible to the arithmetic unit 120. By way of example, but not limited to, computer readable recording media may include computer storage and communication media. Computer storage media are volatile and non-volatile, removable and non-volatile, which can be achieved by any method or technique for storing information such as computer-readable instructions, data structures, program modules or other data. Includes possible media. The recording medium of the computer is RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, DVD or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or desired. It includes, but is not limited to, any other medium that can be used to store information and is accessible to the computer 120. A communication medium is typically any medium that embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal, such as a carrier wave or other transmission method, and delivers information. Also includes. The term "modulated data signal" means a signal having one or more of its characteristics set or altered to encode information in the signal. As an example, the communication medium is a wired network or a directly connected connection. It includes, but is not limited to, any wired and audio, RF (radio radio), infrared, and other radio media. Any combination of the above must be included in the range of computer readable recording media. [0050] FIG. 1 shows the functional configuration of such a handheld arithmetic unit 120, which includes a processor 122, a memory 124, a display 126 and a keyboard 128 (which may be a physical or virtual keyboard). .. Memory 124 generally includes both volatile memory (eg RAM) and non-volatile memory (eg ROM, PCMCIA card, etc.). The operating system 130, like Microsoft® Windows® CE, or any other operating system, resides in memory 124 and runs on processor 122. [0051] One or more application programs 132 are loaded into memory 124 and run on operating system 130. Examples of applications include e-mail programs, scheduling programs, PIM (personal information management) programs, word processing programs, spreadsheet programs, Internet browser programs and others. The handheld personal computer 120 may also include a notification management program 134 that is loaded into memory 124 and runs on processor 122. The notification management program 134 processes, for example, a notification request from the application program 132. [0052] The handheld personal computer 120 has a power supply 136 implemented as one or more batteries. The power supply 136 may further include an external power source that preferentially or recharges the built-in battery, such as an AC adapter or a powered docking cradle. [0053] The example handheld personal computer 120 shown in FIG. 1 has three types of external notification mechanisms: one or more light emitting diodes (LEDs) 140 and an audio generator 144. These devices are directly coupled to power supply 136 and, when activated, the notification mechanism even if the handheld personal computer processor 122 and other components are powered off to save battery power. Try to stay on for the period specified by. The LED 140 is preferably kept on indefinitely until the user operates it. Modern audio generators 144 use too much power for the batteries of today's handheld personal computers, and therefore have a constant duration when other parts of the system are operating or after booting. It is configured to be turned off over. [0054] [Generalized reference decoder] The leaky bucket is a model that conceptualizes the state (or full amount) of the encoder or decoder buffer as a function of time. FIG. 2 illustrates this concept, in which the input data 200 is fed to the enhanced encoder 202, where the data is encoded (as described below) and input to the encoder buffer 204. The encoded data is transmitted through the transmission medium (pipe) 206 to the decoder buffer 208 and then decoded by the decoder 210 into output data such as video or video frames. For the sake of simplicity, the decoder buffer 208 is mainly described. The reason is that the full encoder and decoder buffers are conceptually complementary to each other, that is, the more data there is in the decoder buffer, the less data there is in the encoder buffer and vice versa. .. [0055] The leaky bucket model is characterized by a set of three parameters R, B and F. Here, R is the peak bit rate (bits / second) when a bit enters the decoder buffer 208. In constant bitrate scenarios, R is often the bitrate of the channel and the average bitrate of the video or audio clip, which can be conceptually considered to correspond to the width of the pipe 206. B is the size (in bits) of the bucket or decoder buffer 208 that smoothes the fluctuations in the video bit rate. This buffer size cannot be larger than the physical buffer of the decoder device. F is the initial decoder buffer fullness (also in bits) required before the decoder begins removing bits from its buffer. F is at least as large as the amount of coded data that represents the first frame. The processing time may be considered instantaneous for the purposes of this embodiment, but if this processing time is not taken into account, F and R determine the initial or rising delay D. Here, D = F / R (seconds). [0056] Therefore, in the leaky bucket model, the bits are input to the decoder buffer 208 at rate R until the full level reaches F (ie D seconds), and then the bits b required for the first frame.<sub>0</sub>Is removed (instantly in this example). Bits continue to be input to the buffer at rate R, and the decoder b<sub>1</sub>, B<sub>2</sub>, ..., b<sub>n-1</sub>Bits are removed in a given moment, typically (but not required) every 1 / Ms. Where M is the frame rate of the video. [0057] FIG. 3 is a graph showing the time variation of the decoder buffer full amount with respect to the bitstream stored in the leaky bucket of the parameters (R, B, F). Here, as described above, the i-th bit is b.<sub>i</sub>Is. In Figure 3, encoded video frames are removed from the buffer (generally according to the video frame rate), as represented by the drop in buffer full. And especially the time t<sub>i</sub>In b<sub>i</sub>B is the full amount of the decoder buffer just before removing the bit.<sub>i</sub>And. A general leaky bucket model operates according to the following equation. [0058] [0058] [Number 1]<img file="JP4199973B2_D0001.tif" />[0059] Typically t<sub>i + 1</sub>-t<sub>i</sub>= 1 / M seconds. Where M is the frame rate (frames / second) for the bitstream. [0060] The leaky bucket model of parameters (R, B, F) stores one bitstream if there is no underflow in the decoder buffer 208 (Fig. 2). This is equivalent to the encoder buffer 204 not overflowing, as the encoder and decoder buffer full amounts complement each other. However, the encoder buffer 204 (leaky bucket) is allowed to be empty, or equally, and the decoder buffer 208 is full in that no more bits are transmitted from the encoder buffer 204 to the decoder buffer 208. It is also good. Therefore, the decoder buffer 208 stops receiving bits when it is full. This is why we use the min operator in the second equation above. To complement these, a full decoder buffer 208 means that the encoder buffer 204 is free, as described below for variables of variable bit rate (VBR). [0061] It should be added that a given video stream may be included in the configuration of various leaky buckets. For example, if a video stream is stored in the leaky bucket of parameters (R, B, F), it is also stored in the leaky bucket of the larger buffer (R, B', F). Here, B'is greater than B. Alternatively, it is also stored in a leaky bucket with a higher peak transmission rate (R', B, F) than in this case. Here, R'is greater than R. In addition, for any bitrate R', there is a buffer size to store the (time-limited) video bitstream. In the worst case, that is, when R'approaches 0, the buffer size should be as large as the bitstream itself. In other words, the video bitstream can be transmitted at any rate (regardless of the average rate of the clip) as long as the buffer size is large enough. [0062] Figure 4 shows the peak bit rate R for a given bitstream.<sub>min</sub>The minimum buffer size B obtained using the second equation above for<sub>min</sub>It is a graph of. Here, the desired initial buffer full amount is set at a fixed portion of the total buffer size. The curve in Figure 4 shows that the decoder must at least B to carry the stream at the peak bit rate r.<sub>min</sub>(r) Indicates that bits need to be buffered. Moreover, as can be seen from the graph, higher peak rates require smaller buffer sizes and therefore shorter rise buffer delays. Alternatively, this graph shows that if the size of the decoder buffer is b, the minimum peak rate required to carry this bitstream is the associated R.<sub>min</sub>It shows that it is (b). In addition, for any bitstream (as shown in Figure 4) (R)<sub>min</sub>, B<sub>min</sub>The curves in the set of) are partially linear and convex. [0063] According to one embodiment of the invention, if at least two points of the curve are given by the enhanced encoder 202, the generalized reference decoder 210 selects one point and takes linear interpolation between the two points, or they. Extrapolate the point of (R<sub>min</sub>, B<sub>min</sub>) Slightly larger on the safe side, at some point (R)<sub>interp</sub>, B<sub>interp</sub>) Can be reached. As an important result, you can often safely reduce the buffer size on an order-by-order basis compared to a single leaky bucket that stores bitstreams at an average rate. As a result, the delay is reduced as well. Alternatively, for the same delay, the peak transmission rate may be (probably) reduced by a factor of 4, and the signal-to-noise ratio (SNR) may be (probably) improved by a few dB. [0064] To this end, the leaky bucket parameter 214 corresponding to at least two points on the rate-buffer curve useful for a given video or video clip (eg, reasonably divided over the range of R and / or B). , For example (R<sub>1</sub>, B<sub>1</sub>, F<sub>1</sub>), (R<sub>2</sub>, B<sub>2</sub>, F<sub>2</sub>) ..., (R<sub>N</sub>, B<sub>N</sub>, F<sub>N</sub>) Enhances the encoder 202 by configuring the encoder 202 to generate at least two sets. The enhanced encoder 202 inserts these leaky bucket parameter sets into the initial stream header, or somehow out of band, for a generalized reference decoder 210, along with a number N of these leaky bucket parameters. Provide a set. Two to four buckets are usually sufficient to reasonably represent the RB curve, but even for a few dozen buckets, for example, it is needed to provide this information. Note that the amount of additional bytes (eg 1 byte for N plus 8 bytes per leaky bucket model or parameter set) is negligible when compared to general video or video data. [0065] Furthermore, it is particularly noted that at higher bit rates, the content creator may decide to identify different leaky bucket models at different times within the bitstream. This is useful even if the connection is lost during transmission and restarted in the middle of the bitstream. For example, by giving a leaky bucket model at 15-minute intervals, the decoder can reselect, re-interpolate, or re-extrapolate its operating conditions (eg, buffer size or rate) an appropriate number of times, as desired. You may change it. [0066] The desired value of N can be selected with the encoder (note that if N = 1, the generalized decoder 210 extrapolates points like an MPEG video buffer tester). The encoder selects the value of the leaky bucket in advance and encodes the bitstream by performing rate control so that the constraint of the leaky bucket is satisfied, or encodes the bitstream and then N pieces using the above equation. You can choose to calculate a set of leaky bucket parameters that store the bitstream at different R values, or both. The first method can be applied to live or on-demand transmission, while the others apply to on-demand. [0067] In the case of the present invention, once received by the generalized reference decoder 210, which leaky bucket is used by the decoder 210, taking into account the peak bit rate and / or physical buffer size possible for it. Can be determined if is desirable. Alternatively, the generalized reference decoder 210 seeks linear interpolation between these points or seeks linear extrapolation from these points to search for the appropriate parameter set for a given setting. You may. [0068] Figure 5 shows two leaky bucket parameter sets and their linearly interpolated (R, B) values. For reference, the calculated RB curve is represented by a fine dashed line and is a leaky bucket model (R).<sub>x</sub>, B<sub>x</sub>) And (R<sub>y</sub>, B<sub>y</sub>The values of R and B given in) are represented by stars. (R<sub>x</sub>, B<sub>x</sub>) To (R<sub>y</sub>, B<sub>y</sub>The solid line to) represents the interpolated value. Any set of R or B selected on this solid line keeps the decoder buffer 208 properly (eg, without overflow or underflow). The leaky bucket parameters are represented by the coarse dashed line in FIG. 5, which can be extrapolated from these points, and also any R or B pair chosen on the solid line properly maintains the decoder buffer 208. [0069] The interpolated buffer size B between points k and k + 1 follows the following straight line. [0070] [Number 2]<img file="JP4199973B2_D0002.tif" />[0071] Where R<sub>k</sub><R <R<sub>k + 1</sub>Is. Similarly, the initial decoder buffer full amount F can be linearly interpolated. [0072] [Number 3]<img file="JP4199973B2_D0003.tif" />[0073] Where R<sub>k</sub><R <R<sub>k + 1</sub>Is. [0074] The leaky bucket with parameters (R, B, F) thus obtained is guaranteed to contain a bitstream, because (as can be mathematically proved) the minimum buffer size B.<sub>min</sub>Is convex to both R and F, i.e. any convex combination (R, F) = a (R)<sub>k</sub>, F<sub>k</sub>) + (1-a) (R<sub>k + 1</sub>, F<sub>k + 1</sub>), Minimum buffer size B corresponding to 0 <a <1<sub>min</sub>Is B = aB<sub>k</sub>+ (1-a) B<sub>k + 1</sub>Because it is below. [0075] As mentioned above, R is R<sub>N</sub>If larger than, leaky bucket (R, B)<sub>N</sub>, F<sub>N</sub>) Also contains a bitstream. Where B<sub>N</sub>And F<sub>N</sub>Is R<sub>N</sub>The desired buffer size and the initial decoder buffer full amount at. R is R<sub>1</sub>If less than, upper limit B = B<sub>1</sub>+ (R<sub>1</sub>-R) T can be used. Where T is the time length of the stream in seconds. These (R, B) values beyond the range of N points may be extrapolated. [0076] The decoder does not need to select the parameters of the leaky bucket and extrapolate or extrapolate, but the other way is rather to select the parameters to send a single set to the decoder, thereby the decoder. Uses that single set. Given some information, such as a decoder request, the server can determine the appropriate set of leaky bucket parameters to send to the decoder (by selection, interpolation or extrapolation), and then the decoder , Can be decrypted using only one parameter set. The proxy for the server or decoder can also select, interpolate or extrapolate the leaky bucket information without the decoder even looking at two or more leaky buckets. In other words, instead of the decoder deciding, the server can probably decide while the server and client decoders negotiate the parameters. However, in general, and in the case of the present invention, the proper leaky bucket model is determined based on at least two leaky bucket models. [0077] The value of the RB curve for a given bitstream can be calculated from the time of the highest and lowest fills in the decoder buffer plot, as shown in FIG. More specifically, for the bitstream contained in the leaky bucket of parameters (R, B, F), two times (tM, tm) of the maximum and minimum full amount of the decoder buffer are set in advance or dynamically. Think about it. The full amount may reach the highest and lowest values several times, but consider the set of the largest values (tM, tm) such that tM <tm. Assuming that the leaky bucket can be calculated properly, B is the minimum buffer size containing the bitstream for the values R, F. [0078] [Number 4]<img file="JP4199973B2_D0004.tif" />[0079] Here, b (t) is the number of bits of the frame at time t, and M is the frame rate at frames / second. In this equation, n is the number of frames between time tM and tm, and c is the sum of the bits for these frames. [0080] [0080] This equation can be interpreted as the point of the straight line B (r) where -n / M is the slope of the line with r = R. There is a bitrate range r [R-r1, R + r2] such that the largest pair of values tM and tm stays at the same value. Here, the equation corresponds to a straight line that determines the minimum buffer size associated with the bit rate r. If the bit rate r is outside the above range, at least one of the values tM and / or tm changes. Here, when r> R + r2, the time interval between tM and tm becomes smaller. And the value of n in the new straight line that defines B (r) is also smaller, and the slope of each line is larger (less negative). When r <R-r1, the time interval between tM and tm increases, and the value of n on the straight line defining B (r) also increases. The slope of the line is then smaller (more negative). [0081] A set of values (tM, tm) for a range of bitrates (or associated n values) and some values of c (at least one for any set) are stored in the bitstream header. Therefore, the above equation can be used to obtain a partially linear B (r) curve. In addition, this equation may be used to simplify the calculation of the parameters of the leaky bucket model after the encoder has generated the bitstream. [0082] In an attempt, the bitstream in Figure 5 was generated to generate an average bit rate of 797 Kbps. At a constant transmission rate of 797 Kbps, as is commonly shown in Figure 5, the decoder has a buffer size of approximately 18,000 Kbits (R).<sub>x</sub>, B<sub>x</sub>) Is required. If the initial decoder buffer full is equal to 18,000 Kbits, the rise delay will be about 22.5 seconds. Therefore, this coding (generated without rate control) shifts the bits up to 22.5 seconds in order to achieve essentially the best possible quality for that total coding length. [0083] Figure 5 shows that at a peak transmission rate of 2,500 Kbps (for example, the video bit rate portion of a 2x CD), the decoder has a buffer size of only 2,272 Kbits (R).<sub>y</sub>, B<sub>y</sub>) Is only required, which is reasonable for a consumer hardware device. If the initial buffer full is equal to 2,272 Kbits, the rise delay is only about 0.9 seconds. [0084] Therefore, for this coding, typically two leaky bucket models, eg, (R = 797Kbps, B = 18,000K bits, F = 18,000K bits) and (R = 2,500Kbps, B = 2,272K bits,), The model with F = 2,272 Kbits) is useful. This first leaky bucket parameter set allows video to be transmitted over a channel at a constant bit rate with a delay of approximately 22.5 seconds. For many scenarios, this delay is too great, but probably acceptable, for example for internet streaming of movies. The second set of leaky bucket parameters allows video to be transmitted over a shared network with a peak rate of 2,500 Kbps, or played locally from a 2x CD, with a delay of about 0.9 seconds. This delay of less than 1 second is acceptable for random access playback, which has the same function as a VCR (Video Cassette Recorder). [0085] This advantage is clear given that it occurs when only the first leaky bucket is identified in the bitstream, but not the second leaky bucket. In such cases, the decoder uses a buffer size of 18,000 Kbits, even when playing on a channel with a peak bit rate of 2,500 Kbps, so the delay is F / R = 18,000 Kbits / 2,500 Kbps = 7.2 seconds. Let's go. It can be understood that such a delay is unacceptable for random access playback that has a VCR-like function. However, if the second leaky bucket is further identified, the buffer size drops to 2,272 Kbits at a rate of 2,500 Kbps, and the delay drops to 0.9 seconds as described above. [0086] On the other hand, when only the second leaky bucket is identified and the first leaky bucket is not identified, even a sophisticated decoder uses a much larger buffer than required at a constant transmission rate of 797 Kbps. It is forced to ensure that the buffer does not overflow. That is, B'= B + (R-R') T = 2,272 Kbit + (2,500 Kbps-797 Kbps) × 130 seconds = 223,662 Kbit. Even if possible in such a device with a large amount of memory, this corresponds to an initial delay of 281 seconds or almost 5 minutes, which is well beyond the permissible limit. However, if the first leaky bucket is also identified, at a rate of 797 Kbps, the buffer size drops to 18,000 Kbits, and the delay drops to 22.5 seconds, as described above. [0087] Furthermore, once both leaky buckets are identified, the decoder linearly interpolates between them for any bitrate R between 797 Kbps and 2,500 Kbps (using the interpolation formula above). This allows near minimum buffer size and delay at a given rate. Extrapolation is also lower than 797 Kbps and higher than 2,500 Kbps (shown in Figure 5 with a coarse dashed line), with a single leaky bucket anywhere between 792 Kbps and 2,500 Kbps (including both ends). It is more efficient when compared to the extrapolation used. [0088] As illustrated by the above embodiment, even just two sets of leaky bucket parameters can reduce the buffer size on an order-by-order basis (eg, from 223,662 Kbits to 18,000 Kbits in some cases, and others. In the case of 18,000 Kbits to 2,272 Kbits), the delay at a given peak transmission rate is also reduced on an order-by-order basis (for example, from 281 seconds to 22.5 seconds in some cases, and in other cases). Reduced from 7.2 seconds to 0.9 seconds). [0089] Alternatively, the peak transmission rate for a given decoder buffer size can be reduced. In fact, as is clear from Figure 5, if the RB curve is obtained by interpolating and / or extrapolating multiple leaky buckets, a decoder with a fixed physical buffer size will have no decoder buffer underflow. You can choose the minimum peak transmission rate required to safely decode the bitstream. For example, if the decoder has a fixed buffer with a size of 18,000 Kbits, the peak transmission rate for coding can be as low as 797 Kbps. However, if the second leaky bucket is identified but the first leaky bucket is not identified, the decoder sets the bitrate to R'= R- (B'-B) / T = 2,500Kbps- (18,000K bits). -2,272 Kbit) /130 seconds = Bit rates lower than 2,379 Kbps can be reduced. In this case, compared to using a single leaky bucket, using only two leaky buckets reduces the peak transmission rate by a factor of 4 for the same decoder buffer size. [0090] Having multiple leaky bucket parameters can also improve the quality of video reconstructed at the same average code rate. Consider the case where both leaky buckets can be used for coding. As mentioned above, with this information in the decoder, encode and play with a delay of 22.5 seconds if the peak transmission rate is 797 Kbps and with a delay of 0.9 seconds if the peak transmission rate is 2,500 Kbps. Can be done. [0091] However, if the second leaky bucket cannot be used, the delay increases from 0.9 seconds to 7.2 seconds at 2,500 Kbps. One way to reduce the delay to 0.9 seconds without the benefit of the second leaky bucket is to reduce the buffer size (of the first leaky bucket) from 18,000 Kbits to (0.9 seconds x 2,500 Kbps) = 2,250 Kbits. , Controlling the rate to recode the clip. This ensures that the delay at 797 Kbps is reduced from 22.5 seconds to 2.8 seconds, but at a peak transmission rate of 2,500 Kbps, the delay is only 0.9 seconds. However, as a result, quality (SNR) is also reduced by an estimated amount of a few dB, especially for clips with large dynamic range. [0092] Therefore, by identifying the second leaky bucket, the SNR is probably a number, without changing the average bit rate, by identifying the second leaky bucket, except for a negligible amount of additional bits per clip. It can be increased by dB. This increase in signal-to-noise ratio is noticeable during reproduction for each peak transmission rate. [0093] The advantage of identifying multiple leaky buckets in a generalized reference decoder is for devices that have a single encoding and are transmitted over channels with different peak rates or have different physical buffer sizes. Understood when transmitted. But in reality, this is becoming more and more commonplace. For example, content encoded offline and stored on disk flows through networks with different peak rates and is often played locally. Even in local playback, various drive speeds (eg, 1x CD to 8x DVD) affect the peak transmission rate. In addition, the peak transmission rate over the network connection dramatically follows the speed of limiting links (eg, 100 or 10baseT Ethernet®, T1, DSL, ISDN, modems, etc.) that are generally close to the end user. The buffer capacity of a playback device also changes significantly from a desktop computer having a buffer space of several gigabytes to an electronic device for small consumers having a buffer space of several orders of magnitude smaller. Multiple leaky buckets and the generalized reference decoder proposed by the present invention allow the same bitstream to be transmitted over various channels with minimum rising delay, minimum decoder buffer requirements and maximum possible quality. Is possible. This applies not only to offline encoded video, but also to live video delivered simultaneously to different devices over different channels. In summary, the generalized reference decoder proposed by the present invention adds great flexibility to existing bitstreams. [0094] As can be understood from the details above, a generalized reference decoder improved over the prior art is provided. Generalized reference decoders require only a small amount of information from the encoder (eg, in the header of a bitstream), have varying bandwidths, and / or terminals have varying bitrates and buffer capacities. Provides greater flexibility for bitstream delivery over modern networks with. The reference decoders of the present invention enable these new scenarios while minimizing transmission delays for possible bandwidth, and in addition for delivery to devices with a given physical buffer size limit. Virtually minimize the channel bit rate requirement of. [0095] Although various modifications and alternative configurations are possible in the present invention, the illustrated examples are shown in the drawings and described in detail above. However, there is no intention to limit the invention to the particular form disclosed, and on the contrary, to include all modifications, alternative configurations and all equivalents in the spirit and intent of the invention. It should be understood that there is. [Simple explanation of drawings] FIG. 1 is a block diagram showing an example of a computer system incorporating the present invention. FIG. 2 is a block diagram showing an enhanced encoder for encoding and decoding video or video data according to one aspect of the present invention, a generalized reference decoder, and their respective buffers. FIG. 3 is a graph showing the amount of buffer full contained in the leaky bucket of parameters (R, B, F) over time. FIG. 4 is a graph showing a rate vs. buffer size curve for a typical video clip. FIG. 5: Rate vs. buffer size for a representative video clip with (parameter set) of two leaky bucket models provided in a generalized reference decoder for interpolation and extrapolation according to one aspect of the invention. It is a graph which shows the curve of. [Explanation of symbols] 120 operating environment 122 processor 124 memory 126 display 128 keyboard 130 operating system 132 Application program 134 Notification Management Program 136 power supply 140 LED 144 Audio Generator 200 Input data 202 Enhanced encoder 204 encoder buffer 206 Transmission medium (pipe) 208 decoder buffer 210 Generalized reference decoder 212 Output data 214 Leaky bucket parameters
Every citation, both waysCites: the store holds 4 of 5
| Document | Relation | Office |
|---|---|---|
| JP2001169261A | Cites | Japan |
| JP10294757A | Cites | Japan |
| JP08223385A | Cites | Japan |
| JP04297179A | Cites | Japan |
| STUDY GROUP 16 ,TRANSMISSION OF NON-TELEPHONE SIGNALS VIDEO CODINGFOR LOW BIT RATE COMMUNICATION,DRAFT H.263,ITU-T,1998年 1月27日,p.5,Annex B,URL,http://www.ece.cmu.edu/~ece796/documents/H263.doc | Non-patent | – |
28 members in 8 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 09955731 | United States of America | – | |
| 95573101 | United States of America | A | |
| 2001955731 | – | – | – |
| US20010955731 | – | – | – |
Members28
| Document | Office | Kind | |
|---|---|---|---|
| US2003053416A1 | United States of America | A1 | |
| KR20030025186A | Republic of Korea | A | |
| EP1298938A2 | European Patent Office (EPO) | A2 | |
| CN1426235A | China | A | |
| JP2003179665A | Japan | A | |
| EP1298938A3 | European Patent Office (EPO) | A3 | |
| TW574831B | Taiwan Province of China | B | |
| US2006198446A1 | United States of America | A1 | |
| CN1848964A | China | A | |
| EP1746844A2 | European Patent Office (EPO) | A2 | |
| EP1753248A2 | European Patent Office (EPO) | A2 | |
| EP1746844A3 | European Patent Office (EPO) | A3 | |
| EP1753248A3 | European Patent Office (EPO) | A3 | |
| KR20070097375A | Republic of Korea | A | |
| JP2007329953A | Japan | A | |
| JP4199973B2This record | Japan | B2 | |
| CN100461858C | China | C | |
| US7593466B2 | United States of America | B2 | |
| US7646816B2 | United States of America | B2 | |
| KR100947162B1 | Republic of Korea | B1 | |
| JP4489794B2 | Japan | B2 | |
| KR100999311B1 | Republic of Korea | B1 | |
| CN1848964B | China | B | |
| DE20222026U1 | Germany | U1 | |
| EP1746844B1 | European Patent Office (EPO) | B1 | |
| EP1753248B1 | European Patent Office (EPO) | B1 | |
| EP1298938B1 | European Patent Office (EPO) | B1 | |
| HK1053034B | Hong Kong, China | B |
36 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 |
Numbers
- Publication
- 4199973
- Publication, DOCDB
- 4199973
- Publication, EPODOC
- JP4199973B
- Application
- 273882
- Application, DOCDB
- 2002273882
- Application, EPODOC
- JP20020273882
Titles2
- Japanese
- 映像又はビデオ処理のための一般化されたリファレンスデコーダ
- English
- Generalized reference decoder for video or video processing
Classification
- CPC, 7
- H04N19/44
- G06T1/00
- H04N19/115
- H04N19/149
- H04N19/152
- H04N19/172
- H04N19/61
- IPC, 8
- H04L29 08
- H04L13 08
- H04N7 26
- G06T1 00
- H04N7 173
- H04N7 50
- H04N19 00
- H04N21 438