Video encoder
Abstract
Video coding method comprising: defining a first set of multiple values (Ri1) characteristic of a transmission bit rate for a first access point at a starting point of a video sequence; defining a second set of multiple values (Bi1) characteristic of a buffer size for said first access point; defining a third set of multiple values (Di1) characteristic of a delay until a video sequence is presented for said first access point; defining a fourth set of multiple values (Ri2) characteristic of a transmission bit rate for another access point located after said first access point and then; defining a fifth set of multiple values (Bi2) characteristic of a buffer size for said other access point; defining a sixth set of multiple values (Di2) characteristic of a delay until a video sequence is presented for said other access point; wherein a value of said first set of multiple values, a value of said second set of multiple values, and a value of said third set of multiple values are selected such that said video sequence is free of a memory overflow condition buffer at said first access point; and a value of said fourth set of multiple values, a value of said fifth set of multiple values, and a value of said sixth set of multiple values are selected such that said video sequence is free of a buffer overflow condition in said another access point.

Term
Term ended
Projected expiry passed 26 March 2024, 2.5 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
1 claim: 1 independent, 0 dependent
- 1ES 2 390 596 T3 ES 2 390 596 T3 CLAIMS REIVINDICACIONES Método de codificación de vídeo que comprende:Video encoding method comprising: defining a first set of multiple values (Ri1) characteristic of a transmission bit rate for a first access point at a starting point of a video sequence;definir un primer conjunto de múltiples valores (Ri1) característicos de una velocidad binaria de transmisión para un primer punto de acceso en un punto de comienzo de una secuencia de vídeo;defining a second set of multiple values (Bi1) characteristic of a buffer size for said first access point;definir un segundo conjunto de múltiples valores (Bi1) característicos de un tamaño de memoria tampón para dicho primer punto de acceso;defining a third set of multiple values (Di1) characteristic of a delay until a video sequence is presented for said first access point;definir un tercer conjunto de múltiples valores (Di1) característicos de un retardo hasta que se presenta una secuencia de vídeo para dicho primer punto de acceso;defining a fourth set of multiple values (Ri2) characteristic of a transmission bit rate for another access point located after said first access point and thereafter;definir un cuarto conjunto de múltiples valores (Ri2) característicos de una velocidad binaria de transmisión para otro punto de acceso localizado después de dicho primer punto de acceso y a continuación;defining a fifth set of multiple values (Bi2) characteristic of a buffer size for said other access point;definir un quinto conjunto de múltiples valores (Bi2) característicos de un tamaño de memoria tampón para dicho otro punto de acceso;defining a sixth set of multiple values (Di2) characteristic of a delay until a video sequence is presented for said other access point;wherein a value from said first multi-value set, a value from said second multi-value set, and a value from said third multi-value set such that said video sequence is free from a memory overflow condition are selected buffer in said first access point;Y a value from said fourth multi-value set, a value from said fifth multi-value set, and a value from said sixth multi-value set such that said video sequence is selected from a buffer overflow condition in said another access point. definir un sexto conjunto de múltiples valores (Di2) característicos de un retardo hasta que se presenta una secuencia de vídeo para dicho otro punto de acceso;en el que se seleccionan un valor de dicho primer conjunto de múltiples valores, un valor de dicho segundo conjunto de múltiples valores, y un valor de dicho tercer conjunto de múltiples valores tal que dicha secuencia de vídeo está libre de una condición de desbordamiento de memoria tampón en dicho primer punto de acceso;y se seleccionan un valor de dicho cuarto conjunto de múltiples valores, un valor de dicho quinto conjunto de múltiples valores, y un valor de dicho sexto conjunto de múltiples valores tal que dicha secuencia de vídeo está libre de una condición de desbordamiento de memoria tampón en dicho otro punto de acceso. Método de codificación de vídeo según la reivindicación 1, en el que dicho primer conjunto: de múltiples valores (Ri1), dicho segundo conjunto de múltiples valores (Bi1), y dicho tercer conjunto de múltiples valores (Di1) definen al menos un modelo de cubeta con goteo para una memoria tampón de un decodificador hipotético de referencia. Video coding method according to claim 1, wherein said first set of multiple values (Ri1), said second set of multiple values (Bi1), and said third set of multiple values (Di1) define at least one model of trickle bucket for a reference hypothetical decoder buffer. Método de codificación de vídeo según la reivindicación 2, en el que dicho modelo de cubeta con goteo usa una velocidad binaria fija. Video coding method according to claim 2, wherein said trickle bucket model uses a fixed bit rate. Método de codificación de vídeo según la reivindicación 2, en el que dicho modelo de cubeta con goteo usa una velocidad binaria variable. Video coding method according to claim 2, wherein said trickle bucket model uses a variable bit rate.
103 paragraphs in 6 sections, as filed
ES 2 390 596 T3
DESCRIPTION
Video encoder
Background of the invention
The present invention relates to a hypothetical reference decoder.
A digital video system includes a transmitter and a receiver that assembles video consisting of audio, images, and ancillary components for a coordinated presentation to the user. The transmitting system includes subsystems for receiving and compressing the digital source data (the elementary or application data streams, representing the audio, video, and ancillary data components of a program); multiplexing the data from several elementary data streams into a single transport bit stream; and transmitting the data to the receiver. At the receiver, the transport bit stream is demultiplexed into its constituent elementary data streams. The elementary data streams are decoded, and the audio and video data streams are distributed as synchronized program elements, to the receiver's display subsystem, for display as parts of a coordinated program.
In many video coding standards, a decoder-compatible bitstream is decoded by a hypothetical decoder that is conceptually connected to the output of an encoder, and consists of a decoder buffer, a decoder, and a display unit. This virtual decoder is known as a H.263 Hypothetical Reference Decoder (HRD) and a Video Buffering Verifier (VBV) in MPEG-2. The encoder creates a bit stream in such a way that the hypothetical decoder buffer does not overflow or underflow.
As a result, the amount of data that the receiver can be requested to buffer may exceed its capacity (a memory overflow condition), or its throughput capabilities. Alternatively the receiver may not receive, in a data access unit, all the data in time for decoding and synchronized presentation, which at a certain point in time results in data being lost in the audio and video data streams. as well as inconsistent operation (a memory underflow condition).
In existing hypothetical reference decoders, the video bit stream is received at a given, constant bit rate (usually the average bit / second rate of the stream), and stored in the decoder's buffer until the fill reaches a desired level. Such a desired level is denoted as the decoder's initial buffer fill, and is directly proportional to the transmission or startup delay (buffer). At this point, the decoder instantly removes the bits for the first video image in the sequence, decodes the bits, and displays the image. The bits of the following images are also removed, decoded, and instantly displayed in subsequent time slots.
Traditional hypothetical decoders operate at fixed bit rate, buffer, and initial delay. However, in many of today's video applications (for example, streaming video over the Internet or ATM networks) the available bandwidth varies according to the path of the network (for example, depending on how the user connects to the network: by modem, ISDN, DLS, cable, etc.) and it also fluctuates over time according to network conditions (eg congestion, number of connected users, etc.). In addition, video bitstreams are distributed to a variety of devices with different buffer memory capacities (for example handsets, PDAs, PCs, connection modules, DVD-type players, etc.), and are created for scenarios with different delay requirements (eg low delay direct broadcast, progressive discharge, etc.). As a result, these applications require a more flexible hypothetical reference decoder, which can decode a bit stream at different maximum bit rates, and with different buffer sizes and start-up delays.
Jordi Ribas-Corbera and Philip A. Chou, in a document entitled A Generalized Hypothetical Reference Decoder For H.26L, on September 4, 2001 proposed the modified hypothetical reference decoder. The decoder operates according to N sets of rate and buffer parameters for a given bit stream. Each set characterizes what is known as a leaky bucket model and contains three values (R, B, F), where R is the transmission bit rate, B is the buffer size and F is the fill. decoder initial buffer (F / R is the start-up or initial buffer delay). An encoder can create a video bitstream that is contained in certain N desired trickle buckets, or it can simply compute the N sets of parameters after the bitstream has been generated. The hypothetical reference decoder can interpolate between the trickle bucket parameters, and can operate at any desired maximum bit rate, buffer or delay. For example, given a maximum transmission rate R 'the reference decoder can select the minimum buffer and delay (according to the available data from the trickle bucket) with which it will be possible to decode
ES 2 390 596 T3 the bit stream without experiencing buffer overflow or underflow. Conversely, for a given size of buffer B ', the hypothetical decoder can select and operate at the minimum required maximum baud rate.
There are benefits to using such a generalized reference hypothetical decoder. For example, a content provider can create a bitstream once, and a server can distribute it to multiple devices of different capacities, using a variety of different, maximum transmission rate channels. Either a server and a terminal can negotiate the best trickle bucket for the given network conditions - for example, the one that produces the least startup delay (buffer), or the one that requires the least maximum transmission speed for the data size of the device's buffer.
As described in VCEG-58, sections 2.1-2.4, a trickle cuvette is a model for the state (or fill) of an encoder or decoder buffer, as a function of time. The encoder and decoder buffer fills are mutually complementary. A drip bucket model is characterized by three parameters (R, B, F), where:
R is the maximum bit rate (in bits per second) at which bits enter the decoder's buffer. In constant bit rate scenarios, R is often the bit rate of the channel and the average bit rate of the vldeocllp.
B is the size of the decoder bucket or buffer (in bits) that smooths out fluctuations in video bitrate. This buffer size cannot be larger than the physical buffer of the decoding device.
F is the Initial decoder buffer fill (also in bits) before the decoder begins to remove bits from the buffer. F and R determine the Initial or start-up delay D, where D = F / R seconds.
In a dripping bucket model, bits enter the buffer at the rate R, until the fill level is F (ie, for D seconds), and then bO bits are Instantly removed for the first Image. The bits continue to enter the buffer at the rate R, and the decoder removes b1, b2, ..., bn-1 bits for subsequent Images, at certain times data, typically (but not necessarily) every 1 / M seconds. , where M is the frame rate of the video. Figure 1 illustrates the filling of the decoder buffer over time, of a bit stream that is modulated in a trickle bucket, of parameters (R, B, F).
Let B, be the decoder buffer filling immediately before removing b, bits at Time t,. A generic drip bucket model works according to the following equations:
= min (B.Brbi + RÍ ^ -t,)), 1 = 0,1,2.-. (1)
Typically t, + i, t, = 1 / M seconds, where M is the Image rate (usually in images / second) for the bit stream.
A trickle bucket model with parameters (R, R, F) contains a stream of bits, if there is no underflow of the decoder buffer. Because the fills of the encoder and decoder buffers are complementary to each other, this is equivalent to no encoder buffer overflow. However, the encoder buffer (the dripping bucket) is allowed to empty, or equivalently the decoder buffer to fill up, at which point no more bits are transmitted from the encoder buffer to the buffer. decoder. Thus, the decoder buffer stops receiving bits when it is full, which is why the mln operator is included in equation (1). A full decoder buffer simply means that the encoder buffer is empty.
The following observations can be made.
A given video stream can be contained in many drip buckets. For example, if a video stream is contained in a trickle bucket with parameters (R, B, F), it will also be contained in a trickle bucket with a larger buffer (R, B ', F), B'> B, or in a drip tray with a higher maximum transmission speed (R ', B, F), R'> R.
ES 2 390 596 T3
For any bit rate R 'the system can always find a buffer size that contains the video bit stream (limited in time). In the worst case (R 'close to 0), the buffer size will need to be as large as the bit stream itself. In other words, a video bit stream can be transmitted at any frequency (regardless of the average bit rate of the clip), as long as the buffer size is large enough.
The system is assumed to set the relationship F = a B for all drip cuvettes, where a is some desired fraction of the initial buffer fill. For each value of the maximum bit rate R, the system can find the minimum size of buffer B<sub>min</sub> containing the bit stream, using equation (1). Figure 2 shows the representation of the curve of R - B values.
By inspection, the curve of (R<sub>m</sub>¡N, Bm¡n) for any bit stream (like the one in Figure 2) is piecewise linear, and convex. Therefore, if N points of the curve are provided, the decoder can linearly interpolate the values to arrive at some points (R,<sub>n</sub>ter<sub>P</sub>, B¡nter<sub>P</sub>) that are slightly, but certainly, greater than (R<sub>m</sub>¡N, B<sub>m</sub>¡N). In this way the size of the buffer, and consequently also the delay, can be reduced by an order of magnitude relative to a single trickle bucket containing the bit stream at its average frequency. Alternatively, for the same delay, the maximum transmit frequency can be reduced by a factor of four, or possibly even the signal-to-noise ratio improved by several dB.
MPEG Video Buffer Checker (VBV)
The MPEG Video Buffer Verifier (VBV) can operate in two modes: Constant Bit Rate (CBR) and Variable Bit Rate (VBR). MPEG-1 only supports CBR mode, while MPEG-2 supports both nodes.
The VBV works in CBR mode when the bitstream is contained in a bucket model with dripping parameters (R, B, F) and:
R = Rmax = average bit rate of the flow.
The value of B is stored in the syntactic parameter buffer_memory_size_vbv, using a special size unit (namely 16 x 1024 bit units).
The value of F / R is stored in the syntactic element delay_vbv, associated with the first video image in the sequence, using a special unit of time (namely, number of periods of a 90 kHz clock).
Decoder buffer filling follows the following equations:
Bpn. Bj-bf + ¡= 0, 1,2, ... (2)
The encoder has to ensure that B¡ - b¡ is always greater than or equal to zero, while B¡ is always less than or equal to B. In other words, the encoder ensures that the decoder buffer does not present an overflow not an underflow.
The VBV works in VBR mode when the bit stream is limited in a bucket model with parameter dripping (R, B, F) and:
R = Rmax = peak of the maximum frequency. R<sub>ma</sub>x is greater than the average frequency of the bit stream.
F = B, ie the buffer is initially full.
The value of R is represented in the syntactic parameter buffer_memory_size_vbv, as in the case of CBR.
Decoder buffer filling follows the following equations:
ES 2 390 596 T3 = tnin (B, - hi + JM), i = 0,1,2, ... (3)
The encoder ensures that B¡ - b¡ is always greater than or equal to zero. That is, the encoder must ensure that the decoder's buffer does not have underflow. However, in this VBR case the encoder does not need to ensure that the decoder buffer does not overflow. IF the decoder buffer becomes full, then the encoder buffer is assumed to be empty and therefore bits are no longer transmitted from the encoder buffer to the decoder buffer.
VBR mode is useful for devices that can read data up to the maximum frequency R<sub>ma</sub>x For example, a DVD includes VBR clips in which R<sub>ma</sub>x is approximately 10 Mb / s, which corresponds to the maximum read speed of the disc drive, even though the average frequency of the DVD video stream is only about 4 Mb / s.
Referring to Figure 3A and Figure 3B, decoder buffer filling graphs are shown for some operational bitstreams in CBR and VBR mode.
Broadly speaking, the CBR mode can be considered a special case of VBR where R<sub>max</sub> is the average frequency of the clip.
H.263 Hypothetical Reference Decoder (HRD)
The hypothetical reference model for H.263 is similar to the MPEG VBV CBR mode discussed previously, except for the following.
The decoder inspects the buffer filling at certain time intervals, and decodes an image as soon as all the bits are available for the image. This approach results in a couple of benefits: (a) the delay is minimized because F is usually only slightly greater than the number of bits for the first image, and (b) if skipping is common, the decoder simply waits until the next available image. The latter is also enabled in the low-delay mode of MPEG VBV.
The buffer overflow check is performed after the bits for an image are removed from the buffer. This relaxes the limitation of sending large images once every so often, but there is a maximum value for the largest image.
The H.263 HRD can essentially be mapped to a low-delay, trickle-bucket type of model. Limitations of Previous Hypothetical Reference Decoders
Previously existing hypothetical reference decoders operate at a single point (R, B) on the curve in Figure 2. As a result, these decoders have the following drawbacks:
if the bit rate available on the R 'channel is lower than that available on R (for example, this is common in internet streaming and progressive downloading, or when it is required to stream a VBR clip in MPEG to a frequency lower than maximum), strictly speaking the hypothetical decoder would not be able to decode the bit stream.
If the available bandwidth R 'is greater than R (for example, this is common also for streaming over the Internet, as well as for local playback), the previous hypothetical decoders could work in VBR mode and decode the bit stream. However, if more information on the Frequency-Buffer curve were available, the buffer size and associated startup delay required to decode the bit stream would be significantly reduced.
If the physical size of the buffer in a decoding device is less than R, the device will not be able to decode such a bit stream.
If the buffer size is larger than B, the device will be able to decode the bit stream but the startup delay will be the same.
ES 2 390 596 T3
More generally, a bit stream that has been generated according to a trickle bucket (R, B, F) will usually not be capable of being distributed across different networks of bit rate less than R, and to a variety of devices with buffer sizes smaller than B. In addition, the startup delay will not be minimized.
Hypothetical Generalized Reference Decoder (GHRD)
A hypothetical generalized reference decoder (GHRD) can work given the information of N drip bucket models,
<img file="ES2390596T3_D0001.tif" />
each of which contains the bit stream. Without losing generality, we assume that these dripping buckets are ordered from the lowest to the highest bit rate, that is, R<sub>i</sub> <R<sub>i</sub> + <sub>1</sub>. We also assume that the encoder computes these drip tray models correctly, and therefore B<sub>i</sub> <B<sub>i</sub> + <sub>1</sub>.
The desired value of N can be selected by the encoder. If N = 1, the GHRD is essentially equivalent to the VBV of MPEG. The encoder can choose to: (a) preset trickle bucket values and encode the bit stream with a speed control, ensuring that all drip bucket limitations are satisfied, (b) encode the bit stream, and then use equation (1) to compute a set of trickle buckets containing the stream of bits at N different values of R, or (c) do both. The first approach (a) can be applied to live or on-demand streaming, while (b) or (c) applies only to on-demand streaming.
The number N of drip cuvettes and the drip cuvette parameters (4) are inserted into the bit stream. In this way, the decoder can determine which trickle bucket it wants to use, knowing the maximum bit rate available for it, and / or its physical buffer size. The drip cuvette models at (4), as well as all linearly interpolated or extrapolated models, are of viable use. Figure 4 illustrates a set of N drip cuvette models, and their interpolated or extrapolated values (R, B).
The interpolated buffer size B between points k and k + 1, follows the straight line:
5- {(Rk + rR) / (Rk + rRk)} Bk + {(R-Rk) / (Rk # i-Rk)} Bk<sub>+</sub>go<sub>k</sub><R <R<sub>k + 1</sub>
Similarly, the initial filling of the decoder F buffer can be linearly interpolated:
F = {(Rk<sub>+</sub>iR) / (Rm-R<sub>k</sub>)}F<sub>k</sub> + {(RR<sub>k</sub>) / (R<sub>k + r</sub>R<sub>k</sub>)}F<sub>k + 1</sub> R<sub>k</sub><R <R<sub>k + I</sub>
The resulting trickle bucket with parameters (R, B, F) contains the bit stream, due to which the minimum buffer size B<sub>min</sub> is convex in both R and F, i.e. the minimum buffer size B<sub>min </sub>corresponding to any convex combination (R, F) = a (R<sub>k</sub>, F<sub>k</sub>) + (1 - a) (R<sub>k</sub>+<sub>1</sub>, F<sub>k</sub>+<sub>1</sub>), 0 <a <1, is less than or equal to B = a Bk + (1 - a) Bk + 1.
It is observed that if R is greater than R<sub>N</sub>, the drip tray (R, B<sub>N</sub>, F<sub>N</sub>) will also contain the bit stream, and therefore BN and FN are the buffer size and the initial buffer fill of the decoder, recommended when R> = R<sub>N</sub>. If R is less than R<sub>l</sub> the upper limit B = B + (R<sub>l</sub> - R) T (and the time can be set F = B), where T is the time duration of the flow in seconds. These values (R, B) outside the range of the N points are also shown in figure 4.
The document Working Draft Number 2, revision 0 (WD-2), of the Joint Video Team of ISO / IEC MPEG and ITU-T VCEG, incorporates many of the concepts of the hypothetical reference decoder proposed by Jordi Ribas-Cobera et al., from Microsoft Corporation, incorporated herein by reference. The WD-2 document is similar to the decoder proposed by Jordi Ribas-Cobera et al., From Microsoft Corporation, although the syntax is somewhat modified. Furthermore, WD-2 describes an exemplary algorithm to compute B, and F for a given frequency R.
US Patent Application Publication No. US 2003/0 053 416 A describes the use of two drip tray models. The first set of trickle bucket parameters would allow video transmission on a constant bit rate channel, with a delay of about 22.5 seconds. The second set of trickle bucket parameters would allow video transmission over a shared network, with a maximum speed of 2500 kbps, or it would allow local playback from a 2x CD, with a delay of about 0.9 seconds.
ES 2 390 596 T3
Summary of the invention
According to one aspect of the present invention, a video coding method is provided comprising defining a first set of multiple characteristic values of a transmission bit rate for a first access point at a starting point of a video sequence. ; defining a second set of multiple characteristic values of a buffer size for said first access point; defining a third set of multiple characteristic values of a delay until a video sequence is presented for said first access point; defining a fourth set of multiple characteristic values of a transmission bit rate for another access point located after said first access point and thereafter; defining a fifth set of multiple characteristic values of a buffer size for said other access point; define a sixth set of multiple values characteristic of a delay until a video sequence is presented for said other access point in which a value from said first set of multiple values is selected, a value from said second set of multiple values, and a value of said third set of multiple values such that said video sequence is free from a buffer overflow condition at said first access point; and a value from said fourth multi-value set, a value from said fifth multi-value set, and said sixth multi-value set are selected such that said video sequence is free from a buffer overflow condition at said other point access.
Preferred features of the invention are set out in the dependent claims
Brief description of the drawings
Figure 1 illustrates a decoder buffer fill.
Figure 2 illustrates a R - B curve.
Figures 3A and 3B are representations of decoder buffer filling, for some operational bit streams respectively in CBR and VBR modes.
Figure 4 illustrates a set of N drip cuvette models, and their interpolated or extrapolated values (R, B).
Figure 5 illustrates the decoder's initial buffering Bj, for whatever point the user looks for when the rate is Rj.
Figure 6 illustrates sets of (R, B, F) defined in forward mode, for the particular video stream.
Figure 7 illustrates the initial buffer fill (in bits) for a video segment.
Figure 8 illustrates the selection criteria in a set of 10 points for Figure 7.
Figure 9 illustrates selection criteria.
Figure 10 illustrates reductions in delay.
Detailed description of the preferred embodiment
As previously described, the JVT standard (WD-2) allows the storage of (N> = 1) drip cuvettes, values (R1, B1, F1), ..., (Rn, Bn, Fn) that are contained in the bit stream. These values can be stored in the header. Using Fi as the initial buffer fill and Bi as the buffer size, ensures that the decoder buffer will not underflow when the input stream enters the Ri rate. This will be the case if the user wants to present the encoded video from start to finish. In a typical video-on-demand application, the user may want to search for different parts of the video stream. The point that the user wishes to search can be referred to as the access point. During the process of receiving video data and constructing video images, the amount of data in the buffer fluctuates. Upon consideration, the present inventor comes to the understanding that if the initial buffer fill value Fi (when the channel rate is R,) is used before starting to decode the video from the access point, then the decoder may have an underflow. For example at the access point or something thereafter, the number of bits required for video reconstruction may be larger than the bits currently in the buffer, resulting in underflow and inability to display video images. in a timely manner. Likewise, it can be shown that in a video stream the initial buffer fill value, necessary to ensure that there is no underflow in the decoder, varies as a function of the point the user is looking for. This value is limited by B¡. Therefore, the combination of B and F provided for the entire video sequence will probably not be appropriate if used for an intermediate point in the video, resulting in underflow and thus freezing the images.
ES 2 390 596 T3
Based on this potential for underflow not previously understood, the present inventor then comes to the realization that if only one set of R, B and F values is defined for an entire video segment, then the system should wait until memory buffer B for the corresponding rate R is full, or substantially full (or full above 90%), to start decoding images when a user jumps to an access point. In this way the initial buffer fill will be at maximum, and therefore there is no potential for underflow during subsequent decoding starting from the access point. This can be achieved without any additional change to the existing bit stream, and therefore without impact on existing systems. Therefore, the decoder would use the initial buffering value Bj, for whatever point the user looks for, when the rate is Rj, as shown in Figure 5. Unfortunately, however, this sometimes results in a delay. meaningful until the images are presented, after selecting a different location (for example, an access point) from which to present the video.
The initial buffer fill (F) can likewise be characterized as a delay until the video sequence is presented. The delay is temporary in nature, being related to the time required to achieve the initial buffer fill (F). Delay and / or F can be associated with all video or access points. Similarly, it should be understood that delay can be substituted for F in all embodiments described herein (eg (R, B, delay)). A specific value for the delay can be calculated as delay = F / R, using a special unit of time (90 kHz clock units).
To reduce the potential delay, the present inventor came to the understanding that sets of (R, B, F) can be defined for a particular video stream, at each access point. Referring to Figure 6, these sets of (R, B, F) are preferably defined in forward mode, for the particular video stream. For example, the set of values (R, B, F) can be computed in the previously existing way for the video stream as a whole, in addition to the set of values F for the same values (R, B) as those of the entire video stream, it can be computed in the previously existing way for the video stream, relative to the video stream from position 2 forward, and so on. The same process can be used for the other access points. The access points can be any picture within the video stream, I pictures of the stream, B pictures of the stream, or P pictures of the stream (I, B, and P pictures are typically used in video decoding based on MPEG). Consequently, the user can select one of the access points, and then use the respective Fj for the desired initial fill (assuming that the buffer Bi and the rate Ri remain unchanged), or else a set of two or more than Ri, Bi, Fij.
Index i represents each trickle cuvette, and index j represents each random access point. Assuming that the buffer Bi and the speed Ri remain unchanged, the header stores the multiple set of values (Bi, Ri, F, i), where i = 1, 2, ..., N and F, i represents initial buffer filling. Then, in the access point j, Fij is stored, where j = 2, 3, ... On the other hand, assuming that the buffer Bi and the rate Ri will be modified in each access point j, multiple sets of values of (Rij, Bij, Fij) can be stored in each access point. The benefit of the first case is that you save on the amount of data, since only one multiple set of Fij has been stored in each access point, and the benefit of the last case is that you can adjust the set of values more appropriately for each access point. When using the delay (D) until the video sequence is displayed, instead of the initial filling of the buffer (F), it is possible to carry out the present invention by replacing Fij with Dj. In this case Dij represents the value of the delay. Thus, when it is assumed that the buffer Bi and the rate R remain unchanged, (Bi, Ri, Dii) is stored in the header and Dij is stored in each access point j. When it is assumed that the buffer Bi and the rate R change at each access point j, a multiple set of values (Bij, Rij, Dij) can be stored at the access point.
The sets of R, B, F values for each access point can be located at any appropriate location, such as for example at the beginning of the video stream, together with sets of values (R, B, F) for the entire video stream, or before each access point which avoids the need for an index; or be stored external to the video stream itself, which is especially suitable for a server / client environment.
This technique can be characterized by the following model:
(Ri, Bi, Fi, Mi, fii, tii, ..., Ímii, Ímii) ..., (Rn, Bn, Fn, Mn, fiN, tiN, ..., Ímnn, Ímnn), where fkj denotes the initial buffer fill value at rate Rj at access point tkj (timestamp). The Mj values can be provided as an input parameter, or they can be selected automatically. For example, Mj can include the following options:
(a) Mj can be set to the value equal to the number of access points. In this way the values of fkj can be stored for each access point, at each rate Rj (either at the beginning of the video stream, within the video stream, distributed across the video stream, or with any other location).
(b) Mj can be set equal to zero, otherwise search support is desired.
ES 2 390 596 T3 (c) Mj values can be selected automatically for each speed Rj (described below).
For a given Rj, the system can use an initial buffer fill equal to f,<sub>k</sub> if the user searches for a tkj access point. This occurs when the user selects to start at an access point, or when the system adjusts the user's selection to one of the access points.
It is noted that in the case where a variable bit rate (in bit stream) is used, preferably the initial buffer fill value (or delay) is different from the buffer size (or the delay calculated by the buffer size), although it may be the same. In the case of a variable bit rate in MPEG-2 VBV, the buffer is filled until it is full, that is, F = B (the value of B is represented by vbv_buffer_memory_size).
In the present invention, the initial fill value of the buffer size can be chosen appropriately at each random access point to avoid any underflow or overflow of the buffer. When using the delay until the video sequence is presented, instead of filling the buffer, the delay value is chosen appropriately at each random access point, to avoid any underflow or buffer overflow. Generally, this means that each random access point achieves less delay than completely filling the VBV buffer. Therefore, determining the buffer fill value (or the delay) which is less than the buffer size (or the delay calculated by the buffer size) by the present invention has the advantage of a reduced delay, since it is required to store less data in the buffer before the beginning of the decoding, than in the prior art.
If the system allows the user to jump to any image in the video, as an access point, then the decoding data set would need to be provided for each and every image. If allowable, the resulting data set would be excessively large and consume a significant amount of the available bit rate for the data. A more reasonable approach would be to limit the user to specific access points within the video stream, such as every second, every 10 seconds, every minute, and so on. Although an improvement, the resulting data set can still be somewhat large, resulting in excessive data for devices with limited bandwidth, such as mobile communication devices.
In the case that the user selects a position that is one of the access points with an associated data set, then the initial buffer fill can be equal to max (fkj, f (k + i) j) for a time between tkj and t (k + i) j, especially if the access points are properly selected. This ensures that the system has a set of values that will be free from resulting in an underflow condition, or otherwise reduce the probability of an underflow condition as will be explained below.
Reference is made to Figure 7, to select a set of values that ensures that an underflow condition does not occur (or in another case, that it is reduced) when the selection criteria referred to above has been used. Figure 7 illustrates the initial buffer fill (in bits) for a video segment, where the initial buffer fill is calculated in advance, for increments of 10 seconds. The system then preferably selects an access point at the beginning of the video stream, and an access point at the end of the video segment. Between the beginning and the end of the video segment, the system selects local maxima for inclusion as access points. Additionally, the system can select local minima for inclusion as access points. Preferably, if a limited set of access points is desired the system selects first the local maxima and then the local minima, which helps to ensure that underflow does not occur. Then, if desired, the system can also select intermediate points.
Based on the selection criteria, a set of 10 points can be selected for Figure 7, as indicated in Figure 8. Referring to Figure 9, the 10 selected points are shown by the dashed curve. The resulting initial buffer fill values at all access points are shown by the solid curve. The solid curve illustrates a safe set of values for all access points in the video so that the decoder buffer does not underflow. If there have been extreme fluctuations in the bit rate of the current bit stream, which were not detected in the processing, such as sharp spikes, then an underflow is possible, although this is usually unlikely. The optimal initial buffer fill values at all access points are shown by a dotted curve. A significant reduction in buffering time delay is achieved, in contrast to requiring a full buffer when accessing an access point, as illustrated in Figure 10.
Also, if the bit rate and buffer size remain the same while a different access point is selected, then the modified buffer fill, F.
Contents6
6 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6
42 members in 11 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 404947 | United States of America | – | |
| 40494703 | United States of America | A | |
| 40494703 | United States of America | A | |
| 404947 | – | – | – |
| US20030404947 | – | – | – |
Members42
| Document | Office | Kind | |
|---|---|---|---|
| US2004190606A1 | United States of America | A1 | |
| WO2004088988A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1611747A1 | European Patent Office (EPO) | A1 | |
| EP1611747A4 | European Patent Office (EPO) | A4 | |
| JP2006519517A | Japan | A | |
| CN1826812A | China | A | |
| EP1791369A2 | European Patent Office (EPO) | A2 | |
| EP1791369A3 | European Patent Office (EPO) | A3 | |
| JP2007189731A | Japan | A | |
| US7266147B2 | United States of America | B2 | |
| EP1611747B1 | European Patent Office (EPO) | B1 | |
| PT1611747E | Portugal | E | |
| AT390019T | Austria | T | |
| DE602004012540D1 | Germany | D1 | |
| ES2300757T3 | Spain | T3 | |
| DE602004012540T2 | Germany | T2 | |
| EP1791369B1 | European Patent Office (EPO) | B1 | |
| AT472228T | Austria | T | |
| EP2209319A2 | European Patent Office (EPO) | A2 | |
| CN1826812B | China | B | |
| JP2010166600A | Japan | A | |
| DE602004027847D1 | Germany | D1 | |
| PT1791369E | Portugal | E | |
| CN101854552A | China | A | |
| CN101854553A | China | A | |
| ES2348075T3 | Spain | T3 | |
| EP2209319A3 | European Patent Office (EPO) | A3 | |
| HK1145413A1 | Hong Kong, China | A1 | |
| HK1147374A1 | Hong Kong, China | A1 | |
| USRE43062E | United States of America | E | |
| EP2209319B1 | European Patent Office (EPO) | B1 | |
| JP2012135009A | Japan | A | |
| PT2209319E | Portugal | E | |
| JP5025289B2 | Japan | B2 | |
| ES2390596T3This record | Spain | T3 | |
| PL2209319T3 | Poland | T3 | |
| JP5444047B2 | Japan | B2 | |
| JP5536811B2 | Japan | B2 | |
| CN101854553B | China | B | |
| USRE45983E | United States of America | E | |
| CN101854552B | China | B | |
| USRE48953E | United States of America | E |
Numbers
- Publication
- 2390596
- Publication, DOCDB
- 2390596
- Publication, EPODOC
- ES2390596T
- Application
- 10162205
- Application, DOCDB
- 10162205
- Application, EPODOC
- ES20100162205T
Titles2
- English
- Video encoder
- Spanish
- Codificador de vídeo
Classification
- CPC, 6
- H04N19/00
- H04N21/44004
- H04N19/15
- H04N19/61
- H04N19/44
- H04N21/2401
- IPC, 2
- H04N7 50
- H04N7 26