Method and apparatus for variable accuracy inter-picture timing specification for digital video encoding with reduced requirements fo division operations
Abstract
A method and apparatus for performing motion estimation in a digital video system is disclosed (Fig.1). Specifically, the present invention discloses a system that quickly calculates estimated motion vectors in a very efficient manner (Fig.1, item 160). In one embodiment, a first multiplicand is determined by multiplying a first display time difference between a first video picture and a second video picture by a power of two scale value (Fig.1, items 150, 160). This step scales up a numerator for a ratio (Fig.1, item 120). Next, the system determines a scaled ratio by divi ding that scaled numerator by a second first display time difference between the second video picture and a third video picture. The scaled ratio is then stored calculating motion vector estimations. By storing the scaled ratio, all the estimated motion vectors can be calculated quickly with good precision since the scaled ratio saves significant bits and reducing the scale is performed by simple shifts (Fig.1, item 180).

Term
Term ended
Expired 7 August 2023, 3.1 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
91 claims: 13 independent, 78 dependent
- 1CA 02502004 2011-07-15 The embodiments of the invention in which an exclusive property or privilege is claimed are defined as follows:1. For a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, a method comprising: computing a scaling value that is based on (i) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein an order value for a video picture specifies a display order for the video picture;and computing a motion vector for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the motion vector for the second video picture comprises performing a bit shifting operation.
- 20For a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, a method comprising:computing a scaling value that is based on (i) a power of two value, (ii) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (iii) a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein the order value for the second video picture is encoded in a slice header associated with the second video picture;and computing a motion vector for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the motion vector for the second video picture comprises performing a bit shifting operation.
- 26For a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, a method comprising:computing a scaling value that is based on (i) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein an order value for a video picture is encoded in a slice header of a bitstream associated with the video picture;and computing a particular motion vector for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the particular motion vector comprises performing a division by a power of two value.
- 32A method comprising:computing a scaling value based on (i) a first order difference value between an order value for a first video picture and an order value for a second video picture, and (ii) a particular value that is based on a power of two value and a second order difference value between an order value for a third video picture and the order value for the second video picture, wherein the order value for the second video picture is encoded in a slice header of a bitstream associated with the second video picture;and using said scaling value to compute a motion vector for the second video picture by performing a bit shifting operation.
- 38For a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, the method comprising:computing a scaling value that is (i) inversely proportional to a first order difference value between an order value for the third video picture and an order value for the first video picture and (ii) directly proportional to a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein the order value for the second picture is encoded in a bitstream more than once;and computing a motion vector for the second video picture by multiplying the scaling value by a motion vector for the third video picture and performing a division operation that is based on a power of two value.
- 43The method of 38, wherein the order value is compressed in the bitstream using variable length coding.
- 45For a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, a method comprising:computing a scaling value that is (i) inversely proportional to a first order difference value between an order value for the third video picture and an order value for the first video picture, (ii) directly proportional to a second order difference value between an order value for the second video picture and the order value for the first video picture, and (iii) directly proportional to a power of two value, wherein an order value for a video picture specifies a display order for the video picture;and computing a motion vector for the second video picture by multiplying the scaling value and a motion vector for the third video picture and performing a division operation based on said power of two value.
- 52For a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, a method comprising:computing a scaling value that is based on (i) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein computing the scaling value is subject to a truncation operation, wherein the order value for the second video picture specifies a display order for the second video picture;computing a particular motion vector for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the particular motion vector for the second video picture comprises performing a bit shifting operation.
- 59For a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, a method comprising:subject to a truncation operation, computing a scaling value that is (i) inversely proportional to a first order difference value between an order value for the third video picture and an order value for the first video picture and (ii) directly proportional to a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein the order value for the second video picture is encoded in a slice header of a bitstream associated with the second video picture;and computing a motion vector for the second video picture by multiplying the scaling value by a motion vector for the third video picture and performing a bit shifting operation.
- 65For a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, the method comprising:computing a scaling value that is (i) inversely proportional to a first order difference value between an order value for the third video picture and an order value for the first video picture and (ii) directly proportional to a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein the order value for the second video picture is encoded in a bitstream by using an exponent of a power of two integer;and computing a motion vector for the second video picture by multiplying the scaling value by a motion vector for the third video picture and performing a bit shifting operation.
- 70The method of 69, wherein the order value is compressed in the bitstream using variable length coding. -34CA 02502004 2011-07-15
- 72A method of decoding a bitstream comprising encoded first, second, and third video pictures, the method comprising:receiving an integer value representing an exponent of a particular power of two integer;computing a scaling value that is based on (i) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein the order value for the second video picture is derived from the particular power of two integer;computing a motion vector for the second video picture by multiplying the scaling value by a motion vector for the third video picture and performing a bit shifting operation;and decoding the second video picture by using the computed motion vector.
- 81A method for encoding a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, the method comprising:computing a scaling value that is based on (i) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture;computing a particular motion vector for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the particular motion vector for the second video picture comprises performing a bit shifting operation;encoding the second video picture in a bitstream by using the computed motion vector;and encoding the order value for the second video picture in the bitstream by using an exponent of a power of integer.
Independent claims13
176 paragraphs in 64 sections, as filed
CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
Method and Apparatus for Variable Accuracy Inter-Picture Timing Specification for Digital Video Encoding With Reduced Requirements fo Division Operations
RELATED APPLICATIONS
This patent application claims the benefit of an earlier filing date under title 35, United States Code, Section 120 to the United States Patent Application having serial number 10/313,773 filed on December 6, 2002.
FIELD OF THE INVENTION
The present invention relates to the field of multimedia compression systems. In particular the present invention discloses methods and systems for specifying variable accuracy inter-picture timing with reduced requirements for processor intensive division operation.
BACKGROUND OF THE INVENTION
Digital based electronic media formats are finally on the cusp of largely replacing analog electronic media formats. Digital compact discs (CDs) replaced analog vinyl records long ago. Analog magnetic cassette tapes are becoming increasingly rare. Second and third generation digital audio systems such as Mini-discs and MP3 (MPEG Audio - layer 3) are now taking market share from the first generation digital audio format of compact discs.
-2CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
The video media formats have been slower to move to digital storage and digital transmission formats than audio media. The reason for this slower digital adoption has been largely due to the massive amounts of digital information required to accurately represent acceptable quality video in digital form and the fast processing capabilities needed to encode compressed video. The massive amounts of digital information needed to accurately represent video require very high-capacity digital storage systems and high-bandwidth transmission systems.
However, video is now rapidly moving to digital storage and transmission formats. Faster computer processors, high-density storage systems, and new efficient compression and encoding algorithms have finally made digital video transmission and storage practical at consumer price points. The DVD (Digital Versatile Disc), a digital video system, has been one of the fastest selling consumer electronic products in years.
DVDs have been rapidly supplanting Video-Cassette Recorders (VCRs) as the prerecorded video playback system of choice due to their high video quality, very high audio quality, convenience, and extra features. The antiquated analog NTSC (National Television Standards Committee) video transmission system is currently in the process of being replaced with the digital ATSC (Advanced Television Standards Committee) video transmission system.
Computer systems have been using various different digital video encoding formats for a number of years. Specifically, computer systems have employed different video coder/decoder methods for compressing and encoding or decompressing
-3CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 and decoding digital video, respectively. A video coder/decoder method, in hardware or software implementation, is commonly referred to as a “CODEC”.
Among the best digital video compression and encoding systems used by 5 computer systems have been the digital video systems backed by the Motion Pictures
Expert Group commonly known by the acronym MPEG. The three most well known and highly used digital video formats from MPEG are known simply as MPEG-1, MPEG-2, and MPEG-4. VideoCDs (VCDs) and early consumer-grade digital video editing systems use the early MPEG-1 digital video encoding format. Digital Versatile Discs (DVDs) and the Dish Network brand Direct Broadcast Satellite (DBS) television broadcast system use the higher quality MPEG-2 digital video compression and encoding system. The MPEG4 encoding system is rapidly being adapted by the latest computer based digital video encoders and associated digital video players.
The MPEG-2 and MPEG-4 standards compress a series of video frames or video fields and then encode the compressed frames or fields into a digital bitstream. When encoding a video frame or field with the MPEG-2 and MPEG-4 systems, the video frame or field is divided into a rectangular grid of pixelblocks. Each pixelblock is independently compressed and encoded.
When compressing a video frame or field, the MPEG-4 standard may compress the frame or field into one of three types of compressed frames or fields: Intraframes (I-frames), Unidirectional Predicted frames (P-frames), or Bi-Directional Predicted frames (B-frames). Intra-frames completely independently encode an independent video frame with no reference to other video frames. P-frames define a —4—
CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 video frame with reference to a single previously displayed video frame. B-frames define a video frame with reference to both a video frame displayed before the current frame and a video frame to be displayed after the current frame. Due to their efficient usage of redundant video information, P-frames and B-frames generally provide the best compression.
-5CA 02502004 2011-07-15
SUMMARY OF THE INVENTION
A method and apparatus for performing motion estimation in a video codec is disclosed. Specifically, the present invention discloses a system that quickly calculates estimated motion vectors in a very efficient manner without requiring an excessive number of division operations.
In one embodiment, a first multiplicand is determined by multiplying a first display time difference between a first video picture and a second video picture by a power of two scale value. This step scales up a numerator for a ratio. Next, the system determines a scaled ratio by dividing that scaled numerator by a second first display time difference between said second video picture and a third video picture. The scaled ratio is then stored to be used later for calculating motion vector estimations. By storing the scaled ratio, all the estimated motion vectors can be calculated quickly with good precision since the scaled ratio saves significant bits and reducing the scale is performed by simple shifts thus eliminating the need for time consuming division operations.
In another aspect, the present invention provides for a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, a method comprising: computing a scaling value that is based on (i) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein an order value for a video picture specifies a display order for the video picture; and computing a motion vector
-6CA 02502004 2011-07-15 for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the motion vector for the second video picture comprises performing a bit shifting operation.
In a further aspect, the present invention provides for a sequence of video pictures 5 comprising a first video picture, a second video picture, and a third video picture, a method comprising: computing a scaling value that is based on (i) a power of two value, (ii) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (iii) a second order difference value between an order value for the second video picture and the order value for the first video picture wherein the order value for the second video picture is encoded in a slice header associated with the second video picture; and computing a motion vector for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the motion vector for the second video picture comprises performing a bit shifting operation.
In a still further aspect, the present invention provides a method comprising:
computing a scaling value based on (i) a first order difference value between an order value for a first video picture and an order value for a second video picture, and (ii) a particular value that is based on a power of two value and a second order difference value between an order value for a third video picture and the order value for the second video picture, wherein the order value for the second video picture is encloded in a slice header of a bitstream associated with the second video picture; and using said scaling value to compute a motion vector for the second video picture by performing a bit shifting operation.
-6a—
CA 02502004 2009-08-20
In a further aspect, the present invention provides a method comprising: a.
receiving a plurality of video pictures and at least one order value; and b. computing an implicit B prediction block weighting for a video picture based on said at least one order value.
In a still further aspect, the present invention provides a method comprising: a. decoding a plurality of video pictures by using at least one order value, said order value is for establishing an ordering for reference video picture selection; and c. outputting the decoded video pictures based on the order value.
In a further aspect, the present invention provides a method of decoding coded video data, comprising: calculating, for a current frame of video data to be decoded, a scale factor (Z) based on a ratio of two time differences (Δίι/Δΐ2), the first time difference (Δίι) representing a temporal difference between the current frame and a first reference frame and the second time difference (Δί2) representing a temporal difference between the first reference frame and a second reference frame, the scale factor representing the ratio multiplied by a power of two (Ζ=2<sup>Ν</sup>*(Δί]/Δί2)) subject to rounding, predicting data of pixelblocks of the current frame from data of the reference frames according to predictive coding techniques, further comprising, as part of the prediction, interpolating a motion
-6bCA 02502004 2009-08-20 vector (mvpB) of the pixelblocks from a motion vector (hivref) of a co-located pixelblock in one of the reference frames as mvpB = mvREF*Z/2<sup>N</sup>.
In a still further aspect, the present invention provides a method of decoding coded video data, comprising: calculating, for a current frame of video data to be decoded, a scale factor (Z) based on a ratio of two time differences (Δί[/ΔΪ2), the first time difference (Δίι) representing a temporal difference between the current frame and a first reference frame and the second time difference (Δί2) representing a temporal difference between the first reference frame and a second reference frame, the scale factor representing the ratio multiplied by a power of two (Ζ=2<sup>Ν</sup>*(Δίι/Δΐ<sub>2</sub>)) subject to rounding, predicting data of pixelblocks of the current frame from data of the reference frames according to predictive coding techniques, as part of the prediction, determining whether motion vectors for the pixelblock are to be interpolated from motion vectors of the reference frames, if so, interpolating motion vectors (mv<sub>PB</sub>[j]) of the respective pixelblocks from motion vector (mvREF[ij) of a co-located pixelblock in one of the reference frames as mvp<sub>B</sub> = mv<sub>RE</sub>F*Z/2<sup>N </sup>wherein the scale factor Z is common to motion vector derivations of all pixelblocks i in the current frame.
In a further aspect, the present invention provides a video decoder, comprising: a motion estimator to calculate, for a current frame of video data to be decoded, a scale factor (Z) based on a ratio of two time differences (Atj/At<sub>2</sub>), the first time difference (Δίι )
-6cCA 02502004 2009-08-20 representing a temporal difference between the current frame and a first reference frame and the second time difference (At<sub>2</sub>) representing a temporal difference between the first reference frame and a second reference frame, the scale factor representing the ratio multiplied by a power of two (Ζ=2<sup>Ν</sup>*(Δί<sub>1</sub>/Δί<sub>2</sub>)) subject to rounding, a predictor to predict data of pixelblocks of the current frame from data of the reference frames according to predictive coding techniques, wherein the estimator further interpolates a motion vector (mvp<sub>B</sub>) of the pixelblocks from a motion vector (mv<sub>REf</sub>.) of a co-located pixelblock in one of the reference frames as mvp<sub>B</sub> = mvREF*Z/2<sup>N</sup>.
In a still further aspect, the present invention provides a video decoder, comprising:
a motion estimator to calculate, for a current frame of video data to be decoded, a scale factor (Z) based on a ratio of two time differences (Δΐ]/Δί<sub>2</sub>), the first time difference (Δΐι) representing a temporal difference between the current frame and a first reference frame and the second time difference (Δί<sub>2</sub>) representing a temporal difference between the first reference frame and a second reference frame, the scale factor representing the ratio multiplied by a power of two (Ζ=2<sup>Ν</sup>*(Δΐι/Δΐ<sub>2</sub>)) subject to rounding, a predictor to predict data of pixelblocks of the current frame from data of the reference frames according to predictive coding techniques, as part of the prediction, determining whether motion vectors for the pixelblock are to be interpolated from motion vectors of the reference frames, wherein the estimator further interpolates motion vectors (mvp<sub>B</sub>[ij) of the respective pixelblocks from motion vector (mvREF[ij) of a co-located pixelblock in one of the reference ~6d~
CA 02502004 2009-08-20 frames as mv<sub>PB</sub> = mvREF*Z/2<sup>N</sup>, wherein the scale factor Z is common to motion vector derivations of all pixelblocks i in the current frame.
In a further aspect, the present invention provides computer readable medium 5 storing program instructions that, when executed by a processing device, cause the device to: calculate, for a current frame of video data to be decoded, a scale factor (Z) based on a ratio of two time differences (Δίι/Δί<sub>2</sub>), the first time difference (Af) representing a temporal difference between the current frame and a first reference frame and the second time difference (At<sub>2</sub>) representing a temporal difference between the first reference frame and a second reference frame, the scale factor representing the ratio multiplied by a power of two (Z=2<sup>N</sup>*(At]/At<sub>2</sub>)) subject to rounding, predict data of pixelblocks of the current frame from data of the reference frames according to predictive coding techniques, and as part of the prediction, interpolate a motion vector (mv<sub>PB</sub>) of the pixelblocks from a motion vector (iuvref) of a co-located pixelblock in one of the reference frames as mv<sub>PB</sub> = mvRBF*Z/2<sup>N</sup>.
In a still further aspect, the present invention provides a method of coding a sequence of video data: interpolating motion vectors for a pixelblock in a current frame, the pixelblock to be coded according to bi-directional prediction with reference to a pair of reference frames in the sequence, wherein the interpolation includes for at least one motion vector: generating a scale factor (Z) based on a ratio of time differences (Ati/At<sub>2</sub>) among —6e—
CA 02502004 2009-08-20 the current frame and the pair of reference frames, the first time difference (Δίι) representing a temporal difference between the current frame and a first of the reference frame and the second time difference (Δί<sub>2</sub>) representing a temporal difference between the two reference frames, the scale factor representing the ratio multiplied by a power of two (Ζ=2<sup>Ν</sup>*(Δί]/Δΐ<sub>2</sub>)) subject to rounding, generating a first motion vector for the pixelblock (mvpBi) according to a first derivation based on a motion vector extending between the reference frames (mvRFF), the first derivation being mvpBi = mvRE<sub>F</sub>*Z/2<sup>N</sup>, generating a second motion vector for the pixelblock (mvp<sub>B2</sub>) based on a second derivation mvREF, predicting video data for the pixelblock from video data of the first and second reference frames according to the generated motion vectors mv<sub>PB!</sub> and mv<sub>PB2</sub>, coding a difference between actual video data of the pixelblock and predicted video data of the pixel block, and outputting the coded difference to a channel as coded video data of the pixelblock.
In a further aspect, the present invention provides a method of coding a sequence of video data: coding a first reference frame at a first temporal location in the sequence, coding a second reference frame at a second temporal location in the sequence, interpolating motion vectors for a pixelblock in a third, current frame, the pixelblock to be coded according to bi-directional prediction with reference to the pair of reference frames, wherein the interpolation includes for at least one motion vector: generating a scale factor (Z) based on a ratio of time differences (Δίι/Δΐ<sub>2</sub>) among the current frame and the reference frames, the first time difference (Δίι ) representing a temporal difference between the —6f—
CA 02502004 2009-08-20 current frame and the first reference frame and the second time difference (Δΐ<sub>2</sub>) representing a temporal difference between first and second reference frames, the scale factor representing the ratio multiplied by a power of two (Ζ=2<sup>Ν</sup>*(ΔΤ/Δί<sub>2</sub>)) subject to rounding, generating a first motion vector for the pixelblock (mvp<sub>B]</sub>) according to a first derivation based on a motion vector extending between the reference frames (mv^), the first derivation being mvp<sub>B</sub>i - mvR£F*Z/2<sup>N</sup>, generating a second motion vector for the pixelblock (mvp<sub>B2</sub>) based on a second derivation of hivref, predicting video data for the pixelblock from video data of the first and second reference frames according to the generated motion vectors mv<sub>PB</sub>i and mvp<sub>B</sub>2, coding a difference between actual video data of the pixelblock and predicted video data of the pixel block, and outputting the coded difference to a channel as coded video data of the pixelblock.
In a still further aspect, the present invention provides a video coder, comprising: coding unit comprising a discrete cosine transform unit, a quantization unit to code an output of the discrete cosine transform unit, and an entropy coder to code an output of the quantization unit; motion estimator to interpolate motion vectors (mvp<sub>B</sub>i, mv<sub>PB2</sub>) for a pixelblock in a current frame, the motion vector estimator interpolating the motion vectors mvpBi, mvp<sub>B2</sub> with reference to a motion vector (hivref) extending from between previously-coded pair of reference frames by: generating a scale factor (Z) based on a ratio of time differences (Δίι/Δί<sub>2</sub>) among the current frame and the pair of reference frames, the first time difference (Δίι) representing a temporal difference between the current frame and
-6gCA 02502004 2009-08-20 a first of the reference frames and the second time difference (Δΐ2) representing a temporal difference between the two reference frames, the scale factor representing the ratio multiplied by a power of two (Ζ=2<sup>Ν</sup>*(Δΐι/Δί<sub>2</sub>)) subject to rounding, generating a first motion vector for the pixelblock (mv<sub>PB]</sub>) according to a first derivation based on a motion vector extending between the reference frames (hivref), the first derivation being mv<sub>PB]</sub> = mvR£F*Z/2<sup>N</sup>, generating a second motion vector for the pixelblock (mv<sub>PB</sub>2) based on a second derivation itivref; and a video data predictor to predict video data for the pixelblock from video data of the first and second reference frames according to the generated motion vectors mv<sub>PB</sub>i and mv<sub>PB2</sub>; wherein the coding unit codes a difference between actual video data of the pixelblock and predicted video data of the pixel block, and outputs the coded difference to a channel as coded video data of the pixelblock.
In a further aspect, the present invention provides computer readable medium storing program instructions that, when executed by a processing device, cause the device to: interpolate motion vectors for a pixelblock in a current frame, the pixelblock to be coded according to bi-directional prediction with reference to a pair of reference frames in the sequence, wherein the interpolation includes for at least one motion vector: generating a scale factor (Z) based on a ratio of time differences (Δΐι/Δΐ2) among the current frame and the pair of reference frames, the first time difference (Δΐι) representing a temporal difference between the current frame and a first of the reference frame and the second time difference (Δί<sub>2</sub>) representing a temporal difference between the two reference frames, the —6h—
CA 02502004 2009-08-20 scale factor representing the ratio multiplied by a power of two (ΖΑ2<sup>Ν</sup>*(Δί]/Δΐ<sub>2</sub>)) subject to rounding, generating a first motion vector for the pixelblock (mvpBi) according to a first derivation based on a motion vector extending between the reference frames (iuvref), the first derivation being mvpsi = mvREF*Z/2<sup>N</sup>, generating a second motion vector for the pixelblock (mv<sub>P</sub>B<sub>2</sub>) based on a second derivation iuvref, predict video data for the pixelblock from video data of the first and second reference frames according to the generated motion vectors mvpBi and mv<sub>P</sub>B<sub>2</sub>, code a difference between actual video data of the pixelblock and predicted video data of the pixel block, and output the coded difference to a channel as coded video data of the pixelblock.
In a still further aspect, the present invention provides computer readable medium storing program instructions that, when executed by a processing device, cause the device to: code a first reference frame at a first temporal location in the sequence, code a second reference frame at a second temporal location in the sequence, interpolate motion vectors for a pixelblock in a third, current frame, the pixelblock to be coded according to bidirectional prediction with reference to the pair of reference frames, wherein the interpolation includes for at least one motion vector: generating a scale factor (Z) based on a ratio of time differences (Δί]/Δί<sub>2</sub>) among the current frame and the reference frames, the first time difference (Δίι) representing a temporal difference between the current frame and the first reference frame and the second time difference (Δί<sub>2</sub>) representing a temporal difference between first and second reference frames, the scale factor representing the ratio —6i—
CA 02502004 2009-08-20 multiplied by a power of two (Ζ=2<sup>Ν</sup>*(Δίι/Δί2)) subject to rounding, generating a first motion vector for the pixelblock (mvp<sub>B</sub>i) according to a first derivation based on a motion vector extending between the reference frames (uivref), the first derivation being mvp<sub>B</sub>i = mvREF*Z/2<sup>N</sup>, generating a second motion vector for the pixelblock (mvp<sub>B</sub>2) based on a second derivation of mvRFp, predict video data for the pixelblock from video data of the first and second reference frames according to the generated motion vectors mvp<sub>B</sub>i and mv<sub>PB2</sub>, code a difference between actual video data of the pixelblock and predicted video data of the pixel block, and output the coded difference to a channel as coded video data of the pixelblock.
In a further aspect, the present invention provides a coded video signal created by a method, comprising: interpolating motion vectors for a pixelblock in a current frame, the pixelblock to be coded according to bi-directional prediction with reference to a pair of reference frames in the sequence, wherein the interpolation includes, for at least one motion vector: generating a scale factor (Z) based on a ratio of time differences (Δίι/Δί2) among the current frame and the pair of reference frames, the first time difference (Δίι) representing a temporal difference between the current frame and a first of the reference frame and the second time difference (Δΐ2) representing a temporal difference between the two reference frames, the scale factor representing the ratio multiplied by a power of two (Ζ=2<sup>Ν</sup>*(Δΐι/Δί<sub>2</sub>)) subject to rounding, generating a first motion vector for the pixelblock (mvp<sub>B</sub>i) according to a first derivation based on a motion vector extending between the .-6j~
CA 02502004 2009-08-20 reference frames (mv<sub>RE</sub>F), the first derivation being mvp<sub>B</sub>i = mvREF*Z/2<sup>N</sup>, generating a second motion vector for the pixelblock (mv<sub>PB</sub>2) based on a second derivation mvi®, predicting video data for the pixelblock from video data of the first and second reference frames according to the generated motion vectors mv<sub>P</sub>Bi and mv<sub>PB</sub>2, coding a difference between actual video data of the pixelblock and predicted video data of the pixel block, and outputting the coded difference to a channel as coded video data of the pixelblock.
In a still further aspect, the present invention provides an apparatus comprising: a storage for storing a stream comprising a first video picture, a second video picture, and a third video picture; and at least one processor for: (i) computing a scaling value that is based on (i) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture; and computing a particular motion vector for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the particular motion vector comprises performing a bit shifting operation.
In a further aspect, the present invention provides an apparatus comprising: a storage for storing a stream comprising a first video picture, a second video picture, and a third video picture; and at least one processor for: computing a scaling value that is based on (i) a first order difference value between an order value for the third video picture and
-6kCA 02502004 2011-07-15 an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture;
and computing a particular motion vector for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the particular motion vector comprises performing a division by a power of two value.
In a still further aspect, the present invention provides an apparatus comprising: at least one processor for: computing a scaling value based on (i) a first order difference value between an order value for a first video picture and an order value for a second video picture, and (ii) a particular value that is based on a power of two value and a second order difference value between an order value for a third video picture and the order value for the second video picture; and using said scaling value to compute a motion vector.
In a further aspect, the present invention provides an apparatus comprising: at least one processor for decoding a plurality of video pictures by using at least one order value, said order value is for establishing an ordering for reference video picture selection; and a storage for outputting the decoded video pictures based on the order value.
In a further aspect, the present invention provides for a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, the method comprising: computing a scaling value that is (i) inversely proportional to a first order difference value between an order value for the third video picture and an order value
-61CA 02502004 2011-07-15 for the first video picture and (ii) directly proportional to a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein the order value for the second picture is encoded in a bitstream more than once; and computing a motion vector for the second video picture by multiplying the scaling value by a motion vector for the third video picture and performing a division operation that is based on a power of two value.
In a still further aspect, the present invention provides for a sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, a method comprising: computing a scaling value that is (i) inversely proportional to a first order difference value between an order value for the third video picture and an order value for the first video picture, (ii) directly proportional to a second order difference value between an order value for the second video picture and the order value for the first video picture, and (iii) directly proportional to a power of two value, wherein an order value for a video picture specifies a display order for the video picture; and computing a motion vector for the second video picture by multiplying the scaling value and a motion vector for the third video picture and performing a division operation based on said power of two value.
In a further aspect, the present invention provides a method of decoding a bitstream comprising encoded first, second, and third video pictures, the method comprising:
receiving an integer value representing an exponent of a particular power of two integer; computing a scaling value that is based on (i) a first order difference value between an
-6mCA 02502004 2011-07-15 order value for the third video picture and an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture, wherein the order value for the second video picture is derived from the particular power of two integer; computing a motion vector for the second video picture by multiplying the scaling value by a motion vector for the third video picture and performing a bit shifting operation; and decoding the second video picture by using the computed motion vector.
In a still further aspect, the present invention provides a method for encoding a 10 sequence of video pictures comprising a first video picture, a second video picture, and a third video picture, the method comprising: computing a scaling value that is based on (i) a first order difference value between an order value for the third video picture and an order value for the first video picture, and (ii) a second order difference value between an order value for the second video picture and the order value for the first video picture; computing a particular motion vector for the second video picture based on the scaling value and a motion vector for the third video picture, wherein computing the particular motion vector for the second video picture comprises performing a bit shifting operation; encoding the second video picture in a bitstream by using the computed motion vector; and encoding the order value for the second video picture in the bitstream by using an exponent of a power of integer.
Other objects, features, and advantages of present invention will be apparent from the company drawings and from the following detailed description.
—6n—
CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
BRIEF DESCRIPTION OF THE DRAWINGS
The objects, features, and advantages of the present invention will be apparent to one skilled in the art, in view of the following detailed description in which:
Figure 1 illustrates a high-level block diagram of one possible digital video encoder system.
Figure 2 illustrates a series of video pictures in the order that the pictures 10 should be displayed wherein the arrows connecting different pictures indicate interpicture dependency created using motion compensation.
Figure 3 illustrates the video pictures from Figure 2 listed in a preferred transmission order of pictures wherein the arrows connecting different pictures indicate inter-picture dependency created using motion compensation.
Figure 4 graphically illustrates a series of video pictures wherein the distances between video pictures that reference each other are chosen to be powers of two.
—7—
CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENT
A method and system for specifying Variable Accuracy Inter-Picture Timing in a multimedia compression and encoding system with reduced requirements for division operations is disclosed. In the following description, for purposes of explanation, specific nomenclature is set forth to provide a thorough understanding of the present invention. However, it will be apparent to one skilled in the art that these specific details are not required in order to practice the present invention. For example, the present invention has been described with reference to the MPEG multimedia compression and encoding system. However, the same techniques can easily be applied to other types of compression and encoding systems.
Multimedia Compression and Encoding Overview
Figure 1 illustrates a high-level block diagram of a typical digital video encoder 100 as is well known in the art. The digital video encoder 100 receives an incoming video stream of video frames 105 at the left of the block diagram. The digital video encoder 100 partitions each video frame into a grid of pixelblocks. The pixelblocks are individually compressed. Various different sizes of pixelblocks may be used by different video encoding systems. For example, different pixelblock resolutions include 8x8, 8x4, 16x8, 4x4, etc. Furthermore, pixelblocks are occasionally referred to as ‘macroblocks. ’ This document will use the term pixelblock to refer to any block of pixels of any size.
-8CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
A Discrete Cosine Transformation (DCT) unit 110 processes each pixelblock in the video frame. The frame may be processed independently (an intraframe) or with reference to information from other frames received from the motion compensation unit (an inter-frame). Next, a Quantizer (Q) unit 120 quantizes the information from the Discrete Cosine Transformation unit 110. Finally, the quantized video frame is then encoded with an entropy encoder (H) unit 180 to produce an encoded bitstream. The entropy encoder (H) unit 180 may use a variable length coding (VLC) system.
Since an inter-frame encoded video frame is defined with reference to other nearby video frames, the digital video encoder 100 needs to create a copy of how each decoded frame will appear within a digital video decoder such that inter-frames may be encoded. Thus, the lower portion of the digital video encoder 100 is actually a digital video decoder system. Specifically, an inverse quantizer (Q'<sup>1</sup>) unit 130 reverses the quantization of the video frame information and an inverse Discrete Cosine
Transformation (DCT'<sup>1</sup>) unit 140 reverses the Discrete Cosine Transformation of the video frame information. After all the DCT coefficients are reconstructed from inverse Discrete Cosine Transformation (DCT'<sup>1</sup>) unit 140, the motion compensation unit will use that information, along with the motion vectors, to reconstruct the encoded video frame.
The reconstructed video frame is then used as the reference frame for the motion estimation of the later frames.
The decoded video frame may then be used to encode inter-frames (Pframes or B-frames) that are defined relative to information in the decoded video frame.
Specifically, a motion compensation (MC) unit 150 and a motion estimation (ME) unit
..9..
CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
160 are used to determine motion vectors and generate differential values used to encode inter-frames.
A rate controller 190 receives information from many different 5 components in a digital video encoder 100 and uses the information to allocate a bit budget for each video frame. The rate controller 190 should allocate the bit budget in a manner that will generate the highest quality digital video bit stream that that complies with a specified set of restrictions. Specifically, the rate controller 190 attempts to generate the highest quality compressed video stream without overflowing buffers (exceeding the amount of available memory in a video decoder by sending more information than can be stored) or underflowing buffers (not sending video frames fast enough such that a video decoder runs out of video frames to display).
Digital Video Encoding With Pixelblocks
In some video signals the time between successive video pictures (frames or fields) may not be constant. (Note: This document will use the term video pictures to generically refer to video frames or video fields.) For example, some video pictures may be dropped because of transmission bandwidth constraints. Furthermore, the video timing may also vary due to camera irregularity or special effects such as slow motion or fast motion. In some video streams, the original video source may simply have nonuniform inter-picture times by design. For example, synthesized video such as computer graphic animations may have non-uniform timing since no arbitrary video timing is imposed by a uniform timing video capture system such as a video camera system. A
-10CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 flexible digital video encoding system should be able to handle non-uniform video picture timing.
As previously set forth, most digital video encoding systems partition 5 video pictures into a rectangular grid of pixelblocks. Each individual pixelblock in a video picture is independently compressed and encoded. Some video coding standards, e.g., ISO MPEG or ITU H.264, use different types of predicted pixelblocks to encode video pictures. In one scenario, a pixelblock may be one of three types:
1. I-pixelblock - An Intra (I) pixelblock uses no information from any other video pictures in its coding (it is completely self-defined);
2. P-pixelblock - A unidirectionally predicted (P) pixelblock refers to picture information from one preceding video picture; or
3. B-pixelblock - A bi-directional predicted (B) pixelblock uses information from one preceding picture and one future video picture.
If all the pixelblocks in a video picture are Intra-pixelblocks, then the video picture is an Intra-frame. If a video picture only includes unidirectional predicted macro blocks or intra-pixelblocks, then the video picture is known as a P-frame. If the video picture contains any bi-directional predicted pixelblocks, then the video picture is known as a B-frame. For the simplicity, this document will consider the case where all pixelblocks within a given picture are of the same type.
An example sequence of video pictures to be encoded might be represented as:
fi B2 B3 B4 P5 Ββ B7 Bg B9 P10 Bn P12 B131¼...
-11CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 where the letter (I, P, or B) represents if the video picture is an I-frame, P-frame, or Bfiame and the number represents the camera order of the video picture in the sequence of video pictures. The camera order is the order in which a camera recorded the video pictures and thus is also the order in which the video pictures should be displayed (the display order).
The previous example series of video pictures is graphically illustrated in Figure 2. Referring to Figure 2, the arrows indicate that pixelblocks from a stored picture (I-frame or P-frame in this case) are used in the motion compensated prediction of other pictures.
In the scenario of Figure 2, no information from other pictures is used in the encoding of the intra-frame video picture f. Video picture P<sub>5</sub> is a P-frame that uses video information from previous video picture Ii in its coding such that an arrow is drawn from video picture Ii to video picture P5. Video picture B<sub>2</sub>, video picture B<sub>3</sub>, video picture B4 all use information from both video picture L and video picture P<sub>5</sub> in their coding such that arrows are drawn from video picture f and video picture P<sub>5</sub> to video picture B<sub>2</sub>, video picture B3, and video picture B4. As stated above the inter-picture times are, in general, not the same.
Since B-pictures use information from future pictures (pictures that will be displayed later), the transmission order is usually different than the display order. Specifically, video pictures that are needed to construct other video pictures should be transmitted first. For the above sequence, the transmission order might be:
h P5 B<sub>2</sub> B3 B4 P10 Bô Βγ Bs B9 Pi<sub>2</sub> Bu 1¼ B13 ...
-12CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
Figure 3 graphically illustrates the preceding transmission order of the video pictures from Figure 2. Again, the arrows in the figure indicate that pixelblocks from a stored video picture (I or P in this case) are used in the motion compensated prediction of other video pictures.
Referring to Figure 3, the system first transmits I-frame Ii which does not depend on any other frame. Next, the system transmits P-frame video picture P5 that depends upon video picture fr. Next, the system transmits B-frame video picture B<sub>2</sub> after video picture P5 even though video picture B<sub>2</sub> will be displayed before video picture P5. The reason for this is that when it comes time to decode video picture B<sub>2</sub>, the decoder will have already received and stored the infonnation in video pictures Ii and P5 necessary to decode video picture B<sub>2</sub>. Similarly, video pictures f and P5 are ready to be used to decode subsequent video picture B3 and video picture B<sub>4</sub>. The receiver/decoder reorders the video picture sequence for proper display. In this operation I and P pictures are often referred to as stored pictures.
The coding of the P-frame pictures typically utilizes Motion Compensation, wherein a Motion Vector is computed for each pixelblock in the picture.
Using the computed motion vector, a prediction pixelblock (P-pixelblock) can be formed by translation of pixels in the aforementioned previous picture. The difference between the actual pixelblock in the P-frame picture and the prediction pixelblock is then coded for transmission.
-13CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
P-Pictures
The coding of P-Pictures typically utilize Motion Compensation (MC), wherein a Motion Vector (MV) pointing to a location in a previous picture is computed for each pixelblock in the current picture. Using the motion vector, a prediction pixelblock can be formed by translation of pixels in the aforementioned previous picture.
The difference between the actual pixelblock in the P-Picture and the prediction pixelblock is then coded for transmission.
Each motion vector may also be transmitted via predictive coding. For 10 example, a motion vector prediction may be formed using nearby motion vectors. In such a case, then the difference between the actual motion vector and the motion vector prediction is coded for transmission.
B-Pictures
Each B-pixelblock uses two motion vectors: a first motion vector referencing the aforementioned previous video picture and a second motion vector referencing the future video picture. From these two motion vectors, two prediction pixelblocks are computed. The two predicted pixelblocks are then combined together, using some function, to form a final predicted pixelblock. As above, the difference between the actual pixelblock in the Β-frame picture and the final predicted pixelblock is then encoded for transmission.
As with P-pixelblocks, each motion vector (MV) of a B-pixelblock may be transmitted via predictive coding. Specifically, a predicted motion vector is formed using
-14CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 nearby motion vectors. Then, the difference between the actual motion vector and the predicted is coded for transmission.
However, with B-pixelblocks the opportunity exists for interpolating 5 motion vectors from motion vectors in the nearest stored picture pixelblock. Such motion vector interpolation is carried out both in the digital video encoder and the digital video decoder.
This motion vector interpolation works particularly well on video pictures 10 from a video sequence where a camera is slowly panning across a stationary background.
In fact, such motion vector interpolation may be good enough to be used alone. Specifically, this means that no differential information needs be calculated or transmitted for these B-pixelblock motion vectors encoded using interpolation.
To illustrate further, in the above scenario let us represent the inter-picture display time between pictures i and j as D, j, i.e., if the display times of the pictures are Tj and Tj, respectively, then <sup>D</sup>i,j = Τχ - Tj from which it follows that
Di,k ' Di<sub>;</sub>j + Dj,k
Di,k <sup>=</sup> “Dk,i
Note that Dÿ may be negative in some cases.
-15CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
Thus, if MVy is a motion vector for a P5 pixelblock as referenced to Ii, then for the corresponding pixelblocks in B<sub>2</sub>, B3 and B4 the motion vectors as referenced to Ii and P5, respectively would be interpolated by
<td></td><td> MV<sub>2</sub>,i</td><td> = MV<sub>5</sub>,i*D<sub>2(</sub>i/D<sub>5</sub>,i</td>
<td> 5</td><td> MV5,2</td><td> = ^5,1*05,2/05,3.</td>
<td></td><td> MV<sub>3</sub>,1</td><td> = MV<sub>5</sub>,i*D<sub>3j1</sub>/D<sub>5</sub>,<sub>1</sub></td>
<td></td><td> MV<sub>s</sub>,3</td><td> = MV5,1*D5,3/D<sub>5</sub>,1</td>
<td> 10</td><td> MV<sub>4</sub>,i</td><td> = MV<sub>5</sub>, i*D<sub>4</sub>, i/D<sub>5</sub>,i</td>
<td></td><td> MV<sub>5</sub>,4</td><td> = MV<sub>5</sub>,i*D<sub>5</sub>,4/D<sub>5</sub>,i</td>
Note that since ratios of display times are used for motion vector prediction, absolute display times are not needed. Thus, relative display times may be used for Dy interpicture display time values.
This scenario may be generalized, as for example in the H.264 standard. In the generalization, a P or B picture may use any previously transmitted picture for its motion vector prediction. Thus, in the above case picture B3 may use picture Ii and picture B<sub>2</sub> in its prediction. Moreover, motion vectors maybe extrapolated, not just interpolated. Thus, in this case we would have:
Mt/3,1 = MV2,l*D3<sub>(</sub>i/D2,l
Such motion vector extrapolation (or interpolation) may also be used in the prediction 25 process for predictive coding of motion vectors.
-16CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
Encoding Inter-Picture Display Times
The variable inter-picture display times of video sequences should be encoded and transmitted in a maimer that renders it possible to obtain a very high coding efficiency and has selectable accuracy such that it meets the requirements of a video decoder. Ideally, the encoding system should simplify the tasks for the decoder such that relatively simple computer systems can decode the digital video.
The variable inter-picture display times are potentially needed in a number 10 of different video encoding systems in order to compute differential motion vectors,
Direct Mode motion vectors, and/or Implicit B Prediction Block Weighting.
The problem of variable inter-picture display times in video sequences is intertwined with the use of temporal references. Ideally, the derivation of correct pixel values in the output pictures in a video CODEC should be independent of the time at which that picture is decoded or displayed. Hence, timing issues and time references should be resolved outside the CODEC layer.
There are both coding-related and systems-related reasons underlying the desired time independence. In a video CODEC, time references are used for two purposes:
(1) To establish an ordering for reference picture selection; and (2) To interpolate motion vectors between pictures.
To establish an ordering for reference picture selection, one may simply send a relative position value. For example, the difference between the frame position N in decode order
-17CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 and the frame position M in the display order, i.e., N-M. In such an embodiment, timestamps or other time references would not be required. To interpolate motion vectors, temporal distances would be useful if the temporal distances could be related to the interpolation distance. However, this may not be true if the motion is non-linear.
Therefore, sending parameters other than temporal information for motion vector interpolation seems more appropriate.
In terms of systems, one can expect that a typical video CODEC is part of a larger system where the video CODEC coexists with other video (and audio) CODECs.
In such multi-CODEC systems, good system layering and design requires that general functions, which are logically CODEC-independent such as timing, be handled by the layer outside the CODEC. The management of timing by the system and not by each CODEC independently is critical to achieving consistent handling of common functions such as synchronization. For instance in systems that handle more than one stream simultaneously, such as a video/audio presentation, timing adjustments may sometimes be needed within the streams in order to keep the different streams synchronized. Similarly, in a system that handles a stream from a remote system with a different clock timing adjustments may be needed to keep synchronization with the remote system. Such timing adjustments may be achieved using time stamps. For example, time stamps that are linked by means of “Sender Reports” from the transmitter and supplied in RTP in the RTP layer for each stream may be used for synchronization. These sender reports may take the form of:
Video RTP TimeStamp X is aligned with reference timestamp Y Audio RTP TimeStamp W is aligned with reference timestamp Z
-18CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
Wherein the wall-clock rate of the reference timestamps is known, allowing the two streams to be aligned. However, these timestamp references arrive both periodically and separately for the two streams, and they may cause some needed re-alignment of the two streams. This is generally achieved by adjusting the video stream to match the audio or vice-versa. System handling of time stamps should not affect the values of the pixels being displayed. More generally, system handling of temporal information should be performed outside the CODEC.
A Specific Example
As set forth in the previous section, the problem in the case of non uniform inter-picture times is to transmit the inter-picture display time values Djj to the digital video receiver in an efficient manner. One method of accomplishing this goal is to have the system transmit the display time difference between the current picture and the most recently transmitted stored picture for each picture after the first picture. For error resilience, the transmission could be repeated several times within the picture. For example, the display time difference may be repeated in the slice headers of the MPEG or H.264 standards. If all slice headers are lost, then presumably other pictures that rely on the lost picture for decoding information cannot be decoded either.
Thus, with reference to the example of the preceding section, a system would transmit the following inter-picture display time values:
D54 D2,5 D3,5 Ü4,5 Dio,5 D<sub>6</sub>,10 D7JO D<sub>8j</sub>io D940 Dj2,io Dn<sub>;</sub>i2 Di4<sub>;</sub>i2 D1344 ...
For the purpose of motion vector estimation, the accuracy requirements for the inter25 picture display times Djj may vary from picture to picture. For example, if there is only a
-19CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 single B-frame picture B<sub>6</sub> halfway between two P-frame pictures P5 and P7, then it suffices to send only:
D<sub>7;</sub>s = 2 and Dg,<sub>7</sub> = -1 where the Dÿ inter-picture display time values are relative time values.
If, instead, video picture Bg is only one quarter the distance between video picture P5 and video picture P7 then the appropriate Dÿ inter-picture display time values to send would be:
D<sub>7j</sub>s = 4 and Dg<sub>;7</sub> = — T
Note that in both of the preceding examples, the display time between the video picture
Bg and video picture video picture P<sub>7</sub> (inter-picture display time Dg<sub>:7</sub>) is being used as the display time “unit” value . In the most recent example, the display time difference between video picture P<sub>5</sub> and picture video picture P<sub>7</sub> (inter-picture display time Dg<sub>;7</sub>) is four display time “units” (4 * Dg<sub>j7</sub>) .
Improving Decoding Efficiency
In general, motion vector estimation calculations are greatly simplified if divisors are powers of two. This is easily achieved in our embodiment if Djj (the interpicture time) between two stored pictures is chosen to be a power of two as graphically illustrated in Figure 4. Alternatively, the estimation procedure could be defined to truncate or round all divisors to a power of two.
In the case where an inter-picture time is to be a power of two, the number of data bits can be reduced if only the integer power (of two) is transmitted instead of the full value of the inter-picture time. Figure 4 graphically illustrates a case wherein the
-20CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 distances between pictures are chosen to be powers of two. In such a case, the D3J display time value of 2 between video picture Pi and picture video picture P3 is transmitted as 1 (since 2<sup>1</sup> = 2) and the D73 display time value of 4 between video picture P<sub>7</sub> and picture video picture P3 can be transmitted as 2 (since 2<sup>2</sup> = 4).
Alternatively, the motion vector interpolation of extrapolation operation can be approximated to any desired accuracy by scaling in such a way that the denominator is a power of two. (With a power of two in the denominator division may be performed by simply shifting the bits in the value to be divided.) For example,
D<sub>5</sub>,<sub>4</sub>/D<sub>5rl</sub> ~ Z<sub>5</sub>,<sub>4</sub>/P
Where the value P is a power of two and = P*D<sub>5j</sub>4/D5,i is rounded or truncated to the
I nearest integer. The value of P may be periodically transmitted or set as a constant for the system. In one embodiment, the value of P is set as P = 2<sup>8</sup> = 256.
The advantage of this approach is that the decoder only needs to compute
Z5,4 once per picture or in many cases the decoder may pre-compute and store the Z value. This allows the decoder to avoid having to divide by D54 for every motion vector in the picture such that motion vector interpolation may be done much more efficiently. For example, the normal motion vector calculation would be:
mv<sub>5</sub>,<sub>4</sub> = MV<sub>5</sub>,1*D<sub>5</sub>,<sub>4</sub>/D<sub>5</sub>,1
But if we calculate and store Ζ<sub>5>4</sub> wherein Zs,4 = PW^/D^i then
MV<sub>5(4</sub> = MV<sub>5</sub>,i*Z<sub>5</sub>,<sub>4</sub>/P
But since the P value has been chosen to be a power of two, the division by P is merely a simple shift of the bits. Thus, only a single multiplication and a single shift are required to calculate motion vectors for subsequent pixelblocks once the Z value has been
-21CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 calculated for the video picture. Furthermore, the system may keep the accuracy high by performing all divisions last such that significant bits are not lost during the calculation. In this manner, the decoder may perform exactly the same as the motion vector interpolation as the encoder thus avoiding any mismatch problems that might otherwise arise.
Since division (except for division by powers of two) is a much more computationally intensive task for a digital computer system than addition or multiplication, this approach can greatly reduce the computations required to reconstruct pictures that use motion vector interpolation or extrapolation.
In some cases, motion vector interpolation may not be used. However, it is still necessary to transmit the display order of the video pictures to the receiver/player system such that the receiver/player system will display the video pictures in the proper order. In this case, simple signed integer values for D;j suffice irrespective of the actual display times. In some applications only the sign (positive or negative) may be needed to reconstruct the picture ordering.
The inter-picture times Djj may simply be transmitted as simple signed integer values. However, many methods may be used for encoding the Djj values to achieve additional compression. For example, a sign bit followed by a variable length coded magnitude is relatively easy to implement and provides coding efficiency.
-22CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953
One such variable length coding system that may be used is known as UVLC (Universal Variable Length Code). The UVLC variable length coding system is given hy the code words:
= 1
2 = 0 10 = 0 11 = 0 0 10 0 = 0 0 10 1 = 0 0 110
7= 00111
0 0 1 0 0 0...
Another method of encoding the inter-picture times may be to use arithmetic coding. Typically, arithmetic coding utilizes conditional probabilities to effect a very high compression of the data bits.
Thus, the present invention introduces a simple but powerful method of encoding and transmitting inter-picture display times and methods for decoding those inter-picture display times for use in motion vector estimation. The encoding of inter20 picture display times can be made very efficient by using variable length coding or arithmetic coding. Furthermore, a desired accuracy can be chosen to meet the needs of the video codec, but no more.
The foregoing has described a system for specifying variable accuracy 25 inter-picture timing in a multimedia compression and encoding system. It is
-23CA 02502004 2005-04-11
WO 2004/054257
PCT/US2003/024953 contemplated that changes and modifications may be made by one of ordinary skill in the art, to the materials and arrangements of elements of the present invention without departing from the scope of the invention.
Contents64
5 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10123037B2 | Cited by | United States of America | Applicant |
71 members in 11 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 10313773 | United States of America | – | |
| 31377302 | United States of America | A | |
| 0324953 | United States of America | W | |
| 10313773 | – | – | – |
| PCTUS2003024953 | – | – | – |
| US20020313773 | – | – | – |
| WO2003US24953 | – | – | – |
Members71
| Document | Office | Kind | |
|---|---|---|---|
| EP0036607A1 | European Patent Office (EPO) | A1 | |
| JPS56144423A | Japan | A | |
| US4320572A | United States of America | A | |
| US4394709A | United States of America | A | |
| CA1153477A | Canada | A | |
| EP0036607B1 | European Patent Office (EPO) | B1 | |
| DE3164377D1 | Germany | D1 | |
| US2004017851A1 | United States of America | A1 | |
| US6728315B2 | United States of America | B2 | |
| CA2502004A1 | Canada | A1 | |
| CA2784948A1 | Canada | A1 | |
| WO2004054257A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003258142A1 | Australia | A1 | |
| US2004184543A1 | United States of America | A1 | |
| NO20052764L | Norway | L | |
| KR20050085392A | Republic of Korea | A | |
| BR0316175A | Brazil | A | |
| EP1579689A1 | European Patent Office (EPO) | A1 | |
| CN1714572A | China | A | |
| JP2006509463A | Japan | A | |
| US2007183501A1 | United States of America | A1 | |
| US2007183502A1 | United States of America | A1 | |
| US2007183503A1 | United States of America | A1 | |
| US2007286282A1 | United States of America | A1 | |
| KR100804335B1 | Republic of Korea | B1 | |
| US7339991B2 | United States of America | B2 | |
| CN101242536A | China | A | |
| CN100417218C | China | C | |
| US2009022224A1 | United States of America | A1 | |
| US2009022225A1 | United States of America | A1 | |
| AU2003258142B2 | Australia | B2 | |
| AU2009202255A1 | Australia | A1 | |
| EP1579689A4 | European Patent Office (EPO) | A4 | |
| JP2010154568A | Japan | A | |
| JP2011155678A | Japan | A | |
| US8009736B2 | United States of America | B2 | |
| US8009737B2 | United States of America | B2 | |
| AU2009202255B2 | Australia | B2 | |
| US2011243235A1 | United States of America | A1 | |
| US2011243236A1 | United States of America | A1 | |
| US2011243237A1 | United States of America | A1 | |
| US2011243238A1 | United States of America | A1 | |
| US2011243239A1 | United States of America | A1 | |
| US2011243240A1 | United States of America | A1 | |
| US2011243241A1 | United States of America | A1 | |
| US2011243242A1 | United States of America | A1 | |
| US2011243243A1 | United States of America | A1 | |
| US2011249752A1 | United States of America | A1 | |
| US2011249753A1 | United States of America | A1 | |
| US8077779B2 | United States of America | B2 | |
| US8090023B2 | United States of America | B2 | |
| US8094729B2 | United States of America | B2 | |
| AU2011265362A1 | Australia | A1 | |
| CN101242536B | China | B | |
| US8254461B2 | United States of America | B2 | |
| CA2502004CThis record | Canada | C | |
| US8817880B2 | United States of America | B2 | |
| US8817888B2 | United States of America | B2 | |
| US8824565B2 | United States of America | B2 | |
| US8837603B2 | United States of America | B2 | |
| US8885732B2 | United States of America | B2 | |
| US8934546B2 | United States of America | B2 | |
| US8934547B2 | United States of America | B2 | |
| US8934551B2 | United States of America | B2 | |
| US8938008B2 | United States of America | B2 | |
| US8942287B2 | United States of America | B2 | |
| US8953693B2 | United States of America | B2 | |
| US2015117541A1 | United States of America | A1 | |
| US9554151B2 | United States of America | B2 | |
| US2017094308A1 | United States of America | A1 | |
| US10123037B2 | United States of America | B2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| LapsedLapsedMKLA | MKLA | |
| Examination requestEEER | EEER |
Numbers
- Publication
- 2502004
- Publication, DOCDB
- 2502004
- Publication, EPODOC
- CA2502004
- Application
- 2502004
- Application, DOCDB
- 2502004
- Application, EPODOC
- CA20032502004
Titles2
- English
- METHOD AND APPARATUS FOR VARIABLE ACCURACY INTER-PICTURE TIMING SPECIFICATION FOR DIGITAL VIDEO ENCODING WITH REDUCED REQUIREMENTS FO DIVISION OPERATIONS
- French
- PROCEDE ET APPAREIL DE SPECIFICATION DE MINUTAGE ENTRE IMAGES A PRECISION VARIABLE POUR CODAGE VIDEO NUMERIQUE A EXIGENCES REDUITES POUR DES OPERATIONS DE DIVISION
Classification
- CPC, 24
- H04N19/107
- H04N19/51
- H04N19/52
- H04N19/124
- H04N19/15
- H04N19/152
- H04N19/159
- H04N19/176
- H04N19/196
- H04N19/40
- H04N19/46
- H04N19/463
- H04N19/48
- H04N19/513
- H04N19/577
- H04N19/61
- H04N19/13
- H04N19/132
- H04N19/137
- H04N19/139
- H04N19/146
- H04N19/172
- H04N19/587
- H04N19/625
- IPC, 7
- H04N7 12
- G06K9 36
- G06T9 00
- H04N9 74
- H04N7 26
- H04N7 36
- H04N19 94