EP1369820A2

Spatiotemporal prediction for bidirectionally predictive (B) pictures and motion vector prediction for multi-picture reference motion compensation

Abstract

Several improvements for use with Bidirectionally Predictive (B) pictures within a video sequence are provided. In certain improvements Direct Mode encoding and/or Motion Vector Prediction are enhanced using spatial prediction techniques. In other improvements Motion Vector prediction includes temporal distance and subblock information, for example, for more accurate prediction. Such improvements and other presented herein significantly improve the performance of any applicable video coding system/logic.

EP1369820A2, drawing sheet 1
Sheet 1 of 35

Term

Term ended

Projected expiry passed 27 May 2023, 3.3 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

117 claims: 21 independent, 96 dependent

  1. 1
    A method for use in encoding video data within a sequence of video frames, the method comprising:identifying at least a portion of at least one video frame to be a Bidirectionally Predictive (B) picture;andselectively encoding said B picture using at least spatial prediction to encode at least one motion parameter associated with said B picture.
  2. 16
    A computer-readable medium having computer implementable instructions for configuring at least one processing unit to perform acts comprising:accessing data for a sequence of video frames;identifying at least a portion of at least one video frame to be a Bidirectionally Predictive (B) picture;andselectively encoding said B picture using at least spatial prediction to encode at least one motion parameter associated with said B picture.
  3. 31
    An apparatus for use in encoding video data within a sequence of video frames, the apparatus comprising:logic operatively configured to access video data for a sequence of video frames, identify at least a portion of at least one video frame to be a Bidirectionally Predictive (B) picture, and selectively encode said B picture using at least spatial prediction to encode at least one motion parameter associated with said B picture.
  4. 46
    A method for encoding video data, the method comprising:identifying at least a portion of at least one video frame to be coded in an enhanced direct mode;andencoding said portion in said enhanced direct mode using at least spatial information associated with said portion within said at least one video frame.
  5. 49
    The method as recited in Clam 48, wherein said motion vector prediction includes median prediction.
  6. 52
    A computer-readable medium having computer implementable instructions for configuring at least one processing unit to perform acts comprising encoding video data by identifying at least a portion of at least one video frame to be coded in an enhanced direct mode, and encoding said portion in said enhanced direct mode using at least spatial information associated with said portion within said at least one video frame.
  7. 55
    The computer-readable medium as recited in Clam 54, wherein said motion vector prediction includes median prediction.
  8. 58
    An apparatus comprising:logic operatively configured to encode video data by identifying at least a portion of at least one video frame to be coded in an enhanced direct mode, and encode said portion in said enhanced direct mode using at least spatial information associated with said portion within said at least one video frame.
  9. 61
    The apparatus as recited in Clam 60, wherein said motion vector prediction includes median prediction.
  10. 64
    A method to predict a reference picture in direct mode video encoding, the method comprising:selecting a reference picture from a group comprising a minimum reference picture for a plurality of predictions related to at least a portion of a video frame to be encoded, a median reference picture for said plurality of predictions, and a current reference picture based on a single direction prediction;andencoding said at least one portion of said video frame based on selected reference picture.
  11. 67
    A computer-readable medium having computer implementable instructions for configuring at least one processing unit to perform acts comprising:selecting a reference picture from a group comprising a minimum reference picture for a plurality of predictions related to at least a portion of a video frame to be encoded, a median reference picture for said plurality of predictions, and a current reference picture based on a single direction prediction;andencoding said at least one portion of said video frame based on selected reference picture.
  12. 70
    An apparatus comprising:logic that is operatively configured to select a reference picture from a group comprising a minimum reference picture for a plurality of predictions related to at least a portion of a video frame to be encoded, a median reference picture for said plurality of predictions, and a current reference picture based on a single direction prediction, and encode said at least one portion of said video frame based on selected reference picture.
  13. 73
    A method for use in selecting between temporal prediction, spatial prediction, or both temporal and spatial prediction for encoding at least a portion of at least one video frame in an enhanced direct mode, the method comprising:selecting temporal prediction if at least one motion vector of a collocated portion of said video frame is zero;if surrounding portions within said video frame use different reference pictures than a collocated reference picture, then select spatial prediction only;if a motion flow associated with said portion of said video frame is substantially different than a motion flow associated with a reference picture, then select spatial prediction;if temporal prediction of direct mode is signaled inside an image header, then selecting temporal prediction;andif spatial prediction of direct mode is signaled inside said image header, then selecting spatial prediction.
  14. 76
    A computer-readable medium having computer implementable instructions for configuring at least one processing unit to perform acts comprising:selecting between temporal prediction, spatial prediction, or both temporal and spatial prediction for encoding at least a portion of at least one video frame in an enhanced direct mode, such that: temporal prediction is selected if at least one motion vector of a collocated portion of said video frame is zero,only spatial prediction is selected if surrounding portions within said video frame use different reference pictures than a collocated reference picture,spatial prediction is selected if a motion flow associated with said portion of said video frame is substantially different than a motion flow associated with a reference picture,temporal prediction is selected if temporal prediction of direct mode is signaled inside an image header, andspatial prediction is selected if spatial prediction of direct mode is signaled inside said image header.
  15. 79
    An apparatus comprising:logic operatively configured to select between and employ temporal prediction, spatial prediction, or both temporal and spatial prediction for encoding at least a portion of at least one video frame in an enhanced direct mode, wherein said logic: selects temporal prediction if at least one motion vector of a collocated portion of said video frame is zero,selects only spatial prediction if surrounding portions within said video frame use different reference pictures than a collocated reference picture,selects spatial prediction if a motion flow associated with said portion of said video frame is substantially different than a motion flow associated with a reference picture,selects temporal prediction if temporal prediction of direct mode is signaled inside an image header, andselects spatial prediction if spatial prediction of direct mode is signaled inside said image header.
  16. 82
    A method for use in encoding video data, the method comprising:selecting a reference portion of a future video frame to serve as a B picture to at least one portion of an earlier video frame;using motion vectors associated with said reference frame to calculate motion vectors associated with said at least one portion;andencoding said at least one portion based on said calculated motion vectors associated with said at least one portion.
  17. 88
    A computer-readable medium having computer implementable instructions for configuring at least one processing unit to perform acts comprising:selecting a reference portion of a future video frame to serve as a B picture to at least one portion of an earlier video frame;using motion vectors associated with said reference frame to calculate motion vectors associated with said at least one portion;andencoding said at least one portion based on said calculated motion vectors associated with said at least one portion.
  18. 94
    An apparatus comprising:logic operatively configured to select a reference portion of a future video frame to serve as a B picture to at least one portion of an earlier video frame, use motion vectors associated with said reference frame to calculate motion vectors associated with said at least one portion, and encode said at least one portion based on said calculated motion vectors associated with said at least one portion.
  19. 100
    A method for use in determining motion vectors during video encoding, the method comprising:selecting at least three predictors A, B and C that each uses a different reference picture having an associated temporal distance TRA, TRB, and TRC respectively, and a motion vector MVA, MVB, and MVC;andpredicting a median motion vector MVpred associated with a current reference picture that has a temporal distance equal to TR.
  20. 106
    A computer-readable medium having computer implementable instructions for configuring at least one processing unit to perform acts comprising:selecting at least three predictors A, B and C that each uses a different reference picture having an associated temporal distance TRA, TRB, and TRC respectively, and a motion vector MVA, MVB, and MVC;andpredicting a median motion vector MVpred associated with a current reference picture that has a temporal distance equal to TR.
  21. 112
    An apparatus comprising logic operatively configured to select at least three predictors A, B and C that each uses a different reference picture having an associated temporal distance TRA, TRB, and TRC respectively, and a motion vector MVA, MVB, and MVC, and predict a median motion vector MVpred associated with a current reference picture that has a temporal distance equal to TR.
Independent claims21