EP1538566A2

Method and apparatus for scalable video encoding and decoding

Abstract

A method and apparatus for scalable video and decoding are provided. A method for video coding includes eliminating temporal redundancy in constrained temporal level sequence from a plurality of frames constituting a video sequence input, and generating a bit-stream by quantizing transformation coefficients obtained from the frames whose temporal redundancy has been eliminated. A video encoder for performing the encoding method includes a temporal transformation unit (10), a spatial transformation unit (20), a quantization unit (30), and a bit-stream generation unit (40). A video decoding method is in principle performed inversely to the video coding sequence, wherein decoding is performed by extracting information on encoded frames by receiving bit-streams input and interpreting them.

EP1538566A2, drawing sheet 1
Sheet 1 of 33

Term

Term ended

Projected expiry passed 29 November 2024, 1.8 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

65 claims: 25 independent, 40 dependent

  1. 1
    A method for video coding, the method comprising:(a) eliminating a temporal redundancy in a constrained temporal level sequence from a plurality of frames of a video sequence;and (b) generating a bit-stream by quantizing transformation coefficients obtained from the frames whose temporal redundancy has been eliminated.
  2. 5
    The method as claimed in any preceding claim, wherein temporal levels of the frames have dyadic hierarchal structures.
  3. 6
    The method as claimed in any preceding claim, wherein the constrained temporal level sequence is a sequence of the frames from a highest temporal level to a lowest temporal level and a sequence of the frames from a lowest frame index to a highest frame index in a same temporal level.
  4. 13
    The method as claimed in any one of claims 9 to 12, wherein the elimination of the temporal redundancy from the remaining frame is performed based on at least one frame of a next GOP, whose temporal level is higher than temporal levels of each of the frames currently being processed in step (a).
  5. 14
    The method as claimed in any preceding claim, wherein the constrained temporal level sequence is determined based on a coding mode.
  6. 19
    The method as claimed in any one of claims 15 to 18, wherein the coding mode is determined depending upon an end-to-end delay parameter D, where the constrained temporal level sequence progresses from the frames at a highest temporal level to a lowest temporal level among the frames having frame indexes not exceeding D in comparison to a frame at the lowest temporal level, which has not yet had the temporal redundancy removed, and from the frames at a lowest frame index to a highest frame index in a same temporal level.
  7. 23
    The method as claimed in any one of claims 20 to 22, wherein the elimination of the temporal redundancy from the remaining frame is performed based on the remaining frame.
  8. 25
    The method as claimed in any one of claims 20 to 24, wherein the elimination of the temporal redundancy from the remaining frame is performed based on at least one frame of a next GOP, whose temporal level is higher than a temporal level of each of the frames currently being processed in step (a) and whose temporal distances from each of the frames currently being processed in step (a) are less than or equal to D.
  9. 26
    A video encoder comprising:a temporal transformation unit (10) eliminating a temporal redundancy in a constrained temporal level sequence from a plurality of frames of an input video sequence;a spatial transformation unit (20) eliminating a spatial redundancy from the frames;a quantization unit (30) quantizing transformation coefficients obtained from eliminating the temporal redundancies in the temporal transformation unit (10) and the spatial redundancies in the spatial transformation unit (20);and a bit-stream generation unit (40) generating a bit-stream based on quantized transformation coefficients generated by the quantization unit (30).
  10. 29
    The video encoder as claimed in any one of claims 26 to 28, wherein the spatial transformation encoder eliminates the spatial redundancy of the frames through the wavelet transformation and transmits the frames whose spatial redundancy has been eliminated to the temporal transformation unit (10), and the temporal transformation unit (10) eliminates the temporal redundancy of the frames whose spatial redundancy has been eliminated to generate the transformation coefficients.
  11. 30
    The video encoder as claimed in any one of claims 26 to 29, wherein the temporal transformation unit (10) comprises:a motion estimation unit (12) obtaining motion vectors from the frames ;a temporal filtering unit (14) temporally filtering in the constrained temporal level sequence the frames based on the motion vectors obtained by the motion estimation unit (12);and a mode selection unit (16) determining the constrained temporal level sequence.
  12. 34
    The video encoder as claimed in any one of claims 30 to 33, wherein the mode selection unit (16) determines the constrained temporal level sequence based on a delay control parameter D, where a determined temporal level sequence is a sequence of frames from a highest temporal level to a lowest temporal level among the frames of indexes not exceeding D in comparison to a frame at the lowest level, whose temporal redundancy is not eliminated, and a sequence of the frames from a lowest frame index to a highest frame index in a same temporal level.
  13. 38
    The video encoder as claimed in any one of claims 35 to 37, wherein the elimination of the temporal redundancy from the remaining frame is performed based on the remaining frame.
  14. 40
    The video encoder as claimed in any one of claims 26 to 39, wherein the bit-stream generation unit (40) generates the bit-stream including information on the constrained temporal level sequence.
  15. 41
    The video encoder as claimed in any one of claims 26 to 40, wherein the bit-stream generation unit (40) generates the bit-stream including information regarding sequences of eliminating temporal and spatial redundancies to obtain the transformation coefficients.
  16. 42
    A video decoding method comprising:(a) extracting information regarding encoded frames by receiving and interpreting a bit-stream;(b) obtaining transformation coefficients by inverse-quantizing the information regarding the encoded frames;and (c) restoring the encoded frames through an inverse-temporal transformation of the transformation coefficients in a constrained temporal level sequence.
  17. 46
    The method as claimed in any one of claims 42 to 45, wherein the constrained temporal level sequence is a sequence of the encoded frames from a highest temporal level to a lowest temporal level, and a sequence of the encoded frames from a highest frame index to a lowest frame index in a same temporal level.
  18. 49
    The method as claimed in any one of claims 42 to 48, wherein the constrained temporal level sequence is determined according to coding mode information extracted from the bit-stream input.
  19. 52
    The method as claimed in any one of claims 42 to 51, wherein the redundancy elimination sequence is extracted from the bit-stream.
  20. 53
    A video decoder for restoring frames from a bit-stream, the decoder comprising:a bit-stream interpretation unit (100) interpreting a bit-stream to extract information regarding encoded frames therefrom;an inverse-quantization unit (210) inverse-quantizing the information regarding the encoded frames to obtain transformation coefficients therefrom;an inverse spatial transformation unit (220) performing an inverse-spatial transformation process;and an inverse temporal transformation unit (230) performing an inverse-temporal transformation process in a constrained temporal level sequence,    wherein the encoded frames of the bit-stream are restored by performing the inverse-spatial process and the inverse-temporal transformation process on the transformation coefficients.
  21. 57
    The video decoder as claimed in any one of claims 53 to 56, wherein the constrained temporal is a sequence of the encoded frames from a highest temporal level to a lowest temporal level, and a sequence of the encoded frames from a highest frame index to a lowest frame index in a same temporal level.
  22. 60
    The video decoder as claimed in any one of claims 53 to 59, wherein the bit-stream interpretation unit (100) extracts coding mode information from the bit-stream input and determines the constrained temporal level sequence according to the coding mode information.
  23. 63
    The video decoder as claimed in any one of claims 53 to 62, wherein the redundancy elimination sequence is extracted from the input stream.
  24. 64
    A storage medium recording thereon a program readable by a computer so as to execute a video coding method comprising:eliminating a temporal redundancy in a constrained temporal level sequence from a plurality of frames of a video sequence;and generating a bit-stream by quantizing transformation coefficients obtained from the frames whose temporal redundancy has been eliminated.
  25. 65
    A storage medium recording thereon a program readable by a computer so as to execute a video coding method comprising:extracting information regarding encoded frames by receiving and interpreting a bit-stream;obtaining transformation coefficients by inverse-quantizing the information regarding the encoded frames;and restoring the encoded frames through an inverse-temporal transformation of the transformation coefficients in a constrained temporal level sequence.
Independent claims25