Nova Patents
EP1589765B1

Transformation block optimisation

Abstract

This record has no abstract on file.

EP1589765B1, drawing sheet 1
Sheet 1 of 47

Term

Term ended

Expired 4 October 2016, 10 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

17 claims: 11 independent, 6 dependent

  1. 1
    A method of decoding plural video objects in a video sequence, the method comprising:receiving encoded data for the plural video objects in the video sequence, wherein the plural video objects include a first video object and a second video object;decoding the first video object at a first time in the video sequence;and decoding the second video object at the first time in the video sequence;characterized in that: each of the first and second video objects is divided into plural blocks each having normal size or quarter size, wherein four quarter size blocks (380a-380d) are derived by sub-dividing a normal size block (374) into four equal sub-blocks;the encoded data includes: inter-coded data for the first video object at the first time in the video sequence, wherein the plural blocks of the first video object include one or more normal size blocks (374) each having a multi-dimensional affine transformation representing motion of the block (374) and at least four quarter size blocks (380a-380d) each having a multi-dimensional affine transformation representing motion of the block (380a-380d), and wherein the inter-coded data for the first video object includes the affine transformation for each of the one or more normal size blocks (374) and the affine transformation for each of the at least four quarter size blocks (380a-380d) of the first video object;one or more masks (66,80) that define shape for the first video object;inter-coded data for the second video object at the first time, wherein the plural blocks of the second video object include one or more normal size blocks each having a multi-dimensional affine transformation representing motion of the block (374) and at least four quarter size blocks (380a-380d) each having a multi-dimensional affine transformation representing motion of the block (380a-380d) and wherein the inter-coded data for the second video object includes the affine transformation for each of the one or more normal size blocks (374) and the affine transformation for each of the at least four quarter size blocks (380a-380d) of the second video object;and one or more masks (66,80) that define shape for the second video object;the decoding the first video object at the first time includes, using the affine transformation for each of the one or more normal size blocks (374) and the affine transformation for each of the at least four quarter size blocks (380a-380d) of the first video object to predict pixel values of pixels of the first video object at the first time from a previously decoded version of the first video object, wherein the one or more masks (66,80) that define shape for the first video object indicate which pixels are part of the first video object at the first time;and the decoding the second video object at the first time includes, using the affine transformation for each of the one or more normal size blocks (374) and the affine transformation for each of the at least four quarter size blocks (380a-380d) of the second video object to predict pixel values of pixels of the second video object at the first time from a previously decoded version of the second video object, wherein the one or more masks (60,80) that define shape for the second video object indicate which pixels are part of the second video object at the first time.
  2. 4
    An object-based video decoder adapted to decode encoded data for plural video objects in a video sequence, the object-based video decoder comprising:A receiver for receiving the encoded data for the plural video objects in the video sequence, wherein the plural video objects include a first video object and a second video object, and A processor for decoding the first video object at a first time in the video sequence and decoding the second video object at the first time in the video sequence;characterized in that: each of the first and second video objects is divided into plural blocks each having normal size or quarter size, wherein four quarter size blocks (380a-380d) are derived by sub-dividing a normal size block (374) into four equal sub-blocks;the encoded data includes: inter-coded data for the first video object at the first time in the video sequence, wherein the plural blocks of the first video object include one or more normal size blocks (374) each having a multi-dimensional affine transformation representing motion of the block and at least four quarter size blocks (380a-380d) each having a multi-dimensional affine transformation representing motion of the block, and wherein the inter-coded data for the first video object includes the affine transformation for each of the one or more normal size blocks (374) and the affine transformation for each of the at least four quarter size blocks (380a-380d) of the first video object;one or more masks (66,80) that define shape for the first video object;inter-coded data for the second video object at the first time, wherein the plural blocks of the second video object include one or more normal size blocks (374) each having a multi-dimensional affine transformation representing motion of the block (374) and at least four quarter size blocks (380a-380d) each having a multi-dimensional affine transformation representing motion of the block (380a-380d), and wherein the inter-coded data for the second video object includes the affine transformation for each of the one or more normal size blocks (374) and the affine transformation for each of the at least four quarter size blocks (380a-380d) of the second video object;and one or more masks (66,80) that define shape for the second video object;the decoding the first video object at the first time includes, using the affine transformation for each of the one or more normal size blocks (374) and the affine transformation for each of the at least four quarter size blocks (380a-380d) of the first video object to affine transformation predict pixel values of pixels of the first video object at the first time from a previously decoded version of the first video object, wherein the one or more masks (66,80) that define shape for the first video object indicate which pixels are part of the first video object at the first time;and the decoding the second video object at the first time includes, using the affine transformation for each of the one or more normal size blocks (374) and the affine transformation for each of the at least four quarter size blocks (380a-380d) of the second video object to predict pixel values of pixels of the second video object at the first time from a previously decoded version of the second video object, wherein the one or more masks that define shape for the second video object indicate which pixels are part of the second video object at the first time.
  3. 6
    The method of any one of claims 1 to 3 or 5, wherein the one or more masks (66,80) that define shape for the first video object and the one or more masks (66,80) that define shape for the second video object are binary masks.
  4. 7
    The method of any one of claims 1 to 3 or 5, wherein the one or more masks that define shape for the first video object and the one or more masks (66,80) that define shape for the second video object are multi-bit alpha channel masks.
  5. 8
    The method of any one of claims 1 to 3 or 5 to 7, wherein the affine transformation is coded in terms of pixel coordinates.
  6. 9
    The method of any one of claims 1 to 3 or 5 to 8, wherein each of the plural blocks (374, 380a-380d) is a rectangular array.
  7. 10
    A computer program comprising computer program code means adapted to perform all of the steps of any one of claims 1 to 3 or 5 to 7 when the program is run on a computer.
  8. 14
    The decoder of any one of claims 4, 12 or 13, wherein the one or more masks (66,80) that define shape for the first video object and the one or more masks (66,80) that define shape for the second video object are binary masks.
  9. 15
    The decoder of any one of claims 4, 12 or 13, wherein the one or more masks (66,80) that define shape for the first video object and the one or more masks (66,80) that define shape for the second video object are multi-bit alpha channel masks.
  10. 16
    The decoder of any one of claims 4 or 12 to 15, wherein the affine transformation is coded in terms of pixel coordinates.
  11. 17
    The decoder of any one of claims 4 or 12 to 15, wherein each of the plural blocks (374, 380a-380d) is a rectangular array.