Image encoding method, image decoding method, image encoding device, and image decoding device
3 claims: 3 independent, 0 dependent
- 1複数のピクチャにおける複数のブロックのうちそれぞれのブロックを符号化する符号化方法であって、 符号化対象のカレントブロックを含むピクチャとは異なる第1ピクチャに含まれる第1ブロックの第1動きベクトルから、前記カレントブロックの動きベクトルの符号化に用いられる予測動きベクトルの候補を導出する導出ステップと、 導出された前記候補を候補リストに追加する追加ステップと、 前記候補リストから、前記予測動きベクトルを選択する選択ステップと、 前記カレントブロックの動きベクトルおよび前記カレントブロックの参照ピクチャを用いて前記カレントブロックを符号化し、選択された前記予測動きベクトルを用いて前記動きベクトルを符号化する符号化ステップとを含み、 前記導出ステップでは、 前記カレントブロックの参照ピクチャがロングターム参照ピクチャであるかショートターム参照ピクチャであるか、および、前記第1ブロックの第1参照ピクチャがロングターム参照ピクチャであるかショートターム参照ピクチャであるかを判定し、 前記カレントブロックの参照ピクチャおよび前記第1ブロックの第1参照ピクチャのそれぞれがロングターム参照ピクチャであると判定された場合、前記第1動きベクトルから、時間的距離に基づくスケーリングを行わずに、前記候補を導出し、 前記カレントブロックの参照ピクチャおよび前記第1ブロックの第1参照ピクチャのそれぞれがショートターム参照ピクチャであると判定された場合、前記第1動きベクトルから、時間的距離に基づくスケーリングを行って、前記候補を導出する 符号化方法。
- 2複数のピクチャにおける複数のブロックのうちそれぞれのブロックを符号化する符号化装置であって、 符号化対象のカレントブロックを含むピクチャとは異なる第1ピクチャに含まれる第1ブロックの第1動きベクトルから、前記カレントブロックの動きベクトルの符号化に用いられる予測動きベクトルの候補を導出する導出部と、 導出された前記候補を候補リストに追加する追加部と、 前記候補リストから、前記予測動きベクトルを選択する選択部と、 前記カレントブロックの動きベクトルおよび前記カレントブロックの参照ピクチャを用いて前記カレントブロックを符号化し、選択された前記予測動きベクトルを用いて前記動きベクトルを符号化する符号化部とを備え、 前記導出部は、 前記カレントブロックの参照ピクチャがロングターム参照ピクチャであるかショートターム参照ピクチャであるか、および、前記第1ブロックの第1参照ピクチャがロングターム参照ピクチャであるかショートターム参照ピクチャであるかを判定し、 前記カレントブロックの参照ピクチャおよび前記第1ブロックの第1参照ピクチャのそれぞれがロングターム参照ピクチャであると判定された場合、前記第1動きベクトルから、時間的距離に基づくスケーリングを行わずに、前記候補を導出し、 前記カレントブロックの参照ピクチャおよび前記第1ブロックの第1参照ピクチャのそれぞれがショートターム参照ピクチャであると判定された場合、前記第1動きベクトルから、時間的距離に基づくスケーリングを行って、前記候補を導出する 符号化装置。
- 3処理回路と、前記処理回路がアクセス可能な記憶装置とを備えるシステムであって、 前記処理回路は、前記記憶装置を用いて、複数のピクチャにおける複数のブロックのうちそれぞれのブロックを符号化する動作を実行し、 前記動作は、 符号化対象のカレントブロックを含むピクチャとは異なる第1ピクチャに含まれる第1ブロックの第1動きベクトルから、前記カレントブロックの動きベクトルの符号化に用いられる予測動きベクトルの候補を導出する導出ステップと、 導出された前記候補を候補リストに追加する追加ステップと、 前記候補リストから、前記予測動きベクトルを選択する選択ステップと、 前記カレントブロックの動きベクトルおよび前記カレントブロックの参照ピクチャを用いて前記カレントブロックを符号化し、選択された前記予測動きベクトルを用いて前記動きベクトルを符号化する符号化ステップとを含み、 前記導出ステップでは、 前記カレントブロックの参照ピクチャがロングターム参照ピクチャであるかショートターム参照ピクチャであるか、および、前記第1ブロックの第1参照ピクチャがロングターム参照ピクチャであるかショートターム参照ピクチャであるかを判定し、 前記カレントブロックの参照ピクチャおよび前記第1ブロックの第1参照ピクチャのそれぞれがロングターム参照ピクチャであると判定された場合、前記第1動きベクトルから、時間的距離に基づくスケーリングを行わずに、前記候補を導出し、 前記カレントブロックの参照ピクチャおよび前記第1ブロックの第1参照ピクチャのそれぞれがショートターム参照ピクチャであると判定された場合、前記第1動きベクトルから、時間的距離に基づくスケーリングを行って、前記候補を導出する システム。
Independent claims3
313 paragraphs, as filed
0001The present invention relates to an image coding method for encoding each of a plurality of blocks in a plurality of pictures.
0002There is a technique described in Non-Patent Document 1 as a technique relating to an image coding method for encoding each of a plurality of blocks in a plurality of pictures.
<p num="0003"><nplcit num="1"><text>ISO / IEC 14496-10 "MPEG-4 Part10 Advanced Video Coding"</text></nplcit></p>
<p num="0004"> However, with the conventional image coding method, a sufficiently high coding efficiency may not be obtained.</p><p num="0005"> Therefore, the present invention provides an image coding method capable of improving the coding efficiency in image coding.</p>
<p num="0006"> The coding method according to one aspect of the present invention is a coding method for coding each block among a plurality of blocks in a plurality of pictures, and is a first picture different from the picture including the current block to be coded. A derivation step for deriving a candidate for a predicted motion vector used for encoding the motion vector of the current block from the first motion vector of the first block included in, and an additional step for adding the derived candidate to the candidate list. Then, the current block is encoded by using the selection step for selecting the predicted motion vector from the candidate list, the motion vector of the current block, and the reference picture of the current block, and the selected predicted motion vector is used. In the derivation step, whether the reference picture of the current block is a long-term reference picture or a short-term reference picture, and a first block of the first block, including a coding step for encoding the motion vector. 1 It is determined whether the reference picture is a long-term reference picture or a short-term reference picture, and it is determined that each of the reference picture in the current block and the first reference picture in the first block is a long-term reference picture. If, the candidate is derived from the first motion vector without scaling based on the time distance, and each of the reference picture of the current block and the first reference picture of the first block is a short-term reference picture. If it is determined that the above is true, the candidate is derived from the first motion vector by scaling based on the time distance.</p><p num="0007"> It should be noted that these comprehensive or specific embodiments may be implemented in non-temporary recording media such as systems, devices, integrated circuits, computer programs or computer readable CD-ROMs, systems, devices, methods. , Integrated circuits, computer programs and any combination of recording media.</p>
<p num="0008"> The image coding method of the present invention can improve the coding efficiency in coding an image.</p>
0009<figref num="1">FIG. 1 is a flowchart showing the operation of the image coding apparatus according to the reference example.</figref><figref num="2">FIG. 2 is a flowchart showing the operation of the image decoding device according to the reference example.</figref><figref num="3">FIG. 3 is a flowchart showing the details of the derivation process according to the reference example.</figref><figref num="4">FIG. 4 is a diagram for explaining a co-located block according to a reference example.</figref><figref num="5">FIG. 5 is a block diagram of the image coding apparatus according to the first embodiment.</figref><figref num="6">FIG. 6 is a block diagram of the image decoding device according to the first embodiment.</figref><figref num="7">FIG. 7 is a flowchart showing the operation of the image coding apparatus according to the first embodiment.</figref><figref num="8">FIG. 8 is a flowchart showing the operation of the image decoding apparatus according to the first embodiment.</figref><figref num="9">FIG. 9 is a flowchart showing the details of the derivation process according to the first embodiment.</figref><figref num="10">FIG. 10 is a flowchart showing the details of the derivation process according to the second embodiment.</figref><figref num="11">FIG. 11 is a diagram for explaining the co-located block according to the second embodiment.</figref><figref num="12">FIG. 12 is a flowchart showing the details of the derivation process according to the third embodiment.</figref><figref num="13A">FIG. 13A is a block diagram of the image coding apparatus according to the fourth embodiment.</figref><figref num="13B">FIG. 13B is a flowchart showing the operation of the image coding apparatus according to the fourth embodiment.</figref><figref num="14A">FIG. 14A is a block diagram of the image decoding device according to the fourth embodiment.</figref><figref num="14B">FIG. 14B is a flowchart showing the operation of the image decoding apparatus according to the fourth embodiment.</figref><figref num="15A">FIG. 15A is a diagram showing a first example of the storage position of the parameter indicating the classification of the reference picture.</figref><figref num="15B">FIG. 15B is a diagram showing a second example of the storage position of the parameter indicating the classification of the reference picture.</figref><figref num="15C">FIG. 15C is a diagram showing a third example of the storage position of the parameter indicating the classification of the reference picture.</figref><figref num="16">FIG. 16 is a diagram showing an example of a storage position of a parameter indicating a prediction mode.</figref><figref num="17">FIG. 17 is an overall configuration diagram of a content supply system that realizes a content distribution service.</figref><figref num="18">FIG. 18 is an overall configuration diagram of the digital broadcasting system.</figref><figref num="19">FIG. 19 is a block diagram showing a configuration example of a television.</figref><figref num="20">FIG. 20 is a block diagram showing a configuration example of an information reproduction / recording unit that reads / writes information to / from a recording medium such as an optical disc.</figref><figref num="21">FIG. 21 is a diagram showing a structural example of a recording medium which is an optical disc.</figref><figref num="22A">FIG. 22A is a diagram showing an example of a mobile phone.</figref><figref num="22B">FIG. 22B is a block diagram showing a configuration example of a mobile phone.</figref><figref num="23">FIG. 23 is a diagram showing the structure of the multiplexed data.</figref><figref num="24">FIG. 24 is a diagram schematically showing how each stream is multiplexed in the multiplexed data.</figref><figref num="25">FIG. 25 is a more detailed diagram of how a video stream is stored in a PES packet sequence.</figref><figref num="26">FIG. 26 is a diagram showing the structure of TS packets and source packets in the multiplexed data.</figref><figref num="27">FIG. 27 is a diagram showing a data structure of PMT.</figref><figref num="28">FIG. 28 is a diagram showing an internal configuration of multiplexed data information.</figref><figref num="29">FIG. 29 is a diagram showing an internal configuration of stream attribute information.</figref><figref num="30">FIG. 30 is a diagram showing steps for identifying video data.</figref><figref num="31">FIG. 31 is a block diagram showing a configuration example of an integrated circuit that realizes the moving image coding method and the moving image decoding method of each embodiment.</figref><figref num="32">FIG. 32 is a diagram showing a configuration for switching the drive frequency.</figref><figref num="33">FIG. 33 is a diagram showing steps for identifying video data and switching the drive frequency.</figref><figref num="34">FIG. 34 is a diagram showing an example of a look-up table in which the video data standard and the drive frequency are associated with each other.</figref><figref num="35A">FIG. 35A is a diagram showing an example of a configuration in which the module of the signal processing unit is shared.</figref><figref num="35B">FIG. 35B is a diagram showing another example of the configuration in which the module of the signal processing unit is shared.</figref>
0010(Knowledge on which the present invention is based) The present inventor has found that the following problems arise with respect to the image coding method described in the Background Art column. In the following, the image may be any of a moving image composed of a plurality of pictures, a still image composed of one picture, a part of the pictures, and the like.
0011Recent image coding methods include MPEG-4 AVC / H.264 and HEVC (High Efficiency Video Coding). In these image coding schemes, inter-prediction using encoded reference pictures is available.
0012Further, in these image coding methods, a reference picture called a long-term reference picture may be used. For example, when the reference picture is maintained in the DPB (Decoded Picture Buffer) for a long time, the reference picture may be used as a long-term reference picture.
0013In HEVC, there is a mode called AMVP (Adaptive Motion Vector Prediction) mode. In the AMVP mode, the predicted motion vector obtained by predicting the motion vector of the current block from the motion vector of the adjacent block or the like is used for coding the motion vector of the current block.
0014In addition, time-predicted motion vectors are available in HEVC. The time-predicted motion vector is derived from the motion vector of the co-located block in the encoded co-located picture. The coordinates of the co-located block in the co-located picture correspond to the coordinates of the current block in the current picture to be encoded.
0015Here, the motion vector of the co-located block may be called a co-located motion vector. Further, the reference picture of the co-located block may be called a co-located reference picture. The co-located block is encoded with a co-located motion vector and a co-located reference picture. In addition, co-located may be described as collocated or collocated.
0016Similarly, the motion vector of the current block current motion vector may be referred to as torr. Further, the reference picture of the current block may be referred to as a current reference picture. The current block is encoded using the current motion vector and the current reference picture.
0017The above-mentioned current block and co-located block are Prediction Units (PUs), respectively. A prediction unit is a block of images and is defined as a data unit for prediction. In HEVC, a coding unit (CU) is defined separately from a prediction unit as a coding data unit. The prediction unit is a block within the coding unit. The blocks described below may be replaced with predictive units or coding units.
0018The size of the coding unit and the prediction unit is not constant. For example, a picture may contain multiple coding units of various sizes, and a picture may contain multiple prediction units of various sizes.
0019Therefore, a block that exactly matches the area of the current block may not be defined in the co-located picture. Therefore, in HEVC, the co-located block is selected from a plurality of blocks included in the co-located picture by a predetermined selection method.
0020The time-predicted motion vector is generated by scaling the motion vector of the selected co-located block according to the POC (proof of concept count) distance. POC is an ordinal number assigned to a picture in display order. The POC distance corresponds to the temporal distance between the two pictures. Scaling based on POC distance is also called POC-based scaling. Equation 1 shown below is an arithmetic expression that performs POC-based scaling on the motion vector of the co-located block.
0021pmv = (tb / td) × colmv (Equation 1)
0022Here, colmv is the motion vector of the co-located block. pmv is a time-predicted motion vector derived from the motion vector of the co-located block. tb is the signed POC distance, which is the difference from the current picture to the current reference picture. td is the signed POC distance, which is the difference from the co-located picture to the co-located reference picture.
0023If a valid time-predicted motion vector exists, the time-predicted motion vector is placed in an ordered list of predicted motion vector candidates. The predicted motion vector used to code the current motion vector is selected from the ordered list of predicted motion vector candidates. The selected predicted motion vector is then indicated by the parameters in the coded stream.
0024FIG. 1 is a flowchart showing the operation of the image coding apparatus according to the reference example. In particular, FIG. 1 shows a process of encoding an image by inter-prediction.
0025First, the image encoding device classifies each of the plurality of reference pictures into a short-term reference picture or a long-term reference picture (S101). The image coding device writes information indicating the classification of each of the plurality of reference pictures in the header of the coded stream (S102).
0026Next, the image coding device identifies the current reference picture and the current motion vector by motion detection (S103). The image encoder then derives a predicted motion vector (S104). The details of the derivation process will be described later.
0027Next, the image coding device subtracts the predicted motion vector from the current motion vector to derive the differential motion vector (S105). Next, the image coding device generates a prediction block by performing motion compensation using the current reference picture and the current motion vector (S106).
0028Next, the image coding device subtracts the prediction block from the current block to generate a residual block (S107). Finally, the image coding apparatus encodes the residual block, the differential motion vector, and the reference index indicating the current reference picture to generate a coded stream containing these (S108).
0029FIG. 2 is a flowchart showing the operation of the image decoding device according to the reference example. In particular, FIG. 2 shows a process of decoding an image by inter-prediction.
0030First, the image decoding device acquires the coded stream and parses the header of the coded stream to acquire information indicating the classification of each of the plurality of reference pictures (S201). Further, the image decoding apparatus acquires the residual block, the differential motion vector, and the reference index indicating the current reference picture by analyzing the coded stream (S202).
0031Next, the image decoding device derives a predicted motion vector (S203). The details of the derivation process will be described later. Next, the image decoding device adds the predicted motion vector to the differential motion vector to generate the current motion vector (S204). Next, the image decoding device generates a prediction block by performing motion compensation using the current reference picture and the current motion vector (S205). Finally, the image decoding device adds the prediction block to the residual block to generate the reconstruction block (S206).
0032FIG. 3 is a flowchart showing details of the derivation process shown in FIGS. 1 and 2. The following shows the operation of the image coding device. If the coding is read as decoding, the operation of the image decoding device is the same as the operation of the image coding device.
0033First, the image encoder selects a co-located picture (S301). The image encoder then selects the co-located block in the co-located picture (S302). The image encoder then identifies the co-located reference picture and the co-located motion vector (S303). The image coding device then derives the predicted motion vector according to a derivation method that performs POC-based scaling (S304).
0034FIG. 4 is a diagram for explaining the co-located block used in the derivation process shown in FIG. The co-located block is selected from multiple blocks in the co-located picture.
0035A co-located picture is different from a current picture that contains a current block. For example, a co-located picture is a picture immediately before or after the current picture in the display order. More specifically, for example, the co-located picture is the first reference picture in any of the two reference picture lists used to encode the B picture (bi-predictive coding).
0036The first block containing sample c0 in the co-located picture is the first candidate for the co-located block and is also called the primary co-located block. The second block containing sample c1 in the co-located picture is the second candidate for the co-located block and is also called the secondary co-located block.
0037If the coordinates of the sample tl in the upper left of the current block are (x, y), the width of the current block is w, and the height of the current block is h, the coordinates of the sample c0 are (x + w, y +). h). In this case, the coordinates of sample c1 are (x + (w / 2) -1, y + (h / 2) -1).
0038If the first block is not available, the second block is selected as the co-located block. When the first block is not available, there are cases where the first block does not exist because the current block is at the right or bottom edge of the picture, or the first block is encoded by intra-prediction.
0039Hereinafter, a more specific example of the process of deriving the time prediction motion vector will be described with reference to FIG. 3 again.
0040First, the image encoder selects a co-located picture (S301). The image encoder then selects the co-located block (S302). If the first block containing sample c0 shown in FIG. 4 is available, the first block is selected as the co-located block. If the first block is not available and the second block containing sample c1 shown in FIG. 4 is available, the second block is selected as the co-located block.
0041If an available co-located block is selected, the image encoder sets the time-predicted motion vector as available. If no available co-located block is selected, the image encoder sets the time-predicted motion vector as unavailable.
0042If the time-predicted motion vector is set as available, the image encoder identifies the co-located motion vector as the reference motion vector. The image coding device also identifies the co-located reference picture (S303). Then, the image coding device derives the time prediction motion vector from the reference motion vector by scaling the equation 1 (S304).
0043Through the above processing, the image coding device and the image decoding device derive a time prediction motion vector.
0044However, it may be difficult to derive an appropriate time prediction motion vector due to the relationship between the current picture, the current reference picture, the co-located picture, and the co-located reference picture.
0045For example, when the current reference picture is a long-term reference picture, the time distance from the current reference picture to the current picture may be long. Further, when the co-located reference picture is a long-term reference picture, the time distance from the co-located reference picture to the co-located reference picture may be long.
0046In these cases, POC-based scaling can generate extremely large or small time-predicted motion vectors. As a result, the prediction accuracy deteriorates and the coding efficiency deteriorates. In particular, with a fixed number of bits, an extremely large or small time prediction motion vector is not properly expressed, and the prediction accuracy and the coding efficiency are significantly deteriorated.
0047In order to solve such a problem, the image coding method according to one aspect of the present invention is an image coding method that encodes each of a plurality of blocks in a plurality of pictures, and is a current block to be encoded. A derivation step for deriving a candidate for a predicted motion vector used for encoding the motion vector of the current block from a motion vector of a co-located block, which is a block included in a picture different from the picture containing the above, and the derived step. The current block using the additional step of adding the candidate to the list, the selection step of selecting the predicted motion vector from the list to which the candidate is added, the motion vector of the current block, and the reference picture of the current block. Including a coding step of encoding the motion vector of the current block using the selected predicted motion vector, and in the derivation step, is the reference picture of the current block a long-term reference picture? It is determined whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture, and the reference picture of the current block and the co-located block of the current block. When it is determined that each of the reference pictures is a long-term reference picture, the candidate is derived from the motion vector of the co-located block by the first derivation method that does not perform scaling based on the time distance, and the current block is derived. When it is determined that the reference picture of the co-located block and the reference picture of the co-located block are short-term reference pictures, the motion vector of the co-located block is scaled based on the time distance by the second derivation method. The candidate is derived.
0048As a result, the candidates for the predicted motion vector are appropriately derived without becoming extremely large or extremely small. Therefore, the prediction accuracy can be improved, and the coding efficiency can be improved.
0049For example, in the derivation step, when one of the reference picture of the current block and the reference picture of the co-located block is determined to be a long-term reference picture and the other is determined to be a short-term reference picture. , The candidate is not derived from the motion vector of the co-located block, and it is determined that the reference picture of the current block and the reference picture of the co-located block are long-term reference pictures, respectively, or the current When it is determined that the reference picture of the block and the reference picture of the co-located block are short-term reference pictures, the candidate may be derived from the motion vector of the co-located block.
0050As a result, if the prediction accuracy is expected to be low, the prediction motion vector candidates are not derived from the motion vector of the co-located block. Therefore, deterioration of prediction accuracy is suppressed.
0051Further, for example, in the coding step, information indicating whether the reference picture of the current block is a long-term reference picture or a short-term reference picture, and the reference picture of the co-located block is a long-term reference picture. Information indicating whether the picture is a reference picture or a short-term reference picture may be encoded.
0052As a result, information indicating whether each reference picture is a long-term reference picture or a short-term reference picture is notified from the encoding side to the decoding side. Therefore, the same determination result is obtained on the coding side and the decoding side, and the same processing is performed.
0053Further, for example, in the derivation step, the reference picture of the current block is a long-term reference picture or a short-term reference picture by using the time distance from the reference picture of the current block to the picture including the current block. It is determined whether there is, and the reference picture of the co-located block is a long-term reference picture or a short-term by using the time distance from the reference picture of the co-located block to the picture including the co-located block. It may be determined whether it is a reference picture.
0054Thereby, whether each reference picture is a long-term reference picture or a short-term reference picture is concisely and appropriately determined based on the time distance.
0055Further, for example, in the derivation step, it is determined whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture during the period when the co-located block is encoded. May be good.
0056As a result, it is more accurately determined whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture.
0057Further, for example, in the derivation step, it may be determined whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture during the period when the current block is coded. ..
0058As a result, the information on whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture does not have to be maintained for a long period of time.
0059Further, for example, in the derivation step, when it is determined that the reference picture of the current block and the reference picture of the co-located block are long-term reference pictures, the motion vector of the co-located block is used as the candidate. When it is derived and it is determined that the reference picture of the current block and the reference picture of the co-located block are short-term reference pictures, from the reference picture of the co-located block to the picture including the co-located block. The candidate is derived by scaling the motion vector of the co-located block by using the ratio of the time distance from the reference picture of the current block to the picture including the current block with respect to the time distance of the co-located block. You may.
0060As a result, when the two reference pictures are long-term reference pictures, scaling is omitted and the amount of calculation is reduced. Then, when the two reference pictures are short-term reference pictures, the candidates for the predicted motion vector are appropriately derived based on the time distance.
0061Further, for example, in the derivation step, when it is further determined that the reference picture of the current block is a short-term reference picture and the reference picture of the co-located block is a long-term reference picture, the above-mentioned The second derivation method is performed by selecting another co-located block encoded by referring to the short-term reference picture without deriving the candidate from the co-located block and using the motion vector of the other co-located block. The candidate may be derived by.
0062As a result, a block for deriving a candidate with high prediction accuracy is selected. Therefore, the prediction accuracy is improved.
0063Further, the image decoding method according to one aspect of the present invention is an image decoding method for decoding each of a plurality of blocks in a plurality of pictures, and is a block included in a picture different from the picture including the current block to be decoded. A derivation step for deriving a candidate for a predicted motion vector used for decoding the motion vector of the current block from a motion vector of a co-located block, an additional step for adding the derived candidate to the list, and the candidate From the added list, the motion vector of the current block is decoded using the selection step for selecting the predicted motion vector and the selected predicted motion vector, and the motion vector of the current block and the reference of the current block are referenced. In the derivation step, the derivation step includes a decoding step of decoding the current block using a picture, whether the reference picture of the current block is a long-term reference picture or a short-term reference picture, and the co-located block. When it is determined whether the reference picture of is a long-term reference picture or a short-term reference picture, and it is determined that the reference picture of the current block and the reference picture of the co-located block are long-term reference pictures, respectively. , The candidate is derived from the motion vector of the co-located block by the first derivation method that does not perform scaling based on the time distance, and the reference picture of the current block and the reference picture of the co-located block are shorted, respectively. When it is determined that the picture is a term reference picture, the image decoding method may be used in which the candidate is derived from the motion vector of the co-located block by a second derivation method that performs scaling based on the time distance.
0064As a result, the candidates for the predicted motion vector are appropriately derived without becoming extremely large or extremely small. Therefore, the prediction accuracy can be improved, and the coding efficiency can be improved.
0065For example, in the derivation step, when one of the reference picture of the current block and the reference picture of the co-located block is determined to be a long-term reference picture and the other is determined to be a short-term reference picture. , The candidate is not derived from the motion vector of the co-located block, and it is determined that the reference picture of the current block and the reference picture of the co-located block are long-term reference pictures, respectively, or the current When it is determined that the reference picture of the block and the reference picture of the co-located block are short-term reference pictures, the candidate may be derived from the motion vector of the co-located block.
0066As a result, if the prediction accuracy is expected to be low, the prediction motion vector candidates are not derived from the motion vector of the co-located block. Therefore, deterioration of prediction accuracy is suppressed.
0067Further, for example, in the decoding step, information indicating whether the reference picture of the current block is a long-term reference picture or a short-term reference picture, and the reference picture of the co-located block are long-term references. Information indicating whether the picture is a picture or a short-term reference picture is decoded, and in the derivation step, information indicating whether the reference picture in the current block is a long-term reference picture or a short-term reference picture is used. , Determines whether the reference picture of the current block is a long-term reference picture or a short-term reference picture, and determines whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture. Using the information shown, it may be determined whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture.
0068As a result, information indicating whether each reference picture is a long-term reference picture or a short-term reference picture is notified from the encoding side to the decoding side. Therefore, the same determination result is obtained on the coding side and the decoding side, and the same processing is performed.
0069Further, for example, in the derivation step, the reference picture of the current block is a long-term reference picture or a short-term reference picture by using the time distance from the reference picture of the current block to the picture including the current block. It is determined whether there is, and the reference picture of the co-located block is a long-term reference picture or a short-term by using the time distance from the reference picture of the co-located block to the picture including the co-located block. It may be determined whether it is a reference picture.
0070Thereby, whether each reference picture is a long-term reference picture or a short-term reference picture is concisely and appropriately determined based on the time distance.
0071Further, for example, in the derivation step, it may be determined whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture during the period when the co-located block is decoded. Good.
0072As a result, it is more accurately determined whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture.
0073Further, for example, in the derivation step, it may be determined whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture during the period when the current block is decoded.
0074As a result, the information on whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture does not have to be maintained for a long period of time.
0075Further, for example, in the derivation step, when it is determined that the reference picture of the current block and the reference picture of the co-located block are long-term reference pictures, the motion vector of the co-located block is used as the candidate. When it is derived and it is determined that the reference picture of the current block and the reference picture of the co-located block are short-term reference pictures, from the reference picture of the co-located block to the picture including the co-located block. The candidate is derived by scaling the motion vector of the co-located block by using the ratio of the time distance from the reference picture of the current block to the picture including the current block with respect to the time distance of the co-located block. You may.
0076As a result, when the two reference pictures are long-term reference pictures, scaling is omitted and the amount of calculation is reduced. Then, when the two reference pictures are short-term reference pictures, the candidates for the predicted motion vector are appropriately derived based on the time distance.
0077Further, for example, in the derivation step, when it is further determined that the reference picture of the current block is a short-term reference picture and the reference picture of the co-located block is a long-term reference picture, the above-mentioned The candidate is not derived from the co-located block, another co-located block decoded by referring to the short term reference picture is selected, and the motion vector of the other co-located block is derived by the second derivation method. The candidate may be derived.
0078As a result, a block for deriving a candidate with high prediction accuracy is selected. Therefore, the prediction accuracy is improved.
0079Further, in the content supply method according to one aspect of the present invention, the image data is transmitted from the server in which the image data encoded by the image coding method is recorded in response to a request from an external terminal.
0080It should be noted that these comprehensive or specific embodiments may be implemented in non-temporary recording media such as systems, devices, integrated circuits, computer programs or computer readable CD-ROMs, systems, devices, methods. , Integrated circuits, computer programs or any combination of recording media.
0081Hereinafter, embodiments will be specifically described with reference to the drawings. It should be noted that all of the embodiments described below are comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, the order of steps, etc. shown in the following embodiments are examples, and are not intended to limit the present invention. Further, among the components in the following embodiments, the components not described in the independent claims indicating the highest level concept are described as arbitrary components.
0082(Embodiment 1) FIG. 5 is a block diagram of the image coding apparatus according to the present embodiment. The image coding apparatus 500 shown in FIG. 5 encodes an image block by block and outputs a coded stream containing the coded image. Specifically, the image coding device 500 includes a subtraction unit 501, a conversion unit 502, a quantization unit 503, an entropy coding unit 504, an inverse quantization unit 505, an inverse conversion unit 506, an addition unit 507, and a block memory 508. It includes a picture memory 509, an intra prediction unit 510, an inter prediction unit 511, and a selection unit 512.
0083The subtraction unit 501 outputs a difference image by subtracting the predicted image from the image input to the image coding device 500. The conversion unit 502 generates a plurality of frequency coefficients by frequency-converting the difference image output from the subtraction unit 501. The quantization unit 503 generates a plurality of quantization coefficients by quantizing the plurality of frequency coefficients generated by the conversion unit 502. The entropy coding unit 504 generates a coded stream by encoding a plurality of quantization coefficients generated by the quantization unit 503.
0084The dequantization unit 505 restores a plurality of frequency coefficients by dequantizing a plurality of quantization coefficients generated by the quantization unit 503. The inverse conversion unit 506 restores the difference image by performing inverse frequency conversion of a plurality of frequency coefficients restored by the inverse quantization unit 505. The addition unit 507 restores (reconstructs) the image by adding the predicted image to the difference image restored by the inverse conversion unit 506. The addition unit 507 stores the restored image (reconstructed image) in the block memory 508 and the picture memory 509.
0085The block memory 508 is a memory for storing the image restored by the addition unit 507 for each block. Further, the picture memory 509 is a memory for storing the image restored by the addition unit 507 for each picture.
0086The intra prediction unit 510 refers to the block memory 508 to perform intra prediction. That is, the intra prediction unit 510 predicts the pixel value in the picture from the other pixel values in the picture. As a result, the intra prediction unit 510 generates a prediction image. In addition, the inter-prediction unit 511 refers to the picture memory 509 to perform inter-prediction. That is, the inter-prediction unit 511 predicts the pixel value in the picture from the pixel value in another picture. As a result, the inter-prediction unit 511 generates a prediction image.
0087The selection unit 512 selects either the prediction image generated by the intra prediction unit 510 or the prediction image generated by the inter prediction unit 511, and outputs the selected prediction image to the subtraction unit 501 and the addition unit 507. To do.
0088Although not shown in FIG. 5, the image coding apparatus 500 may include a deblock filter unit. Then, the deblock filter unit may remove the noise near the block boundary by performing the deblock filter process on the image restored by the addition unit 507. Further, the image coding device 500 may include a control unit that controls each process in the image coding device 500.
0089FIG. 6 is a block diagram of the image decoding device according to the present embodiment. The image decoding device 600 shown in FIG. 6 acquires a coded stream and decodes the image block by block. Specifically, the image decoding device 600 includes an entropy decoding unit 601, an inverse quantization unit 602, an inverse conversion unit 603, an addition unit 604, a block memory 605, a picture memory 606, an intra prediction unit 607, an inter prediction unit 608, and a selection unit. It has a part 609.
0090The entropy decoding unit 601 decodes a plurality of encoded quantization coefficients included in the encoded stream. The dequantization unit 602 restores a plurality of frequency coefficients by dequantizing a plurality of quantization coefficients decoded by the entropy decoding unit 601. The inverse conversion unit 603 restores the difference image by performing inverse frequency conversion of a plurality of frequency coefficients restored by the inverse quantization unit 602.
0091The addition unit 604 restores (reconstructs) the image by adding the predicted image to the difference image restored by the inverse conversion unit 603. The addition unit 604 outputs the restored image (reconstructed image). Further, the addition unit 604 stores the restored image in the block memory 605 and the picture memory 606.
0092The block memory 605 is a memory for storing the image restored by the addition unit 604 for each block. Further, the picture memory 606 is a memory for storing the image restored by the addition unit 604 for each picture.
0093The intra prediction unit 607 refers to the block memory 605 to perform intra prediction. That is, the intra prediction unit 607 predicts the pixel value in the picture from the other pixel values in the picture. As a result, the intra prediction unit 607 generates a prediction image. In addition, the inter-prediction unit 608 refers to the picture memory 606 to perform inter-prediction. That is, the inter-prediction unit 608 predicts the pixel values in the picture from the pixel values in the other pictures. As a result, the inter-prediction unit 608 generates a prediction image.
0094The selection unit 609 selects either the prediction image generated by the intra prediction unit 607 or the prediction image generated by the inter prediction unit 608, and outputs the selected prediction image to the addition unit 604.
0095Although not shown in FIG. 6, the image decoding device 600 may include a deblocking filter unit. Then, the deblock filter unit may remove the noise near the block boundary by performing the deblock filter process on the image restored by the addition unit 604. Further, the image decoding device 600 may include a control unit that controls each process in the image decoding device 600.
0096The above coding process and decoding process are performed for each coding unit. The transformation process, the quantization process, the inverse conversion process, and the inverse quantization process are performed for each transformation unit (TU: Transform Unit) in the coding unit. The prediction process is performed for each prediction unit in the coding unit.
0097FIG. 7 is a flowchart showing the operation of the image coding device 500 shown in FIG. In particular, FIG. 7 shows a process of encoding an image by inter-prediction.
0098First, the inter-prediction unit 511 classifies each of the plurality of reference pictures into a short-term reference picture or a long-term reference picture (S701).
0099A long-term reference picture is a reference picture suitable for long-term use. Further, the long-term reference picture is defined as a reference picture to be used longer than the short-term reference picture. Therefore, the long-term reference picture is likely to be maintained in the picture memory 509 for a long time. The long-term reference picture is specified by an absolute POC that does not depend on the current picture. On the other hand, the short term reference picture is specified by the POC relative to the current picture.
0100Next, the entropy encoding unit 504 writes information indicating the classification of each of the plurality of reference pictures in the header of the encoded stream (S702). That is, the entropy encoding unit 504 writes information indicating whether each of the plurality of reference pictures is a long-term reference picture or a short-term reference picture.
0101Next, the inter-prediction unit 511 identifies the reference picture and the motion vector of the current block of the coding target (prediction target) by motion detection (S703). Next, the inter-prediction unit 511 derives the predicted motion vector (S704). The details of the derivation process will be described later.
0102Next, the inter-prediction unit 511 subtracts the predicted motion vector from the current motion vector to derive the differential motion vector (S705). Next, the inter-prediction unit 511 generates a prediction block by performing motion compensation using the current reference picture and the current motion vector (S706).
0103Next, the subtraction unit 501 subtracts the prediction block from the current block (original image) to generate a residual block (S707). Finally, the entropy coding unit 504 encodes the residual block, the differential motion vector, and the reference index indicating the current reference picture to generate a coded stream containing these (S708).
0104FIG. 8 is a flowchart showing the operation of the image decoding apparatus 600 shown in FIG. In particular, FIG. 8 shows a process of decoding an image by inter-prediction.
0105First, the entropy decoding unit 601 acquires the coded stream and parses the header of the coded stream to acquire information indicating the classification of each of the plurality of reference pictures (S801). That is, the entropy decoding unit 601 acquires information indicating whether each of the plurality of reference pictures is a long-term reference picture or a short-term reference picture.
0106Further, the entropy decoding unit 601 acquires the residual block, the differential motion vector, and the reference index indicating the current reference picture by analyzing the coded stream (S802).
0107Next, the inter-prediction unit 608 derives the predicted motion vector (S803). The details of the derivation process will be described later. Next, the inter-prediction unit 608 adds the predicted motion vector to the differential motion vector to generate the current motion vector (S804). Next, the inter-prediction unit 608 generates a prediction block by performing motion compensation using the current reference picture and the current motion vector (S805). Finally, the addition unit 604 adds the prediction block to the residual block to generate the reconstruction block (S806).
0108FIG. 9 is a flowchart showing details of the derivation process shown in FIGS. 7 and 8. The following mainly shows the operation of the inter-prediction unit 511 in FIG. If the coding is read as decoding, the operation of the inter-prediction unit 608 in FIG. 6 is the same as the operation of the inter-prediction unit 511 in FIG.
0109First, the inter-prediction unit 511 selects a co-located picture from a plurality of available reference pictures (S901). The plurality of reference pictures available are encoded pictures and pictures maintained in picture memory 509.
0110Next, the inter-prediction unit 511 selects the co-located block in the co-located picture (S902). Then, the inter-prediction unit 511 identifies the co-located reference picture and the co-located motion vector (S903).
0111Next, the inter-prediction unit 511 determines whether or not any of the current reference picture and the co-located reference picture is a long-term reference picture (S904). Then, when it is determined that either the current reference picture or the co-located reference picture is a long-term reference picture (Yes in S904), the inter-prediction unit 511 derives the predicted motion vector according to the first derivation method. (S905).
0112The first derivation method is a method using a co-located motion vector. More specifically, the first derivation method is a method of directly deriving a co-located motion vector as a predicted motion vector without POC-based scaling. The first derivation method may be a method of deriving a predicted motion vector by scaling the co-located motion vector at a predetermined constant ratio.
0113When it is determined that neither the current reference picture nor the co-located reference picture is a long-term reference picture (No in S904), the inter-prediction unit 511 derives the predicted motion vector according to the second derivation method (S906). .. That is, when it is determined that both the current reference picture and the co-located reference picture are short-term reference pictures, the inter-prediction unit 511 derives the predicted motion vector according to the second derivation method.
0114The second derivation method is a method using a current reference picture, a co-located reference picture, and a co-located motion vector. More specifically, the second derivation method is a method of deriving a predicted motion vector by performing POC-based scaling (Equation 1) on the co-located motion vector.
0115Hereinafter, a more specific example of the process of deriving the time prediction motion vector will be described with reference to FIG. 9 again. The derivation process described above may be modified as follows.
0116First, the inter-prediction unit 511 selects a co-located picture (S901). More specifically, if the slice header parameter slice_type is B and the slice header parameter collapse_from_l0_flag is 0, the picture RefPicList1 [0] is selected as the co-located picture. The picture RefPicList1 [0] is the first reference picture in the ordered reference picture list RefPicList1.
0117If the slice header parameter slice_type is not B, or if the slice header parameter collapse_from_l0_flag is not 0, picture RefPicList0 [0] is selected as the co-located picture. Picture RefPicList0 [0] is the first reference picture in the ordered reference picture list RefPicList0.
0118Next, the inter-prediction unit 511 selects the co-located block (S902). If the first block containing sample c0 shown in FIG. 4 is available, the first block is selected as the co-located block. If the first block is not available and the second block containing sample c1 shown in FIG. 4 is available, the second block is selected as the co-located block.
0119When an available co-located block is selected, the inter-prediction unit 511 sets the time-predicted motion vector as available. If no available co-located block is selected, the inter-prediction unit 511 sets the time-predicted motion vector as unavailable.
0120When the time-predicted motion vector is set as available, the inter-prediction unit 511 specifies the co-located motion vector as the reference motion vector. In addition, the inter-prediction unit 511 identifies the co-located reference picture (S903). When the co-located block has a plurality of motion vectors, that is, when the co-located block is encoded using the plurality of motion vectors, the inter-prediction unit 511 selects the reference motion vector according to a predetermined priority. To do.
0121For example, when the current reference picture is a short-term reference picture, the inter-prediction unit 511 may preferentially select a motion vector indicating a position in the short-term reference picture as a reference motion vector among a plurality of motion vectors. Good.
0122That is, when the motion vector indicating the position in the short-term reference picture exists, the inter-prediction unit 511 selects the motion vector as the reference motion vector. Then, when the motion vector pointing to the position in the short-term reference picture does not exist, the inter-prediction unit 511 selects the motion vector pointing to the position in the long-term reference picture as the reference motion vector.
0123After that, when either the current reference picture or the co-located reference picture is a long-term reference picture (Yes in S904), the inter-prediction unit 511 derives the reference motion vector as the time-predicted motion vector (S905).
0124On the other hand, if neither of the two reference pictures is a long-term reference picture (No in S904), the inter-prediction unit 511 derives a time-predicted motion vector from the reference motion vector by POC-based scaling (S906).
0125As mentioned above, the time prediction motion vector is set as available or unavailable. The inter-prediction unit 511 puts the time-predicted motion vector set as available into the ordered list of predicted motion vector candidates. The ordered list holds various motion vectors as predicted motion vector candidates, not limited to the time predicted motion vector.
0126The inter-prediction unit 511 selects one predicted motion vector from the ordered list and predicts the current motion vector using the selected predicted motion vector. At that time, the inter-prediction unit 511 selects the predicted motion vector closest to the current motion vector or the predicted motion vector capable of coding the current motion vector with the highest coding efficiency from the ordered list. The index corresponding to the selected predicted motion vector is written to the coded stream.
0127By the above processing, the time prediction motion vector is appropriately derived from the co-located motion vector without becoming extremely large or extremely small. Therefore, the prediction accuracy is improved and the coding efficiency is improved.
0128Whether each reference picture is a long-term reference picture or a short-term reference picture may be changed depending on the time. For example, the short-term reference picture may later be changed to a long-term reference picture. Conversely, the long-term reference picture may later be changed to a short-term reference picture.
0129Further, the inter-prediction unit 511 may determine whether the co-located reference picture is a long-term reference picture or a short-term reference picture during the period when the co-located block is coded. Then, the image coding device 500 may have an additional memory for holding the determination result from the coding of the co-located block to the coding of the current block.
0130In this case, it is more accurately determined whether the co-located reference picture is a long-term reference picture or a short-term reference picture.
0131Alternatively, the inter-prediction unit 511 may determine whether the co-located reference picture is a long-term reference picture or a short-term reference picture during the period when the current block is encoded.
0132In this case, the information on whether the co-located reference picture is a long-term reference picture or a short-term reference picture does not have to be maintained for a long period of time.
0133Further, the inter-prediction unit 511 may determine whether the current reference picture is a long-term reference picture or a short-term reference picture by using the time distance from the current reference picture to the current picture.
0134For example, when the time distance from the current reference picture to the current picture is larger than a predetermined threshold value, the inter-prediction unit 511 determines that the current reference picture is a long-term reference picture. Then, when the time distance is equal to or less than a predetermined threshold value, the inter-prediction unit 511 determines that the current reference picture is a short-term reference picture.
0135Similarly, the inter-prediction unit 511 determines whether the co-located reference picture is a long-term reference picture or a short-term reference picture by using the time distance from the co-located reference picture to the co-located reference picture. You may.
0136For example, when the time distance from the co-located reference picture to the co-located picture is larger than a predetermined threshold value, the inter-prediction unit 511 determines that the co-located reference picture is a long-term reference picture. Then, when the time distance is equal to or less than a predetermined threshold value, the inter-prediction unit 511 determines that the co-located reference picture is a short-term reference picture.
0137Then, the inter-prediction unit 608 of the image decoding device 600 determines whether each reference picture is a long-term reference picture or a short-term reference picture in terms of time distance, similarly to the inter-prediction unit 511 of the image coding device 500. It may be judged based on. In this case, the information indicating whether each reference picture is a long-term reference picture or a short-term reference picture does not have to be encoded.
0138Regarding the other processes shown in the present embodiment, each component of the image decoding device 600 is encoded with high coding efficiency by performing the same process as the corresponding component in the image coding device 500. The image is properly decoded.
0139Moreover, the operation shown above can be applied to other embodiments. The configurations and operations shown in this embodiment may be incorporated into other embodiments, and the configurations and operations shown in other embodiments may be incorporated into this embodiment.
0140(Embodiment 2) The configuration of the image coding device and the image decoding device according to the present embodiment is the same as that of the first embodiment. Therefore, these operations according to the present embodiment will be described with reference to the configuration of the image coding device 500 of FIG. 5 and the configuration of the image decoding device 600 of FIG.
0141Further, the image coding apparatus 500 according to the present embodiment performs the operation shown in FIG. 7 as in the first embodiment. Further, the image decoding apparatus 600 according to the present embodiment performs the operation shown in FIG. 8 as in the first embodiment. In the present embodiment, the derivation process of the predicted motion vector is different from that in the first embodiment. The details will be described below.
0142FIG. 10 is a flowchart showing details of the derivation process according to the present embodiment. The inter-prediction unit 511 according to the present embodiment performs the operation shown in FIG. 10 instead of the operation shown in FIG. The following mainly shows the operation of the inter-prediction unit 511 in FIG. If the coding is read as decoding, the operation of the inter-prediction unit 608 in FIG. 6 is the same as the operation of the inter-prediction unit 511 in FIG.
0143First, the inter-prediction unit 511 selects a co-located picture from a plurality of available reference pictures (S1001). Next, the inter-prediction unit 511 selects the co-located block in the co-located picture (S1002). Then, the inter-prediction unit 511 identifies the co-located reference picture and the co-located motion vector (S1003).
0144Next, the inter-prediction unit 511 determines whether or not the current reference picture is a long-term reference picture (S1004). Then, when it is determined that the current reference picture is a long-term reference picture (Yes in S1004), the inter-prediction unit 511 derives the predicted motion vector according to the first derivation method similar to the first embodiment (S1005). ).
0145When it is determined that the current reference picture is not a long-term reference picture (No in S1004), the inter-prediction unit 511 determines whether or not the co-located reference picture is a long-term reference picture (S1006).
0146Then, when it is determined that the co-located reference picture is not a long-term reference picture (No in S1006), the inter-prediction unit 511 derives a predicted motion vector according to the second derivation method similar to the first embodiment (No). S1007). That is, when it is determined that both the current reference picture and the co-located reference picture are short-term reference pictures, the inter-prediction unit 511 derives the predicted motion vector according to the second derivation method.
0147If the co-located reference picture is determined to be a long-term reference picture (Yes in S1006), the inter-prediction unit 511 selects another co-located block in the co-located picture (S1008). In the example of FIG. 10, the block encoded with reference to the short term reference picture is selected as another co-located block.
0148After that, the inter-prediction unit 511 identifies the co-located reference picture and the co-located motion vector corresponding to another co-located block (S1009). Next, the inter-prediction unit 511 derives the predicted motion vector according to the second derivation method using POC-based scaling (S1010).
0149That is, when the reference picture of the current block is a short-term reference picture and the reference picture of the co-located block is a long-term reference picture, the inter-prediction unit 511 calculates the predicted motion vector from the motion vector of the co-located block. Do not derive. In this case, the inter-prediction unit 511 selects another encoded co-located block with reference to the short-term reference picture, and derives a predicted motion vector from the motion vector of the other selected co-located block. ..
0150For example, when the reference picture of the current block is a short-term reference picture and the reference picture of the co-located block is a long-term reference picture, the inter-prediction unit 511 refers to the short-term reference picture and encodes the block. To explore. Then, the inter-prediction unit 511 selects the encoded block as another co-located block by referring to the short-term reference picture.
0151As another example, when the reference picture of the current block is a short-term reference picture and the reference picture of the co-located block is a long-term reference picture, the inter-prediction unit 511 first refers to the short-term reference picture. Search for the encoded block.
0152Then, if there is a block encoded by referring to the short-term reference picture, the inter-prediction unit 511 selects the block as another co-located block. If there is no coded block with reference to the short-term reference picture, then the inter-prediction unit 511 searches for the coded block with reference to the long-term reference picture. Then, the encoded block is selected as another co-located block by referring to the long-term reference picture.
0153Further, for example, first, the inter-prediction unit 511 selects the first block in FIG. 4 as the co-located block. Then, when the current reference picture is a short-term reference picture and the co-located reference picture is a long-term reference picture, the inter-prediction unit 511 newly sets the second block in FIG. 4 as a co-located block. select.
0154In the above example, the inter-prediction unit 511 may select the second block as the co-located block only when the reference picture in the second block of FIG. 4 is a short-term reference picture. Further, the block selected as the co-located block is not limited to the second block in FIG. 4, and other blocks may be selected as the co-located block.
0155FIG. 11 is a diagram for explaining a co-located block according to the present embodiment. FIG. 11 shows samples c0, c1, c2 and c3 in a co-located picture. Samples c0 and c1 in FIG. 11 are equivalent to samples c0 and c1 in FIG. Not only the second block containing the sample c1, but also the third block containing the sample c2 or the fourth block containing the sample c3 may be selected as another co-located block.
0156The coordinates of sample c2 are (x + w-1, y + h-1). The coordinates of sample c3 are (x + 1, y + 1).
0157The inter-prediction unit 511 determines whether or not they are available in the order of the first block, the second block, the third block, and the fourth block. Then, the inter-prediction unit 511 determines the available block as the final co-located block. Examples of unusable blocks include cases where the block does not exist, or where the block is encoded by intra-prediction.
0158When the current reference picture is a short-term reference picture, the inter-prediction unit 511 may determine that the coded block by referring to the long-term reference picture is not available.
0159In the above, an example of the selection method of the co-located block is described, but the selection method of the co-located block is not limited to the example described above. Blocks containing samples other than samples c0, c1, c2 or c3 may be selected as co-located blocks. Moreover, those priorities are not limited to the examples shown in the present embodiment.
0160Hereinafter, a more specific example of the process of deriving the time prediction motion vector will be described with reference to FIG. 10 again. The derivation process described above may be modified as follows.
0161First, the inter-prediction unit 511 selects the co-located picture as in the first embodiment (S1001). Then, the inter-prediction unit 511 selects the first block including the sample c0 shown in FIG. 11 as the co-located block, and identifies the co-located reference picture (S1002 and S1003).
0162Next, the inter-prediction unit 511 determines whether or not the co-located block is available. At that time, if the current reference picture is a short-term reference picture and the co-located reference picture is a long-term reference picture, the inter-prediction unit 511 determines that the co-located block is not available (S1004 and). S1006).
0163If a co-located block is not available, the inter-prediction unit 511 searches for and selects another co-located block available (S1008). Specifically, the inter-prediction unit 511 refers to the short-term reference picture from the second block including the sample c1 in FIG. 11, the third block including the sample c2, and the fourth block including the sample c3. Select the coded block. Then, the inter-prediction unit 511 identifies the reference picture of the co-located block (S1009).
0164When an available co-located block is selected, the inter-prediction unit 511 sets the time-predicted motion vector as available. If no available co-located block is selected, the inter-prediction unit 511 sets the time-predicted motion vector as unavailable.
0165When the time-predicted motion vector is set as available, the inter-predictor 511 identifies the co-located motion vector as the reference motion vector (S1003 and S1009). When the co-located block has a plurality of motion vectors, that is, when the co-located block is encoded using the plurality of motion vectors, the inter-prediction unit 511 performs a predetermined priority as in the first embodiment. Select the reference motion vector according to.
0166Then, when either the current reference picture or the co-located reference picture is a long-term reference picture (Yes in S1004), the inter-prediction unit 511 derives the reference motion vector as the time-predicted motion vector (S1005).
0167On the other hand, when neither the current reference picture nor the co-located reference picture is a long-term reference picture (No in S1004), the inter-prediction unit 511 derives a time prediction motion vector from the reference motion vector by POC-based scaling. (S1007 and S1010).
0168When the time prediction motion vector is set as unavailable, the inter-prediction unit 511 does not derive the time prediction motion vector.
0169As described above, in the present embodiment, when the reference picture of the current block is a short-term reference picture and the reference picture of the co-located block is a long-term reference picture, the motion vector of the co-located block is used. The time-predicted motion vector is not derived.
0170When one of the current reference picture and the co-located reference picture is a long-term reference picture and the other is a short-term reference picture, it is very difficult to derive a time prediction motion vector with high prediction accuracy. .. Therefore, the image coding device 500 and the image decoding device 600 according to the present embodiment suppress the deterioration of the prediction accuracy by the above operation.
0171(Embodiment 3) The configuration of the image coding device and the image decoding device according to the present embodiment is the same as that of the first embodiment. Therefore, these operations according to the present embodiment will be described with reference to the configuration of the image coding device 500 of FIG. 5 and the configuration of the image decoding device 600 of FIG.
0172Further, the image coding apparatus 500 according to the present embodiment performs the operation shown in FIG. 7 as in the first embodiment. Further, the image decoding apparatus 600 according to the present embodiment performs the operation shown in FIG. 8 as in the first embodiment. In the present embodiment, the derivation process of the predicted motion vector is different from that in the first embodiment. The details will be described below.
0173FIG. 12 is a flowchart showing details of the derivation process according to the present embodiment. The inter-prediction unit 511 according to the present embodiment performs the operation shown in FIG. 12 instead of the operation shown in FIG. The following mainly shows the operation of the inter-prediction unit 511 in FIG. If the coding is read as decoding, the operation of the inter-prediction unit 608 in FIG. 6 is the same as the operation of the inter-prediction unit 511 in FIG.
0174First, the inter-prediction unit 511 selects a co-located picture from a plurality of available reference pictures (S1201). Next, the inter-prediction unit 511 selects the co-located block in the co-located picture (S1202). Then, the inter-prediction unit 511 identifies the co-located reference picture and the co-located motion vector (S1203).
0175Next, the inter-prediction unit 511 determines whether or not the current reference picture is a long-term reference picture (S1204). Then, when it is determined that the current reference picture is a long-term reference picture (Yes in S1204), the inter-prediction unit 511 derives the predicted motion vector according to the first derivation method similar to the first embodiment (S1205). ).
0176When it is determined that the current reference picture is not a long-term reference picture (No in S1204), the inter-prediction unit 511 determines whether or not the co-located reference picture is a long-term reference picture (S1206).
0177Then, when it is determined that the co-located reference picture is not a long-term reference picture (No in S1206), the inter-prediction unit 511 derives a predicted motion vector according to the second derivation method similar to the first embodiment (No). S1207). That is, when it is determined that both the current reference picture and the co-located reference picture are short-term reference pictures, the inter-prediction unit 511 derives the predicted motion vector according to the second derivation method.
0178If the co-located reference picture is determined to be a long-term reference picture (Yes in S1206), the inter-prediction unit 511 selects another co-located picture (S1208). Then, the inter-prediction unit 511 selects another co-located block in another co-located picture (S1209). In the example of FIG. 12, the block encoded with reference to the short term reference picture is selected as another co-located block.
0179After that, the inter-prediction unit 511 identifies the co-located reference picture and the co-located motion vector corresponding to another co-located block (S1210). Next, the inter-prediction unit 511 derives the predicted motion vector according to the second derivation method using POC-based scaling (S1211).
0180That is, when the reference picture of the current block is a short-term reference picture and the reference picture of the co-located block is a long-term reference picture, the inter-prediction unit 511 does not derive the predicted motion vector from the co-located block.
0181In this case, the inter-prediction unit 511 selects another co-located picture. Then, the inter-prediction unit 511 selects another co-located block encoded by referring to the short-term reference picture from the other selected co-located picture. The inter-prediction unit 511 derives a predicted motion vector from the motion vector of another selected co-located block.
0182For example, when the current reference picture is a short-term reference picture and the co-located reference picture is a long-term reference picture, the inter-prediction unit 511 selects a picture including a block encoded by referring to the short-term reference picture. Explore. Then, the inter-prediction unit 511 refers to the short-term reference picture and selects a picture including the encoded block as another co-located picture.
0183As another example, if the current reference picture is a short-term reference picture and the co-located reference picture is a long-term reference picture, the interpredictor 511 is first encoded with reference to the short-term reference picture. Search for pictures that contain blocks.
0184Then, if there is a picture including a block encoded by referring to the short-term reference picture, the inter-prediction unit 511 selects the picture as another co-located picture.
0185If there is no picture containing the coded block with reference to the short term reference picture, then the inter-prediction unit 511 searches for the picture containing the coded block with reference to the long term reference picture. Then, the inter-prediction unit 511 refers to the long-term reference picture and selects a picture including the encoded block as another co-located picture.
0186Further, for example, when the picture RefPicList0 [0] is a co-located picture, the picture RefPicList1 [0] is another co-located picture. If picture RefPicList1 [0] is a co-located picture, then picture RefPicList0 [0] is another co-located picture.
0187That is, of the two reference picture lists used in B-picture coding (bi-predictive coding), the first picture in one reference picture list is the co-located picture and one in the other reference picture list. The second picture is another co-located picture.
0188Hereinafter, a more specific example of the process of deriving the time prediction motion vector will be described with reference to FIG. 12 again. The derivation process described above may be modified as follows.
0189First, the inter-prediction unit 511 selects one of the picture RefPicList0 [0] and the picture RefPicList1 [0] as the co-located picture (S1201). Then, the inter-prediction unit 511 selects the first block including the sample c0 shown in FIG. 11 as the co-located block from the selected co-located pictures, and identifies the co-located reference picture (S1202 and). S1203).
0190Next, the inter-prediction unit 511 determines whether or not the co-located block is available. At that time, if the current reference picture is a short-term reference picture and the co-located reference picture is a long-term reference picture, the inter-prediction unit 511 determines that the co-located block is not available (S1204 and S1206).
0191If the co-located block is not available, the inter-prediction unit 511 newly selects an available co-located block. For example, the inter-prediction unit 511 selects the second block including the sample c1 in FIG. 11 as the co-located block. Then, the inter-prediction unit 511 identifies the co-located reference picture.
0192If no available co-located block is selected, the inter-prediction unit 511 selects another co-located picture. At that time, the inter-prediction unit 511 selects the other of the picture RefPicList0 [0] and the picture RefPicList1 [0] as the co-located picture (S1208).
0193Then, the inter-prediction unit 511 selects the first block including the sample c0 shown in FIG. 11 as the co-located block from the selected co-located pictures, and identifies the co-located reference picture (S1209 and). S1210).
0194Next, the inter-prediction unit 511 determines whether or not the co-located block is available. At that time, as in the previous determination, if the current reference picture is a short-term reference picture and the co-located reference picture is a long-term reference picture, the inter-prediction unit 511 can use the co-located block. Judge that it is not.
0195If the co-located block is not available, the inter-prediction unit 511 newly selects an available co-located block (S1209). Specifically, the inter-prediction unit 511 selects the second block including the sample c1 in FIG. 11 as the co-located block. Then, the inter-prediction unit 511 identifies the co-located reference picture (S1210).
0196Finally, when an available co-located block is selected, the inter-prediction unit 511 sets the time-predicted motion vector as available. If no available co-located block is selected, the inter-prediction unit 511 sets the time-predicted motion vector as unavailable.
0197When the time-predicted motion vector is set as available, the inter-predictor 511 identifies the motion vector of the co-located block as the reference motion vector (S1203 and S1210). When the co-located block has a plurality of motion vectors, that is, when the co-located block is encoded using the plurality of motion vectors, the inter-prediction unit 511 performs a predetermined priority as in the first embodiment. Select the reference motion vector according to.
0198Then, when either the current reference picture or the co-located reference picture is a long-term reference picture (Yes in S1204), the inter-prediction unit 511 derives the reference motion vector as the time-predicted motion vector (S1205).
0199On the other hand, when neither the current reference picture nor the co-located reference picture is a long-term reference picture (No in S1204), the inter-prediction unit 511 derives a time prediction motion vector from the reference motion vector by POC-based scaling. (S1207 and S1211).
0200When the time prediction motion vector is set as unavailable, the inter-prediction unit 511 does not derive the time prediction motion vector.
0201As described above, the image coding device 500 and the image decoding device 600 according to the present embodiment select a block suitable for deriving the time prediction motion vector from a plurality of pictures, and from the motion vector of the selected block, Derivation of time prediction motion vector. This improves the coding efficiency.
0202(Embodiment 4) This embodiment confirmably shows the characteristic configuration and the characteristic procedure included in the first to third embodiments.
0203FIG. 13A is a block diagram of the image coding apparatus according to the present embodiment. The image coding device 1300 shown in FIG. 13A encodes each of a plurality of blocks in a plurality of pictures. Further, the image coding device 1300 includes a lead-out unit 1301, an additional unit 1302, a selection unit 1303, and a coding unit 1304.
0204For example, the derivation unit 1301, the addition unit 1302, and the selection unit 1303 correspond to the inter-prediction unit 511 and the like in FIG. The coding unit 1304 corresponds to the entropy coding unit 504 and the like in FIG.
0205FIG. 13B is a flowchart showing the operation of the image coding device 1300 shown in FIG. 13A.
0206The derivation unit 1301 derives a candidate for the predicted motion vector from the motion vector of the co-located block (S1301). The co-located block is a block included in a picture different from the picture containing the current block to be encoded. The predicted motion vector is used to code the motion vector of the current block.
0207In deriving the candidate, the derivation unit 1301 determines whether the reference picture of the current block is a long-term reference picture or a short-term reference picture. Further, the derivation unit 1301 determines whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture.
0208Here, when it is determined that the reference picture of the current block and the reference picture of the co-located block are long-term reference pictures, the derivation unit 1301 selects a candidate from the motion vector of the co-located block by the first derivation method. Derived. The first derivation method is a derivation method that does not perform scaling based on time distance.
0209On the other hand, when it is determined that the reference picture of the current block and the reference picture of the co-located block are short-term reference pictures, the derivation unit 1301 derives a candidate from the motion vector of the co-located block by the second derivation method. To do. The second derivation method is a derivation method that performs scaling based on time distance.
0210Addition 1302 adds the derived candidates to the list (S1302). The selection unit 1303 selects a predicted motion vector from the list to which candidates are added (S1303).
0211The coding unit 1304 encodes the current block using the motion vector of the current block and the reference picture of the current block. Further, the coding unit 1304 encodes the motion vector of the current block using the selected predicted motion vector (S1304).
0212FIG. 14A is a block diagram of the image decoding device according to the present embodiment. The image decoding device 1400 shown in FIG. 14A decodes each of a plurality of blocks in a plurality of pictures. Further, the image decoding device 1400 includes a derivation unit 1401, an additional unit 1402, a selection unit 1403, and a decoding unit 1404.
0213For example, the derivation unit 1401, the addition unit 1402, and the selection unit 1403 correspond to the inter-prediction unit 608 and the like in FIG. The decoding unit 1404 corresponds to the entropy decoding unit 601 and the like in FIG.
0214FIG. 14B is a flowchart showing the operation of the image decoding apparatus 1400 shown in FIG. 14A.
0215The derivation unit 1401 derives a candidate for the predicted motion vector from the motion vector of the co-located block (S1401). The co-located block is a block included in a picture different from the picture containing the current block to be decoded. The predicted motion vector is used to decode the motion vector of the current block.
0216In deriving the candidate, the derivation unit 1401 determines whether the reference picture of the current block is a long-term reference picture or a short-term reference picture. Further, the derivation unit 1401 determines whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture.
0217Here, when it is determined that the reference picture of the current block and the reference picture of the co-located block are long-term reference pictures, the derivation unit 1401 selects a candidate from the motion vector of the co-located block by the first derivation method. Derived. The first derivation method is a derivation method that does not perform scaling based on time distance.
0218On the other hand, when it is determined that the reference picture of the current block and the reference picture of the co-located block are short-term reference pictures, the derivation unit 1401 derives a candidate from the motion vector of the co-located block by the second derivation method. To do. The second derivation method is a derivation method that performs scaling based on time distance.
0219Addition 1402 adds the derived candidates to the list (S1402). The selection unit 1403 selects a predicted motion vector from the list to which candidates are added (S1403).
0220The decoding unit 1404 decodes the motion vector of the current block using the selected predicted motion vector. Further, the decoding unit 1404 decodes the current block using the motion vector of the current block and the reference picture of the current block (S1404).
0221By the above processing, the candidate of the predicted motion vector is appropriately derived from the motion vector of the co-located block without becoming extremely large or extremely small. Therefore, the prediction accuracy can be improved, and the coding efficiency can be improved.
0222When the derivation units 1301 and 1401 determine that one of the reference picture of the current block and the reference picture of the co-located block is a long-term reference picture and the other is a short-term reference picture. , It is not necessary to derive the candidate from the motion vector of the co-located block.
0223In this case, the derivation units 1301 and 1401 further select another co-located block encoded or decoded by referring to the short-term reference picture, and select a candidate from the other co-located block by the second derivation method. It may be derived. Alternatively, in this case, the derivation units 1301 and 1401 may derive candidates by another derivation method. Alternatively, in this case, the derivation units 1301 and 1401 do not have to finally derive the candidates corresponding to the time prediction motion vector.
0224Further, the derivation units 1301 and 1401 use the time distance from the reference picture of the current block to the picture including the current block to determine whether the reference picture of the current block is a long-term reference picture or a short-term reference picture. You may judge.
0225In addition, the derivation units 1301 and 1401 use the time distance from the reference picture of the co-located block to the picture including the co-located block, and the reference picture of the co-located block is a long-term reference picture or a short-term. It may be determined whether it is a reference picture.
0226Further, the derivation units 1301 and 1401 determine whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture during the period when the co-located block is encoded or decoded. May be good.
0227Further, the derivation units 1301 and 1401 may determine whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture during the period when the current block is encoded or decoded. ..
0228Further, the first derivation method may be a method of deriving the motion vector of the co-located block as a candidate. The second derivation method uses the ratio of the time distance from the reference picture of the current block to the picture containing the current block to the time distance from the reference picture of the co-located block to the picture containing the co-located block. A method of deriving candidates by scaling the motion vector of the co-located block may also be used.
0229Further, the encoding unit 1304 further provides information indicating whether the reference picture of the current block is a long-term reference picture or a short-term reference picture, and the reference picture of the co-located block is a long-term reference picture. Information indicating whether it is a short-term reference picture or a short-term reference picture may be encoded.
0230Further, the decoding unit 1404 further indicates information indicating whether the reference picture of the current block is a long-term reference picture or a short-term reference picture, and whether the reference picture of the co-located block is a long-term reference picture. Information indicating whether the picture is a short-term reference picture may be decoded.
0231Then, the derivation unit 1401 may determine whether the reference picture of the current block is a long-term reference picture or a short-term reference picture by using the decoded information. Further, the derivation unit 1401 may determine whether the reference picture of the co-located block is a long-term reference picture or a short-term reference picture by using the decoded information.
0232Further, the information indicating the classification of the reference picture may be stored as a parameter at the following position in the coded stream.
0233FIG. 15A is a diagram showing a first example of the storage position of the parameter indicating the classification of the reference picture. As shown in FIG. 15A, the parameter indicating the classification of the reference picture may be stored in the sequence header. Sequence headers are also called sequence parameter sets.
0234FIG. 15B is a diagram showing a second example of the storage position of the parameter indicating the classification of the reference picture. As shown in FIG. 15B, the parameter indicating the classification of the reference picture may be stored in the picture header. Picture headers are also called picture parameter sets.
0235FIG. 15C is a diagram showing a third example of the storage position of the parameter indicating the classification of the reference picture. As shown in FIG. 15C, the parameter indicating the classification of the reference picture may be stored in the slice header.
0236Further, the information indicating the prediction mode (inter prediction or intra prediction) may be stored as a parameter at the following position in the coded stream.
0237FIG. 16 is a diagram showing an example of a storage position of a parameter indicating a prediction mode. As shown in FIG. 16, this parameter may be stored in the CU header (encoding unit header). This parameter indicates whether the prediction unit within the coding unit was encoded with inter-prediction or intra-prediction. This parameter may be used to determine if a co-located block is available.
0238Further, in each of the above-described embodiments, each component may be configured by dedicated hardware or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory. Here, the software that realizes the image coding apparatus and the like of each of the above embodiments is the following program.
0239That is, this program is an image coding method for encoding each of a plurality of blocks in a plurality of pictures on a computer, and is a block included in a picture different from the picture including the current block to be encoded. A derivation step for deriving a candidate for a predicted motion vector used for encoding the motion vector of the current block from the motion vector of the -located block, an additional step for adding the derived candidate to the list, and the addition of the candidate. The current block is encoded using the motion vector of the current block and the reference picture of the current block, and the predicted motion vector selected is used in the selection step of selecting the predicted motion vector from the list. In the derivation step, which includes a coding step for encoding the motion vector of the current block, whether the reference picture of the current block is a long-term reference picture or a short-term reference picture, and the co-located. It was determined whether the reference picture of the block was a long-term reference picture or a short-term reference picture, and it was determined that the reference picture of the current block and the reference picture of the co-located block were long-term reference pictures, respectively. In this case, the candidate is derived from the motion vector of the co-located block by the first derivation method that does not perform scaling based on the time distance, and the reference picture of the current block and the reference picture of the co-located block are respectively. When it is determined that the picture is a short-term reference picture, the image coding method for deriving the candidate is executed by the second derivation method that performs scaling based on the time distance from the motion vector of the co-located block.
0240In addition, this program is an image decoding method for decoding each of a plurality of blocks in a plurality of pictures on a computer, and is a co-located block which is a block included in a picture different from the picture including the current block to be decoded. A derivation step for deriving a candidate for a predicted motion vector used for decoding the motion vector of the current block from the motion vector of the above, an additional step for adding the derived candidate to the list, and the list to which the candidate is added. From the selection step of selecting the predicted motion vector, the motion vector of the current block is decoded using the selected predicted motion vector, and the motion vector of the current block and the reference picture of the current block are used. In the derivation step, which includes a decoding step of decoding the current block, whether the reference picture of the current block is a long-term reference picture or a short-term reference picture, and the reference picture of the co-located block is long. When it is determined whether the reference picture is a term reference picture or a short term reference picture, and the reference picture of the current block and the reference picture of the co-located block are determined to be long term reference pictures, the co-located The candidate is derived from the motion vector of the block by the first derivation method that does not perform scaling based on the time distance, and the reference picture of the current block and the reference picture of the co-located block are short-term reference pictures, respectively. If it is determined, the image decoding method for deriving the candidate may be executed by the second derivation method that performs scaling based on the time distance from the motion vector of the co-located block.
0241Further, each component may be a circuit. These circuits may form one circuit as a whole, or may be separate circuits. Further, each component may be realized by a general-purpose processor or a dedicated processor.
0242Although the image coding apparatus according to one or more aspects has been described above based on the embodiment, the present invention is not limited to this embodiment. As long as the gist of the present invention is not deviated, various modifications that can be conceived by those skilled in the art are applied to the present embodiment, and a form constructed by combining components in different embodiments is also within the scope of one or more embodiments. May be included within.
0243For example, the image coding / decoding device may include an image coding device and an image decoding device. Further, another processing unit may execute the processing executed by the specific processing unit. Further, the order in which the processes are executed may be changed, or a plurality of processes may be executed in parallel.
0244(Embodiment 5) Each of the above embodiments is performed by recording a program for realizing the configuration of the moving image coding method (image coding method) or the moving image decoding method (image decoding method) shown in each of the above embodiments on a storage medium. It becomes possible to easily carry out the processing shown in the above embodiment in an independent computer system. The storage medium may be a magnetic disk, an optical disk, a magneto-optical disk, an IC card, a semiconductor memory, or the like, as long as it can record a program.
0245Further, an application example of the moving image coding method (image coding method) and the moving image decoding method (image decoding method) shown in each of the above embodiments and a system using the same will be described here. The system is characterized by having an image coding / decoding device including an image coding device using an image coding method and an image decoding device using an image decoding method. Other configurations in the system can be changed appropriately as the case may be.
0246FIG. 17 is a diagram showing the overall configuration of the content supply system ex100 that realizes the content distribution service. The communication service provision area is divided into desired sizes, and fixed radio stations ex106, ex107, ex108, ex109, and ex110 are installed in each cell.
0247This content supply system ex100 is connected to the Internet ex101 via the Internet service provider ex102 and the telephone network ex104, and the base station ex106 to ex110 via the computer ex111, PDA (Personal Digital Assistant) ex112, camera ex113, mobile phone ex114, and game machine ex115. Each device such as is connected.
0248However, the content supply system ex100 is not limited to the configuration shown in FIG. 17, and any element may be combined and connected. Further, each device may be directly connected to the telephone network ex104 without going through the base stations ex106 to ex110, which are fixed radio stations. Further, the devices may be directly connected to each other via short-range radio or the like.
0249The camera ex113 is a device capable of shooting moving images such as a digital video camera, and the camera ex116 is a device capable of shooting still images and moving images such as a digital camera. In addition, the mobile phone ex114 is a GSM (registered trademark) (Global System for Mobile Communications) system, a CDMA (Code Division Multiple Access) system, a W-CDMA (Wideband-Code Division Multiple Access) system, or LTE (Long Term Evolution). The method, HSPA (High Speed Packet Access) mobile phone, PHS (Personal Handyphone System), etc., may be used.
0250In the content supply system ex100, live distribution and the like are possible by connecting the camera ex113 and the like to the streaming server ex103 through the base station ex109 and the telephone network ex104. In live distribution, content (for example, live music video, etc.) taken by a user using the camera ex113 is encoded as described in each of the above embodiments (that is, according to one aspect of the present invention). (Functions as such an image encoding device), and transmits to the streaming server ex103. On the other hand, the streaming server ex103 streams the content data transmitted to the requested client. Clients include a computer ex111, a PDAex112, a camera ex113, a mobile phone ex114, a game machine ex115, and the like, which can decode the coded data. Each device that has received the distributed data decodes and reproduces the received data (that is, functions as an image decoding device according to one aspect of the present invention).
0251The captured data may be encoded by the camera ex113, the streaming server ex103 that performs the data transmission processing, or may be shared with each other. Similarly, the decryption process of the delivered data may be performed by the client, the streaming server ex103, or shared with each other. Further, not only the camera ex113 but also the still image and / or the moving image data taken by the camera ex116 may be transmitted to the streaming server ex103 via the computer ex111. The coding process in this case may be performed by any of the camera ex116, the computer ex111, and the streaming server ex103, or may be shared with each other.
0252Further, these coding / decoding processes are generally performed by the computer ex111 or the LSI ex500 of each device. The LSIex500 may be a single chip or a configuration composed of a plurality of chips. In addition, software for video coding / decoding is embedded in some recording medium (CD-ROM, flexible disk, hard disk, etc.) that can be read by a computer ex111 or the like, and the coding / decoding processing is performed using the software. You may. Further, when the mobile phone ex114 is equipped with a camera, the moving image data acquired by the camera may be transmitted. The moving image data at this time is the data encoded by the LSI ex500 of the mobile phone ex114.
0253Further, the streaming server ex103 may be a plurality of servers or a plurality of computers, and may disperse data for processing, recording, and distribution.
0254As described above, in the content supply system ex100, the client can receive and reproduce the encoded data. In this way, in the content supply system ex100, the client can receive, decode, and reproduce the information transmitted by the user in real time, and even a user who does not have special rights or equipment can realize personal broadcasting.
0255Not limited to the example of the content supply system ex100, as shown in FIG. 18, the digital broadcasting system ex200 also includes at least a moving image coding device (image coding device) or moving image decoding according to each of the above embodiments. Any of the devices (image decoding devices) can be incorporated. Specifically, in the broadcasting station ex201, the multiplexed data in which music data or the like is multiplexed with the video data is transmitted to the satellite ex202 or communication via radio waves. This video data is data encoded by the moving image coding method described in each of the above embodiments (that is, data encoded by the image coding apparatus according to one aspect of the present invention). In response to this, the broadcasting satellite ex202 transmits radio waves for broadcasting, and the radio waves are received by the home antenna ex204 capable of receiving satellite broadcasting. A device such as a television (receiver) ex300 or a set-top box (STB) ex217 decodes and reproduces the received multiplexed data (that is, functions as an image decoding device according to one aspect of the present invention).
0256In addition, the reader / recorder ex218 also reads and decodes the multiplexed data recorded on the recording medium ex215 such as DVD and BD, or encodes the video signal on the recording medium ex215 and, in some cases, multiplexes and writes the music signal. It is possible to implement the moving image decoding device or the moving image coding device shown in each of the above embodiments. In this case, the reproduced video signal is displayed on the monitor ex219, and the video signal can be reproduced in another device or system by the recording medium ex215 in which the multiplexed data is recorded. Further, a moving image decoding device may be mounted in a set-top box ex217 connected to a cable ex203 for cable TV or an antenna ex204 for satellite / terrestrial broadcasting, and this may be displayed on a TV monitor ex219. At this time, the moving image decoding device may be incorporated in the television instead of the set-top box.
0257FIG. 19 is a diagram showing a television (receiver) ex300 using the moving image decoding method and the moving image coding method described in each of the above embodiments. The TV ex300 acquires or outputs multiplexed data in which audio data is multiplexed on video data via an antenna ex204 or a cable ex203 that receives the above broadcast, and a tuner ex301 that outputs the received multiplexed data. Alternatively, the modulation / demodulation unit ex302 that modulates the multiplexed data to be transmitted to the outside, and the video data and audio data that separate the demodulated multiplexed data into video data and audio data, or are encoded by the signal processing unit ex306. It is provided with a multiplexing / separation unit ex303 that multiplexes the data.
0258Further, the television ex300 includes an audio signal processing unit ex304 and a video signal processing unit ex305 (an image coding device or an image according to one aspect of the present invention) that decodes each of the audio data and the video data or encodes the respective information. It has a signal processing unit ex306 having a function as a decoding device), a speaker ex307 for outputting a decoded audio signal, and an output unit ex309 having a display unit ex308 such as a display for displaying the decoded video signal. Further, the television ex300 has an interface unit ex317 having an operation input unit ex312 or the like for receiving input of user operation. Further, the television ex300 has a control unit ex310 that controls each unit in an integrated manner, and a power supply circuit unit ex311 that supplies electric power to each unit. In addition to the operation input unit ex312, the interface unit ex317 has a bridge ex313 connected to an external device such as a reader / recorder ex218, a slot unit ex314 for mounting a recording medium ex216 such as an SD card, and an external recording such as a hard disk. It may have a driver ex315 for connecting to media, a modem ex316 for connecting to a telephone network, and the like. The recording medium ex216 is capable of electrically recording information by a non-volatile / volatile semiconductor memory element that is stored. Each part of the TV ex300 is connected to each other via a synchronization bus.
0259First, a configuration in which the television ex300 decodes and reproduces the multiplexed data acquired from the outside by the antenna ex204 or the like will be described. The television ex300 receives a user operation from the remote controller ex220 or the like, and separates the multiplexed data demodulated by the modulation / demodulation unit ex302 by the multiplexing / separation unit ex303 based on the control of the control unit ex310 having a CPU or the like. Further, the television ex300 decodes the separated audio data by the audio signal processing unit ex304, and decodes the separated video data by the video signal processing unit ex305 using the decoding method described in each of the above embodiments. The decoded audio signal and video signal are output to the outside from the output unit ex309, respectively. When outputting, it is advisable to temporarily store these signals in buffers ex318, ex319, etc. so that the audio signal and the video signal are reproduced in synchronization. Further, the television ex300 may read the multiplexed data from the recording media ex215 and ex216 such as a magnetic / optical disk and an SD card, not from broadcasting or the like. Next, a configuration in which the television ex300 encodes an audio signal or a video signal and transmits it to the outside or writes it to a recording medium or the like will be described. The TV ex300 receives a user operation from a remote controller ex220 or the like, encodes an audio signal by the audio signal processing unit ex304 based on the control of the control unit ex310, and outputs a video signal by the video signal processing unit ex305 according to each of the above embodiments. It is encoded using the coding method described in. The encoded audio signal and video signal are multiplexed by the multiplexing / separating unit ex303 and output to the outside. When multiplexing, it is advisable to temporarily store these signals in buffers ex320, ex321, etc. so that the audio signal and the video signal are synchronized. It should be noted that a plurality of buffers ex318, ex319, ex320, and ex321 may be provided as shown in the figure, or one or more buffers may be shared. Furthermore, in addition to the figures shown, system overflow and underflow can be detected even between the modulation / demodulation section ex302 and the multiplexing / separation section ex303, for example.
0260In addition to acquiring audio data and video data from broadcasting and recording media, the TV ex300 has a configuration that accepts AV input from a microphone or camera, and performs encoding processing on the data acquired from them. May be good. Although the TV ex300 has been described here as a configuration capable of the above-mentioned coding processing, multiplexing, and external output, these processing cannot be performed, and only the above-mentioned reception, decoding processing, and external output are possible. It may be a configuration.
0261When reading or writing the multiplexed data from the recording medium with the reader / recorder ex218, the decoding process or the coding process may be performed with either the TV ex300 or the reader / recorder ex218, or the TV ex300. The reader / recorder ex218 may share the work with each other.
0262As an example, FIG. 20 shows the configuration of the information reproduction / recording unit ex400 when reading or writing data from an optical disc. The information reproduction / recording unit ex400 includes the elements ex401, ex402, ex403, ex404, ex405, ex406, and ex407 described below. The optical head ex401 irradiates the recording surface of the recording medium ex215, which is an optical disk, with a laser spot to write information, detects the reflected light from the recording surface of the recording medium ex215, and reads the information. The modulation recording unit ex402 electrically drives the semiconductor laser built in the optical head ex401 and modulates the laser beam according to the recorded data. The reproduction / demodulation unit ex403 amplifies the reproduction signal by electrically detecting the reflected light from the recording surface by the photodetector built in the optical head ex401, separates and demodulates the signal component recorded on the recording medium ex215, and is necessary. Information is played back. The buffer ex404 temporarily holds the information for recording on the recording medium ex215 and the information reproduced from the recording medium ex215. The disk motor ex405 rotates the recording medium ex215. The servo control unit ex406 moves the optical head ex401 to a predetermined information track while controlling the rotational drive of the disc motor ex405, and performs laser spot tracking processing. The system control unit ex407 controls the entire information reproduction / recording unit ex400. For the above read / write processing, the system control unit ex407 uses various information held in the buffer ex404, and new information is generated / added as necessary, and the modulation recording unit ex402 and the reproduction / demodulation unit are used. This is achieved by recording and reproducing information through the optical head ex401 while coordinating the ex403 and the servo control unit ex406. The system control unit ex407 is composed of, for example, a microprocessor, and executes those processes by executing a read / write program.
0263In the above, the optical head ex401 has been described as irradiating the laser spot, but it may be configured to perform higher density recording using near-field light.
0264FIG. 21 shows a schematic diagram of the recording medium ex215, which is an optical disc. A guide groove (groove) is spirally formed on the recording surface of the recording medium ex215, and address information indicating an absolute position on the disk is recorded in advance on the information track ex230 by changing the shape of the groove. This address information includes information for specifying the position of the recording block ex231, which is a unit for recording data, and the recording block is specified by playing back the information track ex230 and reading the address information in a device that performs recording or playback. Can be done. Further, the recording medium ex215 includes a data recording area ex233, an inner peripheral area ex232, and an outer peripheral area ex234. The area used for recording user data is the data recording area ex233, and the inner circumference area ex232 and the outer circumference area ex234 arranged on the inner circumference or the outer circumference from the data recording area ex233 are used for specific purposes other than recording user data. Used. The information reproduction / recording unit ex400 reads / writes encoded audio data, video data, or multiplexed data obtained by multiplexing those data with respect to the data recording area ex233 of such a recording medium ex215.
0265In the above description, an optical disc such as a single-layer DVD or BD has been described as an example, but the present invention is not limited to these, and an optical disc having a multi-layer structure and capable of recording other than the surface may be used. In addition, an optical disc with a structure that performs multidimensional recording / playback, such as recording information at the same location on a disc using light of various different wavelength colors and recording different layers of information from various angles. It may be.
0266Further, in the digital broadcasting system ex200, it is also possible to receive data from the satellite ex202 or the like by the car ex210 having the antenna ex205 and reproduce the moving image on the display device such as the car navigation ex211 of the car ex210. As the configuration of the car navigation ex211, for example, among the configurations shown in FIG. 19, a configuration including a GPS receiver can be considered, and the same can be considered for a computer ex111, a mobile phone ex114, and the like.
0267FIG. 22A is a diagram showing a mobile phone ex114 using the moving image decoding method and the moving image coding method described in the above embodiment. The mobile phone ex114 has an antenna ex350 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex365 capable of taking images and still images, an image captured by the camera unit ex365, an image received by the antenna ex350, etc. The camera is provided with a display unit ex358 such as a liquid crystal display that displays the decoded data. The mobile phone ex114 further includes a main body having an operation key unit ex366, an audio output unit ex357 such as a speaker for outputting audio, an audio input unit ex356 such as a microphone for inputting audio, and a captured video. In the memory unit ex367 that stores encoded data such as still images, recorded audio, received video, still images, mail, etc. or decoded data, or in the interface unit with the recording media that also stores data. It has a certain slot part ex364.
0268Further, a configuration example of the mobile phone ex114 will be described with reference to FIG. 22B. The mobile phone ex114 has a power supply circuit unit ex361, an operation input control unit ex362, and a video signal processing unit ex355, as opposed to a main control unit ex360 that collectively controls each part of the main body unit having a display unit ex358 and an operation key unit ex366. , Camera interface unit ex363, LCD (Liquid Crystal Display) control unit ex359, modulation / demodulation unit ex352, multiplexing / separation unit ex353, audio signal processing unit ex354, slot unit ex364, memory unit ex367 are connected to each other via bus ex370. ing.
0269When the call end and the power key are turned on by the user's operation, the power circuit unit ex361 activates the mobile phone ex114 in an operable state by supplying power to each unit from the battery pack.
0270The mobile phone ex114 converts the voice signal picked up by the voice input unit ex356 in the voice call mode into a digital voice signal by the voice signal processing unit ex354 based on the control of the main control unit ex360 having a CPU, ROM, RAM, etc. , This is subjected to spectrum diffusion processing by the modulation / demodulation unit ex352, digital-analog conversion processing and frequency conversion processing by the transmission / reception unit ex351, and then transmitted via the antenna ex350. In addition, the mobile phone ex114 amplifies the received data received via the antenna ex350 in the voice call mode, performs frequency conversion processing and analog-to-digital conversion processing, and the modulation / demodulation unit ex352 performs spectrum reverse diffusion processing to perform the audio signal processing unit. After converting to an analog audio signal with ex354, this is output from the audio output unit ex357.
0271Further, when the e-mail is transmitted in the data communication mode, the text data of the e-mail input by the operation of the operation key unit ex366 of the main body unit is sent to the main control unit ex360 via the operation input control unit ex362. The main control unit ex360 performs spread spectrum processing on the modulation / demodulation unit ex352, digital-to-analog conversion processing and frequency conversion processing on the transmission / reception unit ex351, and then transmits the text data to the base station ex110 via the antenna ex350. .. When receiving an e-mail, the received data is processed in almost the reverse manner and output to the display unit ex358.
0272When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex355 compresses the video signal supplied from the camera unit ex365 by the moving image coding method shown in each of the above embodiments. It is encoded (that is, functions as an image coding device according to one aspect of the present invention), and the encoded video data is transmitted to the multiplexing / separating unit ex353. In addition, the audio signal processing unit ex354 encodes the audio signal picked up by the audio input unit ex356 while the camera unit ex365 is capturing images, still images, etc., and sends the encoded audio data to the multiplexing / separation unit ex353. To do.
0273The multiplexing / separating unit ex353 multiplexes the encoded video data supplied from the video signal processing unit ex355 and the encoded audio data supplied from the audio signal processing unit ex354 by a predetermined method, and is obtained as a result. The multiplexed data is subjected to spectrum diffusion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex352, digital-analog conversion processing and frequency conversion processing by the transmission / reception unit ex351, and then transmitted via the antenna ex350.
0274Decrypts the multiplexed data received via the antenna ex350 when receiving video file data linked to a homepage, etc. in data communication mode, or when receiving e-mail with video and / or audio attached. In order to do so, the multiplexing / separating unit ex353 separates the multiplexed data into a bit stream of video data and a bit stream of audio data, and processes the video data encoded via the synchronization bus ex370. While supplying to the unit ex355, the encoded audio data is supplied to the audio signal processing unit ex354. The video signal processing unit ex355 decodes the video signal by decoding by the moving image decoding method corresponding to the moving image coding method shown in each of the above embodiments (that is, the image according to one aspect of the present invention). (Functions as a decoding device), the video and still images included in the moving image file linked to the homepage are displayed from the display unit ex358 via the LCD control unit ex359. Further, the audio signal processing unit ex354 decodes the audio signal, and the audio output unit ex357 outputs the audio.
0275Further, the terminals such as the mobile phone ex114 are referred to as transmission / reception terminals having both a encoder and a decoder, as well as a transmission terminal having only an encoder and a receiving terminal having only a decoder, similar to the television ex300. There are three possible implementation formats. Furthermore, the explanation was given that the digital broadcasting system ex200 receives and transmits multiplexed data in which music data and the like are multiplexed with video data, but data in which character data and the like related to video are multiplexed in addition to audio data. It may be the video data itself instead of the multiplexed data.
0276As described above, it is possible to use the moving image coding method or moving image decoding method shown in each of the above-described embodiments for any of the above-mentioned devices / systems, and by doing so, in each of the above-described embodiments. The described effect can be obtained.
0277Further, the present invention is not limited to the above-described embodiment, and various modifications or modifications can be made without departing from the scope of the present invention.
0278(Embodiment 6) The moving image coding method or apparatus shown in each of the above embodiments is appropriately switched between the moving image coding method or apparatus conforming to different standards such as MPEG-2, MPEG4-AVC, and VC-1 as necessary. By doing so, it is also possible to generate video data.
0279Here, when a plurality of video data conforming to different standards are generated, it is necessary to select a decoding method corresponding to each standard when decoding. However, since it is not possible to identify which standard the video data to be decoded conforms to, there arises a problem that an appropriate decoding method cannot be selected.
0280In order to solve this problem, the multiplexed data in which audio data or the like is multiplexed with the video data is configured to include identification information indicating which standard the video data conforms to. The specific configuration of the multiplexed data including the video data generated by the moving image coding method or the apparatus shown in each of the above embodiments will be described below. The multiplexed data is a digital stream in the MPEG-2 transport stream format.
0281FIG. 23 is a diagram showing the structure of the multiplexed data. As shown in FIG. 23, the multiplexed data is obtained by multiplexing one or more of a video stream, an audio stream, a presentation graphics stream (PG), and an interactive graphics stream. The video stream shows the main and sub-video of the movie, the audio stream (IG) shows the main audio part of the movie and the sub-audio that mixes with the main audio, and the presentation graphics stream shows the subtitles of the movie. Here, the main image indicates a normal image displayed on the screen, and the sub image is an image displayed on a small screen in the main image. In addition, the interactive graphics stream shows an interactive screen created by arranging GUI components on the screen. The video stream is encoded by the moving image coding method or device shown in each of the above embodiments, or a moving image coding method or device conforming to the conventional standards such as MPEG-2, MPEG4-AVC, and VC-1. ing. Audio stream is Dolby AC-3, Dolby Digital It is encoded by a method such as Plus, MLP, DTS, DTS-HD, or linear PCM.
0282Each stream contained in the multiplexed data is identified by a PID. For example, 0x1011 for video streams used for movie footage, 0x1100 to 0x111F for audio streams, 0x1200 to 0x121F for presentation graphics, and 0x1400 to 0x141F for interactive graphics streams. 0x1B00 to 0x1B1F are assigned to the video stream used for the secondary video, and 0x1A00 to 0x1A1F are assigned to the audio stream used for the secondary audio to be mixed with the main audio.
0283FIG. 24 is a diagram schematically showing how the multiplexed data is multiplexed. First, the video stream ex235 composed of a plurality of video frames and the audio stream ex238 composed of a plurality of audio frames are converted into PES packet sequences ex236 and ex239, respectively, and converted into TS packets ex237 and ex240, respectively. Similarly, the data of the presentation graphics stream ex241 and the interactive graphics ex244 are converted into PES packet strings ex242 and ex245, respectively, and further converted into TS packets ex243 and ex246. The multiplexing data ex247 is configured by multiplexing these TS packets into a single stream.
0284Figure 25 shows in more detail how the video stream is stored in the PES packet sequence. The first stage in FIG. 25 shows the video frame sequence of the video stream. The second row shows the PES packet sequence. As shown by arrows yy1, yy2, yy3, yy4 in FIG. 25, a plurality of Video Presentation Units I picture, B picture, and P picture in the video stream are divided into pictures and stored in the payload of the PES packet. .. Each PES packet has a PES header, and the PES header stores PTS (Presentation Time-Stamp), which is the display time of the picture, and DTS (Decoding Time-Stamp), which is the decoding time of the picture.
0285FIG. 26 shows the format of the TS packet that is finally written to the multiplexed data. The TS packet is a 188-byte fixed-length packet composed of a 4-byte TS header containing information such as a PID that identifies the stream and a 184-byte TS payload that stores data. The PES packet is divided and stored in the TS payload. To. In the case of BD-ROM, a 4-byte TP_Extra_Header is added to the TS packet to form a 192-byte source packet, which is written to the multiplexed data. Information such as ATS (Arrival_Time_Stamp) is described in TP_Extra_Header . ATS indicates the transfer start time of the TS packet to the PID filter of the decoder. As shown in the lower part of FIG. 26, source packets are lined up in the multiplexed data, and the number incremented from the beginning of the multiplexed data is called SPN (source packet number).
0286In addition to each stream such as video, audio, and subtitles, TS packets included in the multiplexed data include PAT (Program Association Table), PMT (Program Map Table), and PCR (Program Clock Reference). The PAT indicates what the PID of the PMT used in the multiplexed data is, and the PID of the PAT itself is registered as 0. The PMT has the PID of each stream such as video, audio, and subtitles included in the multiplexed data and the attribute information of the stream corresponding to each PID, and also has various descriptors related to the multiplexed data. Descriptors include copy control information that instructs whether to allow or disallow copying of multiplexed data. PCR corresponds to ATS in which the PCR packet is transferred to the decoder in order to synchronize ATC (Arrival Time Clock), which is the time axis of ATS, with STC (System Time Clock), which is the time axis of PTS / DTS. Has STC time information.
0287FIG. 27 is a diagram illustrating the data structure of PMT in detail. At the beginning of the PMT, a PMT header that describes the length of the data included in the PMT is placed. Behind it, a plurality of descriptors related to the multiplexed data are arranged. The copy control information and the like are described as descriptors. After the descriptor, a plurality of stream information about each stream included in the multiplexed data is arranged. The stream information is composed of a stream descriptor in which the stream type, the PID of the stream, and the attribute information of the stream (frame rate, aspect ratio, etc.) are described in order to identify the compression codec of the stream. There are as many stream descriptors as there are streams in the multiplexed data.
0288When recording on a recording medium or the like, the multiplexed data is recorded together with the multiplexed data information file.
0289As shown in FIG. 28, the multiplexed data information file is the management information of the multiplexed data, has a one-to-one correspondence with the multiplexed data, and is composed of the multiplexed data information, the stream attribute information, and the entry map.
0290As shown in FIG. 28, the multiplexed data information is composed of a system rate, a playback start time, and a playback end time. The system rate indicates the maximum transfer rate of the multiplexed data to the PID filter of the system target decoder described later. The ATS interval included in the multiplexed data is set to be less than or equal to the system rate. The playback start time is the PTS of the first video frame of the multiplexed data, and the playback end time is set by adding the playback interval of one frame to the PTS of the video frame at the end of the multiplexed data.
0291As the stream attribute information is shown in FIG. 29, the attribute information for each stream included in the multiplexed data is registered for each PID. Attribute information has different information for each video stream, audio stream, presentation graphics stream, and interactive graphics stream. The video stream attribute information includes what kind of compression codec the video stream was compressed with, what the resolution of the individual picture data that makes up the video stream is, what the aspect ratio is, and the frame rate. It has information such as how much it is. The audio stream attribute information includes what compression codec the audio stream was compressed with, how many channels the audio stream contains, what language it supports, what the sampling frequency is, and so on. Has the information of. This information is used for initializing the decoder before the player plays it.
0292In the present embodiment, the stream type included in the PMT is used among the above-mentioned multiplexed data. When the multiplexed data is recorded on the recording medium, the video stream attribute information included in the multiplexed data information is used. Specifically, in the moving image coding method or apparatus shown in each of the above embodiments, the moving image coding shown in each of the above embodiments is applied to the stream type or video stream attribute information included in the PMT. Provide steps or means to set unique information indicating that the video data is generated by the method or device. With this configuration, it becomes possible to distinguish between the video data generated by the moving image coding method or the apparatus shown in each of the above embodiments and the video data conforming to other standards.
0293Further, FIG. 30 shows the steps of the moving image decoding method in the present embodiment. In step exS100, the stream type included in the PMT or the video stream attribute information included in the multiplexed data information is acquired from the multiplexed data. Next, in step exS101, it is determined whether or not the stream type or the video stream attribute information indicates that it is the multiplexed data generated by the moving image coding method or apparatus shown in each of the above embodiments. To do. Then, when it is determined that the stream type or the video stream attribute information is generated by the moving image coding method or apparatus shown in each of the above embodiments, in step exS102, each of the above implementations is performed. Decoding is performed by the moving image decoding method shown in the form. If the stream type or video stream attribute information indicates that it conforms to the conventional standards such as MPEG-2, MPEG4-AVC, and VC-1, in step exS103, the conventional method is used. Decoding is performed by a moving image decoding method that conforms to the standard.
0294In this way, by setting a new eigenvalue in the stream type or the video stream attribute information, it is possible to determine whether or not the video stream decoding method or device shown in each of the above embodiments can be used for decoding. You can judge. Therefore, even when multiplexed data conforming to different standards is input, an appropriate decoding method or device can be selected, so that decoding can be performed without causing an error. Further, the moving image coding method or device shown in the present embodiment, or the moving image decoding method or device can be used for any of the above-mentioned devices and systems.
0295(Embodiment 7) The moving image coding method and apparatus, moving image decoding method and apparatus shown in each of the above embodiments are typically realized by an LSI which is an integrated circuit. As an example, Fig. 31 shows the configuration of LSI ex500 integrated into one chip. The LSI ex500 includes the elements ex501, ex502, ex503, ex504, ex505, ex506, ex507, ex508, and ex509 described below, and each element is connected via the bus ex510. When the power supply is on, the power supply circuit unit ex505 starts up in an operable state by supplying power to each unit.
0296For example, when performing coding processing, the LSI ex500 is AV based on the control of the control unit ex501 having the CPU ex502, the memory controller ex503, the stream controller ex504, the drive frequency control unit ex512, and the like. AV signal is input from microphone ex117, camera ex113, etc. by I / O ex509. The input AV signal is temporarily stored in an external memory ex511 such as SDRAM. Based on the control of the control unit ex501, the accumulated data is appropriately divided into a plurality of times according to the processing amount and the processing speed and sent to the signal processing unit ex507, and the signal processing unit ex507 encodes the audio signal and / or the video. The signal is coded. Here, the video signal coding process is the coding process described in each of the above embodiments. The signal processing unit ex507 further performs processing such as multiplexing the encoded audio data and the encoded video data in some cases, and outputs the stream I / O ex506 to the outside. This output multiplexed data is transmitted to the base station ex107 or written to the recording medium ex215. It is advisable to temporarily store data in the buffer ex508 so that it will be synchronized when multiplexing.
0297In the above description, the memory ex511 has been described as an external configuration of the LSI ex500, but it may be a configuration included inside the LSI ex500. The buffer ex508 is not limited to one, and may have multiple buffers. Further, the LSI ex500 may be integrated into one chip or a plurality of chips.
0298Further, in the above, it is assumed that the control unit ex501 has a CPU ex502, a memory controller ex503, a stream controller ex504, a drive frequency control unit ex512, and the like, but the configuration of the control unit ex501 is not limited to this configuration. For example, the signal processing unit ex507 may be configured to further include a CPU. By providing a CPU inside the signal processing unit ex507, the processing speed can be further improved. Further, as another example, the CPU ex502 may be configured to include a signal processing unit ex507 or, for example, an audio signal processing unit that is a part of the signal processing unit ex507. In such a case, the control unit ex501 is configured to include a signal processing unit ex507 or a CPU ex502 having a part thereof.
0299Although it is referred to as LSI here, it may be referred to as IC, system LSI, super LSI, or ultra LSI depending on the degree of integration.
0300Further, the method of making an integrated circuit is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after the LSI is manufactured, or a reconfigurable processor that can reconfigure the connection and settings of circuit cells inside the LSI may be used.
0301Furthermore, if an integrated circuit technology that replaces an LSI appears due to advances in semiconductor technology or another technology derived from it, it is naturally possible to integrate functional blocks using that technology. There is a possibility of adaptation of biotechnology.
0302(Embodiment 8) When decoding the video data generated by the moving image coding method or device shown in each of the above embodiments, the video data conforming to the conventional standards such as MPEG-2, MPEG4-AVC, and VC-1 is decoded. It is conceivable that the processing amount will increase as compared with the case. Therefore, in LSIex500, it is necessary to set the drive frequency higher than the drive frequency of CPUex502 when decoding video data conforming to the conventional standard. However, when the drive frequency is increased, there arises a problem that power consumption increases.
0303In order to solve this problem, a moving image decoding device such as a television ex300 or LSI ex500 is configured to identify which standard the video data conforms to and switch the drive frequency according to the standard. FIG. 32 shows the configuration ex800 in the present embodiment. The drive frequency switching unit ex803 sets the drive frequency high when the video data is generated by the moving image coding method or device shown in each of the above embodiments. Then, the decoding processing unit ex801 that executes the moving image decoding method shown in each of the above embodiments is instructed to decode the video data. On the other hand, when the video data is video data conforming to the conventional standard, as compared with the case where the video data is generated by the moving image coding method or the apparatus shown in each of the above embodiments. Set the drive frequency low. Then, the decoding processing unit ex802, which conforms to the conventional standard, is instructed to decode the video data.
0304More specifically, the drive frequency switching unit ex803 is composed of the CPU ex502 and the drive frequency control unit ex512 in FIG. 31. Further, the decoding processing unit ex801 that executes the moving image decoding method shown in each of the above embodiments and the decoding processing unit ex802 that conforms to the conventional standard correspond to the signal processing unit ex507 of FIG. 31. CPUex502 identifies which standard the video data conforms to. Then, the drive frequency control unit ex512 sets the drive frequency based on the signal from the CPU ex502. Further, the signal processing unit ex507 decodes the video data based on the signal from the CPU ex502. Here, for the identification of the video data, for example, it is conceivable to use the identification information described in the sixth embodiment. The identification information is not limited to the information described in the sixth embodiment, and may be any information that can identify which standard the video data conforms to. For example, it is possible to identify which standard the video data conforms to based on an external signal that identifies whether the video data is used for a television or a disc. In some cases, identification may be based on such an external signal. Further, it is conceivable that the drive frequency in the CPUex 502 is selected based on, for example, a lookup table in which the video data standard as shown in FIG. 34 and the drive frequency are associated with each other. The lookup table is stored in the buffer ex508 or the internal memory of the LSI, and the CPU ex502 can select the drive frequency by referring to this lookup table.
0305FIG. 33 shows the steps to implement the method of this embodiment. First, in step exS200, the signal processing unit ex507 acquires identification information from the multiplexed data. Next, in step exS201, CPUex502 identifies whether or not the video data is generated by the coding method or apparatus shown in each of the above embodiments based on the identification information. When the video data is generated by the coding method or device shown in each of the above embodiments, in step exS202, the CPU ex502 sends a signal for setting the drive frequency high to the drive frequency control unit ex512. Then, the drive frequency control unit ex512 sets a high drive frequency. On the other hand, if it is shown that the video data conforms to the conventional standards such as MPEG-2, MPEG4-AVC, and VC-1, CPUex502 drives the signal to set the drive frequency low in step exS203. Send to frequency control unit ex512. Then, in the drive frequency control unit ex512, the drive frequency is set to be lower than that in the case where the video data is generated by the coding method or the apparatus shown in each of the above embodiments.
0306Further, the power saving effect can be further enhanced by changing the voltage applied to the LSI ex500 or the device including the LSI ex500 in conjunction with the switching of the drive frequency. For example, when the drive frequency is set low, it is conceivable to set the voltage applied to the LSI ex500 or the device including the LSI ex500 lower than when the drive frequency is set high.
0307Further, as the driving frequency setting method, when the processing amount at the time of decoding is large, the driving frequency may be set high, and when the processing amount at the time of decoding is small, the driving frequency may be set low. Not limited to the method. For example, when the amount of processing for decoding video data conforming to the MPEG4-AVC standard is larger than the amount of processing for decoding video data generated by the moving image coding method or device shown in each of the above embodiments. It is conceivable to reverse the setting of the drive frequency as described above.
0308Further, the method of setting the drive frequency is not limited to the configuration in which the drive frequency is lowered. For example, when the identification information indicates that it is the video data generated by the moving image coding method or the apparatus shown in each of the above embodiments, the voltage applied to the LSI ex500 or the apparatus including the LSI ex500 is set high. However, if it indicates that the video data conforms to the conventional standards such as MPEG-2, MPEG4-AVC, and VC-1, it is also possible to set the voltage applied to the LSI ex500 or the device including the LSI ex500 low. Be done. Further, as another example, when the identification information indicates that it is the video data generated by the moving image coding method or the apparatus shown in each of the above embodiments, the drive of the CPU ex502 is stopped. If it is shown that the video data conforms to the conventional standards such as MPEG-2, MPEG4-AVC, and VC-1, there is a margin in processing, so the drive of CPUex502 should be paused. Is also possible. Even when the identification information indicates that it is the video data generated by the moving image coding method or the apparatus shown in each of the above embodiments, if there is a margin in the processing, the CPU ex502 is temporarily driven. It is also possible to stop it. In this case, it is conceivable to set the stop time shorter than in the case of indicating that the video data conforms to the conventional standards such as MPEG-2, MPEG4-AVC, and VC-1.
0309In this way, power saving can be achieved by switching the drive frequency according to the standard to which the video data conforms. Further, when the LSI ex500 or the device including the LSI ex500 is driven by using the battery, the life of the battery can be extended along with the power saving.
0310(Embodiment 9) A plurality of video data conforming to different standards may be input to the above-mentioned devices / systems such as televisions and mobile phones. In this way, in order to enable decoding even when a plurality of video data conforming to different standards are input, the signal processing unit ex507 of LSI ex500 needs to support the plurality of standards. However, if the signal processing unit ex507 corresponding to each standard is used individually, there arises a problem that the circuit scale of the LSI ex500 becomes large and the cost increases.
0311In order to solve this problem, a decoding processing unit for executing the moving image decoding method shown in each of the above embodiments, and decoding conforming to conventional standards such as MPEG-2, MPEG4-AVC, and VC-1. The configuration is such that a part is shared with the processing unit. An example of this configuration is shown in ex900 of FIG. 35A. For example, the moving image decoding method shown in each of the above embodiments and the moving image decoding method conforming to the MPEG4-AVC standard are processed in processing such as entropy coding, dequantization, deblocking filter, and motion compensation. Some of the contents are common. For common processing content, the decoding processing unit ex902 corresponding to the MPEG4-AVC standard is shared, and for other processing content specific to one aspect of the present invention that does not correspond to the MPEG4-AVC standard, a dedicated decoding processing unit is used. A configuration using ex901 is conceivable. In particular, since one aspect of the present invention is characterized by inter-prediction, for example, for inter-prediction, a dedicated decoding processing unit ex901 is used, and other entropy decoding, deblocking filter, and dequantization are used. It is conceivable to share the decoding processing unit for any or all of the processing. Regarding the sharing of the decoding processing unit, regarding the common processing content, the decoding processing unit for executing the moving image decoding method shown in each of the above embodiments is shared, and the processing content peculiar to the MPEG4-AVC standard is shared. May be configured to use a dedicated decoding processing unit.
0312In addition, another example of partially sharing the processing is shown in ex1000 in Fig. 35B. In this example, a dedicated decoding processing unit ex1001 corresponding to the processing content peculiar to one aspect of the present invention, a dedicated decoding processing unit ex1002 corresponding to the processing content peculiar to other conventional standards, and one aspect of the present invention. It is configured to use the common decoding processing unit ex1003 corresponding to the processing contents common to the moving image decoding method according to the above and the moving image decoding method of other conventional standards. Here, the dedicated decoding processing units ex1001 and ex1002 are not necessarily specialized in one aspect of the present invention or processing contents peculiar to other conventional standards, but can execute other general-purpose processing. May be good. It is also possible to implement the configuration of this embodiment with LSI ex500.
0313As described above, the LSI circuit scale can be reduced by sharing the decoding processing unit for the processing contents common to the moving image decoding method according to one aspect of the present invention and the moving image decoding method of the conventional standard. Moreover, it is possible to reduce the cost.
0314The present invention can be used, for example, in television receivers, digital video recorders, car navigation systems, mobile phones, digital cameras, digital video cameras, and the like.
0315500, 1300 image encoder 501 subtraction part 502 Conversion unit 503 Quantization unit 504 Entropy encoding unit 505, 602 Inverse quantization unit 506, 603 Inverse converter 507, 604 Addition part 508, 605 block memory 509, 606 picture memory 510, 607 Intra Prediction Department 511, 608 Inter Prediction Department 512, 609, 1303, 1403 selection 600, 1400 image decoder 601 Entropy decoding unit 1301, 1401 Derivation part 1302, 1402 additional part 1304 Encoding unit 1404 Decryptor
41 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2004215215A | Cites | Japan |
| JP2003333602A | Cites | Japan |
| Yunfei Zheng(外2名),Unified Motion Vector Predictor Selection for Merge and AMVP,Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 JCTVC-E396,米国,ITU-T,2011年 3月23日,p.1-5 | Non-patent | – |
55 members in 14 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 61552863 | United States of America | – | |
| 201161552863 | United States of America | P |
Members55
| Document | Office | Kind | |
|---|---|---|---|
| CA2836243A1 | Canada | A1 | |
| US2013107965A1 | United States of America | A1 | |
| WO2013061546A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201334562A | Taiwan Province of China | A | |
| MX2013012223A | Mexico | A | |
| JP5383958B2 | Japan | B2 | |
| JP2014042291A | Japan | A | |
| JP2014042292A | Japan | A | |
| CN103688545A | China | A | |
| AR088538A1 | Argentina | A1 | |
| KR20140092758A | Republic of Korea | A | |
| EP2773111A1 | European Patent Office (EPO) | A1 | |
| US8861606B2 | United States of America | B2 | |
| US2014314153A1 | United States of America | A1 | |
| US2014348245A1 | United States of America | A1 | |
| JPWO2013061546A1 | Japan | A1 | |
| EP2773111A4 | European Patent Office (EPO) | A4 | |
| US9088799B2 | United States of America | B2 | |
| US2015208088A1 | United States of America | A1 | |
| US9191679B2 | United States of America | B2 | |
| US9357227B2 | United States of America | B2 | |
| US2016241872A1 | United States of America | A1 | |
| TWI563831B | Taiwan Province of China | B | |
| JP6083607B2 | Japan | B2 | |
| JP6089302B2This record | Japan | B2 | |
| CN103688545B | China | B | |
| JP2017099001A | Japan | A | |
| US9699474B2 | United States of America | B2 | |
| BR112013027350A2 | Brazil | A2 | |
| CN107071471A | China | A | |
| US2017264915A1 | United States of America | A1 | |
| JP6260920B2 | Japan | B2 | |
| MY164782A | Malaysia | A | |
| BR112013027350A8 | Brazil | A8 | |
| US9912962B2 | United States of America | B2 | |
| US2018152726A1 | United States of America | A1 | |
| US10045047B2 | United States of America | B2 | |
| US2018316933A1 | United States of America | A1 | |
| KR101935976B1 | Republic of Korea | B1 | |
| EP2773111B1 | European Patent Office (EPO) | B1 | |
| US10567792B2 | United States of America | B2 | |
| CN107071471B | China | B | |
| US2020154133A1 | United States of America | A1 | |
| PL2773111T3 | Poland | T3 | |
| ES2780186T3 | Spain | T3 | |
| CA2836243C | Canada | C | |
| US10893293B2 | United States of America | B2 | |
| US2021092440A1 | United States of America | A1 | |
| BR112013027350B1 | Brazil | B1 | |
| US11356696B2 | United States of America | B2 | |
| US2022248049A1 | United States of America | A1 | |
| US11831907B2 | United States of America | B2 | |
| US2023412835A1 | United States of America | A1 | |
| US12132930B2 | United States of America | B2 | |
| US2025024069A1 | United States of America | A1 |
13 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Notification of change in applicantJAPANESE INTERMEDIATE CODE: A711A711 | A711 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Notification of change in applicantJAPANESE INTERMEDIATE CODE: A711A711 | A711 |
Numbers
- Publication
- 6089302
- Application
- 206666
Titles2
- Japanese
- 符号化方法および符号化装置
- English
- Coding method and coding device
Classification
- CPC, 8
- H04N19/52
- H04N19/51
- H04N19/56
- H04N19/58
- H04N19/61
- H04N19/513
- H04N19/172
- H04N19/176
- IPC, 1
- H04N19 513
