Digital signal encoding device, digital signal decoding device, digital signal arithmetic encoding method and digital signal arithmetic decoding method
Abstract
In the bit stream syntax of the segment video image compression data of the segment structured video image compression data, each segment of the video image compression data is multiplexed as the segment header of each segment video image compression data: segment start code; register The reset flag indicates whether to reset the register value indicating the state of the word code of the arithmetic coding processing program in the next transmission unit; and the initial register value, which should be used only when the register reset flag indicates <do not perform reset>. The register value used at the beginning of the arithmetic coding and decoding of the next transmission unit is the register value at that moment.

Term
Term ended
Projected expiry passed 10 April 2023, 3.5 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
20 claims: 8 independent, 12 dependent
- 1一种数字信号编码装置,将数字信号以既定单位进行分割而执行压缩编码,其特征在于:包括算术编码部,以算术编码来压缩既定单位的数字信号,而相当算术编码部是将表现在某传输单位的编码为完毕的时刻的算术编码状态的信息,作为次一传输单位的数据的一部分进行复用。
- 2如权利要求1所述的数字信号编码装置,其中,上述算术编码部是将既定单位的数字信号以基于与含于1个或多个相邻接的传输单位的信号间的依存关系,来决定编码符号的发生概率而执行算术编码。
- 3如权利要求2所述的数字信号编码装置,其中,上述算术编码部是以计数被编码的符号的出现频度来学习上述发生概率。
- 4如权利要求1所述的数字信号编码装置,其中,所谓上述表现算术编码状态的信息是指:寄存器复位标志,表示以显示算术编码处理程序的寄存器值的复位的有无;及初始寄存器值,仅于不复位寄存器值的场合时进行附加。
- 5一种数字信号编码装置,将数字信号以既定单位进行分割而执行压缩编码,其特征在于:包括算术编码部,以算术编码来压缩既定单位的数字信号,而相当算术编码部是将既定单位的数字信号以基于与含于1个或多个相邻接的传输单位的信号间的依存关系,来决定编码符号的发生概率,同时以计数被编码的符号的出现频度来学习上述发生概率,并将表现在某传输单位的编码为完毕的时刻的发生概率学习状态的信息,作为次一传输单位的数据的一部分进行复用。
- 6如权利要求5所述的数字信号编码装置,其中,所谓表现发生概率学习状态的信息是表示将与形成编码符号的发生概率的变动原因的其他信息的依存关系进行模型化的文脉模型状态的信息。
- 7如权利要求1所述的数字信号编码装置,其中,上述数字信号是视频图像信号,而上述传输单位是由视频图像帧内的1个至多个微块所构成的片段。
- 8如权利要求1所述的数字信号编码装置,其中,上述数字信号是视频图像信号,而上述传输单位是根据含于上述片段内的编码数据的种别而被再构成的编码数据单位。
- 9如权利要求1所述的数字信号编码装置,其中,上述数字信号是视频图像信号,而上述传输单位是视频图像帧。
- 10一种数字信号解码装置,将被压缩编码的数字信号以既定单位进行接收而执行解码,其特征在于:包括算术解码部,将既定单位的压缩数字信号以基于算术编码的程序进行解码,而该算术解码部是在某传输单位的解码开始时,以基于表现做为该传输单位数据的一部分而被复用的算术编码状态的信息,来执行解码动作的初始化。
- 11如权利要求10所述的数字信号解码装置,其中,上述算术解码部是在解码既定单位的压缩数字信号之际,基于与包含于1个或多个相邻接的传输单位的信号间的依存关系,来决定解码符号的发生概率而执行解码。
- 12如权利要求10所述的数字信号解码装置,其中,上述算术解码部是以计数被解码的符号的出现频度来学习上述发生概率。
- 13一种数字信号解码装置,将被压缩编码的数字信号以既定单位进行接收而执行解码,其特征在于:包括算术解码部,将既定单位的压缩数字信号以基于算术编码的程序进行解码,而该算术解码部是在某传输单位的解码开始时,以基于表现做为该传输单位数据的一部分而被复用的符号发生概率学习状态的信息,而执行使用在解码该传输单位的发生概率的初始化,同时于解码既定单位的压缩数字信号之际,以基于与含于1个或多个相邻接的传输单位的信号间的依存关系来决定解码符号的发生概率,并以计数被解码的符号的出现频度来学习上述发生概率而执行解码。
- 14如权利要求10所述的数字信号解码装置,其中,上述数字信号是视频图像信号,而上述传输单位是由视频图像帧内的1个至多个微块所构成的片段。
- 15如权利要求10所述的数字信号解码装置,其中,上述数字信号是视频图像信号,而上述传输单位是根据包含于上述片段内的编码数据的种别而被再构成的编码数据单位。
- 16如权利要求10所述的数字信号解码装置,其中,上述数字信号是视频图像信号,而上述传输单位是视频图像帧。
- 17一种数字信号算术编码方法,为在将数字信号以既定单位进行分割而执行压缩编码之际的数字信号算术编码方法,其特征在于:在将既定单位的数字信号以算术编码进行压缩之际,将表现在某传输单位的编码为完毕的时刻的算术编码状态的信息,作为次一传输单位的数据的一部分进行复用。
- 18一种数字信号算术编码方法,为在将数字信号以既定单位进行分割而执行压缩编码之际的数字信号算术编码方法,其特征在于:在将既定单位的数字信号以算术编码进行压缩之际,将既定单位的数字信号以基于与含于1个或多个相邻接的传输单位的信号间的依存关系,来决定编码符号的发生概率,同时以计数被编码的符号的出现频度来学习上述发生概率,并将表现在某传输单位的编码为完毕的时刻的发生概率学习状态的信息,作为次一传输单位的数据的一部分进行复用。
- 19一种数字信号算术解码方法,为在将数字信号以既定单位进行分割而执行压缩编码之际的数字信号算术编码方法,其特征在于:在将既定单位的压缩数字信号以基于算术编码的程序进行解码之际,在某传输单位的解码开始时,以基于表现做为该传输单位数据的一部分而被复用的算术编码状态的信息,而执行解码动作的初始化。
- 20一种数字信号算术解码方法,为在将数字信号以既定单位进行分割而执行压缩编码之际的数字信号算术编码方法,其特征在于:在将既定单位的压缩数字信号以基于算术编码的程序进行解码之际,在某传输单位的解码开始时,以基于表现做为该传输单位数据的一部分而被复用的算术编码状态的信息,而执行使用在解码该传输单位的发生概率的初始化,同时于解码既定单位的压缩数字信号之际,以基于与含于1个或多个相邻接的传输单位的信号间的依存关系来决定解码符号的发生概率,并以计数被解码的符号的出现频度来学习上述发生概率而执行解码。
Independent claims20
173 paragraphs, as filed
Digital signal coding device, digital signal decoding device, digital signal arithmetic coding method and digital signal arithmetic decoding method
Technical field
The present invention relates to a digital signal coding device, a digital signal decoding device, a digital signal arithmetic coding method, and a digital signal arithmetic decoding method used in video image compression coding technology and compressed video image data transmission technology.
Background technique
In the known international standard video image coding methods such as MPEG and ITU-T H.26x, Huffman coding is used as entropy coding. Although Huffman coding can provide the most suitable coding performance when each information source symbol is required to be expressed as an independent character code, on the one hand, the shape of a signal such as a video image signal changes locally. There is a problem that the optimality cannot be guaranteed when the so-called probability of occurrence of the information source symbol fluctuates.
In this case, the following scheme can be adopted: dynamically adapting to the occurrence probability of each information source symbol, combining multiple symbols and expressing them in one character code as arithmetic coding.
By quoting Mark Nelson, "Arithmetic Coding + Statistical Modeling = Data Compress partl-Arithmetic Coding", Dr. Dobb's Journal, February 1991, the idea of arithmetic coding is simply explained. Here is to consider using alphabetic characters as the information source of the information source symbol, and thinking about the arithmetic coding of the so-called "BILLGATES" information.
At this time, the occurrence probability of each character is defined as shown in FIG. 1. Moreover, as shown in the value range of the graph, only one area defined on the probability number line of the interval [0, 1] is determined.
Second, enter the encoding process. First, the character "B" is encoded, but this is equivalent to the range [0.2, 0.3] on the straight line of the selected probability number. Therefore, the character "B" becomes a value corresponding to the upper limit (High) and the lower limit (Low) of a set of value ranges [0.2, 0.3].
Secondly, when encoding "1", the value range [0.2, 0.3] selected in the "B" encoding is changed to [0, 1] interval, and the interval of [0.5, 0.6] is selected. In short, the processing program of arithmetic coding is equivalent to the squeeze of the value range of the execution probability number straight line.
As long as this process is repeated for each character, as shown in FIG. 2, the arithmetic coding result of "BILLGATES" is represented by the Low value <0.2572167752> at the time when the character "S" is encoded.
In the decoding process, it is also possible to consider the opposite process.
First, it is checked that the encoding result <0.2572167752> is the value range assigned to the character on the line corresponding to the probability number, and "B" is obtained.
After that, after subtracting the Low value of "B", division is performed in the range to obtain <0.572167752>. As a result, the character "I" corresponding to the interval [0.5, 0.6] can be decoded. Hereinafter, by repeating this process, "BILL GATES" can be decoded.
Through the above processing, if arithmetic coding is performed, even the coding of a very long message can be mapped to one character code at the end. However, from the actual implementation, it is impossible to deal with the infinite decimal point precision, and the encoding and decoding procedures require multiplication and division operations to increase the computational load. For example, the implementation of the use of integer registers as a character code expression Floating decimal point calculations are performed by approximating the above Low value by a power of two, and replacing multiplication and division operations with shift operations. If it is based on arithmetic coding, it is ideally suitable for entropy coding of the occurrence probability of information source symbols through the above-mentioned procedure. In particular, when the probability of occurrence changes dynamically, the table of FIG. 1 is appropriately updated to track the change of the probability of occurrence, and a higher coding efficiency than Huffman coding can be obtained.
Since the known digital signal arithmetic coding method and digital signal arithmetic decoding method are constructed as described above, when transmitting an entropy-encoded video image signal, it is usually in order to remove the video image caused by the transmission error. Disturbance is suppressed to a minimum, and each frame of the video image is divided into partial regions, and the majority of them are transmitted in units that can be resynchronized (for example, MPEG-2 segment structure).
Therefore, in Huffman coding, although each coding target symbol is to be mapped to a character code of integer bit length, and only the character code equivalent to the set is defined as the transmission unit, in arithmetic coding , Because it is not only necessary to explicitly interrupt the special symbols of the encoding program, but also when the encoding is restarted, the learning processing program of the probability of occurrence of the symbols up to this point is once reset, and the bits that can be determined by the code need to be discharged, so there will be Incurring the possibility of lowering the coding efficiency before and after the interruption. Furthermore, if the arithmetic coding process is to encode without resetting in 1 video image frame, and when it has to be divided into small units such as packet data during transmission, the decoding process of a certain packet is just perfect. One packet of data cannot be implemented, and there is a problem that the video image quality is significantly degraded when a packet loss caused by transmission errors and delays occurs.
Summary of the invention
This is because the present invention is made to solve the above-mentioned problems, and aims to obtain a digital signal coding device and a digital signal arithmetic coding method that can ensure error tolerance while improving the coding efficiency of arithmetic coding.
Moreover, the present invention is to obtain a digital signal decoding device that can decode correctly even when the encoding device continues to be encoded without resetting the arithmetic encoding state or symbol occurrence probability learning state of the previous transmission unit on the encoding device side. And digital signal arithmetic decoding method as the purpose.
The digital signal coding device and digital signal arithmetic coding method of the present invention is to compress the digital signal of a predetermined transmission unit by arithmetic coding, and use the information of the arithmetic coding state when the coding of a certain transmission unit is completed. , Multiplexed as part of the data of the next transmission unit, or based on the dependence relationship between the signal contained in one or more adjacent transmission units, to determine the probability of occurrence of the coded symbol, and at the same time it is counted The occurrence frequency of the coded symbols is used to learn the occurrence probability, and the information that can express the occurrence probability learning state at the time when the encoding of a certain transmission unit is completed is multiplexed as part of the next transmission unit data.
Therefore, coding can be continued without resetting the previous arithmetic coding state or symbol occurrence probability learning state. Therefore, it is possible to ensure error tolerance and implement coding that improves the coding efficiency of arithmetic coding.
Furthermore, the digital signal decoding device and digital signal arithmetic decoding method of the present invention perform decoding based on information representing the state of arithmetic coding multiplexed as part of the transmission unit data when the decoding of a certain transmission unit starts. The initialization of the action, or when the decoding of a certain transmission unit starts, is performed based on the information representing the learning state of the probability of occurrence of symbols multiplexed as part of the data of the transmission unit, and executes the information used in decoding the probability of occurrence of the transmission unit Initialization, when decoding the compressed digital signal of a predetermined transmission unit, the probability of the occurrence of the decoded symbol is determined based on the dependence relationship with the signal contained in one or more adjacent transmission units, and it can be counted. The frequency of appearance of the decoded symbol is learned to perform decoding by learning the probability of occurrence.
Therefore, even when the encoding device is continuously performing encoding without resetting the arithmetic encoding state or the symbol occurrence probability learning state of the previous transmission unit, it has an effect of accurately decodability.
BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1 is an explanatory diagram showing the occurrence probability of each character when the characters called "BILLGATES" are arithmetic-coded.
Fig. 2 is an explanatory diagram showing the result of arithmetic coding in the case of arithmetic coding of the so-called "BILLGATES" characters.
Fig. 3 is a diagram showing the structure of a video image coding device (digital signal coding device) according to the first embodiment of the present invention.
Fig. 4 is a diagram showing the structure of a video image decoding device (digital signal decoding device) according to the first embodiment of the present invention.
FIG. 5 is a configuration diagram showing the internal structure of the arithmetic coding unit 6 in FIG. 3.
FIG. 6 is a flowchart showing the processing content of the arithmetic coding unit 6 in FIG. 5.
Fig. 7 is an explanatory diagram showing an example of a context model.
Fig. 8 is an explanatory diagram showing an example of a context model for motion vectors.
Fig. 9 is an explanatory diagram for explaining the segment structure.
FIG. 10 is an explanatory diagram showing an example of a bit data stream generated by the arithmetic coding unit 6.
FIG. 11 is an explanatory diagram showing an example of another bit data stream generated by the arithmetic coding unit 6.
FIG. 12 is an explanatory diagram showing an example of another bit data stream generated by the arithmetic coding unit 6.
FIG. 13 is a configuration diagram showing the internal configuration of the arithmetic decoding unit 27 in FIG. 4.
FIG. 14 is a flowchart showing the processing content of the arithmetic decoding unit 27 in FIG. 13.
Fig. 15 is a structural diagram showing the internal structure of the arithmetic coding unit 6 in the second embodiment.
FIG. 16 is a flowchart showing the processing content of the arithmetic coding unit 6 in FIG. 15.
FIG. 17 is an explanatory diagram for explaining the learning state of the context model.
Fig. 18 is an explanatory diagram showing an example of a bit stream generated by the arithmetic coding unit 6 of the second embodiment.
Fig. 19 is a structural diagram showing the internal structure of the arithmetic decoding unit 27 of the second embodiment.
FIG. 20 is a flowchart showing the processing content of the arithmetic decoding unit 27 of FIG. 19.
Fig. 21 is an explanatory diagram showing an example of a bit stream generated by the arithmetic coding unit 6 of the third embodiment.
Specific embodiments of the invention
Hereinafter, in order to explain the present invention in more detail, the best mode for carrying out the present invention will be described with reference to the drawings.
First Embodiment In the first embodiment, arithmetic coding is applied to a video image coding method in which a video image frame is equally divided into a rectangular area of 16×16 pixels (hereinafter, referred to as microblocks) to perform coding. It uses the examples disclosed by D. Marpe and others in "VideoCompression Using Context-Based Adaptive Arithmetic Coding", International Conference on Image Processing 2001.
Fig. 3 is a diagram showing the structure of a video image encoding device (digital signal encoding device) according to the first embodiment of the present invention. In the figure, the motion detection unit 2 uses the reference image 4 stored in the frame memory 3a to input the video image signal. 1 Detect the motion vector in units of micro-blocks 5. The motion compensation unit 7 obtains the temporal prediction image 8 based on the motion vector 5 detected by the motion detection unit 2. The subtractor 51 obtains the difference between the input video image signal 1 and the temporal prediction image 8 and outputs the difference as the temporal prediction residual signal 9.
The spatial prediction unit 10a refers to the input video image signal 1 and performs prediction from the vicinity of the space in the same video image frame to generate the spatial prediction residual signal 11. The coding model determination unit 12 is derived from: a motion prediction model that encodes the temporal prediction residual signal 9; as a skip model when the motion vector 5 is zero and there is no temporal prediction residual signal 9 component; and the spatial prediction residual Among the internal models in which the difference signal 11 is encoded, a model that can encode equivalent microblocks most efficiently is selected, and the encoded model information 13 is output.
The orthogonal transform unit 15 performs orthogonal transform on the encoding target signal selected by the encoding model determination unit 12 and outputs orthogonal transform coefficient data. The quantization unit 16 performs the quantization of the orthogonal transform coefficient data at the granularity indicated by the quantization step parameter 23 determined by the encoding control unit 22.
The inverse quantization unit 18 performs inverse quantization of the orthogonal transform coefficient data 17 output from the quantization unit 16 at the granularity indicated by the quantization step parameter 23. The inverse orthogonal transform unit 19 performs inverse orthogonal transform on the orthogonal transform coefficient data inverse quantized by the inverse quantization unit 18. The switching unit 52 selects and outputs the temporal prediction image 8 output from the motion compensation unit 7 or the spatial prediction image 20 output from the spatial prediction unit 10a based on the coding model information 13 output from the coding model determination unit 12. The adder 53 adds the output signal of the switching unit 52 and the output signal of the inverse orthogonal transform unit 19 to generate the locally decoded image 21, and stores the locally decoded image 21 as the reference image 4 in the frame memory 3a.
The arithmetic coding unit 6 performs entropy coding of coding target data such as motion vector 5, coding model information 13, spatial prediction model 14, and orthogonal transform coefficient data 17, and uses the coding result as video image compression data 26, It is output from the transmission buffer 24. The encoding control unit 22 controls the encoding model determination unit 12, the quantization unit 16, the inverse quantization unit 18, and the like.
4 is a diagram showing the structure of a video image decoding device (digital signal decoding device) according to the first embodiment of the present invention. In the figure, the arithmetic decoding unit 27 performs entropy decoding processing to decode: motion vector 5, coding model information 13 , Spatial prediction model 14, orthogonal transform coefficient data 17, and quantization step parameters 23, etc. The inverse quantization unit 18 inversely quantizes the orthogonal transform coefficient data 17 and the quantization step parameters 23 decoded by the arithmetic decoding unit 27. The inverse orthogonal transform unit 19 performs inverse orthogonal transform on the inversely quantized orthogonal transform coefficient data 17 and the quantization step parameters 23 and locally decodes them.
The motion compensation unit 7 uses the motion vector 5 decoded by the arithmetic decoding unit 27 to restore the temporal prediction image 8. The spatial prediction unit 10b restores the spatial prediction image 20 from the spatial prediction model 14 decoded by the arithmetic decoding unit 27.
The switching unit 54 selects and outputs the temporal prediction image 8 or the spatial prediction image 11 based on the coding model information 13 decoded by the arithmetic decoding unit 27. The adder 55 adds the local decoded signal as the output signal of the inverse orthogonal transform unit 19 and the output signal of the switching unit 54 to output the decoded image 21. In addition, the decoded image 21 is stored in the frame memory 3b used in the generation of the predicted image of the following frame.
Next, the action will be explained.
First, the outline of the operation of the video image encoding device and the video image decoding device will be explained.
(1) Overview of the operation of the video image encoding device The input video image signal 1 is input in units of each video image frame divided into micro-blocks, and the motion detection unit 2 of the video image encoding device is used and stored in the frame memory 3a With reference to image 4, the motion vector 5 is detected in units of microblocks.
The motion compensation unit 7 obtains the temporal prediction image 8 based on the motion vector 5 as soon as the motion detection unit 2 detects the motion vector 5.
The subtractor 51 receives the temporal prediction image 8 from the motion compensation unit 7, and obtains the difference between the input video image signal 1 and the temporal prediction image 8, and outputs the difference as the temporal prediction residual signal 9 to the encoding Model determination unit 12.
On the one hand, the spatial prediction unit 10a generates a spatial prediction residual signal 11 by referring to the input video image signal 1 and performing prediction from the nearby area of the space within the same video image frame as long as an input video image signal 1 is input.
The coding model determination unit 12 is derived from: a motion prediction model that encodes the temporal prediction residual signal 9; as a skip model when the motion vector 5 is zero and there is no component of the temporal prediction residual signal 9; and the spatial prediction Among the internal models in which the residual signal 11 is encoded, a model that encodes the corresponding microblock with the best efficiency is selected, and the encoding model information 13 is output to the arithmetic encoding unit 6. Also, when the motion prediction model is selected, the temporal prediction residual signal 9 is output to the orthogonal transform unit 15 as the encoding target signal, and when the internal model is selected, the spatial prediction residual signal 11 It is output to the orthogonal transform unit 15 as an encoding target signal.
When the motion prediction model is selected, the motion vector 5 is output from the motion detection unit 2 to the arithmetic coding unit 6 as the encoding target information. When the internal model is selected, the intra prediction model 14 is the encoding target information. The spatial prediction unit 10a is output to the arithmetic coding unit 6.
As long as the orthogonal transform unit 15 receives the encoding target signal from the encoding model determination unit 12, the encoding target signal is orthogonally transformed and the orthogonal transform coefficient data is output to the quantization unit 16.
The quantization unit 16 is designed to execute the orthogonal transformation coefficient data at the granularity indicated by the quantization step parameter 23 determined by the encoding control unit 22 as long as it receives the orthogonal transformation coefficient data from the orthogonal transformation unit 15. Quantization.
In addition, the encoding control unit 22 adjusts the quantization step parameter 23 to achieve a balance between encoding rate and quality. Generally speaking, after arithmetic coding, the occupancy of the coded data stored in the transmission buffer 24 just before transmission is confirmed at regular intervals, and the quantization step parameter 23 is executed according to the buffer margin 25. Parameter adjustment. For example, when the buffer margin 25 is large, in addition to suppressing the encoding rate, when the buffer margin 25 has margin, the encoding rate can also be increased to improve the quality.
As long as the inverse quantization unit 18 receives the orthogonal transform coefficient data 17 from the quantization unit 16, it will perform the inverse quantization of the orthogonal transform coefficient data 17 at the granularity indicated by the quantization step parameter 23.
The inverse orthogonal transform unit 19 performs inverse orthogonal transform on the orthogonal transform coefficient data inverse quantized by the inverse quantization unit 18.
The switching unit 52 selects and outputs the temporal prediction image 8 output from the motion compensation unit 7 or the spatial prediction image 20 output from the spatial prediction unit 10a based on the coding model information 13 output from the coding model determination unit 12. That is, when the coding model information 13 is a display motion prediction model, the temporal prediction image 8 output from the motion compensation unit 7 is selected for output, and when the coding model information 13 is a display internal model, the spatial prediction is selected The spatial prediction image 20 output by the unit 10a is output.
The adder 53 adds the output signal of the switching unit 52 and the output signal of the inverse orthogonal transform unit 19 to generate the local decoded image 21. In addition, the locally decoded image 21 is stored in the frame memory 3a as the reference image 4 in order to be used for the motion prediction of the following frame.
The arithmetic coding unit 6 performs entropy coding of coding target data such as motion vector 5, coding model information 13, spatial prediction model 14, and orthogonal transform coefficient data 17, according to a program described later, and uses the coding result as a video The image compression data 26 is output from the transmission buffer 24.
(2) Operation summary of the video image decoding device The arithmetic decoding unit 27 decodes the motion vector 5 and the encoding model information 13 by performing the entropy decoding process described later as soon as it receives the video image compression data 26 from the video image encoding device. , Spatial prediction model 14, orthogonal transform coefficient data 17, and quantization step parameters 23, etc.
The inverse quantization unit 18 inversely quantizes the orthogonal transform coefficient data 17 decoded by the arithmetic decoding unit 27 and the quantization step parameters 23, and the inverse orthogonal transform unit 19 inversely quantizes the orthogonal transform coefficients. The data 17 and the quantization step parameters 23 are subjected to inverse orthogonal transformation to perform local decoding.
The motion compensation unit 7 uses the motion vector 5 decoded by the arithmetic decoding unit 27 to restore the temporal prediction image 8 when the coding model information 13 decoded by the arithmetic decoding unit 27 is a display motion prediction model.
The spatial prediction unit 10b restores the spatial prediction image 20 from the spatial prediction model 14 decoded by the arithmetic decoding unit 27 when the coding model information 13 decoded by the arithmetic decoding unit 27 shows an internal model.
Here, the difference between the spatial prediction unit 10a on the video image encoding device side and the spatial prediction unit 10b on the video image decoding device side is the type of all spatial prediction models obtained for the former, and includes the most efficient selection The process of determining the spatial prediction model 14, and the latter is limited to the process of generating the spatial prediction image 20 from the provided spatial prediction model 14.
The switching unit 54 selects the temporal prediction image 8 restored by the motion compensation unit 7 or the spatial prediction image 11 restored by the spatial prediction unit 10b based on the coding model information 13 decoded by the arithmetic decoding unit 27, and selects it The image is output to the adder 55 as a predicted image.
When the adder 55 receives the predicted image from the switching unit 54, it adds the predicted image and the local decoded signal output from the inverse orthogonal transform unit 19 to obtain the decoded image 21.
In addition, the decoded image 21 is stored in the frame memory 3b in order to be used for the generation of the predicted image of the following frame. The difference between the frame memories 3a and 3b is only the difference between so-called being mounted on a video image encoding device and a video image decoding device.
(3) Arithmetic coding and decoding processing Hereinafter, the arithmetic coding and decoding processing, which is the gist of the present invention, will be described in detail. The encoding process is executed in the arithmetic encoding section 6 of FIG. 3, and the decoding process is executed in the arithmetic decoding section 27 of FIG. 4.
FIG. 5 is a configuration diagram showing the internal structure of the arithmetic coding unit 6 in FIG. 3. In the figure, the arithmetic coding unit 6 includes: a context model determining unit 28, which determines the motion vector 5, coding model information 13, spatial prediction model 14, and orthogonal transform coefficient data 17, which are target data to be coded. The context model defined by the data type (described later); the binarization unit 29 converts n-carry data into binary data according to the binarization rule determined for each encoding target data type; the occurrence probability generation unit 30. Provide the probability of occurrence of the value (0 or 1) of each binarized sequence bin after binarization; the encoding unit 31 performs arithmetic coding based on the generated probability of occurrence; and the transmission unit generating unit 35 notifies the interruption of the arithmetic The sequence of encoding, and at the same time the sequence constitutes the data that becomes the transmission unit.
FIG. 6 is a flowchart showing the processing content of the arithmetic coding unit 6 in FIG. 5.
1) Context model determination processing (step ST1) The context model is to model the dependence relationship with other information that is the cause of the change in the probability of occurrence of the information source (code) symbol, and switch to the corresponding dependence relationship The state of occurrence probability becomes a code that is more suitable for the actual occurrence probability of the symbol.
Fig. 7 is an explanatory diagram explaining the concept of the context model. Also, in Fig. 7, the information source symbol is taken as a binary bit. The selection branch of ctx from 0 to 2 in FIG. 7 is defined by the fact that the probability of occurrence of the information source symbol using the ctx is imagined and changed according to the situation.
With regard to the video image coding in the first embodiment, the value of ctx can be switched according to the dependency between the coded data of a certain macroblock and the coded data of the surrounding macroblocks.
Fig. 8 is an explanatory diagram showing an example of a context model for motion vectors. Fig. 8 is a micro-block disclosed in "Video Compression Using Context-Based Adaptive Arithmetic Coding", International Conference on Image Processing 2001, about D. Marpe and others. Take the context model of the motion vector as an example.
In FIG. 8, the motion vector of the block C is used as the coding target. To be precise, the prediction difference value mvdk(C) of the motion vector of the block C is predicted from the vicinity. And ctx_mvd(C, k) is the context model.
The motion vector prediction difference value displayed in block A as mvdk(A) and the motion vector prediction difference value displayed in block B as mvdk(B) are used in the context model switching evaluation value ek(C) Definition.
The evaluation value ek(C) shows the deviation of the nearby motion vector. Generally speaking, when the deviation is small, mvdk(C) will decrease. Conversely, when ek(C) is large, mvdk(C) will decrease. There is a tendency for mvdk(C) to become larger.
Therefore, the symbol occurrence probability of mvdk(C) is best adapted based on ek(C). The change setting of the occurrence probability is a context model, and it can be said that there are three types of occurrence probability changes in this situation.
In addition, for each of the encoding target data such as the encoding model information 13, the spatial prediction model 14, and the orthogonal transform coefficient data 17, a context model is defined in advance, and the arithmetic encoding unit 6 of the video image encoding device is The arithmetic decoding unit 27 of the video image decoding device is shared. The context model determination unit 28 of the arithmetic coding unit 6 shown in FIG. 5 executes a process of selecting a predetermined model based on the type of the encoding target data.
In addition, since the process of selecting an arbitrary occurrence probability change from the context model is equivalent to the occurrence probability generation process of 3) below, it will be described here.
2) Binary processing (step ST2) The context model is to perform binary serialization of the encoding target data in the binarization unit 29, and is determined based on each bin (binary position) of the binary sequence. The rule of binarization is to convert into a variable-length binary sequence based on the approximate distribution of the acquired values of each coded data. Binaryization may still be performed by performing arithmetic coding on the coding target data originally obtained with n-bits, and coding in units of bins. Since the number of linear divisions of the probability number can be reduced, the calculation can be simplified. Therefore, it has the advantage of making the context model slim.
3) Occurrence probability generation processing (step ST3) In the above-mentioned 1) and 2) processing procedures, the binarization of the target data for multi-value encoding and the setting of the context model for each bin have been completed, and encoding is prepared. Since each context model includes changes to each occurrence probability of 0/1, the occurrence probability generation unit 30 refers to the context model determined in step ST1 to execute the generation process of the occurrence probability of 0/1 in each bin.
Fig. 8 shows an example of the evaluation value ek(C) selected as the probability of occurrence. The occurrence probability generation unit 30 determines the evaluation value selected as the probability of occurrence as shown in ek(C) in Fig. 8, and based on this, Among the selected branches of the referenced context model, it is determined which probability change will be used for the current code.
4) Coding processing (steps ST3 to ST7) Because through 3), the probability of occurrence of each value 0/1 on the straight line of the probability number required by the arithmetic coding processing program can be obtained, so it is based on the processing program cited in the conventional example. Arithmetic coding is performed in the coding unit 31 (step ST4).
Furthermore, the actual code value (0 or 1) 32 is fed back to the occurrence probability generation unit 30, and the occurrence frequency of 0/1 is calculated in order to update the occurrence probability variation part of the context model used (step ST5).
For example, when the encoding process of 100 bins is executed using the occurrence probability change in a certain specific context model, the occurrence probability of 0/1 in the occurrence probability change is 0.25 and 0.75, respectively. Here, as long as 1 is coded with the same occurrence probability change, the occurrence frequency of 1 is updated, and the occurrence probability of 0/1 changes to 0.247 and 0.752. With this mechanism, efficient coding suitable for the actual occurrence probability can be executed.
In addition, the new arithmetic code 33 of the encoded value (0 or 1) 32 generated by the encoding unit 31 is sent to the transmission unit generating unit 35, as described in 6) below, as the arithmetic code 33 constituting the transmission unit Data is multiplexed (step ST6).
Then, it is judged whether or not the encoding process is completed for the entire binary sequence bin of one encoding target data (step ST7), and if it is not completed yet, the process returns to step ST3, and the processes following the generation process of the occurrence probability in each bin are executed. On the other hand, if it is the end, the process proceeds to the transmission unit generation process described next.
5) Transmission unit generation processing (steps ST8 to ST9) Although arithmetic coding is to transform a sequence of multiple encoding target data into one character code, because the video image signal is subjected to motion prediction between frames, or is performed in frame Unit display, so you need to use the frame as a unit to generate a decoded image to update the frame memory. Therefore, it is necessary to be able to clearly determine the so-called frame unit gap on the compressed data that is arithmetic coded. Furthermore, for the purpose of multiplexing with other media such as sound and audio, and packet transmission, it is also necessary to The finer units in the frame distinguish the compressed data for transmission. In this example, a segment structure, that is, a unit in which a plurality of micro-blocks are grouped in a post-scanning order, can generally be cited.
FIG. 9 is an explanatory diagram explaining the segment structure.
The rectangle enclosed by the dotted line is equivalent to a micro block. Generally, the segment structure is handled as a unit of resynchronization during decoding. As an extreme example, in order to map the fragment data into a package for IP transmission as usual. RTP (Real Time Transport Protocol) is mostly used for the IP transmission of real-time media that does not allow transmission delays such as video images. In most cases, the RTP packet provides the time stamp to the header part, and the segment data of the video image is mapped and transmitted in the loading part. For example, in "RTP PayloadFormat for MPEG-4 Audio/Visual Streams" and RFC 3016 of Kikuchi and others, it is stipulated that the compressed data of MPEG-4 video images shall be processed in units of MPEG-4 fragments (video image packets). Mapped to the method of RTP loading.
Because RTP is transmitted as a UDP packet, there is generally no resend control. In the case of packet loss, there may be cases where the fragment data cannot be completely delivered to the decoding device. If the subsequent segment data is to be encoded depending on the information of the discarded segment, it will not be able to be decoded normally even if it is assumed to have been delivered to the decoding device normally.
Therefore, any segment needs to be decoded normally from the beginning without any dependencies. For example, generally speaking, if you encounter encoding that executes Slice5, do not execute encoding that uses the information of the microblock group of Slice3 at the top and Slice4 at the left.
On the other hand, in order to improve the efficiency of arithmetic coding, it is better to adapt it to the probability of occurrence of symbols based on the surrounding conditions, or to continue the division processing program of the probability number straight line. For example, in order to encode Slice5 completely independently of Slice4, when the arithmetic coding of the final microblock of Slice4 ends, the register value of the word code that can be expressed in the arithmetic coding cannot be maintained, but in Slice5, the register is reset to the initial state. After the code is opened again. Therefore, the correlation existing between the end of Slice4 and the beginning of Slice5 cannot be used, resulting in a decrease in coding efficiency. In short, it is generally designed to improve the resistance to unexpected loss of segment data due to transmission errors and the like at the expense of a reduction in coding efficiency.
In the transmission unit generating unit 35 of the first embodiment, a method and an apparatus for improving the adaptability of the design are provided. In other words, when the probability of loss of segment data due to transmission errors or the like is extremely low, it is possible to actively use the segment data without constantly cutting off the dependency relationship between the segments related to arithmetic coding.
On the one hand, when the possibility of fragment data loss is high, the dependency between the fragments can be cut off, and the coding efficiency in the transmission unit can be adaptively controlled.
In short, the transmission unit generation unit 35 in the first embodiment receives the transmission unit instruction signal 36 at a timing that distinguishes the transmission unit as a control signal inside the video image encoding device, and is based on the transmission unit instruction signal 36. The input timing distinguishes the character code of the arithmetic code 33 input from the encoding unit 31 to generate the data of the transmission unit.
Specifically, the transmission unit generation unit 35 multiplexes the arithmetic code 33 of the coded value 32 as transmission unit constituent bits one by one (step ST6), and at the same time, judges that it is only included in the transmission unit through the transmission unit instruction signal 36. If the encoding of the partial data of the obtained macro block is completed (step ST8), if it is determined that all the encodings in the transmission unit are not completed, the process returns to step ST1 and executes the following processing of context model determination.
Conversely, when it is judged that all the encodings in the transmission unit are completed, the transmission unit generating unit 35 adds the following two pieces of information as the header information of the next transmission unit data (step ST9).
1. In the next transmission unit, add the linear division status of the probability number, that is, display whether to reset the "register reset flag" that can represent the register value of the arithmetic coding processing program expressed as a character code. In addition, the register reset flag always instructs <reset> to be set in the transfer unit that is generated first.
2. Only when the register reset flag of 1. above shows <No reset>, it is used as the register value at the beginning of the arithmetic encoding and decoding of the next transmission unit, and it is added as the register value at the beginning of the arithmetic encoding and decoding of the next transmission unit. The "initial register value" of the register value at the moment. In addition, this initial register value is the initial register value 34 input from the encoding unit 31 to the transmission unit generating unit 35 as shown in FIG. 5.
FIG. 10 is an explanatory diagram showing an example of a bit data stream generated by the arithmetic coding unit 6.
As shown in FIG. 10, in addition to the segment start code, the data of each segment of video image compression data and the segment header (referred to as the segment header in the drawing) as the header of each segment of the video image compression data, Setting: the register reset flag of 1. above; and the initial register value. This value is multiplexed only when the register reset flag of 1. above shows <no reset>.
As described above, based on the two additional information, even when the fragment just before is missing, it is initialized by using the register reset flag contained in the header data of the fragment and the register initialization as the initial register value. The value becomes the encoding that maintains the continuity of the arithmetic code even between segments, and the encoding efficiency can be maintained.
Also, in FIG. 10, although the segment header data and the segment video image compression data are multiplexed on the same data stream, as shown in FIG. 11, the segment header data is in the form of another data stream. While being transmitted offline, the segment video image compression data can also be constructed by adding the ID information of the corresponding segment header data. In the same figure, it is shown that the data stream is transmitted according to the IP protocol. It also shows that the header data part is transmitted by the more reliable TCP/IP, and the video image compression data part is transmitted by the low-latency RTP/UDP /IP to transmit example. According to the header and the separate transmission format of the transmission unit based on the configuration of FIG. 11, the data transmitted by RTP/UDP/IP may not necessarily be divided into so-called fragmented data units.
In a segment, basically, although it is necessary to reset all the dependencies (context model) with the video image signal in the nearby area, so that decoding can be performed separately in the segment, but this will lead to video encoding Decrease in efficiency.
As shown in Figure 11, if TCP/IP can be used to transmit the initial register state, the video image signal itself is encoded using each context model in the frame, and it can also be divided during the RTP packetization stage. Arithmetic coded data to be transmitted. Therefore, according to this structure, since the arithmetic coding processing program can be stably obtained regardless of the condition of the line, it is possible to transmit a bit data stream that performs coding that is not restricted by the fragment structure while maintaining high error tolerance. .
In addition, as shown in FIG. 12, the syntax of whether to use the register reset flag and the initial register value may be configured to be displayed in a higher layer. In FIG. 12, it is shown on the header information given in the unit of the sequence of the video image composed of multiple video image frames, and the multiplexed register reset can indicate whether to use the register reset flag and the syntax of the initial register value. Examples of control flags.
For example, when it is judged that the quality of the circuit is deteriorated and the register reset is performed through the video image sequence to enable stable video image transmission, the register reset control flag is set to indicate < through the video image sequence, and The beginning of the fragment always resets the value of the register >. At this time, the multiplexing at the slice level of the register reset flag and the initial register value that are multiplexed in units of slices becomes unnecessary.
Therefore, when a specific transmission condition (error rate of the line, etc.) is continued, if the register reset can be controlled in units of video image sequence, the overhead information transmitted in units of fragments can be reduced. Needless to say, the register reset control flag can also be represented by the Nth frame, the N+1th frame, etc., to add header information of any video image frame in the video image sequence.
FIG. 13 is a configuration diagram showing the internal configuration of the arithmetic decoding unit 27 in FIG. 4.
The arithmetic decoding unit 27 of the video image decoding device includes: a transmission unit decoding initialization unit 37, which performs the initialization of the arithmetic decoding process based on the additional information about the arithmetic coding process program contained in the header for each received transmission unit; The context model determination unit 28 specifies the form of the decoding target data such as the motion vector 5, the coding model information 13, the spatial prediction model 14, and the orthogonal transform coefficient data 17, based on the processing program of arithmetic decoding, and determines the respective formats of the decoding target data and the video image. A context model commonly defined by the encoding device; the binarization unit 29 generates a binarization rule determined based on the form of the decoding target data; the occurrence probability generation unit 30 provides each based on the binarization rule and the context model The probability of occurrence of bin (0 or 1); and the decoding unit 38 performs arithmetic decoding based on the generated probability of occurrence, and decodes the motion vector from the resultant binary sequence and the above-mentioned binarization rule. 5. Coding model information 13. Data such as spatial prediction model 14, orthogonal transform coefficient data 17, etc.
FIG. 14 is a flowchart showing the processing content of the arithmetic decoding unit 27 in FIG. 13.
6) Transmission unit decoding initialization processing (step ST10) As shown in FIG. 10, based on the register reset flag and the initial register value 34, the initialization of the arithmetic decoding start state in the decoding unit 38 is executed (step ST10). The register reset flag indicates that it is multiplexed in the transmission unit of each segment, etc., and shows whether the register value of the arithmetic coding processing program is reset; and when the register value is reset, the initial register value 34 is not used.
7) Context model determination processing, binarization processing, and occurrence probability generation processing Although these processing procedures are respectively executed by the context model determination unit 28, the binarization unit 29, and the occurrence probability generation unit 30 shown in FIG. 13 , But because it is the same as the context model determination process ST1, the binarization process ST2, and the occurrence probability generation process ST3 shown in the processing procedures 1) to 3) on the video image encoding device side, the same step numbers are provided respectively, These descriptions are omitted.
8) Arithmetic decoding processing (step ST11) Since the probability of occurrence of the bin to be decoded from this point has been determined by the processing procedures up to 7), in the decoding unit 38, according to the arithmetic decoding processing shown in the conventional example The program restores the value of bin (step ST11), counts the frequency of occurrence of 0/1 and updates the probability of occurrence of bin (step ST5) in the same way as the processing on the video image encoding device side. Whether the value of bin decoded by comparing the predetermined binary sequence patterns is confirmed (step ST12).
If the value of the decoded bin is indeterminate compared with the binary sequence pattern determined by the binarization rule, the processing following the 0/1 occurrence probability generation processing in each bin of step ST3 is executed again (step ST3, ST11, ST5, ST12).
On the one hand, when it is confirmed that the binary sequence pattern determined by the binarization rule matches and the value of each decoded bin is determined, the data value indicated by the matched pattern is output as the decoded data value. If all the transmission units such as segments have not been decoded yet (step ST13), in order to decode all the transmission units, it is necessary to repeatedly execute the processing following the context model determination processing of step ST1.
It is obvious from the above that according to the first embodiment, when video image compression data is transmitted by dividing the transmission unit into a segment, etc., the addition can be expressed as the segment header data and the arithmetic coding processing program is displayed. The register value has a reset flag for resetting and the initial register value 34, so coding can be performed without cutting off the continuity of the coding process of arithmetic coding, and the coding efficiency can be improved while the resistance to transmission errors is improved. , Making its decoding feasible.
In addition, in the first embodiment, although the segment structure is assumed as the transmission unit, the present invention can be applied even if a video image frame is used as the transmission unit.
Second embodiment. In the second embodiment, another aspect of the arithmetic coding unit 6 and the arithmetic decoding unit 27 will be described. In the second embodiment, it is characterized in that not only the register value indicating the state of the character code of the arithmetic coding processing program, but also the learning state of the occurrence probability change in the context model, that is, the comparison of the bin in the occurrence probability generation unit 30 The learning state of the occurrence probability change in the context model from the occurrence probability update processing of is also reused in the segment header.
For example, in FIG. 8 described in the first embodiment, in order to improve the efficiency of arithmetic coding of block C, for example, the information of the motion vector of block B located in the upper part of block C is determined as the occurrence probability change. To use. Therefore, for example, assuming that the block C and the block B are located in different segments, it is necessary to prohibit the use of the information of the block B in the occurrence probability determination processing program.
This situation means that the coding efficiency of adaptation to the probability of occurrence based on the context model will be reduced.
Therefore, in the second embodiment, since the method and apparatus for improving the adaptability of the design are provided, in the case where the probability of loss of segment data due to transmission errors and the like is extremely low, it is not always possible to cut off the arithmetic coding The inter-fragment dependency relationship can be actively used. In addition, when the possibility of fragment data loss is high, the inter-fragment dependency relationship can be cut off, and the coding efficiency of the transmission unit can be adaptively controlled.
Fig. 15 is a structural diagram showing the internal structure of the arithmetic coding unit 6 in the second embodiment.
The arithmetic coding unit 6 in this second embodiment is different from the arithmetic coding unit 6 in the first embodiment shown in FIG. 5 only in that the occurrence probability generation unit 30 will be used as the context of the multiplexed segment header. The state 39 of the model is passed to the transmission unit generation unit 35.
FIG. 16 is a flowchart showing the processing content of the arithmetic coding unit 6 in FIG. 15.
Compared with the flowchart of FIG. 6 in the first embodiment, it is obvious that the difference is that the context model state 39 of the 0/1 occurrence probability generation process in each bin of step ST3 is generated based on the occurrence probability. The learning state 39 of the occurrence probability change in the context model derived from the occurrence probability update processing of the bin of the unit 30 is also the same as the register value of the binary arithmetic coding processing in step ST4, and it is only the transmission unit generation unit in step ST9. The header of the sub-transmission unit in 35 constitutes only a point that is reused in the segment header in processing.
FIG. 17 is an explanatory diagram for explaining the learning state of the context model. Use Fig. 17 to explain the meaning of the state 39 of the context model.
Figure 17 is a case where there are n microblocks in the k-th transmission unit, and for each microblock, a context model ctx used for only 1 degree is defined, and the occurrence probability of each microblock ctx changes. .
The so-called state 39 of the context model continues to the next transmission unit, which means that the final state of the k-th transmission unit ctxk(n-1) is the ctx of the k+1-th transmission unit as shown in FIG. The initial state of ctxk+1(n-1)=0,1,2, the probability of occurrence of 0, 1 po, p1 and the value of 0, 1 at ctxk(n-1)=0, 1, 2 The probability of occurrence po and p are equal. Therefore, in the transmission unit generating unit 35, the data showing the state of ctxk(n-1) is transmitted as part of the header information in the k+1th transmission unit.
Fig. 18 is an explanatory diagram showing an example of a bit stream generated by the arithmetic coding unit 6 of the second embodiment.
In this second embodiment, the same segment start code, register reset flag, and initial register value as in the first embodiment shown in FIG. 10 are added to the segment header data of each segment of video image compression data, and the same as the previous one. Information about the status of the context model of the fragment.
However, in the second embodiment, not only the register reset flag is made to include the presence or absence of multiplexing of the initial register value, but also the presence or absence of multiplexing of the context model state data.
Also, as information indicating whether the context model status data is multiplexed or not, it is needless to say that not only the register reset flag can be set, but also other flags can be set to form it.
Moreover, although the above-mentioned first embodiment can be described, in FIG. 18, although the segment header data and the segment video image compression data are multiplexed on the same data stream, the segment header is a separate data stream. The shape is transmitted on the line, and the compressed data can also be constructed by attaching the ID information of the corresponding fragment header data.
Fig. 19 is a structural diagram showing the internal structure of the arithmetic decoding unit 27 of the second embodiment. The difference between the arithmetic decoding unit 27 of the second embodiment and the arithmetic decoding unit 27 of the first embodiment shown in FIG. 13 is that the transmission unit decoding initialization unit 37 converts the text of the segment just before multiplexed by the segment header. The state 39 of the context model is passed to the occurrence probability generation unit 30, and becomes a point of constitution that continues the state of the context model from the segment just before.
FIG. 20 is a flowchart showing the processing content of the arithmetic decoding unit 27 of FIG. 19.
Compared with the flowchart of FIG. 14 in the first embodiment, it is obvious that the difference is that in the decoding initialization process of each transmission unit in step ST10, the context model state decoded from the segment header is 39 Refer to the process of step ST3, that is, the process of outputting the context model determined in step ST1 to the process of generating the probability of occurrence of 0/1 in each bin, which is used for the probability of occurrence of 0/1 of the occurrence probability generating unit 30 The point of generation processing.
Also, regarding the status of the context model delivered to the fragment header, it becomes the overhead of the fragment header when the number of context models is extremely large, so it is also possible to select a context model that significantly contributes to coding efficiency. , The state is multiplexed to form it.
For example, since the motion vector and orthogonal transform coefficient data account for a large proportion of the total number of symbols, it is possible to consider a configuration that continues only the state of the context model. Furthermore, the type of context model of the continuation state can be explicitly reused in the bit data stream to construct it, or the state continuation can be selectively executed for only important context models according to the local conditions of the video image.
It is obvious from the above that according to the second embodiment, when the video image compression data is divided into smaller transmission units to transmit video image compression data, because it can be added: as the segment header data, it represents the register that displays the arithmetic coding processing program The register reset flag for whether the value is reset; the initial register value 34; and the information indicating the context model state of the fragment just before, and the coding can be performed without cutting off the continuity of the arithmetic coding coding processing program, which can be improved Error resistance to transmission errors while maintaining coding efficiency.
In addition, although the segment structure is assumed as the transmission unit in the second embodiment, the present invention can be applied even if a video image frame is used as the transmission unit.
In particular, in the second embodiment, information indicating the context model state of the segment immediately before is added. Therefore, for example, in FIG. 8, even if the block B located at the block C and the block B immediately before the block C is different For fragments, the occurrence probability determination processing program of block C can also use the context model state of block B, which can improve the coding efficiency of the adaptation of the occurrence probability based on the context model. In short, according to transmission errors, etc., when the probability of loss of segment data is extremely low, it is not necessary to constantly cut off the dependency relationship between the segments related to arithmetic coding, and it can be positively added until the context model state of the previous segment. In addition, when the possibility of fragment data loss is high, the context model state of the immediately preceding fragment is not used, and the dependency between the fragments is cut off, which becomes an adaptive control of the coding efficiency of the transmission unit.
Also, in the case of the second embodiment, although it has been explained that the bit stream syntax shown in FIG. 18 is used for each segment data, the addition of the register reset flag and the initial register value in the first embodiment is parallel. However, the information indicating the context model status of each data of the immediately preceding segment is added as the segment header data. However, the register reset flag and initial register value of the above-mentioned first embodiment are not added and omitted. Information that only indicates the context model state of each data of the immediately preceding segment can be added as segment header data, and it has nothing to do with whether it is set in parallel with the addition of the register reset flag and initial register value of the first embodiment. Needless to say, even if the context model state reset flag is OFF, that is, information indicating the context model state of each data just before the fragment is added only when the reset is not performed to set the context model state reset flag (refer to Figure 21), but it can also be used for decoding.
Third Embodiment In the third embodiment, a description will be given of an example in which transmission units are separately grouped in the form of encoded data and constituted by a data division format.
For example, take the video image coding method design draft published in the Joint Video Team (JVT) of ISO/IEC MPEG and ITU-T VCEG as an example to show the data classification method of Working DraftNumber2, Revision3, and JVT-B118r3. The segment structure shown in FIG. 9 is a unit, and the data unit constructed by grouping data of a specific form is a method in which the number of microblocks existing only in the segment data is transmitted in the form of segment data. The data format of segment data formed as a data unit formed by grouping includes, for example, the following data formats of 0-7.
0 TYPE_HEADER image (frame) or segment header 1 TYPE_MBHEADER microblock header information (coding model information, etc.) 2 TYPE_MVD motion vector 3 TYPE_CBP CBP (effective orthogonal transform coefficient distribution in the microblock) 4 TYPE-2x2DC orthogonal transform coefficient data (1) 5 TYPE_COEFF_Y Orthogonal transform coefficient data (2) 6 TYPE_COEFF_C Orthogonal transform coefficient data (3) 7 TYPE_EOS Data stream end identification information For example, in the TYPE_MVD segment of data format 2, only the internal micro data will be collected. The block number part and the data of the motion vector information are transmitted as segment data.
Therefore, when the TYPE_MVD data of the k+1 segment is decoded after the TYPE_MVD data of the k-th segment, if only the state of the context model of the motion vector at the end of the k-th segment is pre-determined Multiplexed as the header of the segment that sends the TYPE_MVD data of the k+1th segment, it can continue to be used in the context model learning state of the arithmetic coding of the motion vector.
Fig. 21 is an explanatory diagram showing an example of a bit stream generated by the arithmetic coding unit 6 of the third embodiment. In FIG. 21, for example, when the motion vector in the case of a TYPE_MVD segment of data format 2 is multiplexed as segment data, the segment header is added with a segment start code and a data format indicating TYPE_MVD ID, the context model state reset flag, and information indicating the state of the context model for the motion vector of the just preceding segment.
Also, for example, when only the orthogonal transform coefficient data (2) of the orthogonal transform coefficient data (2) of the TYPE_COEFF_Y of the data format 5 is multiplexed as segment data, the segment header is appended The segment start code and the data form ID indicating TYPE_COEFF_Y, the context model status reset flag, and the information indicating the context model status for the orthogonal transform coefficient data of the segment just before.
Also, in the same figure, although the fragment header data and compressed data are multiplexed on the same data stream, the fragment header is transmitted online in the form of another data stream. In the compressed data, it can also be attached. The ID information of the corresponding segment header data is constructed.
In addition, in the arithmetic coding unit 6 of the third embodiment, the transmission unit generation unit 35 executes the reconstruction of the macro block data in the segment according to the rules of the above-mentioned data classification method in the configuration of FIG. The ID information of the type of the form and the learning state of the context model corresponding to each data form are multiplexed and constituted.
In addition, in the arithmetic decoding unit 27 in the third embodiment, in the configuration of FIG. 19, the transmission unit decoding initialization unit 37 notifies the context model determination unit 28 of the data type type ID multiplexed in the segment header. The context model to be used may be determined, and the context model learning state may be notified to the occurrence probability generation unit 30, and the context model learning state 39 may be continued between segments to perform arithmetic decoding.
It is obvious from the above that according to the third embodiment, even when the video image signal is divided into transmission units grouped in a predetermined data format and compression coding is performed, the data belonging to the transmission unit will be When the video image signal is arithmetic-coded, since the symbol occurrence probability learning state of the transmission unit grouped in the previous predetermined data format is not reset and the encoding is continued, even if it is grouped in the predetermined data format In this case, it is also possible to implement coding that improves the coding efficiency of arithmetic coding while ensuring error tolerance.
In addition, in the third embodiment, although the data format type of each segment structure is exemplified as the transmission unit, the transmission of each data type type in the unit of the video image frame can also be applied. this invention.
In addition, the third embodiment shown in FIG. 21 illustrates an example of bit data stream syntax. In each segment data of each data format, a context model state reset flag is added, and the flag indicates that it is just when it is OFF. The information of the context model status of each data of the previous fragment is used as the fragment header data. However, as in the case of the example of the bit stream syntax of the second embodiment shown in FIG. 18, in the fragment data of each data format, It can also be used in parallel with the addition of the register reset flag and the initial register value, and the context model state reset flag and the context model state information of each data of the immediately preceding fragment when the flag is OFF are used as the fragment header data. It does not matter whether it is set in parallel with the addition of the register reset flag and the initial register value. Needless to say, the context model state reset flag can be omitted, and the context model representing each data of the fragment just before is always added. The status information can also be used for decoding.
In addition, although in the above first to third embodiments, the video image data is taken as an example as a digital signal, the present invention is not limited to this, and not only the digital signal of the video image data, but also the digital signal Audio digital signals, still image digital signals, text digital signals, and multimedia digital signals in which these signals are arbitrarily combined are also applicable.
Furthermore, although in the above first and second embodiments, segments are cited as the transmission unit of digital signals, in the third embodiment, in addition to the form of data and the division of data types, etc. within the segment are listed. The predetermined transmission unit has been described as an example, but in the present invention, it is not limited to this. It is also possible to use one picture (picture) formed by collecting a plurality of fragments, that is, one video image frame unit as the predetermined transmission unit, and Assuming the use of storage systems other than communications, it goes without saying that not only the predetermined transmission unit, but also the predetermined storage unit is also possible.
As described above, the digital signal encoding device according to the present invention is suitable for applications where it is necessary to ensure error tolerance and improve the encoding efficiency of arithmetic encoding when compressing video image signals for transmission.
19 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| CN106445890A | Cited by | China | Search report |
96 members in 11 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 1241142002 | Japan | – | |
| 2002124114 | Japan | A | |
| 2002124114 | Japan | A | |
| 0304578 | Japan | W | |
| 0304578 | Japan | W | |
| 1241142002 | – | – | – |
| JP20020124114 | – | – | – |
| WO2003JP04578 | – | – | – |
Members96
| Document | Office | Kind | |
|---|---|---|---|
| TW200306118A | Taiwan Province of China | A | |
| CA2449924A1 | Canada | A1 | |
| CA2554143A1 | Canada | A1 | |
| CA2632408A1 | Canada | A1 | |
| CA2685312A1 | Canada | A1 | |
| CA2686438A1 | Canada | A1 | |
| CA2686449A1 | Canada | A1 | |
| CA2756577A1 | Canada | A1 | |
| CA2756676A1 | Canada | A1 | |
| CA2807566A1 | Canada | A1 | |
| CA2809277A1 | Canada | A1 | |
| WO03092168A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003236062A1 | Australia | A1 | |
| KR20040019010A | Republic of Korea | A | |
| EP1422828A1 | European Patent Office (EPO) | A1 | |
| US2004151252A1 | United States of America | A1 | |
| CN1522497AThis record | China | A | |
| TWI222834B | Taiwan Province of China | B | |
| JP2005347780A | Japan | A | |
| KR20050122288A | Republic of Korea | A | |
| EP1422828A4 | European Patent Office (EPO) | A4 | |
| US2006109149A1 | United States of America | A1 | |
| KR100585901B1 | Republic of Korea | B1 | |
| JP3807342B2 | Japan | B2 | |
| US7095344B2 | United States of America | B2 | |
| EP1699137A2 | European Patent Office (EPO) | A2 | |
| CN1878002A | China | A | |
| KR100740381B1 | Republic of Korea | B1 | |
| US2007205927A1 | United States of America | A1 | |
| CN101060622A | China | A | |
| US2007263723A1 | United States of America | A1 | |
| US7321323B2 | United States of America | B2 | |
| CA2449924C | Canada | C | |
| US7388526B2 | United States of America | B2 | |
| EP1699137A3 | European Patent Office (EPO) | A3 | |
| US2008158027A1 | United States of America | A1 | |
| US7408488B2 | United States of America | B2 | |
| SG147308A1 | Singapore | A1 | |
| US7518537B2 | United States of America | B2 | |
| US2009153378A1 | United States of America | A1 | |
| CA2554143C | Canada | C | |
| CN100566179C | China | C | |
| CN101626244A | China | A | |
| CN101626245A | China | A | |
| CN101686059A | China | A | |
| CN1522497B | China | B | |
| SG158846A1 | Singapore | A1 | |
| SG158847A1 | Singapore | A1 | |
| CN101815217A | China | A | |
| USRE41729E | United States of America | E | |
| CN101841710A | China | A | |
| HK1140324A1 | Hong Kong, China | A1 | |
| US2010315270A1 | United States of America | A1 | |
| US7859438B2 | United States of America | B2 | |
| EP2288034A1 | European Patent Office (EPO) | A1 | |
| EP2288035A1 | European Patent Office (EPO) | A1 | |
| EP2288036A1 | European Patent Office (EPO) | A1 | |
| EP2288037A1 | European Patent Office (EPO) | A1 | |
| HK1144632A1 | Hong Kong, China | A1 | |
| EP2293450A1 | European Patent Office (EPO) | A1 | |
| EP2293451A1 | European Patent Office (EPO) | A1 | |
| EP2306651A1 | European Patent Office (EPO) | A1 | |
| US7928869B2 | United States of America | B2 | |
| EP2315359A1 | European Patent Office (EPO) | A1 | |
| US2011095922A1 | United States of America | A1 | |
| US2011102210A1 | United States of America | A1 | |
| US2011102213A1 | United States of America | A1 | |
| US2011115656A1 | United States of America | A1 | |
| CA2632408C | Canada | C | |
| US2011148674A1 | United States of America | A1 | |
| US7994951B2 | United States of America | B2 | |
| US8094049B2 | United States of America | B2 | |
| US2012044099A1 | United States of America | A1 | |
| SG177782A1 | Singapore | A1 | |
| CN101815217B | China | B | |
| US8188895B2 | United States of America | B2 | |
| SG180068A1 | Singapore | A1 | |
| CN101626244B | China | B | |
| US8203470B2 | United States of America | B2 | |
| CN101626245B | China | B | |
| CN101841710B | China | B | |
| US8354946B2 | United States of America | B2 | |
| SG186521A1 | Singapore | A1 | |
| SG187281A1 | Singapore | A1 | |
| SG187282A1 | Singapore | A1 | |
| CN101686059B | China | B | |
| SG190454A1 | Singapore | A1 | |
| CN101060622B | China | B | |
| CA2686438C | Canada | C | |
| CA2686449C | Canada | C | |
| CA2685312C | Canada | C | |
| CA2756577C | Canada | C | |
| CA2756676C | Canada | C | |
| CA2807566C | Canada | C | |
| CA2809277C | Canada | C | |
| US8604950B2 | United States of America | B2 |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Expiry of patent termCX01 | CX01 | |
| Transfer of patent rightTR01 | TR01 | |
| Grant of patent or utility modelGrantedC14 | C14 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 1522497
- Publication, DOCDB
- 1522497
- Publication, EPODOC
- CN1522497
- Application
- 38005166
- Application, DOCDB
- 03800516
- Application, EPODOC
- CN20038000516
Titles3
- Chinese
- 数字信号编码装置、数字信号解码装置、数字信号算术编码方法及数字信号算术解码方法
- English
- Digital signal coding device, digital signal decoding device, digital signal arithmetic coding method and digital signal arithmetic decoding method
- Chinese
- 数字信号编码装置、数字信号解码装置、 数字信号算术编码方法及数字信号算术解码方法
Classification
- CPC, 17
- H03M7/40
- H04N7/52
- H03M7/4006
- H04N21/2381
- H04N21/4363
- H04N21/4381
- H04N19/105
- H04N19/107
- H04N19/124
- H04N19/13
- H04N19/137
- H04N19/152
- H04N19/174
- H04N19/176
- H04N19/46
- H04N19/61
- H04N19/70
- IPC, 14
- G06T9 00
- H03M7 40
- H04N7 24
- H04N7 52
- H04N19 00
- H04N19 105
- H04N19 13
- H04N19 174
- H04N19 46
- H04N19 51
- H04N19 593
- H04N19 625
- H04N19 70
- H04N19 91