Computational complexity and precision control in transform-based digital media codec
Abstract
digital media decoding method, digital media encoding method, computer readable storage media, decoding device and encoding device the present invention relates to a digital media decoding method, a digital media encoding method , and a digital media encoder / decoder that includes signaling in various ways related to computational complexity and decoding accuracy. the encoder can send a syntax element indicating arithmetic precision (for example, using 16 or 32 bit operations) of the transform operations performed in decoding. the encoder can signal whether or not scaling is applied to the decoder output, which allows a wider dynamic range of intermediate data, but adds to computational complexity due to the scaling operation.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
11 claims: 5 independent, 6 dependent
- 1Digital media decoding method characterized by the fact that it comprises the steps of:1. Método de decodificação de mídia digital caracterizado pelo fato de que compreende as etapas de: receber um fluxo de bits de mídia digital comprimido em um decodificador de mídia digital;receiving a bit stream of compressed digital media into a digital media decoder;analisar um primeiro elemento sintático e um segundo elemento sintático a partir do fluxo de bits de mídia digital comprimido o primeiro elemento sintático sinaliza um grau de precisão aritmética a ser usado para cálculos de transformada durante processamento de dados de mídia digital;e o segundo elemento sintático sinaliza se uma escala deve ser usada no decodificador de mídia digital;analyzing a first syntactic element and a second syntactic element from the bitstream of compressed digital media the first syntactic element signals a degree of arithmetic precision to be used for transform calculations during digital media data processing;and the second syntactic element signals whether a scale should be used in the digital media decoder;decode the bit stream of compressed digital media, including the application of a two-stage reverse overlap transform (ILT) that includes a first-stage reverse transform that regenerates the DC coefficients followed by a second-stage reverse transform applied to the DC and AC coefficients decoded from the bit stream of compressed digital media and using the precision and the signed arithmetic scale, where the scale includes a rounded division of the output in the digital media decoder;and issue a reconstructed image. decodificar o fluxo de bits da mídia digital compactada, incluindo a aplicação de uma transformada sobreposta inversa em dois estágios (ILT) que inclui uma transformada inversa de primeiro estágio que regenera os coeficientes DC seguidos por uma transformada inversa de segundo estágio aplicada aos coeficientes DC e coeficientes AC decodificados a partir do fluxo de bits de mídia digital compactada e que usa a precisão e a escala aritmética sinalizada, em que a escala inclui uma divisão arredondada da saída no decodificador de mídia digital;e emitir uma imagem reconstruída.
- 6Digital media encoding method characterized by the fact that it comprises:6. Método de codificação de mídia digital caracterizado pelo fato de que compreende: receber dados de mídia digital em um codificador de mídia digital;receive digital media data in a digital media encoder;make a decision whether to use the lowest precision arithmetic for transform calculations when processing digital media data;tomar uma decisão se usa a aritmética de precisão mais baixa para cálculos de transformada durante processamento de dados de mídia digital;represent the decision if the lowest precision arithmetic is used for transform calculations with a first syntactic element in an encoded bit stream, where the syntactic element is operable to communicate the decision to a digital media decoder representar a decisão se usa a aritmética de precisão mais baixa para cálculos de transformada com um primeiro elemento sintático em um fluxo de bits codificado, em que o elemento sintático é operável para comunicar a decisão para um decodificador de mídia digital Petition 870190108373, of 10/25/2019, p. 37/43 Petição 870190108373, de 25/10/2019, pág. 37/43 3/4 and where the lowest precision arithmetic is a 16-bit arithmetic precision;3/4 e em que a aritmética precisão mais baixa é uma precisão aritmética de 16 bits;making a decision whether to scale the digital media input data before transforming the encoding;tomar uma decisão se aplica uma escala dos dados de entrada de mídia digital antes de transformar a codificação;representar a decisão se aplica a escala com um segundo elemento sintático no fluxo de bits codificado;representing the decision is scaled with a second syntactic element in the encoded bit stream;codificar os dados de mídia digital recebidos em um fluxo de bits compactado, incluindo a aplicação de uma transformada sobreposta direta de dois estágios (LT) que inclui uma transformada de primeiro estágio seguida por uma transformada de segundo estágio nos coeficientes DC a partir da transformada de primeiro estágio e que usa a precisão e escala aritmética decidida, em que a escala inclui uma pré-multiplicação dos dados de mídia digital de entrada;e gerar o fluxo de bits codificado. encode the received digital media data into a compressed bit stream, including the application of a direct two-stage overlap (LT) transform that includes a first stage transform followed by a second stage transform in the DC coefficients from the transform of first stage and that uses precision and decided arithmetic scale, where the scale includes a pre-multiplication of the input digital media data;and generate the encoded bit stream.
- 9Computer-readable storage media characterized by the fact that it comprises the method according to claims 1 to 8. 9. Mídia de armazenamento legível por computador caracterizado pelo fato de que compreende o método conforme as reivindicações 1 a 8. Petition 870190108373, of 10/25/2019, p. 38/43 Petição 870190108373, de 25/10/2019, pág. 38/43 4/4 4/4
Independent claims5
235 paragraphs in 1 section, as filed
Invention Patent Descriptive Report for DIGITAL MEDIA DECODING METHOD, DIGITAL MEDIA CODING METHOD, COMPUTER-READABLE STORAGE MEDIA, DECODING DEVICE AND CODING DEVICE.
BACKGROUND
Block transform based coding
[001] Transform encoding is a compression technique used in many digital media compression systems (for example, audio, image and video). Uncompressed digital image and video are typically represented or captured as samples of image elements or colors at locations in an image or video frame arranged on a two-dimensional (2D) network. This is referred to as a spatial domain representation of the image or video. For example, a typical format for images consists of a sample stream of the 24-bit color image element arranged as a network. Each sample is a number representing color components at a pixel location on the network within a color space, such as RGB, or YIQ, among others. Various image and video systems can use several different colors, time resolutions and sampling space. Similarly, digital audio is typically represented as a stream of audio signal sampled over time. For example, a typical audio format consists of a stream of 16-bit amplitude samples of an audio signal taken at regular time intervals.
[002] Digital audio, image and video signals can assume considerable transmission and storage capacity. Transform encoding reduces the size of digital audio, images and video by transforming the spatial domain representation of the signal into a frequency domain representation (or other domain of
Petition 870190108373, of 10/25/2019, p. 5/43
2/31 similar transform), and thus reducing the resolution of certain frequency components generally less noticeable from the transform domain representation. This generally produces much less noticeable degradation of the digital signal compared to reducing the color or spatial resolution of images or video in the spatial domain, or audio in the time domain.
[003] More specifically, an encoder / decoder system based on a typical block transform 100 (also called a codec) shown in Figure 1 divides the pixels of the uncompressed digital image into two-dimensional blocks of fixed size (Xi, Xn), each possibly overlapping the other blocks. A linear transform 120 to 121 that does spatial frequency analysis is applied to each block, which converts the spaced samples within the block to a set of frequency coefficients (or transform) usually representing the digital signal strength in corresponding frequency bands over the block range. For understanding, transform coefficients can be selectively quantized 130 (that is, reduced in resolution, such as dropping less significant bits of the coefficient values or otherwise mapping values into a set of higher resolution number at a lower resolution) , and also encoded in variable length or entropy 130 in a compressed data stream. In decoding, the transform coefficients inversely transform 170 to 171 to almost reconstruct the original spatial / color sample of the image / video signal (reconstructed blocks Xi, Xn).
[004] The transform of block 120 to 121 can be defined as a mathematical operation in a vector x of size N. More often, the operation is a linear multiplication, producing the output of transform domain y = Mx, M being the matrix of transformed. When the input data is arbitrarily long, it
Petition 870190108373, of 10/25/2019, p. 6/43
3/31 is segmented into vectors of size N and the block transform is applied to each segment. For the purpose of data compression, reversible block transforms are shown. In other words, the matrix M is invertible. In multiple dimensions (for example, for image and video), block transforms are typically implemented as separable operations. Matrix multiplication is applied separately across each dimension of the data (that is, both rows and columns).
[005] For understanding, the transform coefficients (vector components y) can be selectively quantized (ie reduced in resolution, such as abandoning less significant bits of the coefficient values or otherwise mapping values in a set of numbers higher resolution to lower resolution), and also encoded in variable length or entropy in a compressed data stream.
[006] When decoding at decoder 150, the inverse of these operations (dequantization / entropy decoding 160 and inverse block transform 170-171) is applied to the decoder side 150, as shown in Figure 1. While reconstructing the data, the matrix reverse M '<sup>1 </sup>(inverse transform 170 to 171) is applied as a multiplier to the transform domain data. When applied to transform domain data, the inverse transform almost reconstructs digital media from spatial domain or frequency domain.
[007] In many coding applications based on the block transform, the transform is reversibly desirable to support both lossy and lossless compressions depending on the quantization factor. Without quantization (usually represented as a quantization factor of 1) for example, a codec using a reversible transform can exactly reproduce the input data in the decoding. However, the request for re
Petition 870190108373, of 10/25/2019, p. 7/43
4/31 versatility in these applications restricts the choice of transforms by which the codec can be designated.
[008] Many video and image compression systems, such as MPEG Media and Windows, among others, use transforms based on the Discrete Cosine Transform (DCT). DCT is known to have favorable energy-condensing properties that result in almost optimal data compression. In these compression systems, the inverse DCT (IDCT) is used in the reconstruction loops in both the encoder and the decoder of the compression system for individual image reconstruction blocks. Quantization
[009] Quantization is the primary mechanism for most image and video codecs to control compressed image quality and compression radius. According to a possible definition, quantization is a term used a non-reversible approximation mapping function commonly used for lossy compression, in which there is a specified set of possible output values, and each member of the set of possible output values has an associated set of input values that result in the selection of that specific output value. A variety of quantization techniques have been developed, including vector or scalar quantization, uniform or non-uniform, with or without a death zone, and adaptive or non-adaptive.
[0010] The quantization operation is essentially a division polarized by a quantization parameter QP that is performed on the encoder. The multiplication or inverse quantization operation is a multiplication by the QP performed in the decoder. These processes together introduce a loss in the original transform coefficient data, which appears as errors or compression artifacts in the decoded image.
Petition 870190108373, of 10/25/2019, p. 8/43
5/31
summary
[0011] The following detailed description presents tools and techniques for controlling computational complexity and decoding accuracy with a digital media codec. In one aspect of the techniques, the encoded signals one of the scale or non-scale precision modes to be used in the decoder. In scale precision mode, the input image is pre-multiplied (for example, by 8) in the encoder. The output at the decoder is also scaled through rounding division. In a non-scale precision mode, no such scale operations are applied. In non-scale precision mode, the encoder and decoder can die with a lower dynamic range for transform coefficient, and thus have less computational complexity.
[0012] In one aspect of the techniques, the codec can also signal the precision required to perform transform operations for the decoder. In an implementation, an element of the bitstream syntax signals employs a less accurate arithmetic operation for the transform in the decoder.
[0013] This summary is provided to introduce a selection of concepts in a simplified way which is further described below in the Detailed Description. This summary is not intended to identify key characteristics or essential characteristics of the claimed subject material, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Additional features and advantages of the invention will be made apparent from the following detailed description of the modalities that proceed with reference to the attached drawings.
Brief Description of Drawings
[0014] Figure 1 is the block diagram of a codec based on the conventional block transform in the prior art.
Petition 870190108373, of 10/25/2019, p. 9/43
6/31
[0015] Figure 2 is a flow diagram of a representative encoder incorporating the standard block encoding.
[0016] Figure 3 is a flow diagram of a representative decoder incorporating the standard block encoding.
[0017] Figure 4 is a diagram of the inverse wound transform including a core transform and post-filter operation (overlay) in an implementation of the encoder / decode representative of Figures 2 and 3.
[0018] Figure 5 is a diagram of identification of the input data points for transform operations.
[0019] Figure 6 is a block diagram of a suitable computing environment to implement the media encoder / decoder of Figures 2 and 3.
Detailed Description
[0020] The following description refers to the techniques for controlling computational precision and complexity of a transform-based digital media codec. The following description describes an exemplary implementation of the techniques in the context of a digital media compression system or codec. The digital media system encodes digital media data in a compressed form for transmission or storage, and decodes the data for playback or other processing. For the sake of illustration, this exemplary compression system incorporating computational complexity and precision control is a video or image compression system. Alternatively, the techniques can also be incorporated into compression systems or codecs for other digital media data. Computational complexity techniques and precision control do not require the digital media compression system to encode compressed digital media data in a specific encoding format.
Petition 870190108373, of 10/25/2019, p. 10/43
7/31
1. Encoder I Decoder
[0021] Figures 2 and 3 are a generalized diagram of the process used in an encoder 200 and a decoder 300 of representative two-dimensional (2D) data. The diagrams present a generalized or simplified illustration of a compression system incorporating the 2D data encoder and decoder that implement compression using computational complexity and precision control techniques. In alternative compression systems that use control techniques, the processes are additional or smaller than those illustrated in this representative encoder and decoder can be used for the compression of 2D data. For example, some encoders / decoders may also include color conversion, color formats, scalable encoding, lossless encoding, macroblock modes, etc. The compression system (encoder and decoder) can provide lossy and / or lossless compression of 2D data, depending on the quantization that can be based on a quantization parameter ranging from lossless to lossy.
[0022] The 2D data encoder 200 produces a compressed bit stream 220 which is a more compact representation (for typical input) of 2D data 210 presented as input to the encoder. For example, the input of 2D data can be an image, a frame of a video sequence, or other data having two dimensions. The 2D data encoder divides a frame of the input data into blocks (usually illustrated in Figure 2 as partitioning 230), which in the illustrated implementation are non-overlapping 4x4 pixel blocks that form a regular pattern across the plane of the frame. These blocks are grouped into clusters, called macroblocks, which are 16<sup>χ</sup> 16 pixels in size in this representative encoder. In turn, macroblocks are grouped into structured
Petition 870190108373, of 10/25/2019, p. 11/43
8/31 regular ras called plates. The plates also form a regular pattern over the image, such that the plates in a horizontal line are of uniform height and aligned, and the plates in a vertical column are of uniform width and aligned. In the representative encoder, the plates can be any arbitrary size that is a multiplication of 16 in the horizontal and / or vertical direction. Alternative encoder implementations can divide the image into blocks, macroblocks, plates, or other units of other sizes and structures. [0023] A direct overlay operator 240 is applied to each edge between the blocks, after which each 4x4 block is transformed using a block transform 250. This block transform 250 can be the reversible, free-scale 2D transform described by Srinivasan, US Patent Application No. 11 / 015,707, entitled, Reversible Transform For Lossy And Lossless 2 -D Data Compression, filed on December 17, 2004. The overlay operator 240 may be the reversible overlay operator described by Tu et al., US Patent Application No. 11 / 015,148, entitled, Reversible Overlap Operator for Efficient Lossless Data Compression, filed on December 17, 2004; and by Tu et al., US Patent Application No. 11 / 035,991, entitled, Reversible 2-Dimensional Pre- / PostFiltering For Lapped Biorthogonal Transform, deposited on January 14, 2005. Alternatively, the discrete cosine transform or other block transforms and overlap operators can be used. Subsequent to the transform, the DC 260 coefficient of each 4x4 transform block is subjected to a similar processing chain (plate, direct overlap, followed by 4x4 block transform). The resulting DC transform coefficients and AC transform coefficients are quantized 270, entropy encoded 280 and packaged 290.
[0024] The decoder performs the reverse process. On the side
Petition 870190108373, of 10/25/2019, p. 12/43
9/31 of the decoder, the bits of transform coefficients are extracted 310 from their respective packages, from which the coefficients are decoded by themselves 320 and dequantized 330. The DC 340 coefficients are regenerated by applying an inverse transform, and the plane of DC coefficients is overlaid inversely using a smoothing operator applied across the edges of the DC block. Subsequently, the entire data is regenerated by applying the 4X4 350 reverse transform to the DC coefficients, and the decoded DC coefficients 342 of the bit stream. Finally, the block edges in the resulting image planes are filtered in 360 reverse overlay. This produces a reconstructed 2D data output.
[0025] In an exemplary implementation, encoder 200 (Figure 2) compresses an input image into the compressed bit stream (for example, a file), and decoder 300 (Figure 3) reconstructs the original input or an approximation of it , based on whether the encoding employed is lossy or lossless. The encoding process involves the application of a direct coiled transform (LT) discussed below, which is implemented with reversible two-dimensional pre- / post-filtering also described more fully below. The decoding process involves the application of the inverse winding transform (ILT) using reversible two-dimensional pre- / post-filtering.
[0026] The illustrated LT and ILT are inversely related to each other, in an exact sense, and therefore can be collectively referred to as reversible rolled transforms. As a reversible transform, the LT / ILT pair can be used for lossless image compression. [0027] The input data 210 compressed by the illustrated encoder 200 / decoder 300 can be images of various color formats (for example, color image format RGB / YUV4: 4: 4, YUV4: 2: 2 or YUV4: 2 : 0). Typically, input image always
Petition 870190108373, of 10/25/2019, p. 13/43
10/31 has a luminance component (Y).
[0028] If it is an RGB / YUV4: 4: 4, YUV4: 2: 2 or YUV4: 2: 0 image, the image also has chrominance components, such as a U component and a V component. The color planes separate images or components may have different spatial resolutions. In the case of an input image in the YUV 4: 2: 0 color format, for example, the U and V components are half the width and length of the y component.
[0029] As discussed above, encoder 200 plates the input image or macroblock image. In an exemplary implementation, the encoder 200 plates the input image in 16X16 pixel areas (called macroblocks) in the Y channel (which can be 16x16, 16x8 or 8x8 areas in the U and V channels depending on the color format). Each macroblock color plane is plated in regions or blocks of 4X4 pixels. Therefore, a macroblock is made up of several color formats as follows for this exemplary encoder implementation:
1. For a grayscale image, each macroblock contains 16 blocks of 4x4 (Y) luminance.
2. For a YUV4: 2: 0 color image, each macroblock contains 16 4x4 Y blocks, and each 4 4x4 chrominance blocks (U and V).
3. For a color image of YUV4: 2: 2 format, each macroblock contains 16 4x4 Y blocks, and each 8 4x4 chrominance blocks (U and V).
4. For an RGB or YUV4: 4: 4 color image, each macroblock contains 16 blocks from each Y, U and V channel.
[0030] Likewise, after the transformation, a macroblock in this representative encoder 200 / decoder 300 has three frequency sub-bands: a DC sub-band (DC macroblock), a
Petition 870190108373, of 10/25/2019, p. 14/43
11/31 low-pass sub-band (low-pass macroblock), and the high-pass sub-band (high-pass macroblock). In the representative system, the low-pass and / or high-pass sub-bands in the bit stream - these sub-bands can be abandoned entirely.
[0031] Additionally, compressed data can be packaged within the bit stream in one of two orders: special order and frequency order. For the spatial order, different sub-bands of the same macroblock within a plate are ordered together, and the resulting bit stream from each plate is recorded within a package. For the frequency order, the same sub-band from different macroblocks within a board is grouped together, and so the bit stream of a board is recorded in three packets: a DC board pack, a pass-through packet low, and a high-pass card pack. In addition, there may be other layers.
[0032] Thus, for the representative system, an image is organized in the following dimensions:
Spatial dimension: Macrobloc plate;
Frequency Dimension: DC | low pass | high pass; and Channel dimension: Luminance | Chrominance_0 | Chrominance_l ... (for example, as Y | U | V).
[0033] The lines above signify a hierarchy, while the vertical bars signify partitioning.
[0034] Although the representative system organizes compressed digital media data in space, frequency and channel dimensions, the flexible quantization approach described here can be applied to alternative encoder / decoder systems that organize their data along smaller, additional or other dimensions. For example, the flexible quantization approach can be applied to coding using a larger number of frequency bands, other color channel formats (for example, YIQ, RGB,
Petition 870190108373, of 10/25/2019, p. 15/43
12/31 etc.), additional image channels (for example, for stereo viewing or other multiple camera arrangements).
2. Inverse and Coiled Transformed Core
Overview
[0035] In one implementation, of the encoder 200 / decoder 300, the reverse transform on the decoder side takes the form of a two-level coiled transform. The steps are as follows:
• An inverse core transform (ICT) is applied to each 4x4 block corresponding to the reconstructed DC and low-pass coefficients arranged in a flat arrangement known as the DC plane.
• A post-filtering operation is optionally applied to the 4x4 areas evenly matched around the blocks in the DC plane. Additionally, a post-filter is applied to the 2x4 and 4x2 boundary areas, and the four 2x2 corner areas are left untouched.
• The resulting arrangement contains DC coefficients of the 4x4 blocks corresponding to the first transform level. The DC coefficients are copied (figuratively) in a larger arrangement, and the reconstructed high-pass coefficients populated in the remaining positions.
• An ICT is applied to each 4x4 block.
• a post-filter operation is optionally applied to the 4x4 areas eventually hit around the blocks in the DC plane. Additionally, a post-filter is applied to the 2x4 and 4x2 boundary areas, and the four 2x2 curve areas are left untouched.
[0036] This process is shown in Figure 4.
[0037] The application of the post-filters is governed by an OVERLAP_INFO syntactic element in the compressed bit stream 220. OVERLAP_INFO can acquire three values:
Petition 870190108373, of 10/25/2019, p. 16/43
13/31 • If OVERLAP_INFO = 0, no post-filtering is performed.
• If OVERLAP_INFO = 1, only external post-filtering is performed.
• If OVERLAP_INFO = 2, both internal and external post-filtering are performed
Inverse Core Transform
[0038] The core transform (CT) is inspired by the conventionally known Discrete Cosine Transform 4x4 (DCT), it is still fundamentally different. The first different key is that DCT is linear whereas CT is non-linear. The second different key is that due to the fact that it is defined in real numbers, DCT is not a lossy operation for the entire space. The CT is defined in whole numbers and is without loss in space. The third different key is that the DCT 2D is a separable operation. The CT is not separable by the project.
[0039] The entire inverse transform process can be recorded as the cascade of three 2x2 elementary transform operations, which are:
• Hadamard Transform 2x2: T_h • Reverse Rotation ID: InvT_odd • 2D Reverse Rotation: InvT_odd_odd
[0040] These transforms are implemented as non-separable operations and are first described, followed by the description of the entire ICT.
Hadamard Transform 2D 2x2 T h
[0041] The encoder / decoder implements the Hadamard 2D 2x2 T_h transform as shown in the pseudocode table. R is a rounding factor that can be set to 0 or 1 only. T_h is involuntary (that is, two applications of T_h in a vector of
Petition 870190108373, of 10/25/2019, p. 17/43
14/31 data [abed] successful in recovering original values of [abcd], R provided is unloaded between applications). T_h is T_h itself.
<img file="BRPI0807465B1_D0001.tif" />
Reverse rotation 1D InvT odd
[0042] The inverse without loss of T_odd is defined by the pseudocode in the following table.
lnvT_odd (a, b, c, d) {d / a - c;
d (b »1) 'c ((a + η» 1);
a ™ ({3 * b + 4) »- 3}; b + “((3 * to + 4)» 3); c $ 3 * d + 4) »3); d + - ((3 * and + 4) »3);
c - = ((b + 1) »1); d = ((a + 1) »1) -d: b +» c;
a - d;
2D Inverse Rotation InvT odd odd
[0043] 2D reverse rotation lnvT_odd_odd is defined by the pseudocode in the following table.
lnvT_odd_odd (a, b, c, d) {l | ÍÍH | BÍ · cb;
a = (Í1 ~ d »1);
o + - (t2 c »1);
a - ((b '3 + 3) »3); b * - f (a * 3 + 3) »-2) ;; a- = ((b<sup>x</sup> 3 + 4)»:3);
b t2; .. is you:
c + ~ b; . ..
d ~~ a:
. b = -b '::
Petition 870190108373, of 10/25/2019, p. 18/43
15/31
ICP operations
[0044] The correspondence between 2X2 data and the previously listed pseudocode is shown in Figure 5. Color coding using gray levels to indicate the four data points is introduced here, to facilitate the description of the transform in the next section.
[0045] The ICT 4X4 2D point is constructed using T_h, T_odd inverse and T_odd_odd inverse. Note that the inverse T_h is T h itself. ICT consists of two stages, which are shown in the following pseudocode. Each stage consists of four 2x2 transforms that can be done in any arbitrary sequence, or concurrently, within the stage.
[0046] If the input data block is [_<sup>mn 0</sup> P\
4x4_IPCT_lstStage () and 4x4_IPCT_2ndStage () are defined as:
4x4 IPCT (a ... p) {
T_h (a, c, i, k);
int. odd (b, d, j.!):
InvT oadíe, mgo);
fnvl odd oddíf hn, p) T_h (a, d, m, p);
__________________________________
T h (c, b, o, n);
[0047] The 2x2_ICT function is the same as T_h.
Post-filtering overview
[0048] Four operators determine the post-filters used in the reverse wound transform. These are:
• 4x4 post filter • 4 point post filter • 2x2 post filter • 2 point post filter
[0049] The post-filter uses T_h, lnvT_odd_odd, invScale and invRotate. invRotate and invScale are defined in the tables below, respectivelyPetition 870190108373, of 10/25/2019, p. 19/43
16/31 te.
<img file="BRPI0807465B1_D0002.tif" />
invScale (a, b) {
4x4 post filter
[0050] Primarily, the 4X4 post-filter is applied to all the joints of the blocks (areas that eventually hit the 4 blocks) in all color planes when OVERLAPJNFO is 1 or 2. Also, the 4x4 filter is applied to all junctions of blocks in the DC plane for all planes when OVERLAPJNFO is 2, and only for the luma plane when OVERLAPJNFO is 2 and the color format is YUV 4: 2: 0 or YUV 4: 2: 2.
[0051]
If the input data block is L<sup>mn 0</sup> shovel, the after filter
4x4, 4x4Post filter (a, b, c, d, e, f, g, h, i, j, k, 1, m, n, o, p), is defined as in the following table:
<img file="BRPI0807465B1_D0003.tif" />
4-point post-filter
Petition 870190108373, of 10/25/2019, p. 20/43
17/31
[0052] 4-point linear filters are applied to the border, hitting 2X4 and 4X2 areas at the image boundary. If the input data is [abc d], the 4-point, 4-post-filter (a, b, c, d) post-filter are defined in the following table '' at® d * 3'í
.................:
2X2 post filter
[0053] The post-filter 2X2 is applied in areas of adjustment blocks in the DC plane for the data chroma channels YUV 4: 2: 0 and YUV 42: 2.
If the input data is, the post-filter 2x2 2x2Post-filter (a, b, c, d), is defined as the following table:
X ___________<sup>;</sup> ______________<sup>;</sup> d η »íwwwwxíwsjwwwwwwwwwwwwwwsíw ·;
....................: _j} ~ b
2-point post-filter
[0054] The 2-point post-filter is applied to the 2X1 and 1X2 limit samples that hit the blocks. The 2-point post-filter, post-filter2 (a, b) is defined in the following table
Petition 870190108373, of 10/25/2019, p. 21/43
18/31
<img file="BRPI0807465B1_D0004.tif" />
[0055] The signaling of the precision required to perform transform operations of the rolled transform described above can be performed in the header of a compressed image structure. In the exemplary implementation, LONG_WORD_FLAG and NO_SCALED_FLAGS are syntactic elements transmitted in the compressed bit stream (for example, in the image header) to signal precision and computational complexity to be applied by the decoder.
3. Speech accuracy and length
[0056] The exemplary encoder / decoder performs integer operations. Additionally, the exemplary encoder / decoder supports lossless encoding and decoding. However, the primary machine precision required by the exemplary encoder / decoder is an integer.
[0057] However, integer operations defined in the exemplary encoder / decoder allow rounding errors for lossy encoding. These errors are small by design, however they cause abandonment in the rate distortion curve. For the improved coding performance ratio by reducing rounding errors, the exemplary encoder / decoder defines secondary machine precision. In this mode, the input is pre-multiplied by 8 (that is, left shifted by 3 bits) and the final output is divided by 8 with rounding (that is, right shifted by 3 bits). These operations are carried out on the encoder screen and at the rear of the decoder, and are largely invisible to the rest of the processes. Additionally, the quantization levels are scaled in the same way as a flow created with the primary machine precision and decoded using the second machine precision (and vice versa
Petition 870190108373, of 10/25/2019, p. 22/43
19/31 sa) produces an acceptable image.
[0058] Secondary machine precision cannot be used when lossless compression is desired. The machine precision used to create a compressed file is explicitly marked in the header.
[0059] Secondary machine precision is equivalent to using scale arithmetic in the codec, and therefore this mode is referred to as Scale. Primary machine precision is referred to as non-scale.
[0060] The exemplary encoder / decoder is designed to provide good encoding and decoding speed. A design objective of the exemplary encoder / decoder is that the data values in both the encoder and decoder do not exceed 16 signed bits for an 8-bit input. (However, the intermediate operation within a transform stage can exceed this figure). This holds true for both machine precision modes.
[0061] Conversely, when secondary machine precision is chosen, the range of intermediate values is expanded by 8 bits. Since the primary machine precision prevents a pre-multiplication by 8, its range expansion is 8 - 3 = 5 bits.
[0062] The first exemplary encoder / decoder uses two different word lengths for intermediate values. These lengths are 16 and 32 bits.
[0063] Second Bitstream Syntax and Semantics Example [0064] The second bitstream syntax and semantics example is hierarchical and is comprised of the following layers: Image, Plate, Macrobloco and Block.
[0065] Some elements of the bit stream selected from the second bit stream syntax and semantics example are defined below:
Petition 870190108373, of 10/25/2019, p. 23/43
20/31
Long Word Flag (LONG_WORD_FLAG) (1 bit)
[0066] LONG_WORD_FLAG is a 1-bit syntax element and specifies whether 16-bit integers can be used for transform computations. In a second example of bitstream syntax, if LONG_WORD_FLAG = = 0 (FALSE), 16-bit integers and arrays can be used for the external stage of transform computations (intermediate operations within the transform (such as (3 * a + l) »l) are developed with greater accuracy). If LONG_WORD_FLAG = = TRUE, 32-bit integers and arrays must be used for transform computations.
[0067] Note: 32-bit arithmetic can be used to decode an image considering the value of LONG_WORD_FLAG. This element of syntax can be used by the decoder to choose the most efficient word length for implementation.
[0068] No scale arithmetic flag (NO_SCALED_FLAG) (1 bit)
[0069] NO_SCALED_FLAG is a 1-bit syntax element that specifies whether the transform uses scale. If NO_SCALED_FLAG == 1, the scale must not be performed. If NO_SCALED_FLAG == 0, scale must be used. In this case, the scale must be performed by rounding the final stage output (color conversion) by 3 bits.
[0070] Note: NO_SCALED_FLAG should be set to TRUE if lossless encoding is desired, even if lossless encoding is used only for a sub-region of an image. Lossy encoding can use both modes.
[0071] Note: The rate distortion performance for lossy encoding is superior for lossy encoding when scaling is used (ie, NO_SCALED_FLAG == FALSE), especially in
Petition 870190108373, of 10/25/2019, p. 24/43
21/31
Low QPs.
4. Signaling and Use of Long Word Flag
[0072] An exemplary image format for the representative encoder / decoder to support a wide range of pixel formats, including high dynamic range and wide musical scale formats. Supported data types include unsigned integer, fixed point float, and floating point float. The supported bit depths include eight, 16, 24 and 32 bits per color channel. The exemplary image format allows for lossless compression of images using up to 24 bits per color channel, and lossy image compression using up to 32 bits per color channel.
[0073] At the same time, the exemplary image format was designed to provide high quality images and compression efficiency and allows for low complexity encoding and decoding implementations.
[0074] To support low complexity implementation, the transformed into an exemplary image format was designed to minimize expansion in the dynamic range. The two-stage transform increases the dynamic range by only five bits. Therefore, if the image bit depth is eight bits per color channel, 1 arithmetic bits may be sufficient to perform all transform operations on the decoder. For other bit depths, higher precision arithmetic may be required for transform operations.
[0075] The computational complexity of decoding a specific bit stream can be reduced if the precision required for transform operations is known in the decoder. This information can be signaled to a decoder using a syntax element (for example, a 1-bit flag in a
Petition 870190108373, of 10/25/2019, p. 25/43
22/31 image). Described signaling techniques and syntax elements can reduce computational complexity in decoding bit streams.
[0076] In an exemplary implementation, the 1-bit syntax element LONG_WORD_FLAG is used. For example, if LONG_WORD_FLAG = = FALSE, 16-bit integers and arrays can be used for the external stage of transform computations and if LONG_WORD_FLAG == TRUE, 32-bit integers and arrays must be used for transform computations.
[0077] In an encoder / decoder implementation, transform operations in place can be performed on 16-bit width words, but intermediate operations within the transform (such as computing the 3 * a product for an elevation step given by b + = (3 * a + l) »l)) are performed with high accuracy (for example, 18 bits or more precision). However, in this example, the intermediate transform values a and b by themselves can be stored in 16-bit integers.
[0078] 32-bit arithmetic can be used to decode an image considering the values of the LONG_WORD_FLAG element. The LONG_WORD_FLAG element can be used by the encoder / decoder to choose the most efficient word length. For example, an encoder may choose to set the LONG_WORD_FLAG element to FALSE if it can verify that the 16-bit and 32-bit precision steps produce the same output value.
5. NO SCALED FLAG signage and use
[0079] An example of a representative encoder / decode image format supports a wide pixel range, including high dynamic range and wide range formats at the same time
Petition 870190108373, of 10/25/2019, p. 26/43
23/31 time, the representative encoder / decoder design optimizes image quality and compression efficiency and allows for low complexity and decoding implementations.
[0080] As discussed above, the representative encoder / decoder uses two transforms based on a hierarchical block, where all transform steps are integer operations. The small rounding errors present in these integer operations result in loss of compression efficiency during lossy compression. To combat this problem, an implementation of the representative encoder / decoder defines two different precision modes for decoder operations: the scale mode and the non-scale mode.
[0081] In the scale precision mode, the input image is pre-multiplied by 8 (that is, shifted to the left by 3 bits), in the encoder, and the final output in the decoder is divided by 8 with rounding (this is, shifted to the left by 3 bits). Operation in scale accuracy mode minimizes rounding errors, and results in rate distortion performance.
[0082] In the non-scale precision mode, there is no such scale. An encoder or decoder operating in non-scale precision mode, does not have to deal with a small dynamic range for transform coefficients, and thus has low computational complexity. However, there is a small penalty on the compression efficiency to operate in this mode. Lossless coding (without quantification, that is, setting the Quantification Parameter or QP to 1) can only use the scale precision mode to ensure reversibility.
[0083] The precision mode used by the encoder to create a compressed file is explicitly flagged in the image header of the compressed bit stream 220 (figure 2) using the
Petition 870190108373, of 10/25/2019, p. 27/43
24/31
NO_SCALED_FLAG. It is recommended that the decoder 300 uses the same precision mode for its operations.
[0084] NO_SCALED_FLAG is a syntactic element in the image header that specifies the precision mode as follows:
[0085] If NO_SCALED_FLAG == TRUE, the non-scale mode should be used for decoder operation.
[0086] If NO_SCALED_FLAG == FALSE, scale should be used. In this case, the scale mode should be used for operation by rounding down the final stage output (color conversion) to 3 bits. The rate distortion performance for lossy encoding is superior when the non-scale mode is used (ie, NO_SCALED_FLAG == FALSE), especially at low QPs. However, the computation complexity is lower when non-scale mode is used for two reasons:
[0087] The smaller dynamic range expansion in non-scale mode means that shorter words can be used for transform computations, especially in conjunction with LONG_WORD_FLAG. In VLSI implementations, the reduced dynamic range expansion means that the port logic that implements the most significant bits can be de-energized.
[0088] The scaling mode requires addition and shift of bit to right in 3 bits (implement rounding divided by 8) on the decoder side. On the decoder side, it requires a bit shift to the left by 3 bits. This is slightly more essentially computational than the full scale mode.
[0089] Additionally, the non-scale mode allows for the compression of more significant bits than the scale mode. For example, the non-scale mode allows lossy compression (and decompression) of up to 27 significant bits per example, using 32 arithmetic bits. In contrast, the scale mode allows the same for al
Petition 870190108373, of 10/25/2019, p. 28/43
25/31 guns 24 bits. This is due to the three additional bits of dynamic range introduced by the scaling process.
[0090] The data values in the decoder do not exceed 16 bits signaled for an 8-bit input for both precision modes. (however, intermediate operation within a transform stage can exceed this figure).
[0091] Note: NO_SCALED_FLAG is set to TRUE by the encoder, if lossy encoding (QP = I) is desired, even if lossy encoding is required for only a sub-region of an image.
[0092] The encoder can use any mode for lossy compression. It is recommended that the decoder use the precision mode signaled by NO_SCALED_MODE for its operations. However, the quantification levels are scaled such that a flow created with the precision scale mode and decoded using the non-scale precision mode (and vice versa) produces an acceptable image and many cases.
6. Scale Arithmetic for Increased Accuracy
[0093] In an implementation of the representative encoder / decoder, the transforms (including color conversion) are integer transformed and implemented through a series of elevation steps. In these elevation steps, zeroing errors damage transform performance. For lossy compression cases, to minimize the damage from zeroing errors and thus maximize transform performance, input data for a transform must be shifted to the left by many bits. However, another desired feature is if the input image is 8 bits, then the transform output should be within 16 bits. So the number of bits shifted to the left cannot be large. The representative decoder implements a
Petition 870190108373, of 10/25/2019, p. 29/43
26/31 scale arithmetic technique to achieve both objectives. The scale arithmetic technique maximizes transform performance by minimizing damage from zeroing errors, and further limits the output of each transform step to be within 16 bits if the input image is 8 bits. This makes simple 16-bit implementation possible.
[0094] The transforms used in the representative encoder / decode are integer transformed and implemented by elevation steps. Most transform steps involve shifting to the right, which introduces zeroing errors. A transform generally involves many elevation steps, and accumulated zeroing errors visibly damage transform performance.
[0095] One way to reduce the damage from truncation errors is to shift the input data left before the transform in the decoder, and shift the same number of bits after the transform (combined with quantization) to the right in the decoder. As described above, the representative encoder / decode has a two-stage transform structure: optional first stage overlay + first stage CT + optional second stage overlay + second stage CT. Experiments have shown that shifting to the left in 3 bits is necessary to minimize truncation errors. So that in cases of loss, before color conversion, input data can be shifted to the left in 3 bits, that is, multiplied or scaled by an 8 fact (for example, for the scale mode described above). [0096] However, color conversion and transforms do not contain data. If input data is shifted in 3 bits, the output of the second stage 4 x 4 DCT has a dynamic range of 17 bits, if the input data is 8 bits (output of each other transformed
Petition 870190108373, of 10/25/2019, p. 30/43
27/31 is still within 16 bits) this is undesirable as it avoids 16 bit implementation, a highly desirable feature. To achieve this, before the second stage 4 x 4 CT, input data is shifted to the right and so that the output is also within 16 bits. Since the second stage 4 x 4 CT applies to only 1/6 of the data (the DC transform coefficients of the first stage DCT), and the data is already scaled by the first transform, so that the damage zeroing errors here are minimal.
[0097] So, in cases with loss for 8-bit images, on the encoder side, input is shifted to the left by 3 bits before color conversion, and shifted to the right by 1 bit before the second stage 4 x 4 CT. On the decoder side, the input is shifted to the left by 1 bit before the first stage 4 x 4 IDCT, and shifted to the right by 3 bits after color conversion.
7. Computing Environment
[0098] The processing techniques described above for computational complexity and precision signaling in a digital media codec can be performed in any of a variety of digital media encoding and / or decoding systems, including among other examples, computers ( various form factors, including server, desktop, laptop, handheld devices, etc.); digital media recorders and players; image and video capture devices (such as cameras, scanners, etc.) communication equipment (such as telephones, mobile phones, conference equipment, etc.); display, printing or other presentation devices; and etc. computational complexity and precision signaling techniques in a digital media codec can be implemented in hardware circuit assemblies, in firmware that controls digital media processing hardware, as well as in software
Petition 870190108373, of 10/25/2019, p. 31/43
28/31 of communication that runs inside a computer or other computer environment, as shown in figure 6.
[0099] Figure 6 illustrates a generalized example of a computing environment (600) in which described modalities can be implemented. The computing environment (600) is not intended to suggest any limitations as well as the scope of use or functionality of the invention, as the present invention can be implemented in general purpose or special purpose computing environments. [00100] With reference to figure 6, the computing environment (600) includes at least one processing unit (610) and memory (620). In figure 6, this most basic configuration (630) is included within a dotted line. The processing unit (610) executes computer-readable instructions and can be a virtual processor. In a multiprocessing system, multiple processing units execute instructions executable by computer, to increase processing power. The memory (620) can be volatile memory (for example, registers, cache, RAM), non-volatile memory (for example, ROM, EEPROM, flash memory, etc.), or some combination of the two. The memory (620) stores software (680) that implements the encoding / decoding of digital media described with computational complexity and precision signaling techniques.
[00101] A computing environment may have additional characteristics. For example, the computing environment (600) includes storage (640), one or more input devices (650), one or more output devices (660), and one or one or more communication connections (670). An interconnect mechanism (not shown) such as a bus, controller, or network interconnects the components of the computing environment (600). Typically, the operating system software (not shown) provides an operating environment
Petition 870190108373, of 10/25/2019, p. 32/43
29/31 for other software running in the computing environment (600), and coordinates activities of the components of the computing environment. (600).
[00102] Storage (640) can be removable or non-removable, and includes magnetic disks, magnetic tapes or cassettes, CDROMs, CD-RWs, DVDs, or any other medium that can be used to store information and that can be accessed within computing environment (600). The storage (640) stores instructions for the software (680) implementing the encoding / decoding of digital media described with computational complexity and precision signaling techniques.
[00103] Input devices (650) can be a touch input device such as a keyboard, mouse, pen, or indicator, a voice input device, a scanning device, or other device that provides input to the environment communication (600). For audio, the input devices (650) can be a sound card or similar device that accepts audio input in analog or digital form from a microphone or microphone array, or CD-ROM player that provides audio samples for the computing environment. The computing devices (660) can be a display, printer, speakers, CD recorder, or other device that provides output from the computing environment (600).
[00104] Communication connections (670) capable of communicating over a communication medium to another computing entity. The communication medium carries information such as computer-readable instructions, copied audio or video information, or other data in a modulated data signal. A modulated data signal is a signal that has one or more of the characteristics established or altered in such a way as to encode information in the signal. For example, and not limitation, communication media includes techniques
Petition 870190108373, of 10/25/2019, p. 33/43
30/31 wired or wireless implemented with an electric, optical, RF, infrared, acoustic, or other carrier.
[00105] The encoding / decoding of digital media described with flexible quantification techniques in this document can be described in the general context of computer-readable media. Computer-readable media is any available media that can be accessed within a communication environment. For example, and not limitation, with the communication environment (600), the computer-readable media includes memory (620), storage (640), communication media, and combinations of any of the above
[00106] The encoding / decoding of digital media described with computational complexity and precision signaling techniques in this document can be described in the general context of instructions executable by computer, such as those included in program modules, being executed in a computing environment on a real target or virtual processor. Program modules generally include routines, programs, libraries, objects, classes, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The functionality of the program modules can be combined or divided between program modules as desired in various modalities. Computer executable instructions for program modules can be executed within a local or distributed computing environment.
[00107] For presentation purposes, the detailed description uses terms as it determines, generates, adjusts, and applies to describe computer operations in a computing environment. These terms are high-level abstractions for operations performed by a computer, and should not be confused with acts performed by a human being. The actual computer operations corresponding to these
Petition 870190108373, of 10/25/2019, p. 34/43
31/31 terms vary depending on the implementation.
[00108] In view of the many possible modalities to which the principles of the invention can be applied, all such modalities are claimed as invention as many as may be within the scope and spirit of the following claims and equivalents thereof.
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
28 members in 11 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 60891031 | United States of America | – | |
| 89103107 | United States of America | P | |
| 89103107 | United States of America | P | |
| 11772076 | United States of America | – | |
| 77207607 | United States of America | A | |
| 77207607 | United States of America | A | |
| 2008054473 | United States of America | W | |
| 2008054473 | United States of America | W | |
| 11772076 | – | – | – |
| 60891031 | – | – | – |
| PCTUS2008054473 | – | – | – |
| US20070772076 | – | – | – |
| US20070891031P | – | – | – |
| WO2008US54473 | – | – | – |
Members28
| Document | Office | Kind | |
|---|---|---|---|
| US2008198935A1 | United States of America | A1 | |
| WO2008103766A2 | World Intellectual Property Organization (WIPO) | A2 | |
| TW200843515A | Taiwan Province of China | A | |
| WO2008103766A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20090115726A | Republic of Korea | A | |
| EP2123045A2 | European Patent Office (EPO) | A2 | |
| CN101617539A | China | A | |
| IL199994A0 | Israel | A0 | |
| IL199994D0 | Israel | D0 | |
| JP2010519858A | Japan | A | |
| HK1140341A | Hong Kong, China | A | |
| HK1140341A1 | Hong Kong, China | A1 | |
| RU2009131599A | Russian Federation | A | |
| CN101617539B | China | B | |
| EP2123045A4 | European Patent Office (EPO) | A4 | |
| JP5457199B2 | Japan | B2 | |
| JP2014078952A | Japan | A | |
| BRPI0807465A2 | Brazil | A2 | |
| RU2518417C2 | Russian Federation | C2 | |
| KR20150003400A | Republic of Korea | A | |
| TWI471013B | Taiwan Province of China | B | |
| US8942289B2 | United States of America | B2 | |
| KR101507183B1 | Republic of Korea | B1 | |
| KR101550166B1 | Republic of Korea | B1 | |
| IL199994A | Israel | A | |
| BRPI0807465A8 | Brazil | A8 | |
| EP2123045B1 | European Patent Office (EPO) | B1 | |
| BRPI0807465B1This record | Brazil | B1 |
6 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent or certificate of addition of invention granted [chapter 16.1 patent gazette]GrantedPRAZO DE VALIDADE: 10 (DEZ) ANOS CONTADOS A PARTIR DE 26/05/2020, OBSERVADAS AS CONDICOES LEGAIS.B16A | B16A | |
| Decision: intention to grant [chapter 9.1 patent gazette]B09A | B09A | |
| Preliminary requirement: requests with searches performed by other patent offices: procedure suspended [chapter 6.21 patent gazette]B06U | B06U | |
| Objections, documents and/or translations needed after an examination request according [chapter 6.6 patent gazette]B06F | B06F | |
| Others concerning applications: alteration of classificationB15K | B15K | |
| Requested transfer of rights approvedB25A | B25A |
Numbers
- Publication
- PI0807465
- Publication, DOCDB
- PI0807465
- Publication, EPODOC
- BRPI0807465
- Application
- 7465
- Application, DOCDB
- PI0807465
- Application, EPODOC
- BR2008PI07465
Titles2
- Portuguese
- MÉTODO DE DECODIFICAÇÃO DE MÍDIA DIGITAL, MÉTODO DE CODIFICAÇÃO DE MÍDIA DIGITAL, MÍDIA DE ARMAZENAMENTO LEGÍVEL POR COMPUTADOR, DISPOSITIVO DE DECODIFICAÇÃO E DISPOSITIVO DE CODIFICAÇÃO
- English
- DIGITAL MEDIA DECODING METHOD, DIGITAL MEDIA ENCODING METHOD, COMPUTER LEGIBLE STORAGE MEDIA, DECODING DEVICE AND ENCODING DEVICE
Classification
- CPC, 7
- H04N21/6547
- H04N19/70
- H04N19/122
- H04N19/156
- H04N19/46
- H04N19/60
- H04N19/42
- IPC, 5
- H04N21 6547
- H04N19 122
- H04N19 156
- H04N19 46
- H04N19 60