Coding of significance maps and transform coefficient blocks
18 claims: 11 independent, 7 dependent
- 1Zastrzeżenia patentowe 1. Urządzenie do dekodowania mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji ze strumienia danych obejmujące:dekoder (250) skonfigurowany do pozyskiwania mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji, i następnie wartości znaczących współczynników transformacji wewnątrz bloku współczynników transformacji ze strumienia danych, z, przy pozyskiwaniu mapy istotności, sekwencyjnym pozyskiwaniem elementów składni pierwszego typu ze strumienia danych przez adaptacyjne względem kontekstu dekodowanie entropijne, elementów składni pierwszego typu wskazujących, czy dla powiązanych pozycji wewnątrz bloku współczynników transformacji w odpowiadającej pozycji położony jest znaczący czy nieznaczący współczynnik transformacji;oraz moduł wiązania (250) skonfigurowany do sekwencyjnego wiązania sekwencyjnie pozyskiwanych elementów składni pierwszego typu z pozycjami bloku współczynników transformacji we wstępnie ustalonej kolejności skanowania spośród pozycji bloku współczynników transformacji, przy czym dekoder jest skonfigurowany do użycia, w adaptacyjnym względem kontekstu dekodowaniu entropijnym elementów składni pierwszego typu, konteksty które są indywidualnie wybrane dla każdego z elementów składni pierwszego typu zależnie od liczby pozycji przy których według poprzednio pozyskanych i powiązanych elementów składni pierwszego typu położone są znaczące współczynniki transformacji, w sąsiedztwie pozycji z którą powiązany jest bieżący element składni pierwszego typu.
- 2Urządzenie według zastrz. 1, przy czym dekoder (250) jest skonfigurowany w taki sposób, że sąsiedztwo pozycji, z którą powiązany jest odpowiedni element składni pierwszego typu, może zawierać jedynie pozycje bezpośrednio przylegające do lub pozycje albo bezpośrednio przylegające do, albo oddzielone od pozycji, z którą powiązany jest odpowiedni element składni pierwszego typu, maksymalnie w jednej pozycji w kierunku pionowym i/lub jednej pozycji w kierunku poziomym, przy czym rozmiar zegara współczynników transformacji jest równy lub większy od 8x8 pozycji.
- 3Urządzenie według zastrz. 1 albo 2, przy czym dekoder (250) jest dalej skonfigurowany do mapowania liczby pozycji, w których zgodnie z uprzednio pozyskanymi i powiązanymi elementami składni pierwszego typu znajdują się znaczące współczynniki transformacji, w sąsiedztwie pozycji z którymi powiązany jest odpowiedni element składni pierwszego typu, do indeksu kontekstu z wstępnie ustalonego zestawu możliwych indeksów kontekstu do ważenia z pewną liczbą dostępnych pozycji w sąsiedztwie pozycji z którą odpowiedni element składni pierwszego typu jest powiązany.
- 4Urządzenie według dowolnego z zastrz. od 1 do 3, przy czym moduł wiązania (252) jest skonfigurowany w taki sposób, że wstępnie ustalona kolejność skanowania zależy od pozycji znaczących współczynników transformacji wskazanych przez uprzednio pozyskane i powiązane elementy składni pierwszego typu.
- 5Urządzenie według zastrz. 4, przy czym dekoder (250) dodatkowo jest skonfigurowany do rozpoznawania, w oparciu o informację w strumieniu danych, i niezależnie od liczby pozycji nieznaczących współczynników transformacji wskazanych przez uprzednio pozyskane i powiązane elementy składni pierwszego typu, czy w pozycji, z którą powiązany jest bieżąco pozyskiwany element składni pierwszego typu, który wskazuje, że w tej pozycji znajduje się znaczący współczynnik transformacji, znajduje się ostatni znaczący współczynnik transformacji w bloku współczynników transformacji.
- 6Urządzenie według zastrz. 4 albo 5, przy czym dekoder (250) dodatkowo jest skonfigurowany do pozyskiwania, pomiędzy elementami składni pierwszego typu wskazującymi, że w odpowiedniej powiązanej pozycji znajduje się znaczący współczynnik transformacji, i bezpośrednio następującymi elementami składni pierwszego typu, elementami składni drugiego typu ze strumienia bitów wskazującymi, dla powiązanych pozycji, w których znajduje się znaczący współczynnik transformacji, czy odpowiednią powiązaną pozycją jest ostatni znaczący współczynnik transformacji w bloku współczynników transformacji.
- 7Urządzenie według dowolnego z zastrz. 1 do 6, przy czym dekoder (250) dodatkowo jest skonfigurowany do sekwencyjnego pozyskiwania, po pozyskaniu wszystkich pierwszego typu elementów składni bloku współczynników transformacji, wartości znaczących współczynników transformacji wewnątrz bloku współczynników transformacji ze strumienia danych za pomocą adaptacyjnego względem kontekstu dekodowania entropijnego, przy czym moduł (252) wiązania jest skonfigurowany do sekwencyjnego wiązania sekwencyjnie pozyskiwanych wartości z pozycjami znaczących współczynników transformacji we wstępnie ustalonej kolejności skanowania współczynników spośród pozycji bloku współczynników transformacji, zgodnie z którą blok współczynników transformacji jest skanowany w podblokach (322) bloku (256) współczynników transformacji z użyciem kolejności (320) skanowania pod-bloku, z, pomocniczym, skanowaniem pozycji współczynników transformacji w pod-blokach (322) w kolejności (324) pod-skanowania pozycji, przy czym dekoder jest skonfigurowany do użycia, w sekwencyjnym adaptacyjnym względem kontekstu dekodowaniu entropijnym wartości znaczących wartości współczynników transformacji, wybranego zbioru pewnej liczby kontekstów z wielu zbiorów pewnej liczby kontekstów, przy czym wybór wybranego zbioru jest realizowany dla każdego pod-bloku w zależności od wartości współczynników transformacji wewnątrz pod-bloku bloku współczynników transformacji, już przebytych w kolejności (320) skanowania pod-bloku, lub wartości współczynników transformacji współumieszczonego pod-bloku w jednakowym rozmiarze bloku uprzednio zdekodowanym współczynników transformacji.
- 8Urządzenie według dowolnego z zastrz. 4 do 7, przy czym moduł (252) wiązania dodatkowo jest skonfigurowany do sekwencyjnego wiązania sekwencyjnie pozyskiwanych pierwszego typu elementów składni z pozycjami bloku współczynników transformacji wzdłuż sekwencji pod-ścieżek rozciągających się pomiędzy pierwszą parą sąsiednich boków bloku współczynników transformacji, wzdłuż których usytuowane są odpowiednio pozycje najniższej częstotliwości w kierunku poziomym i pozycje najwyższej częstotliwości w kierunku pionowym, i drugą parą sąsiednich boków bloku współczynników transformacji, wzdłuż których usytuowane są odpowiednio pozycje najniższej częstotliwości w kierunku pionowym i pozycje najwyższej częstotliwości w kierunku poziomym, z pod-ścieżkami mającymi rosnącą odległość od pozycji najniższej częstotliwości zarówno w kierunku pionowym jak i poziomym i przy czym moduł (252) wiązania jest skonfigurowany do wyznaczania kierunku (300, 302) wzdłuż którego sekwencyjnie pozyskiwane pierwszego typu elementy składni są powiązane z pozycjami bloku współczynników transformacji, w oparciu o pozycje znaczących współczynników transformacji wewnątrz poprzednich pod-skanowań.
- 9Urządzenie według dowolnego z zastrz. 1 do 8, przy czym blok współczynników transformacji odnosi się do zawartości mapy głębi.
- 10Dekoder na bazie transformacji skonfigurowany do dekodowania bloku współczynników transformacji za pomocą urządzenia (150) do dekodowania mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji ze strumienia danych, według dowolnego z zastrz. 1 do 9, i do wykonywania (152) transformacji z dziedziny widmowej do dziedziny przestrzennej do bloku współczynników transformacji.
- 11Dekoder przewidujący obejmujący dekoder na bazie transformacji (150, 152) skonfigurowany do dekodowania bloku współczynników transformacji za pomocą urządzenia do dekodowania mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji ze strumienia danych, według dowolnego z zastrz. 1 do 9, i do wykonywania transformacji z dziedziny widmowej do dziedziny przestrzennej do bloku współczynników transformacji by uzyskać resztkowy blok;moduł predykcji (156) skonfigurowany do dostarczania predykcji dla bloku macierzy próbek informacji reprezentujących przestrzennie próbkowaną informację dotyczącą sygnału;i moduł łączenia (154) skonfigurowany do łączenia predykcji bloku oraz bloku resztkowego w celu rekonstrukcji macierzy próbek informacji.
- 12Urządzenie do kodowania mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji w strumieniu danych, urządzenie skonfigurowane do kodowania mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji, i dalej wartości znaczących współczynników transformacji wewnątrz bloku współczynników transformacji w strumieniu danych, z, przy kodowaniu mapy istotności, sekwencyjnym kodowaniem elementów składni pierwszego typu do strumienia danych przez adaptacyjne względem kontekstu kodowanie entropijne, elementów składni pierwszego typu wskazujących, czy dla powiązanych pozycji wewnątrz bloku współczynników transformacji w odpowiadającej pozycji położony jest znaczący czy nieznaczący współczynnik transformacji, przy czym urządzenie jest ponadto skonfigurowane do sekwencyjnego kodowania elementów składni pierwszego typu do strumienia danych we wstępnie ustalonej kolejności skanowania spośród pozycji bloku współczynników transformacji, przy czym urządzenie jest skonfigurowane do użycia, w adaptacyjnym względem kontekstu kodowaniu entropijnym elementów składni pierwszego typu, konteksty które są indywidualnie wybrane dla każdego z elementów składni pierwszego typu zależnie od liczby pozycji przy których według poprzednio pozyskanych i powiązanych elementów składni pierwszego typu położone są znaczące współczynniki transformacji i z którym uprzednio kodowane elementy składni pierwszego typu są powiązane, w sąsiedztwie pozycji, z którą powiązany jest bieżący element składni pierwszego typu.
- 13Urządzenie według zastrz. 12, przy czym blok współczynników transformacji odnosi się do zawartości mapy głębi.
- 14Sposób dekodowania mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji ze strumienia danych, obejmujący:pozyskiwanie mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji, a następnie wartości znaczących współczynników transformacji wewnątrz bloku współczynników transformacji ze strumienia danych, z, przy pozyskiwaniu mapy istotności, sekwencyjnym pozyskiwaniem elementów składni pierwszego typu ze strumienia danych za pomocą adaptacyjnego względem kontekstu dekodowania entropijnego;elementy składni pierwszego typu wskazujące, czy dla powiązanych pozycji wewnątrz bloku współczynników transformacji w odpowiedniej pozycji położony jest znaczący czy nieznaczący współczynnik transformacji;i sekwencyjne wiązanie sekwencyjnie pozyskiwanych elementów składni pierwszego typu z pozycjami bloku współczynników transformacji we wstępnie ustalonej kolejności skanowania współczynników spośród pozycji bloku współczynników transformacji, przy czym, w sekwencyjnym adaptacyjnym względem kontekstu dekodowaniu entropijnym elementów składni pierwszego typu, używane są konteksty, które są indywidualnie wybrane dla każdego z elementów składni pierwszego typu zależnie od liczby pozycji przy których według poprzednio pozyskanych i powiązanych elementów składni pierwszego typu położone są znaczące współczynniki transformacji, w sąsiedztwie pozycji z którą powiązany jest bieżący element składni pierwszego typu.
- 15Sposób kodowania mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji do strumienia danych, sposób obejmujący kodowanie mapy istotności wskazującej pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji, a następnie wartości znaczących współczynników transformacji wewnątrz bloku współczynników transformacji do strumienia danych, z, w kodowaniu mapy istotności, sekwencyjnym kodowaniem elementów składni pierwszego typu do strumienia danych za pomocą adaptacyjnego względem kontekstu kodowania entropijnego, elementy składni pierwszego typu wskazujące, czy dla powiązanych pozycji wewnątrz bloku współczynników transformacji w odpowiedniej pozycji położony jest znaczący czy nieznaczący współczynnik transformacji, przy czym sekwencyjne kodowanie elementów składni pierwszego typu do strumienia danych jest wykonywane we wstępnie ustalonej kolejności skanowania spośród pozycji bloku współczynników transformacji, i w adaptacyjnym względem kontekstu kodowaniu entropijnym każdy z elementów składni pierwszego typu, używane są konteksty, które są indywidualnie wybrane z elementów składni pierwszego typu w zależności od liczby pozycji, przy których położone są znaczące współczynniki transformacji i z którymi uprzednio kodowane elementy składni pierwszego typu są powiązane, w sąsiedztwie pozycji z którą powiązany jest bieżący element składni pierwszego typu.
- 16Strumień danych zawierający w sobie mapę istotności wskazującą pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji, przy czym mapa istotności wskazująca pozycje znaczących współczynników transformacji wewnątrz bloku współczynników transformacji jest kodowana do strumienia danych, z następującymi wartościami znaczących współczynników transformacji wewnątrz bloku współczynników transformacji, przy czym wewnątrz mapy istotności, elementy składni pierwszego typu są kodowane sekwencyjnie do strumienia danych za pomocą adaptacyjnego względem kontekstu kodowania entropijnego, elementy składni pierwszego typu wskazujące, czy dla powiązanych pozycji wewnątrz bloku współczynników transformacji w odpowiedniej pozycji położony jest znaczący czy nieznaczący współczynnik transformacji, przy czym elementy składni pierwszego typu są kodowane sekwencyjnie do strumienia danych we wstępnie określonej kolejności skanowania spośród pozycji bloku współczynników transformacji, i elementy składni pierwszego typu są adaptacyjnie względem kontekstu kodowane entropijne do strumienia danych z użyciem kontekstów, które są indywidualnie wybrane dla elementów składni pierwszego typu w zależności od liczby pozycji przy których położone są znaczące współczynniki transformacji i z którymi uprzednio kodowane elementy składni pierwszego typu do strumienia danych są powiązane, w sąsiedztwie pozycji z którą powiązany jest bieżący element składni pierwszego typu.
- 17Strumień danych według zastrz. 16, w którym blok współczynników transformacji odnosi się do zawartości mapy głębi.
- 18Odczytywalny komputerowo cyfrowy nośnik danych mający zapisany na nim program komputerowy mający kod programu do realizacji, gdy jest uruchomiony w komputerze, sposobu określonego w zastrz. 14 albo 15. GE Video Compression, LLC, Stany Zjednoczone Ameryki Pełnomocnik:ΕΡ 2 559 244 BI Ζ-16443/17 1/9 FIG 1 ΕΡ 2 559 244 BI Ζ-16443/17 FIG 2Β FIG 20 ΕΡ 2 559 244 BI 3/9 Ζ-16443/17 FIG 3 ΕΡ 2 559 244 BI Ζ-16443/17 4/9 100 FIG 4 ΕΡ 2 559 244 BI Ζ-16443/17 5/9 152 FIG 5 EP 2 559 244 BI Z-16443/17 6/9 FIG 6 ΕΡ 2 559 244 BI Ζ-16443/17 7/9 250 256 FIG 7 ΕΡ 2 559 244 BI Ζ-16443/17 8/9 Ί FIG 10 ΕΡ 2 559 244 BI Ζ-16443/17 9/9 256 FIG 11
Independent claims18
133 paragraphs in 1 section, as filed
Description
[0001] The present application relates to the encoding of significance maps indicating the positions of significant transform coefficients within blocks of transform coefficients and to the encoding of such blocks of transform coefficients. Such encoding may, for example, be used in image and video encoding, for example.
[0002] In conventional video encoding, images of video sequences are typically broken down into blocks. Blocks or color components of blocks are predicted either by motion-compensated prediction or intra prediction. The blocks can be of various sizes and can be either square or rectangular. All samples of a block or color component of a block are predicted using the same set of prediction parameters, such as reference indexes (identifying the reference image in an already encoded image set), motion parameters (defining a measure of the movement of blocks between the reference image and the live image), parameters to determine interpolation filter, intra prediction modes etc. The motion parameters may be represented by displacement vectors with horizontal and vertical components, or by higher order motion parameters such as affine motion parameters consisting of 6 components. It is also possible that more than one set of prediction parameters (such as reference indexes and traffic parameters) are associated with one block. Here, for each set of prediction parameters, one intermediate prediction signal for a block or color component of a block is generated, and the final prediction signal is built from the weighted sum of the intermediate prediction signals. The weighting parameters and potentially also the deviation constant (which is added to the weighted sum) can either be constants for the reference picture or picture or set of reference pictures, or they can be included in the prediction parameter set for the respective block. Similarly, still pictures are also often decomposed into blocks and the blocks are predicted using an intra prediction method (which may be a spatial intra prediction method or a simple intra prediction method that predicts the DC component of the block). In the case of a corner, the prediction signal may also be zero.
[0003] The difference between the original blocks or the color components of the original blocks and the corresponding prediction signals, also called the residual signal, is usually transformed and quantized. A two-dimensional transform is applied to the residual signal and the resulting transform coefficients are quantized. To perform this transform coding, blocks or color components of blocks for which a given set of prediction parameters has been used may be further partitioned before applying a transform. Transform blocks may be equal to or less than the blocks that are used for prediction. It is also possible for a transform block to include more than one of the blocks that are used for prediction. Different transform blocks in a still image or video sequence image can be of different sizes, and the transform blocks can represent square or rectangular blocks.
[0004] The resulting quantized transform coefficients, also called transform coefficient levels, are then transmitted using entropy coding techniques. Accordingly, a block of transform coefficient levels is typically mapped to a vector (i.e. an ordered set) of transform coefficient values using a scan, where different scans may be used for different blocks. A zig-zag scan is often used. For blocks that contain only samples of one field of an interleaved frame (these blocks may be blocks within coded fields or blocks of fields within coded frames), it is also common to use various scans designed specifically for the field blocks. A commonly used entropy coding algorithm for encoding the resulting ordered sequence of transform coefficients is run-level coding. Typically, a large number of levels of transform coefficients are zero and a set of consecutive levels of transform coefficients that are zero can be efficiently represented by encoding the number of consecutive levels of transform coefficients that are zero (series). For the remaining (non-zero) transformation coefficients, the actual level is coded. There are different alternatives for coding a series of levels. The series before the non-zero coefficient and the level of the non-zero transform coefficient may be coded together using a single symbol or code word. Special symbols are often included for the end of the block that is sent after the last non-zero transform coefficient. Or, it is possible to encode the number of non-zero levels of the transform coefficients first, and depending on this number, the levels and series are encoded.
[0005] A slightly different approach is used in the highly efficient H.264 entropy coding of CABAC. Here, the coding of the transform coefficient levels is divided into three steps. In a first step, a syntax binary coded_block_flag is transmitted for each transform block that signals whether the transform block contains significant levels of transform coefficients (i.e., transform coefficients that are non-zero). If this syntax element indicates that significant levels of transform coefficients are present, a binary-valued significance map is encoded that determines which of the transform coefficient levels are nonzero. Then, in reverse scan order, the values of the non-zero levels of the transform coefficients are encoded. The significance map is coded as follows. For each coefficient in the scan order, a binary significant_coeff_flag syntax is encoded which determines whether the corresponding transform coefficient level is not equal to zero. If significant_coeff_flag is one, ie if there is a non-zero transform coefficient level at a given scan entry, the next binary last_significant_coeff_flag in the syntax is encoded. This container indicates whether the current significant transform coefficient level is the last significant transform coefficient level within a block or whether there will be successive significant transform coefficient levels in the scan order. If last_significant_coeff_flag indicates that no further significant transform coefficients will occur, no further syntax elements are coded to determine the significance map for the block. In the next step, values of significant levels of transformation coefficients are coded, the locations of which within the block are already determined by the significance map. The values of significant levels of transform coefficients are encoded in the reverse scan order by using the following three syntax elements. The binary coeff_abs_greater_one element of the syntax indicates whether the absolute value of the significant transform coefficient level is greater than one. If the syntax binary coeff_abs_greater_one indicates that the absolute value is greater than one, another syntax coeff_abs_level_minus_one element is sent that specifies the absolute value of the transform coefficient level minus one. Finally, a syntax binary coeff_sign_flag that specifies the sign of a transform coefficient value is encoded for each level of a significant transform coefficient. Again, note that the syntax items that are associated with the severity map are encoded in the scan order, while the syntax items that are associated with the actual values of the transform coefficient levels are encoded in the reverse scan order, allowing more appropriate context models to be used.
[0006] In H.264 CABAC entropy coding, all syntax elements for the levels of transform coefficients are coded using binary probability modeling. The non-binary coeff_abs_level_minus_one element of the syntax is binarized first, i.e. it is mapped to a sequence of binary decisions (containers) and these containers are sequentially encoded. The binary significant_coeff_flag, last_significant_coeff_flag, coeff_abs_greater_one, and coeff_sign_flag syntax are directly encoded. Each coded container (including binary syntax elements) is associated with a context. Context represents a probability model for a class of encoded containers. The measure associated with the probability for one of the two potential container values is estimated for each context based on the container values already coded with the corresponding context. For several containers associated with transform encoding, the context that is used for encoding is selected based on already transmitted syntax elements or based on the position within the block.
[0007] Significance maps define significance information (the transform coefficient level is not zero) for a scan item. In H.264 CABAC entropy encoding, for a 4x4 block size, a separate context is used for each scan item for encoding binary significant_coeff_flag and last_significant_coeff_ilag syntax when different contexts are used for significant_coeff_flag and last_significant_coeff_flag scan items. For 8x8 blocks, the same context model is used for four consecutive scan items, resulting in 16 context models for significant_coeff_flag and an additional 16 context models for last_significant_coeff_flag. This way of modeling context for significant_coeff_flag and last_significant_coeff_flag has some disadvantages for large block sizes. On the other hand, if each scan item is associated with a separate context model, the number of context models increases significantly when blocks larger than 8x8 are encoded. Such an increased number of context models leads to a slow adaptation of the probability estimation and usually to the inaccuracy of the probability estimation, both aspects having a negative impact on the coding efficiency. On the other hand, assigning the context model to the number of consecutive scan positions (as is the case for 8x8 blocks in H.264) is also not optimal for larger block sizes, since non-zero transform coefficients are usually clustered in certain areas of the transform block (the areas depend on from the main structures inside the respective residual signal blocks).
[0008] After the significance map has been encoded, the blocks are processed in the reverse scan order. If the scan position is significant, ie the coefficient is not zero, the coeff_abs_greater_one binary element of the syntax is transmitted. Initially, a second context model of the corresponding set of context models is selected for the coeff_abs_greater_one element of the syntax. If the encoded value of any coeff_abs_greater_one element of the syntax inside the block is equal to unity (i.e. the absolute factor is greater than 2) context modeling switches back to the first context model in the set and uses this context model until the end of the block. In another case (all the encoded coeff_abs_greater_one values inside the block are zero and the corresponding absolute coefficient levels are zero), the context model is chosen depending on the number of coeff_abs_greater_one elements of the zero syntax that have already been encoded / decoded in the inverse scan of the block in question. The choice of the context model for the coeff-abs_greater_one syntax can be summarized by the following equation, where the index of Ct + and the current context model is selected based on the index Ct + ι of the previous context model and the value of the previously coded coeff_abs_greater_one syntax, which is represented by bint in equation. For the first coeff_abs_reater_one element of the syntax inside the block, the context model index is set to Ct = 1.
C,<sub>+</sub>fC<sub>vol</sub>, bin,)
<img file="PL2559244T3_D0001.tif" />
for for bin<sub>vol</sub> = 1 bin, = 0
[0009] The second syntax element for coding the absolute levels of transform coefficients, coeff_abs_level_minus_one, is only coded when coeff_abs_greater_one of the syntax for the same scan position is equal to one. The non-binary element coeff_abs_level_minus_one of the syntax is binarized into a sequence of containers, and for the first container of this binarization; the context model index is selected as described below. The other binarization containers are encoded with fixed contexts. The context for the first binarization container is selected as follows. For the first coeff_abs_level_minus_one element of the syntax, the first context model is selected from the set of context models for the first container of the coeff_abs_level_minus_one syntax, the corresponding context model index is selected equal to Ct = O. For each successive first container of the coeff_abs_level_minus_one syntax, context modeling switches to the next context model in the set, with the number of context models in the set being limited to 5. The choice of the context model can be expressed as the following formula, where the index Ct + ι of the current context model is selected based on the Ct index of the previous context model. As mentioned above, for the first coeff_abs_level_minus_one element of the syntax inside the block, the context model index is chosen equal to Ct = O. Note that different sets of context models are used for the coeff_abs_greater_one and coeff_abs_level_minus_one syntax elements.
C,<sub>4</sub>., (C 1) = min (C 1 + 1.4)
[0010] In the case of large blocks, this method has some drawbacks. The choice of the first context model for coeff_abs_greater_one (which is used if coeff_abs_greater_one is 1 coded for blocks) is usually done too early and the last context model for coeff_abs_level_minus_one is reached too early because the number of significant coefficients is greater than in small blocks. Thus, most of the coeff_abs_greater_one and coeff_abs_level_minus_one containers are encoded with one context model. However, these bins typically have different probabilities, and hence using one context model for a large number of bins has a negative impact on coding efficiency.
[0011] While large blocks in general increase the computational overhead for performing spectral distribution transformations, the ability to efficiently code both small and large blocks would allow better coding efficiency when coding a sample array such as pictures or sample patterns representing other spatially sampled information signals such as depth maps etc. The reason for this is the relationship between spatial and spectral resolution during the transformation of the sample system within blocks: the larger the blocks, the higher the spectral resolution of the transformation. In general, it would be advantageous to be able to apply the individual transform locally on the sample system such that within the region of such an individual transformation, the spectral composition of the sample system does not change much. For small blocks, they ensure that the content inside the blocks is relatively constant. On the other hand, if the blocks are too small, the spectral resolution is low and the ratio between insignificant and significant transform coefficients is decreased.
[0012] Accordingly, it would be advantageous to have a coding scheme that allows efficient coding for blocks of transform coefficients, even when they are large, and their significance maps.
[0013] Reference EP1487113A2 describes how a first one-bit symbol (CBP4) is transmitted for each block of transform coefficients. If CBP4 indicates that the corresponding block contains significant factors, the significance picture is encoded by transmitting a one-bit symbol (SIG) for each factor in the scan sequence. If a given symbol is "one, a further one-bit symbol (LAST) is transmitted to indicate whether the current significant factor is the last factor within a block or whether further significant factors will follow. The positions of the significant transform coefficients included in the block are determined and coded for each block in a first scan operation followed by a second scan operation performed in the reverse order.
[0014] The reference EP1768415A1 describes the coding and decoding of video data with an improved coding efficiency. After transforming the pixel data into the frequency domain, only a predetermined subset of the transform coefficients are scanned and encoded. In this way, prior knowledge of the location of regularly zero transform coefficients can be used in such a way as to reduce redundancy in encoded video data. The information about the locations of the regularly zero coefficients may be signaled to the decoder either explicitly or by default. The decoder decodes a subset of the transform coefficients, uses the signaled information about the location of regular zero coefficients to invertly scan the decoded transform coefficients, and performs the inverse transform of the transform coefficients back to the pixel block.
[0015] The reference US2003 / 128753A1 discloses and provides an optimal scanning method for encoding / decoding a picture signal. In a method for encoding an image signal by a discrete cosine transform, at least one of a plurality of reference blocks is selected. The scanning order in which the blocks to be encoded and the reference blocks are scanned is generated, and the blocks to be encoded are scanned in the order of the generated scan order. At least one selected reference block is temporally or spatially adjacent to the block to be coded. When the blocks to be encoded are scanned, the probability that non-zero coefficients will occur is obtained from the at least one selected reference block, and the scanning order is determined as the descending order, starting with the highest probability. Here, the scan order is generated as a zigzag scan order if the probabilities are identical. The optimal scanning method increases the efficiency of signal compression.
[0016] Reference GB2264605A describes how a signal is scanned by a plurality of scanners according to a plurality of different schemas, and a scan scheme selection module determines which scan scheme provides the most efficient encoding result, for example for runlength coding. The selected signal is multiplexed with a signal that identifies the selected scan scheme, preferably after variable length encoding. As described, the signal is an image signal that has undergone a discrete cosine transform after the medium has been compensated. Eight different scanning or sampling schemes are disclosed.
[0017] The article "Compression of Sparse Matrices by Arithmetic Coding, Bell T. et al., PROCEEDINGS OF THE DATA COMPRESSION CONFERENCE (DCC '98), March 30, 1998, pages 23-32, ΧΡ010276609, IEEE, Los Alamitos, CA, USA, DOI: 10.1109 / DCC. 1998.672126, ISBN: 978-0-8186-8406-7, describes matrix compression in which most words are fixed constants (usually zero), usually called sparse matrices. The performance of the existing methods is assessed and it is considered how arithmetic coding can be applied in case of a problem in obtaining better compression.
[0018] Data Compression: The Complete Reference (passage) by Salomon D. et al., 1998, Springer, New York, NY, USA, 002270343, ISBN: 978-0-387-98280-9, pp. 69- 84, describes various aspects of arithmetic coding.
[0019] The article "Context-based Arithmetic Coding Reexamined for DCT Video Compression, Zhang L. et al., PROCEEDINGS OF THE 2007 IEEE INTERNATIONAL SYMPOSIUM ON CIRCUITS AND SYSTEMS (ICASP 2007), May 1, 2007, pages 3147-3150, ΧΡ031181972, IEEE, Piscataway, NJ, USA, ISBN: 978-1-4244-0920-4, presents a new context modeling technique for arithmetic coding of DCT coefficients in video compression. A key feature of the new technique is to include all previously coded coefficient modules in the DCT block in context modeling. This enables the use of high-order redundancy Markov operations in the DCT domain with several conditioning states by adaptive arithmetic coding. Additionally, the context weighting technique is used to further improve coding efficiency. The complexity of the new arithmetic coding scheme is slightly lower than the complexity of Context-based Adaptive Binary Arithmetic Coding (CABAC) in H.264.
[0020] The article "An overview of the basie principles of the Q-Coder adaptive binary arithmetic coder, Pennebaker WB et al., IBM JOURNAL OF RESEARCH AND DEVELOPMENT, vol. 32, No. 6, Nov. 1988, pages 717-726, ΧΡ000111384, IBM Corporation, New York, NY, USA, ISSN: 0018-8646 introduces Q-Coder as a new form of adaptive binary arithmetic coding. Part of the binary arithmetic coding of this technique follows from the underlying concepts introduced by Rissanen, Pasco, and Langdon, but extends the coding conventions to resolve the conflict between optimal software and hardware implementation. Additionally, a reliable form of probability estimation is used in which the probability estimation is derived only from the interval renormalization procedures that are part of the arithmetic coding method. A brief explanation of the concept of arithmetic coding is provided followed by compatible optimal hardware and software coding structures and symbol probability estimates from the interval renormalization procedure.
[0021] The article "Context-based Arithmetic Encoding of 2D Shape Sequences, Brady N. et al., PROCEEDINGS OF THE IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP 1997), vol. October 1, 26, 1997, pages 29-32, ΧΡ010254100, IEEE, Los Alamitos, CA, USA, DOI: 10.1109 / ICIP. 1997.647376, ISBN: 978-0-8186-8183-7, describes a new method of shape encoding in object-based video sequences. Context-based arithmetic encoding as used in JBIG has been applied inside the block-based structure and has been further enhanced to ensure efficient use of temporal prediction.
[0022] Article "Context-based adaptive binary arithmetic coding in JVT / H.26L, Marpe D. et al., PROCEEDINGS OF THE IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP 2002), vol. September 2, 22, 2002, pages 513-516, ΧΡ010608021, IEEE, Los Alamitos, CA, USA, ISBN: 978-0-7803-7622-9, describes context-based adaptive binary arithmetic encoding in JVT / H.26L, shows a new adaptive entropy coding scheme for video compression. It uses the adaptive arithmetic coding technique to better adjust the first-order entropy of the coded symbols and to track non-stationary symbol statistics. Additionally, remaining symbol redundancies will be used by context modeling to further reduce bit rate. A new approach to coding the transformation coefficients and a method of viewing a table for probability estimation and arithmetic coding is presented. This new approach has been integrated into the current JVT Test Model (JM) to demonstrate performance gains and has been incorporated as part of the current JVT / H.26L draft.
[0023] Therefore, it is an object of the present invention to provide a coding scheme for coding transform coefficient blocks and, respectively, significance maps indicating positions of significant transform coefficients within blocks of transform coefficients, such that the coding efficiency is increased.
[0024] This object is achieved by the subject of the independent patent claims.
[0025] According to a first aspect of the present application, it is stated in the application that a higher coding efficiency for coding a significance map indicating positions of significant transform coefficients within a block of transform coefficients can be obtained if the scanning order through which the sequentially acquired syntax elements indicate for related positions within transformation coefficient block, Whether a significant or insignificant transform coefficient block is located at the appropriate position are sequentially related to the positions of the block of transform coefficients, among the positions of the block of transform coefficients depends on the positions of the significant transform coefficients indicated by previously related syntax elements. In particular, the inventors found that in typical sample system contents such as image, video or depth map contents, significant transform coefficients mainly form clusters on one side of the block of transform coefficients corresponding to either non-zero frequencies in the vertical direction and low frequencies in the horizontal direction, or vice versa. , such that taking into account the positions of the significant transform coefficients indicated by previously related syntax elements allows control to further cause such a scan, that the probability of reaching the last significant transform coefficient within the block of transform coefficients beforehand is increased relative to the procedure by which the scanning order is predetermined irrespective of the positions of the significant transform coefficients indicated by previously related syntax elements so far. This is especially true for larger blocks, although it is also true for small blocks.
[0026] According to an embodiment of the present application, an entropy decoder is configured to obtain information from the data stream that enables it to recognize whether the significant transform coefficient currently indicated by the currently associated syntax element is the last significant transform coefficient regardless of its exact position within a block of transform coefficients. wherein the entropy decoder is configured not to wait for the next syntax element in the case where the current syntax element relates to such last significant transform factor. This information may include the number of significant transform coefficients within the block. Alternatively, the second syntax elements are interleaved with the first syntax element, the second syntax elements indicating, for related items where a significant transform coefficient is located, whether it is the last transform coefficient in the block of transform coefficients or not.
[0027] According to an embodiment, the binding module adjusts the scanning order depending on the positions of the significant transform coefficients indicated so far only at predetermined positions within the transform coefficient block. For example, several sub-paths that intersect the mutually disjoint subsets of positions within a block of transform coefficients extend substantially diagonally from one pair of sides of the block of transform coefficients corresponding to the minimum frequency along the first direction and the highest frequency along the second direction, respectively. to the opposite pair of sides of the block of transform coefficients corresponding to zero frequency along the second direction and maximum frequency along the first direction, respectively. In this case, the binding module is configured to select the scan order such that the sub-paths follow the order of the sub-paths in which the distance of the sub-paths to the DC position inside the transform coefficient block increases monotonically, each sub-path runs without breaks along the direction of the run, and for each sub-track the direction, along which the subpath runs is selected by the constraint modulus depending on the positions of the significant transformation coefficients that have already run in the previous sub-paths. Thus, the probability is increased that the last sub-path on which the last significant transform coefficient is located is traversed in a direction such that it is more likely that the last significant transform coefficient lies within the first half of this last sub-path than inside it. the other half, thus allowing you to reduce the number of syntax pointing elements, is there a significant or insignificant transformation coefficient in the appropriate position The result is especially valuable for large blocks of transformation coefficients.
According to a further aspect of the present application, the present application is based on the finding that a significance map indicating positions of significant transform coefficients within a block of transform coefficients can be encoded more efficiently if the above-mentioned syntax elements indicating for related positions within the block of transform coefficients are a given item has a significant or insignificant transformation coefficient, are context-adaptive, entropy-decoded using contexts that are selected individually for each of the syntax elements depending on the number of significant transform coefficients in the vicinity of the given syntax element, indicated as significant by any of the preceding syntax elements. In particular, the inventors have found that as the size of blocks of transform coefficients increases, significant transform coefficients are somehow related in certain areas within a block of transform coefficients such that a context adaptation that is not only sensitive to the number of significant transform coefficients that are have been run in pre-determined scanning sequences so far, but also takes into account the vicinity of significant transform coefficients, leading to better context adaptation and therefore increases the coding efficiency, entropy coding.
[0029] Of course, the two aspects outlined above can be advantageously combined.
[0030] Moreover, in accordance with yet another aspect of the present application, the application is based on the finding that coding efficiency in coding blocks of transform coefficients can be increased, when a significance map indicating positions of significant transform coefficients within a block of significant transform coefficients precedes the encoding of the actual values of significant transform coefficients within a block of transform coefficients, and if a predetermined scanning order among the block positions of transform coefficients used to sequentially associate the sequence of significant transform coefficient values with the significant coefficient positions scans the block transform coefficients in sub-blocks using the sub-block scan order of the sub-blocks with, the auxiliary, scan of the position of the transform coefficients within the sub-blocks in the coefficient scan order, and if a selected set of context numbers from multiple sets of context numbers is used for sequential context adaptive entropy decoding of significant values of transformation coefficients, wherein the selection of the selected set depends on the value of the transform coefficients within a sub-block of the transform coefficients already traveled in the scan order of the subblock or the selection depends on the value of the transformation coefficients of the co-located subblock in the previously decoded block of transform coefficients. In this way, the context adaptation is very suitable for the above-sketched property of significant clustered transform coefficients in certain areas within a block of transform coefficients, in particular when large blocks of transform coefficients are considered. In other words, the values can be scanned in sub-blocks and the contexts selected based on the statistics of the sub-blocks.
[0031] Again, even this second aspect may be combined with any of the predefined aspects of the present application or both.
[0032] Preferred embodiments of the present application are described below with reference to the figures, among which
Fig. 1 shows a block diagram of an encoder according to an embodiment;
Figures 2a-2c schematically show various further subdivisions of the sample system, such as the image, into blocks.
Fig. 3 shows a block diagram of a decoder according to an embodiment.
Fig. 4 shows a block diagram of an encoder according to an embodiment of the present application in more detail;
Fig. 5 shows a block diagram of a decoder according to an embodiment of the present application in more detail;
Fig. 6 schematically shows the transformation of a spatial domain block into a spectral domain;
Fig. 7 shows a block diagram of a significance map decoding apparatus and significant transform coefficients of a block of transform coefficients according to an embodiment;
Fig. 8 schematically shows a further subdivision of the scanning sequence into subpaths and their different course directions;
Fig. 9 schematically shows neighborhood definitions for certain scan positions within a transform block according to an embodiment;
Fig. 10 schematically shows potential neighborhood definitions for certain scan positions within transform blocks lying at the boundary of a transform block;
Fig. 11 shows a potential scan of transform blocks according to a further embodiment of the present application.
[0033] It should be noted that in the description of the figures, elements appearing in several of the figures are designated with the same reference sign in each of the figures, and repetition of these elements as to their functionality is avoided to avoid unnecessary duplication. However, the functionalities and descriptions provided with respect to one figure also apply to the other figures unless expressly stated to the contrary.
[0034] Fig. 1 shows an example of an encoder 10 in which aspects of the present application may be implemented. The encoder encodes an array of information samples into the data stream. The information sample array may represent any type of spatially sampled information signal. For example, the sample system 20 may be a still image or a video image. Accordingly, the information samples may correspond to brightness values, color values, luma values, chroma values, etc. However, the information samples may also be depth values in the case of a sample system 20 being a depth map generated, for example, by the time of a light sensor or the like.
[0035] The encoder 10 is a block-based encoder. Ie, encoder 10 encodes an array of 20 samples into the data stream 30 in units of blocks 40. The encoding in units of blocks 40 does not necessarily mean that encoder 10 encodes these blocks 40 completely independently of each other. Rather, encoder 10 may use the reconstruction of previously encoded blocks to extrapolate or predict the intra of the remaining blocks, and may use block granularity to set encoding parameters, i.e. for setting the way in which each area of the sample array corresponding to the respective block is encoded.
[0036] Next, the encoder 10 is a transform encoder. Ie, the encoder 10 encodes blocks 40 by transforming the information samples within each block 40 from the spatial domain to the spectral domain. A two-dimensional transform such as DCT or FFT or the like may be used. Preferably, the blocks 40 are square or rectangular in shape.
[0037] The further subdivision of the sample array 20 into blocks 40 shown in Fig. 1 is for illustrative purposes only. Fig. 1 shows an array of 20 samples, further divided into a regular two-dimensional array of square or rectangular blocks 40 which are non-overlappingly adjacent to each other. The size of the blocks 40 may be predetermined. Ie, the encoder 10 may not send block size information for blocks 40 within the data stream 30 to the decoding side. For example, the decoder can expect a predetermined block size.
[0038] However, several alternatives are possible. For example, blocks can overlap. The overlap may, however, be limited to the extent that each block has a portion not overlapping with any other block, or such that each sample of blocks overlaps, at most, one block of adjacent blocks adjacent to the current block in a predetermined direction. The latter means that the left and right adjacent blocks may overlap the running block in such a way as to fully cover the current block, but they cannot overlap each other, and the same is true for neighbors in vertical and diagonal directions.
[0039] As a further alternative, the division of the sample pattern 20 into blocks 40 may be matched to the contents of the sample pattern 20 by the encoder 10 with the subdivision information on subdivision applied sent to the decoder via the 30 bit stream.
Figures 2a to 2c show various examples of further subdividing the sample array 20 into blocks 40. Figure 2a shows a further subdivision of the sample array 20 into blocks 40 of different sizes based on a quadtree, the representative blocks being labeled 40a, 40b, 40c and 40d with increasing size. According to the further breakdown in Fig. 2a, the sample pattern 20 is first partitioned into a regular two-dimensional tree block pattern 40d, which in turn has individual subdivision information associated therewith, according to which a certain tree block 40d may or may not be further subdivided according to a quad tree structure. The tree block to the left of block 40d is illustratively further broken down into smaller blocks according to the quad tree structure. The encoder 10 may perform one two-dimensional transform for each of the blocks shown by the solid and broken lines in Fig. 2a. In other words, encoder 10 can transform circuit 20 into further block division units.
[0041] Instead of a quadtree further split, a more general tree based multiple further split may be used and the number of child nodes per hierarchy level may vary between different levels of the hierarchy.
[0042] Fig. 2b shows another example of a further subdivision. According to Fig. 2b, the sample array 20 is first divided into macroblocks 40b arranged in a regular two-dimensional non-overlapping and contiguous configuration, each macroblock 40b having associated further partitioning information according to which the macroblock is not further subdivided, or if subdivided further, is subdivided further in a regular two-dimensional manner into sub-blocks of equal size, so as to obtain different subdivision granularities for the different macroblocks. The result is a further division of the sample array 20 into blocks 40 of different sizes with representatives of different sizes labeled 40a, 40b and 40a '. As in Fig. 2a, the encoder 10 performs a two-dimensional transform on each of the blocks shown in Fig. 2b with solid and broken lines. Fig. 2c will be discussed later.
[0043] Fig. 3 shows a decoder 50 capable of decoding the data stream 30 generated by the encoder 10 to reconstruct a reconstructed version 60 of the sample system 20. Decoder 50 acquires from the data stream 30 a block of transform coefficients for each of blocks 40 and reconstructs the reconstructed version 60 by performing an inverted transform on each of the blocks of transform coefficients.
[0044] The encoder 10 and the decoder 50 may be configured to perform entropy encoding / decoding to insert transform coefficient blocks information into and extract this information from the data stream, respectively. Details in this regard will be described later. It should be noted that the data stream 30 does not necessarily include the blocks of transform coefficients for all blocks 40 of the sample pattern. Rather, a subset of blocks 40 may be bitstream encoded in other ways. For example, encoder 10 may decide to refrain from introducing a block of transform coefficients for a certain block of blocks 40 and introduce alternate encoding parameters to the bitstream 30 instead of allowing the decoder 50 to predict or otherwise fill the corresponding block in the reconstructed version 60. For example, the encoder 10 may perform texture analysis to locate blocks within a sample pattern 20 that can be filled on the decoder side by the decoder by texture synthesis and indicate this within the bitstream accordingly.
[0045] As will be discussed in the figures below, blocks of transform coefficients do not necessarily represent the spectral domain representations of the original information samples of the corresponding block 40 of the sample system. Rather, such a block of transform coefficients may represent a spectral domain representation of the prediction remainder of the corresponding block 40. Fig. 4 shows an embodiment for such an encoder. The encoder in Fig. 4 it comprises a transform stage 100, an entropy encoder 102, an inverse transform stage 104, a prediction module 106 and a subtraction module 108 as well as an addition module 110. The subtract module 108, the transform stage 100 and the entropy encoder 102 are connected in series in the mentioned order between the input 112 and the output 114 of the encoder of Fig. 4. The inverse-transform stage 104, the adder 110 and the prediction module 106 are connected in said order between the output of the transform stage 100 and the inverting input of the subtract module 108, the output of the prediction module 106 also being connected to a further input of the adder module 110.
[0046] The encoder of Fig. 4 shows a predictive block coder based on a transform. Ie, blocks of sample array 20 input at input 112 are predicted from previously coded and reconstructed portions of the same sample array or a previously encoded and reconstructed other sample array that may precede or follow the current sample array. The prediction is performed by the prediction module 106. The subtraction module 108 subtracts the prediction from such original block, and the transform step 100 performs a two-dimensional transform on the prediction residuals. The two-dimensional transformation itself or the following measure within the transformation step 100 can lead to quantization of the transform coefficients within blocks of the transform coefficients. The quantized blocks of transform coefficients are losslessly coded, for example, by entropy coding within an entropy coder 102, and the resulting data stream is output 114. The inverse transform step 104 reconstructs the quantized remnant, and the adder 110 in turn combines the reconstructed remnant with the corresponding prediction to obtain reconstructed information samples based on which the prediction module 106 can predict the above-mentioned currently coded prediction blocks. The prediction module 106 may use various prediction modes such as the intra prediction modes and the intra prediction modes for prediction blocks and prediction parameters are passed to an entropy encoder 102 for input into the data stream.
[0047] Ie, according to the Fig. 4 embodiment, the blocks of transform coefficients represent a spectral representation of the remainder of a sample system, rather than actual information samples thereof.
It should be noted that there are several alternatives to the embodiment of Fig. 4, some of which are described in the introductory part of the specification, the description of which is hereby introduced into the description of Fig. 4. For example, the prediction generated by the prediction module 106 may not be entropy coded. Rather, the side information may be communicated to the decoding side by a different encoding scheme.
[0049] Fig. 5 shows a decoder capable of decoding a data stream generated by the encoder of Fig. 4. The decoder of Fig. 5 includes an entropy decoder 150, an inverse transform step 152, an adder 154, and a prediction module 156. The entropy decoder 150, the inverse transform step 152 and the adder 154 are connected in series between the input 158 and the output 160 of the decoder of Fig. 5 in said order. Another output of the entropy decoder 150 is connected to the prediction module 156, which in turn is connected between the output of the addition module 154 and its next input. The entropy decoder 150 acquires from the data stream input to the decoder of Fig. 5 on input 158, blocks of transform coefficients, with an inverted transform being applied to blocks of transform coefficients in step 152 to obtain a residual signal. The residual signal is combined with the prediction from the prediction module 156 in the addition module 154 such as to obtain a reconstructed block of the reconstructed version of the sample pattern at the output 160. Based on the reconstructed versions, the prediction module 156 generates the predictions, thus rebuilding the predictions made by the sample module. 106 prediction on the encoder side. To obtain the same predictions as used at the encoder side, the prediction module 156 uses the prediction parameters which the entropy decoder 150 also obtains from the data stream at input 158.
[0050] It should be noted that in the above-described embodiments, the spatial granularity with which the prediction and residual transformation are performed need not be equal. This is shown in Fig. 2c. This figure shows a further breakdown, for the prediction blocks, the prediction granularity in solid lines and the residual granularity in dashed lines. As can be seen, the further divisions may be chosen independently by the encoder. More specifically, the syntax of the data stream may make it possible to define a further residual split independently of the further prediction split. Alternatively, the further division of the remainder may be an extension of the further division of the prediction such that each block of the remainder is either equal to or is a proper subset of the prediction block. This is shown, for example, in Fig. 2a and Fig. 2b, in which again the prediction granularity is shown in solid lines and the residual granularity is shown in dashed lines. These, in Figs. 2a-2c, all the blocks having an associated reference symbol would be residual blocks for which one two-dimensional transform would be performed, while the larger blocks surround with a solid line the dotted-line blocks 40a, for example, would be prediction blocks, for whose prediction parameters are set individually.
[0051] The above embodiments have in common that a block of samples (residual or original) is to be transformed on the encoder side into a block of transform coefficients which in turn is to be inversely transformed into the reconstructed sample block on the decoder side. This is illustrated in Fig. 6. Fig. 6 shows a sample block 200. In the case of Fig. 6, this block 200 is illustratively square and has a size of 4x4 samples 202. The samples 202 are regularly spaced in the horizontal x direction and vertical y direction. By means of the aforementioned two-dimensional transform T, block 200 is transformed into the spectral domain, namely to block 204 of transform coefficients 206, transform block 204 is the same size as block 200.
This transform block 204 has as many transform coefficients 206 as the block 200 has samples in both the horizontal and vertical directions. However, since the transform T is a spectral transform, the positions of the transform coefficients 206 within transform block 204 do not correspond to the spatial positions but rather to the spectral components of the contents of block 200. In particular, the horizontal axis of the transformation block 204 corresponds to the axis along which the spectral frequency in the horizontal direction increases monotonic, while the vertical axis corresponds to the axis along which the spectral frequency in the vertical direction increases monotonically, the transformation coefficient of the DC component being located at the corner-w. in this case, visually in the upper left corner - block 204, such that in the lower right corner is a transform factor 206 corresponding to the highest frequency in both the horizontal direction and the vertical direction. Disregarding the spatial direction, the spatial frequency to which a certain transformation factor 206 belongs generally increases from the upper left corner to the lower right corner. Using the inverted transformation of T<sup>1</sup> transform block 204 is re-transformed from the spectral domain to the spatial domain so as to re-acquire a copy 208 of block 200. In the event that quantization / transformation losses were not introduced, the reconstruction would be perfect.
[0052] As already mentioned above, it can be seen in Fig. 6 that the larger block size of the block 200 increases the spectral resolution of the resulting spectral representation 204. On the other hand, quantization noise tends to spread throughout block 208, and therefore sharp and extremely local objects within block 200 tend to deviate the re-transformed block from the original block 200 due to quantization noise. The main advantage of using larger blocks, however, is that the ratio between the number of significant, i.e. non-zero (quantized) transform coefficients on the one hand, and the number of insignificant transform coefficients on the other hand, may be reduced for larger blocks compared to smaller blocks, thus allowing for greater encoding efficiency. In other words, often, significant transformation factors, i.e. transform coefficients, not quantized to zero, are sparsely distributed in transform block 204. For this reason, according to the embodiments described in more detail below, the positions of the significant transform coefficients are signaled within the data stream by means of a significance map. Separately from it, the values of significant transformation coefficients, i.e. transform coefficient levels, in the case of quantized transform coefficients, are transferred within the data stream.
[0053] Accordingly, according to an embodiment of the present application, an apparatus for decoding such a significance map from a data stream or for decoding a significance map along the respective values of the transform coefficients from the data stream may be implemented as shown in Fig. 7, each of the entropy decoders mentioned in above, namely, decoder 50 and entropy decoder 150, may include the device of Fig. 7.
[0054] The apparatus of Fig. 7 comprises an entropy map / ratio decoder 250 and a binding module 252. The entropy map / ratio decoder 250 is connected to the input 254 on which syntax elements representing the map significance and significant values of the transform coefficients are input. As will be described in more detail below, there are various possibilities regarding the order in which the syntax elements describing the significance map on the one hand and the significant values of the transform coefficients on the other hand are input into the entropy map / coefficient decoder 250. Significance map syntax elements can precede the appropriate levels or they can be interleaved. However, it is tentatively assumed that the syntax elements representing the significance map precede the values (levels) of the significant transform coefficients, such that the entropy map / coefficient decoder 250 first decodes the significance map and then the transform coefficient levels of the significant transform coefficients.
[0055] When the entropy map / coefficient decoder 250 sequentially decodes syntax elements representing the significance map and significant values of the transform coefficients, the binding module 252 is configured to bind the sequentially decoded syntax elements / values to the positions within transform block 256. The scanning order in which the binding module 252 associates the sequentially decoded syntax elements representing the significance map and the levels of significant transform coefficients with the positions of transform block 256 conforms to the one-dimensional scanning order among the positions of transform blocks 256, which is identical to the sequence used on the coding side for input these elements into the data stream. As also will be outlined in more detail below, the scan order for the significance map syntax elements may or may not be the same order used for the significant coefficient values.
The coefficient map entropy decoder 250 may acquire the transform block information 256 available so far generated by the binding module 252 up to the syntax / level element to be decoded on-the-fly, to set the probability estimation context for decoding the entropy level syntax element to be decoded. be decoded on-the-fly as indicated by the dashed line 258. For example, binding module 252 may record information so far gathered from sequentially linked syntax elements, such as the same levels, or whether or not significant transform coefficients are present at the respective positions, or whether or not nothing is known about the corresponding position of transform block 256 wherein the map / coefficient entropy decoder 250 accesses memory. The memory just mentioned is not shown in Fig. 7, but the reference symbol 256 may also indicate this memory, since the memory or the register buffer will be dedicated to the pre-information obtained so far by the binding module 252 and the entropy decoder 250. Correspondingly, Fig. 7 represents with the items marked with a cross the significant transform coefficients obtained from the previously decoded syntax elements representing the significance map, and "1 will mean that the level of a significant transform factor, a significant transform factor at the corresponding position has already been decoded and is 1. In the case of significance map syntax elements preceding the significant values in the data stream, the cross will be recorded in memory 256 at position "1 (this situation would represent the entire significance map) before entering" 1 after decoding the corresponding value.
[0057] The following description focuses on specific embodiments of coding blocks of transform coefficients or significance maps, which embodiments can be easily translated into the embodiments described above. In these embodiments, a syntax binary coded_block_flag may be transmitted for each transform block that signals whether the transform block includes any level of a significant transform coefficient (i.e., transform coefficients that are not zero). If this syntax element indicates that levels of significant transform coefficients are present, the significance map is encoded, ie only then. The significance map determines, as mentioned above, which of the transformation coefficient levels contains non-zero values. The coded significance map comprises coding of binary significant_coeff_flag syntax elements, each of which determines, for suitably related coefficient positions, whether the corresponding transform coefficient level is not equal to zero. The coding is performed in a scan order that may change during the coding of the significance map depending on the positions of significant coefficients so far identified as significant, as described below in more detail. Additionally, the coding of the significance maps comprises encoding the last_significant_coeff_flag binary elements of the syntax distributed in the significant_coeff_flag sequence at its positions where significant_coeff_flag signals a significant factor. If the significant_coeff_flag container is unity, ie if there is a non-zero transform coefficient level in this scan entry, the next syntax last_significant_coeff_flag is encoded. This container indicates whether the current significant transform coefficient level is the last significant transform coefficient level within a block or whether successive levels of significant transform coefficients will follow in the scan order. If last_significant_coeff_flag indicates that no further significant transform coefficients will occur, no further syntax elements are coded for determining the significance map for the block. Alternatively, the number of significant coefficient positions may be signaled within the data stream before encoding the sequence significant_coeff_flag. In the next step, the values of the levels of significant transformation coefficients are encoded. As described above, alternatively, transmitting the levels may be interleaved with transmitting the significance map. The values of the significant transform coefficient levels are encoded in the next scan sequence, examples of which are described below. The following three syntax elements are used. The binary coeff_abs_greater_one element of the syntax indicates whether the absolute value of the significant transform coefficient level is greater than one. If the syntax binary coeff_abs_greater_one indicates that this absolute value is greater than one, another syntax coeff_abs_level_minus_one element is sent that specifies the absolute value of the transform coefficient level minus one. Finally, a syntax binary coeff_sign_flag that specifies the sign of a transform coefficient value is encoded for each level of a significant transform coefficient.
[0058] The embodiments described below make it possible to further reduce the bit rate and therefore increase the coding efficiency. To do this, these embodiments use a specific context modeling approach for the syntax elements related to the transform coefficient. In particular, the new selection of the context model is used for the significant_coeff_flag, last_significant_coeff_flag, coeff_abs_greater_one and coeff_abs_level_minus_one syntax elements. In addition, adaptive scan switching during coding / decoding of the significance map (determining the locations of non-zero levels of transform coefficients) is described. As to the meaning of the necessarily mentioned syntax elements, please refer to the introductory part of this application.
[0059] The coding of the significant_coeff_flag and last_significant_coeff_flag of the syntax that define the significance map is improved by adaptive scanning and new context modeling based on the defined neighborhood of the already coded scan positions. These new concepts lead to a more efficient significance map coding (i.e. a reduction in the corresponding bit rate), especially for large block sizes.
[0060] One aspect of the exemplary embodiments outlined below is that the scanning order (i.e. mapping a block of transform coefficient values to an ordered set (vector) of transform coefficient levels) is matched when encoding / decoding a significance map based on the values already encoded / decoded. syntax elements for the significance map.
[0061] In a preferred embodiment, the scanning order is adaptively switched between two or more preset scan schemes. In a preferred embodiment, switching may only take place at certain predetermined scan positions. In a further preferred embodiment of the invention, the scanning order is adaptively switched between two predetermined scan schemes. In a preferred embodiment, switching between the two preset scan patterns may only take place at certain preset scan positions.
[0062] An advantage of switching between scan schemes is a reduced bit rate which is a result of fewer coded syntax elements. As an intuitive example and with reference to Fig. 6, it is often the case that significant values of the transform coefficients - especially for large transform blocks - are clustered on one of block boundaries 270, 272, since remnant blocks mainly contain horizontal and vertical structures. When primarily using the zigzag scan 274, there is a probability of about 0.5 that the last diagonal zigzag scan sub-scan where the last significant factor is encountered starts from the side on which the significant factors are not focused. In this case, a large number of syntax elements must be encoded for transform coefficient levels equal to zero before reaching the last non-zero transform coefficient value. This can be avoided if diagonal sub-scans always start on the side on which the levels of the significant transform coefficients are focused.
[0063] More details for a preferred embodiment of the invention are described below.
[0064] As mentioned above, also with large block sizes it is advantageous to keep the number of context models relatively small in order to allow quick adaptation of the context models and ensure high coding efficiency. Hence, a given context should be used for more than one scan item. But the concept of assigning the same context to a number of consecutive scan positions, as is done for 8x8 blocks in H.264, is usually not appropriate because levels of significant transform coefficients are usually clustered in certain areas of transform blocks (this clustering may result from some dominant structures that are usually present, for example, in debris blocks). For the design of the context selection, one has to take into account the observation mentioned above that the levels of significant transform coefficients are often clustered in certain areas of a transform block. The concepts by which this observation can be used are described below.
[0065] In one preferred embodiment, the large transform block (e.g. greater than 8x8) is divided into a number of rectangular sub-blocks (e.g. into 16 sub-blocks), each of the sub-blocks associated with a separate context model for significant_coeff_flag and last_significant_coeff_flag encoding (with different context models being used for significant_coeff_flag and last_significant_coeff_flag). Division into sub-blocks may be different for significant_coeff_flag and last_significant_coeff_flag. The same context model can be used for all scan items that are in a given sub-block.
[0066] In another preferred embodiment, a large transform block (e.g., greater than 8x8) may be partitioned into a number of rectangular and / or non-rectangular sub-areas, each of these sub-areas associated with a separate context model for encoding significant_coeff_flag and / or last_significant_coeff_flag. . The subdivision may be different for significant_coeff_flag and last_significant_coeff_flag. The same context model is used for all scan positions that fall within a given subarea.
[0067] In a further preferred embodiment, the context model for the significant_coeff_flag and / or last_significant_coeff_flag encoding is selected based on the already encoded symbols in a predetermined spatial neighborhood of the current scan position. The preferred neighborhood may be different for different scanning positions. In a preferred embodiment, the context model is selected based on a number of levels of significant transform coefficients in a predetermined spatial neighborhood of the current scan position where only the already coded significance indications are encountered.
[0068] More details of a preferred embodiment of the invention are described below.
[0069] As mentioned above, for large block sizes, conventional context modeling encodes a large number of containers (which typically have different probabilities) with one single context model for the coeff_abs_greater_one and coeff_abs_level_minus_one syntax elements. To avoid this drawback in the case of a large block size, large blocks may, according to an exemplary embodiment, be broken into small square or rectangular sub-blocks of a given size and separate context modeling is applied for each sub-block. In addition, multiple context model sets may be used, with one of the context model sets being selected for each sub-block based on the analysis of the statistics of the previously coded sub-blocks. In a preferred embodiment of the invention, a number of transform coefficients greater than 2 (i.e. coeff_abs_level_minus_l> 1) in a previously coded subblock of the same block is used to derive the context model set for the current sub-block. These context modeling improvements for the coeff_abs_greater_one and coeff_abs_level_minus_one syntax elements lead to more efficient coding of both syntax elements, especially for large block sizes. In a preferred embodiment, the block size for the sub-block is 2x2. In another preferred embodiment, the sub-block size is 4x4.
[0070] In a first step, a block larger than the predetermined size may be divided into smaller sub-blocks of a predetermined size. The coding operation of absolute transform coefficient levels maps square or rectangular blocks of sub-blocks to an ordered set (vector) of sub-blocks by scanning, where different scans may be applied to different blocks. In a preferred embodiment, the sub-blocks are processed using a zigzag scan; the levels of the transform coefficients within the sub-block are converted in an inverse zigzag scan, i.e. with the scan loading from the transform coefficient belonging to the highest frequency in the vertical and horizontal directions to the coefficient associated with the lowest frequency in both directions. In another preferred embodiment of the invention, inverse zigzag scanning is used to encode the sub-blocks and to encode the levels of transform coefficients within the sub-blocks. In another preferred embodiment of the invention, the same adaptive scan that is used to encode the significance map (see above) is used to process the entire block of transform coefficient levels.
[0071] Breaking up large transform blocks into sub-blocks avoids the problem of using only one context model for most large transform block containers. Within the sub-blocks, prior art context modeling (as defined in H.264) or fixed context may be used depending on the actual size of the sub-blocks. Additionally, the statistics (in terms of probability modeling) for such sub-blocks are different from the statistics for a transform block of the same size. This property can be used by extending the set of context models for the coeff_abs_greater_one and coeff_abs_level_minus_one syntax elements. Multiple sets of context models may be provided and for each sub-block one of these context model sets may be selected based on the statistics of the previously coded sub-block in the current transform block or in the previously coded transform blocks. In a preferred embodiment of the invention, a selected set of context models is obtained based on statistics of previously coded sub-blocks in the same block. In another preferred embodiment of the invention, a selected set of context models is obtained based on statistics of previously coded sub-blocks in the same block. In a preferred embodiment, the number of context model sets is selected to be 4, while in another preferred embodiment, the number of context model sets is selected to be 16. In a preferred embodiment, the statistics that are used to derive the set of context models are the number of absolute levels of transform coefficients greater than 2 in the previously coded sub-blocks. In another preferred embodiment, the statistics that are used to derive the set of context models is the difference between the number of significant coefficients and the number of levels of transformation coefficients with an absolute value greater than 2.
[0072] The significance map coding may be implemented as outlined below, namely by adaptive scanning sequence switching.
[0073] In a preferred embodiment, the scanning order for the significance map coding is matched by switching between two predetermined scan schemes. Switching between scan schemes can only be performed at certain preset scan positions. The decision whether the scan scheme is toggled depends on the values of the already encoded / decoded Significance Map syntax elements. In a preferred embodiment, both of the predetermined scan schemes define the diagonal scan scans, similar to the zigzag scan pattern. The scan patterns are shown in Fig. 8. Both scan patterns 300 and 302 consist of a number of diagonal sub-scans for the diagonals from lower left to upper right or vice versa. A scan of diagonal sub-scans (not shown in the figure) is made from top left to bottom right for both predetermined scan schemes. But the scan for diagonal sub-scans is different (as shown in the figure). For the first scan pattern 300, diagonal sub-scans are scanned from bottom left to top right (left illustration in Fig. 8), and for the second scan pattern 302, diagonal sub-scans are scanned from top right to bottom left (right illustration in Fig. 8). In an embodiment, the significance map encoding starts with a second scan pattern. When encoding / decoding syntax elements, the number of significant values of the transform coefficients is counted by the two counters ci and C2. The first numerator ci counts the number of significant transform coefficients that appear in the lower-left part of the transform block; that is, the numerator is incremented by one when the level of a significant transform coefficient is encoded / decoded for which the horizontal x-coordinate inside the transform block is smaller than the vertical y-coordinate. The second numerator C2 counts the number of significant transform coefficients that are located in the upper right part of the transform block; i.e., the numerator is incremented by one when the level of a significant transform coefficient is encoded / decoded for which the horizontal x-coordinate inside the transform block is larger than the vertical y-coordinate. Adaptation of the counters can be performed by the binding module 252 of Fig. 7 and can be described by the following formulas, where t is the index of the scan position and both counters are initialized to zero:
CjCr + l) | cjtj,
In other cases
X> Γ
In other cases
[0074] At the end of each diagonal sub-scan, a decision is made by the binding module 252 whether to use the first or the second of the predefined scan patterns 300, 302 for the next diagonal sub-scan. This decision is based on the values of the ci and C2 counters. When the numerator for the lower left part of the transform block is greater than the numerator for the lower left part, a scan scheme is applied which scans diagonal sub-scans from bottom left to top right; otherwise (the numerator for the lower left part of the transformation block is less than or equal to the numerator for the lower left part), a scan scheme is used which scans diagonal scans from top right to bottom left. This decision can be expressed by the following formula:
<sub>f</sub> top right to bottom left, ci <C2 l bottom left to top right, ci> C2
[0075] It should be noted that the above-described embodiment of the invention can be easily applied to other scanning schemes. For example, the scanning scheme that is used for field macroblocks in H.264 can also be decomposed into subscans. In a further preferred embodiment, a given but arbitrary scan scheme is decomposed into sub-scans. For each of the sub-scans, two scan patterns are defined: one from bottom left to top right and one from top right to bottom left (as base scan directions). Additionally, two counters are introduced that count the number of significant coefficients in the first portion (near the lower left boundary of the transform blocks) and the second portion (near the upper right boundary of the transform blocks) within the sub-scans. Finally, at the end of each sub-scan, a decision is made (based on the values of the counters) whether the next sub-scan will be performed from bottom left to top right or top right to bottom left.
[0076] The following are embodiments of a method for modeling contexts by an entropy decoder 250.
[0077] In one preferred embodiment, context modeling for significant_coeff_flag is performed as follows. For 4x4 blocks, context modeling is performed as specified in H.264. For 8x8 blocks, a transform block is decomposed into 16 sub-blocks of 2x2 samples, each of these sub-blocks is associated with a separate context. It should be noted that the concept can also be extended to larger block sizes, a different number of sub-blocks as well as non-rectangular sub-areas as described above.
[0078] In a further preferred embodiment, the selection of the context model for larger transform blocks (e.g. for blocks larger than 8x8) is based on the number of already encoded significant transform coefficients in a predetermined neighborhood (within a transform block). An example of a neighborhood definition that corresponds to a preferred embodiment of the invention is shown in Fig. 9. The crosses surrounded by circles are available neighbors that are always included in the scoring, and the crosses in triangles are neighbors that are scored depending on the current scan position and the current scan direction):
• If the current scan position is inside the left corner 304 2x2, a separate context model is used for each scan position (Fig. 9, left illustration) • If the current scan position is not inside the left 2x2 corner and is not in the first row or first column of the transform block, the neighbors shown right in Fig. 9 are used to estimate the number of significant transform coefficients in the vicinity of the current scan position "x with nothing around them.
• If the current 'x scan position with nothing around it is in the first row of a transform block, the neighbors defined in the right illustration in Fig. 10 are used.
• If the current position "x scan" is in the first column of the block, the neighbors specified in the left illustration of Fig. 10 are used.
In other words, the decoder 250 may be configured to sequentially acquire the significance map syntax elements using context-adaptive entropy decoding by using contexts that are individually selected for each of the significance map syntax elements as a function of the number of positions according to significant transformation factors are found in previously acquired and related syntax elements of the significance map, the items being limited to the adjacent items ("x on the right side of Fig. 9 and on both sides of Fig. 10 and any of the highlighted items on the left side of Fig. 9) to which the corresponding current significance map syntax element is associated. As shown, the adjacent item to which the corresponding current syntax element is associated may only contain items either directly adjacent to or separate from the item to which the corresponding significance map syntax element is associated, at a maximum in one vertical position and / or one position horizontally. Alternatively, only entries immediately adjacent to the corresponding current syntax element may be included. At the same time, the block size of the transform coefficients may be equal to or greater than 8x8 positions.
[0080] In a preferred embodiment, the context model that is used to encode a given significant_coeff_flag is selected depending on the number of already encoded levels of significant transform coefficients in a defined neighborhood. Here, the number of available context models may be less than the potential value for the number of levels of significant transform coefficients in the defined neighborhood. The encoder and decoder may include a table (or other mapping mechanism) for mapping the number of levels of significant transform coefficients in a defined neighborhood to the context model index.
[0081] In a further preferred embodiment, the index of the selected context model depends on the number of levels of significant transform coefficients in the defined neighborhood and one or more additional parameters as the type of neighborhood or scan position used or a quantized value for the scan position.
[0082] For the last_significant_coeff_flag encoding, similar context modeling as for significant_coeff_flag can be used. However, the probability measure for last_significant_coeff_flag mainly depends on the distance of the current scan position from the upper-left corner of the transform block. In a preferred embodiment, the context model for the last_significant_coeff_flag encoding is selected based on the scan diagonal over which the current scan position lies (i.e. is selected based on x + y, where x and y represent the horizontal and vertical locations of the scan positions within the transform block, respectively, in the case of the above embodiment of Fig. 8, or based on how many sub-scans between the current sub-scan and the upper left DC entry (such as sub-scan index minus 1)). In a preferred embodiment of the invention, the same context is used for different values of x + y. Distance measure, i.e. x + y, or sub-scan index, is mapped to a set of context models in some way (e.g., by quantization x + y or sub-scan index), where the number of potential values for the distance measure is greater than the number of context models available for the la encoding st_sig nificant_coeff_flag.
[0083] In a preferred embodiment, different context modeling schemes are applied to different sizes of transform blocks.
[0084] The coding of the absolute levels of the significant transform coefficients is described below.
[0085] In one preferred embodiment, the size of the sub-blocks is 2x2, and context modeling within the sub-blocks is turned off, i.e., one single context model is used for all transform coefficients in a 2x2 sub-block. Only blocks larger than 2x2 can be subdivided. In a further preferred embodiment of this invention, the size of the sub-blocks is 4x4 and the context modeling within the sub-blocks is performed as in H.264; only blocks larger than 4x4 are subdivided.
As for the scanning order, a zigzag scan 320 is used in the preferred embodiment to scan the sub-blocks 322 of the transform block 256 i.e., in a substantially increasing frequency direction, while the transform coefficients within the sub-block are scanned in inverted zigzag scan 326. (Fig. 11). In another preferred embodiment of the invention, both the sub-blocks 322 and the levels of transform coefficients within the sub-blocks 322 are scanned using an inverted zigzag scan (as shown in Fig. 11 where arrow 320 is inverted). In another preferred embodiment, the same adaptive scan for the significance map coding is used to process the transform coefficient levels, the adaptation decision being the same such that exactly the same scan is used for both the significance map coding and the coding of the transform coefficient levels values. Note that the scan itself usually does not depend on the selected statistics, the number of context model sets, or the decision to enable or disable context modeling within sub-blocks.
[0087] Next, embodiments for modeling the context for the coefficient levels are described.
[0088] In a preferred embodiment, context modeling for a sub-block is similar to context modeling for 4x4 blocks in H.264 as described above. The number of context models that are used in encoding the coeff_abs_greater_one element of the syntax and the first container of the coeff_abs_level_minus_one element of the syntax is five, with, for example, using different context model sets for the two syntax elements. In another preferred embodiment, context modeling within sub-blocks is turned off and only one predefined context model is used within each sub-block. For both embodiments, the context model for sub-block 322 is selected from a predetermined number of context model sets. The selection of the context model set for sub-block 322 is based on some statistics for one or more already coded sub-blocks. In a preferred embodiment, the statistics used to select a set of context models for a sub-block are taken from one or more already coded sub-blocks in the same block 256. A method of using the statistics to derive a selected set of context models is described below. In another embodiment, statistics are taken from the same sub-block in the previously coded block, with the same block size as block 40a and 40a 'in Fig. 2b. In another preferred embodiment of the invention, statistics are taken from the defined adjacent sub-block in the same block, which depends on the selected scan for the sub-blocks. Also, it should be noted that the source of the statistics should be independent of the scanning order and the method of generating statistics for retrieving the set of context models.
[0089] In a preferred embodiment, the number of context model sets is four, while in another preferred embodiment, the number of context model sets is 16. Typically, the number of context model sets is not constant and should be matched according to the selected statistics. In a preferred embodiment, the set of context models for sub-block 322 is derived based on the number of absolute levels of transform coefficients greater than two in one or more already coded sub-blocks. An index for a set of context models is determined by mapping a number of absolute levels of transformation coefficients greater than two in the reference sub-block or sub-blocks onto a set of predetermined context model indexes. This mapping may be performed by quantizing the number of absolute levels of transformation coefficients greater than two or by a predetermined table. In a further preferred embodiment, the set of context models for a sub-block is derived based on the difference between the number of significant levels of transform coefficients and the number of absolute levels of transform coefficients greater than two in one or more already coded sub-blocks. An index for a set of context models is determined by mapping this difference to a set of indexes of predefined context models. This mapping may be performed by quantizing the difference between the number of levels of significant transform coefficients and the number of absolute levels of transform coefficients greater than two, or by a predetermined table.
[0090] In another preferred embodiment, when the same adaptive scan is used to process the absolute levels of the transform coefficients and the significance map, partial sub-block statistics in the same block may be used to derive a set of context models for the current sub-block, or if available, statistics of previously encoded sub-blocks in previously encoded transform blocks may be used. This means, for example, that instead of using the absolute number of absolute levels of transform coefficients greater than two in the sub-block (s) for deriving the context model, the number of already encoded absolute levels of transform coefficients greater than two multiplied by the ratio of the number of transform coefficients in sub-blocks are used for deriving the context model. -block (s) and number of already coded transform coefficients in the sub-block (s); or instead of using the difference between the number of levels of significant transformation coefficients and the number of absolute levels of transformation coefficients greater than two in the sub-block (s), the difference between the number of already encoded levels of significant transform coefficients and the number of already encoded absolute transform coefficient levels greater than two is used multiplied by the ratio of the number of transform coefficients in the sub-block (s) and the number of already encoded transform coefficients in the sub-block (sub-block) blocks).
[0091] For modeling context within sub-blocks, the inversion of prior art context modeling from H.264 may generally be used. This means that when the same adaptive scan is applied to process the absolute levels of the transform coefficients and the significance map, the transform coefficient levels are essentially encoded in the forward scan order, rather than the inverted scan order as in H.264. Hence, switching the context model has to be properly adjusted. According to one embodiment, the coding of the transform coefficient levels starts with a first context model for the syntax coeff_abs_greater_one and coeff_abs_level_minus_one and is switched to the next context model in the set when two coeff_abs_greater_one syntax elements have been coded zero since the last context model switch. In other words, the choice of context depends on the number of coeff_abs_greater_one syntax elements already encoded greater than zero in the scan order. The number of context models for coeff_abs_greater_one and coeff_abs_level_minus_one may be the same as in H.264.
[0092] Thus, the above embodiments can be applied in the field of digital signal processing, and in particular in image and video decoders and encoders. In particular, the above embodiments enable the coding of syntax elements related to transform coefficients in block-based image and video codecs with improved context modeling for the syntax elements related to transform coefficients that are encoded with an entropy encoder that uses probability modeling. Compared to the prior art, an increased coding efficiency is achieved, in particular for large transform blocks.
[0093] While some aspects have been described in the context of an apparatus, it is obvious that these aspects also represent a description of a corresponding method, wherein the block or device corresponds to a method step or a feature of a method step. Likewise, the aspects described in the context of a method step also represent a description of a corresponding block or position or feature of the corresponding device.
[0094] The inventive encoded signal for representing a transform block or a significance map, respectively, may be stored on a digital storage medium or may be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0095] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. Implementation may be via digital storage media, e.g., floppy disks, DVDs, Blue-Ray disks, CDs, ROMs, PROM, EPROM, EEPROM, or FLASH, having electronically readable control signals stored thereon that interact (or are capable of such interaction) with the programmed computer system, so that an appropriate method is performed. Hence, the digital storage medium can be computer readable.
[0096] Some of the embodiments of the invention include a data carrier having electronically readable control signals that are operable to interact with a programmable computer system such that one of the methods described herein is performed.
[0097] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operable to perform one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a machine-readable medium.
[0098] Other embodiments include a computer program for performing one of the methods described herein stored on a machine-readable medium.
[0099] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program product runs on a computer.
[0100] A further embodiment of the inventive method is, therefore, a data medium (or a digital storage medium or a computer readable medium) containing a computer program stored thereon for performing one of the methods described herein.
[0101] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may, for example, be configured to be transmitted over a data link, such as the Internet.
[0102] Another embodiment comprises processing means, for example a computer or programmable logic device, configured or adapted to perform one of the methods described herein.
[0103] Another embodiment includes a computer in which the computer program for performing one of the methods described herein is installed.
[0104] In some embodiments, a programmable logic device (e.g., a user programmable logic table) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a user programmable logic table may interact with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0105] The above described embodiments are merely illustrative for the principles of the present invention. It should be understood that modifications and variations to the systems and details described herein are apparent to those skilled in the art. It is the intention, therefore, to be limited only by the scope of the following patent claims and not by the specific details presented for the description and explanation of the embodiments herein.
GE Video Compression, LLC, United States of America
Proxy:
ΕΡ 2 559 244 BI
Z-16443/17
26 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26
255 members in 21 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 2010054822 | European Patent Office (EPO) | W | |
| 10159766 | European Patent Office (EPO) | A | |
| 2011055644 | European Patent Office (EPO) | W |
Members255
| Document | Office | Kind | |
|---|---|---|---|
| WO2011128303A2 | World Intellectual Property Organization (WIPO) | A2 | |
| TW201142759A | Taiwan Province of China | A | |
| WO2011128303A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20130006678A | Republic of Korea | A | |
| CN102939755A | China | A | |
| EP2559244A2 | European Patent Office (EPO) | A2 | |
| US2013051459A1 | United States of America | A1 | |
| JP2013524707A | Japan | A | |
| EP2693752A1 | European Patent Office (EPO) | A1 | |
| KR20140071466A | Republic of Korea | A | |
| JP2014131286A | Japan | A | |
| KR20150009602A | Republic of Korea | A | |
| KR101502495B1 | Republic of Korea | B1 | |
| TW201537954A | Taiwan Province of China | A | |
| JP2015180071A | Japan | A | |
| CN105187829A | China | A | |
| US2016080742A1 | United States of America | A1 | |
| KR101605163B1 | Republic of Korea | B1 | |
| KR101607242B1 | Republic of Korea | B1 | |
| KR20160038063A | Republic of Korea | A | |
| JP5911027B2 | Japan | B2 | |
| US9357217B2 | United States of America | B2 | |
| BR112012026388A2 | Brazil | A2 | |
| TWI545525B | Taiwan Province of China | B | |
| US2016309188A1 | United States of America | A1 | |
| US2016316212A1 | United States of America | A1 | |
| EP2693752B1 | European Patent Office (EPO) | B1 | |
| JP6097332B2 | Japan | B2 | |
| PT2693752T | Portugal | T | |
| DK2693752T3 | Denmark | T3 | |
| KR101739808B1 | Republic of Korea | B1 | |
| KR20170058462A | Republic of Korea | A | |
| ES2620301T3 | Spain | T3 | |
| PL2693752T3 | Poland | T3 | |
| TWI590648B | Taiwan Province of China | B | |
| US9699467B2 | United States of America | B2 | |
| JP2017130941A | Japan | A | |
| EP2559244B1 | European Patent Office (EPO) | B1 | |
| HUE032567T2 | Hungary | T2 | |
| PT2559244T | Portugal | T | |
| DK2559244T3 | Denmark | T3 | |
| EP3244612A1 | European Patent Office (EPO) | A1 | |
| TW201742457A | Taiwan Province of China | A | |
| ES2645159T3 | Spain | T3 | |
| LT2559244T | Lithuania | T | |
| HRP20171669T1 | Croatia | T1 | |
| NO2559244T3 | Norway | T3 | |
| PL2559244T3This record | Poland | T3 | |
| SI2559244T1 | Slovenia | T1 | |
| US9894368B2 | United States of America | B2 | |
| RS56512B1 | Serbia | B1 | |
| US2018084261A1 | United States of America | A1 | |
| CY1119639T1 | Cyprus | T1 | |
| US9998741B2 | United States of America | B2 | |
| US10021404B2 | United States of America | B2 | |
| EP3244612B1 | European Patent Office (EPO) | B1 | |
| US2018220135A1 | United States of America | A1 | |
| US2018220136A1 | United States of America | A1 | |
| US2018220137A1 | United States of America | A1 | |
| CN102939755B | China | B | |
| HUE037038T2 | Hungary | T2 | |
| KR20180095951A | Republic of Korea | A | |
| KR20180095952A | Republic of Korea | A | |
| KR20180095953A | Republic of Korea | A | |
| CN108471534A | China | A | |
| CN108471537A | China | A | |
| CN108471538A | China | A | |
| US2018255306A1 | United States of America | A1 | |
| US2018295369A1 | United States of America | A1 | |
| CN105187829B | China | B | |
| TWI640190B | Taiwan Province of China | B | |
| KR101914176B1 | Republic of Korea | B1 | |
| KR20180119711A | Republic of Korea | A | |
| US10123025B2 | United States of America | B2 | |
| CN108777790A | China | A | |
| CN108777791A | China | A | |
| CN108777792A | China | A | |
| CN108777793A | China | A | |
| CN108777797A | China | A | |
| DK3244612T3 | Denmark | T3 | |
| US10129549B2 | United States of America | B2 | |
| PT3244612T | Portugal | T | |
| TR2018015695T4 | Türkiye | T4 | |
| TR201815695T4 | Türkiye | T4 | |
| CN108881910A | China | A | |
| CN108881922A | China | A | |
| ES2692195T3 | Spain | T3 | |
| US10148968B2 | United States of America | B2 | |
| EP3410716A1 | European Patent Office (EPO) | A1 | |
| JP6438986B2 | Japan | B2 | |
| JP2018201202A | Japan | A | |
| JP2018201203A | Japan | A | |
| JP2018201204A | Japan | A | |
| CN109151485A | China | A | |
| EP3435674A1 | European Patent Office (EPO) | A1 | |
| PL3244612T3 | Poland | T3 | |
| US2019037221A1 | United States of America | A1 | |
| KR101951413B1 | Republic of Korea | B1 | |
| KR20190019221A | Republic of Korea | A | |
| HUE040296T2 | Hungary | T2 |
Numbers
- Application
- 11713791
Titles2
- English
- CODING OF SIGNIFICANCE MAPS AND TRANSFORM COEFFICIENT BLOCKS
- Polish
- Kodowanie map istotności i bloków współczynników transformacji
Classification
- CPC, 15
- H04N19/18
- H04N19/124
- H04N19/129
- H04N19/13
- H04N19/136
- H04N19/139
- H04N19/176
- H04N19/46
- H04N19/50
- H04N19/51
- H04N19/59
- H04N19/61
- H04N19/70
- H04N19/91
- H04N19/122
- IPC, 14
- H04N19 50
- H04N19 124
- H04N19 129
- H04N19 13
- H04N19 136
- H04N19 139
- H04N19 176
- H04N19 18
- H04N19 46
- H04N19 51
- H04N19 59
- H04N19 61
- H04N19 70
- H04N19 91
