Variable length coding table selection based on block type statistics for refinement coefficient coding
Abstract
This disclosure describes techniques for encoding an enhancement layer in a scalable video encoding scheme (SVC). The techniques can be used in variable length coding of refinement coefficients for an improvement layer of an SVC scheme. According to this disclosure, a method may comprise determining first statistics associated with a first type of video block, determining second statistics associated with a second type of video block, selecting a first variable length coding table (VLC) from a plurality of VLC tables to be used in encoding the first type of video block based on the first statistics, select a second VLC table from the plurality of VLC tables to be used in encoding the second type of video block based on the second statistics, encoding video blocks of the first type based on the first VLC table, and encoding blocks second type of video based on the second VLC table.

Term
1.3 yearsleft in the term
Expires 4 January 2028.
- Priority
- Filed
- Granted
- Today
- Expires
24 claims: 3 independent, 21 dependent
- 1CLAIMS REIVINDICAÇÕES 1. Method of encoding a scalable video encoding scheme (SVC) enhancement layer, the method comprising:1. Método de codificar uma camada de melhoramento de um esquema de codificação de video escalonável (SVC), o método compreendendo: determinar primeiras estatísticas associadas com um primeiro tipo de bloco de vídeo;determine first statistics associated with a first type of video block;determinar segundas estatísticas associadas com um segundo tipo de bloco de vídeo;determine second statistics associated with a second type of video block;selecionar uma primeira tabela de codificação de comprimento variável (VLC) a partir de uma pluralidade de tabelas de VLC a ser usada em codificação do primeiro tipo de bloco de vídeo com base nas primeiras estatísticas;selecting a first variable length coding table (VLC) from a plurality of VLC tables to be used in encoding the first type of video block based on the first statistics;selecionar uma segunda tabela de VLC a partir da pluralidade de tabelas de VLC a ser usada em codificação do select a second VLC table from the plurality of VLC tables to be used in coding the base na segunda tabela de VLC. based on the second VLC table.
- 9Device encoding a scalable video encoding scheme (SVC) enhancement layer, the device comprising:9. Dispositivo que codifica uma camada de melhoramento de um esquema de codificação de vídeo escalonável (SVC), o dispositivo compreendendo: a statistics module that determines first statistics associated with a first type of video block and determines second statistics associated with a second type of video block;um módulo de estatísticas que determina primeiras estatísticas associadas com um primeiro tipo de bloco de vídeo e determina segundas estatísticas associadas com um segundo tipo de bloco de vídeo;a table selection module that selects a first variable length encoding table (VLC) from a plurality of VLC tables to be used in encoding the first type of video block based on the first statistics and selecting a second table VLC from the plurality of VLC tables to be used in encoding the second type of video block based on the second statistics;and an encoding module that encodes video blocks of the first type based on the first VLC table and um módulo de seleção de tabela que seleciona uma primeira tabela de codificação de comprimento variável (VLC) a partir de uma pluralidade de tabelas de VLC a ser usada em codificação do primeiro tipo de bloco de vídeo com base nas primeiras estatísticas e seleciona uma segunda tabela de VLC a partir da pluralidade de tabelas de VLC a ser usada em codificação do segundo tipo de bloco de vídeo com base nas segundas estatísticas;e um módulo de codificação que codifica blocos de vídeo do primeiro tipo com base na primeira tabela de VLC e 4/8 codifica blocos de vídeo do segundo tipo com base na segunda tabela de VLC. 4/8 encodes video blocks of the second type based on the second VLC table.
- 18Computer readable medium comprising instructions that when running on a video encoding device make the device encode a scalable video encoding scheme (SVC) enhancement layer, in which the instructions make the device:18. Meio legível por computador compreendendo instruções que quando em execução em um dispositivo de codificação de vídeo fazem o dispositivo codificar uma camada de melhoramento de um esquema de codificação de vídeo escalonável (SVC), em que as instruções fazem o dispositivo: determinar primeiras estatísticas associadas com um primeiro tipo de bloco de vídeo;determine first statistics associated with a first type of video block;determinar segundas estatísticas associadas com um segundo tipo de bloco de vídeo;determine second statistics associated with a second type of video block;selecionar uma primeira tabela de codificação de comprimento variável (VLC) a partir de uma pluralidade de tabelas de VLC a ser usada em codificação do primeiro tipo de bloco de vídeo com base nas primeiras estatísticas;selecting a first variable length coding table (VLC) from a plurality of VLC tables to be used in encoding the first type of video block based on the first statistics;selecionar, uma segunda tabela de VLC a partir da pluralidade de tabelas de VLC a ser usada em codificação do segundo tipo de bloco de vídeo com base nas segundas estatísticas;selecting, a second VLC table from the plurality of VLC tables to be used in encoding the second type of video block based on the second statistics;Scalable (SVC), the device comprising: escalonável (SVC), o dispositivo compreendendo: means for determining statistics that determine first statistics associated with a first type of video block and determines second statistics associated with a second type of video block;meios para determinar estatísticas que determinam primeiras estatísticas associadas com um primeiro tipo de bloco de vídeo e determina segundas estatísticas associadas com um segundo tipo de bloco de vídeo;means for selecting which select a first variable length coding table (VLC) from a plurality of VLC tables to be used in encoding the first type of video block based on the meios para selecionar que selecionam uma primeira tabela de codificação de comprimento variável (VLC) a partir de uma pluralidade de tabelas de VLC a ser usada em codificação do primeiro tipo de bloco de vídeo com base nas 7/8 first statistics and select a second VLC table from the plurality of VLC tables to be used in encoding the second type of video block based on the second statistics;and means for encoding that encode video blocks of the first type based on the first VLC table and encode video blocks of the second type based on the second VLC table. 7/8 primeiras estatísticas e selecionam uma segunda tabela de VLC a partir da pluralidade de tabelas de VLC a ser usada em codificação do segundo tipo de bloco de vídeo com base nas segundas estatísticas;e meios para codificar que codificam blocos de vídeo do primeiro tipo com base na primeira tabela de VLC e codificam blocos de vídeo do segundo tipo com base na segunda tabela de VLC.
Independent claims3
178 paragraphs in 8 sections, as filed
(54) Title: SELECTION OF VARIABLE LENGTH CODING TABLE BASED ON BLOCK TYPE STATISTICS FOR REFINING COEFFICIENT (30) Unionist Priority: 10/05/2007 us 11 / 868,017, 05/01/2007 US 60 / 883,741.05 / 01/2007 US 60 / 883,741.05 / 10/2007 US 11 / 868,017 (73) Holder (s): Qualcomm Incorporated (72) Inventor (s): Hyukjune Chung, Marta Karczewicz, Phoom Sagetong (74 ) Attorney (s): Montaury Pimenta, Machado & Lioce (86) International Request: pct us2008050261 de 04/01/2008 (57) Abstract: selection of coding table of VARIABLE LENGTH BASED ON STATISTICS OF TYPE OF BLOCK FOR CODING OF REFINING COEFFICIENT. This disclosure describes techniques for encoding an enhancement layer in a scalable video encoding scheme (SVC). The techniques can be used in variable length coding of refinement coefficients for an improvement layer of an SVC scheme. According to this disclosure, a method may comprise determining first statistics associated with a first type of video block, determining second statistics associated with a second type of video block, selecting a first variable length coding table (VLC) from a plurality of VLC tables to be used in encoding the first type of video block based on the first statistics, select a second VLC table from the plurality of VLC tables to be used in encoding the second type of video block based on the second statistics, encoding video blocks of the first type based on the first VLC table, and encoding blocks second type of video based on the second VLC table.
(87) International Publication: wo 2008 / 086i97de 17/07/2008
<img file="BRPI0806304A2_D0001.tif" />
<img file="BRPI0806304A2_D0002.tif" />
ΡΙ0806304-4
SELECTION OF VARIABLE LENGTH CODING TABLE BASED ON BLOCK TYPE STATISTICS FOR CODING
REFINING COEFFICIENT
This order claims to benefit from the following US Provisional Order, the entire contents of which is incorporated herein by reference: US Provisional Order No. 60,883,741, filed on January 5, 2007.
TECHNICAL FIELD
This disclosure relates to digital video encoding and, more particularly, variable length encoding (VLC) of transform coefficients in layers to improve a scalable video encoding scheme (SVC).
FUNDAMENTALS
Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless communication devices, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording device, video game devices, video game consoles, satellite radio or cell phones, and the like. Digital video devices implement video compression techniques, such as MPEG-2, MPEG-4, or H.264 / MPEG-4, Part 10, Advanced Video Encoding (AVC), to more efficiently transmit and receive digital video . Video compression techniques perform spatial and temporal prediction to reduce or remove the redundancy inherent in video sequences.
In video encoding, video compression often includes spatial prediction, motion estimation and motion compensation. Intra-coding depends on spatial prediction to reduce or remove the
2/40 spatial redundancy between video blocks within a given video frame. Inter-coding depends on the temporal prediction to reduce or remove the temporal redundancy between the video blocks of successive video frames of a video sequence. For inter-coding, a video encoder performs motion estimation to follow the movement of coincident video blocks between two or more adjacent frames. Motion estimation generates motion vectors, which indicate the displacement of video blocks relative to corresponding prediction video blocks in one or more reference frames. Motion compensation uses motion vectors to generate the prediction video blocks of a frame of reference. After motion compensation, a residual video block is formed by subtracting the prediction video block from the original video block to be encoded.
The video encoder generally applies transform, quantization and transform coefficient encoding processes to further reduce the bit rate associated with residual block communication. The coding of the transform coefficients of the residual blocks may involve the application of variable length codes to further compress the coefficients produced by the transform and quantization operations. For example, a variable length coding table (VLC) can be used to match different sets of coefficients to variable length codewords in a way that promotes coding efficiency. The different VLC tables can be used for different video content. A video decoder performs reverse VLC operations to reconstruct the coefficients, and then inversely transforms the coefficients to
3/40 reconstruct the video information. The video decoder can decode the video information based on the motion information and residual information associated with the video blocks.
Some video encodings make use of scalable techniques. For example, scalable video encoding (SVC) refers to video encoding in which a base layer and one or more scalable enhancement layers are used. For SVC, a base layer typically loads video data with a base level of quality. One or more enhancement layers carry additional video data to support higher spatial, temporal and / or SNR levels. The base layer can be transmitted in a way that is more reliable than the transmission of enhancement layers. Enhancement layers can add spatial resolution to frames in the base layer, or they can add additional frames to increase the total frame rate. In one example, the most reliable portions of a modulated signal can be used to transmit the base layer, while the less reliable portions of the modulated signal can be used to transmit the enhancement layers. The improvement layers can define the different types of coefficients, referred to as significant coefficients and refinement coefficients.
SUMMARY
In general, this disclosure describes techniques for encoding an enhancement layer in a scalable video encoding scheme (SVC). The techniques provide for the selection of variable length coding tables (VLC) during the coding and decoding processes. The techniques can be used in coding blocks for transform coefficients, and can be
4/40 particularly useful in VLC of block refinement coefficients of an improvement layer of an SVC scheme. The refinement coefficients refer to coefficients of an improvement layer for which the corresponding coefficients of a preceding layer in the SVC scheme had values other than zero. The VLC of refinement coefficients can be performed separately from VLC of significant coefficients, which refer to coefficients of the improvement layer for which the corresponding coefficients of a previous layer in the SVC scheme had values of zero.
According to the techniques of this disclosure, VLC tables are selected for different types of video blocks, for example, intra and inter blocks. Tables can be selected once for each video information frame, or they could be selected once for other types of encoding units (such as, once per video information slice or once per FGS layer of a frame) . The VLC tables for different types of video blocks can be selected based on the statistics associated with the previously encoded blocks. For example, a VLC table for intra blocks can be selected based on the statistics associated with the previously coded intra blocks. Similarly, a VLC table for inter blocks can be selected based on the statistics associated with the previously coded inter blocks. In one example, the statistics for each type of video block can comprise a relationship of the number of refinement coefficients in the previously coded blocks that had the same signal value relative to the number of refinement coefficients in the previously coded blocks that had a value of reversed sign. Based on this
5/40 ratio, VLC tables can be selected to encode the refinement coefficients associated with the blocks of a given frame, and when the next frame is found, the ratio can be calculated again to facilitate VLC table selections for this picture.
In one example, this disclosure provides a method of encoding an enhancement layer for an SVC scheme, the method comprising determining first statistics associated with a first type of video block, determining second statistics associated with a second type of video block, selecting a VLC table from a plurality of VLC tables to be used in encoding the first type of video block based on the first statistics, select a second VLC table from the plurality of VLC tables to be used in encoding the second type of video block based on the second statistics, encoding the video blocks of the first type based on the first VLC table and encoding blocks second type of video based on the second VLC table.
In another example, this disclosure provides a device that encodes an enhancement layer for an SVC scheme, the device comprising a statistics module that determines first statistics associated with a first type of video block and determines second statistics associated with a second type of video. video block, a table selection module that selects a first VLC table from a plurality of VLC tables to be used in encoding the first type of video block based on the first statistics and selects a second VLC table from the plurality VLC tables to be used in encoding the second type of video block based on the
6/40 second statistics and an encoding module that encodes video blocks of the first type based on the first VLC table and encodes video blocks of the second type based on the second VLC table.
In another example, this disclosure provides a device that encodes an enhancement layer for an SVC scheme, the device comprising means for determining statistics that determine first statistics associated with a first type of video block and determine second statistics associated with a second type of video block, means for selecting to select a first VLC table from a plurality of VLC tables to be used in encoding the first type of video block based on the first statistics and selecting a second VLC table from the plurality of video tables VLC to be used in encoding the second type of video block based on the second statistics and means for encoding that encode video blocks of the first type based on the first VLC table and encode video blocks of the second type based on the second VLC table.
The techniques described in this disclosure can be implemented in hardware, software, firmware or any combination thereof. If implemented in software, the software can run on one or more processors, such as a microprocessor, application specific integrated circuit (ASIC), field programmable port arrangement (FPGA) or digital signal processor (DSP). The software that performs the techniques can be initially stored on a computer-readable medium and loaded and run on the processor.
Consequently, this disclosure also includes a computer-readable medium that includes instructions
7/40 that when running on a video encoding device it causes the device to encode an enhancement layer of an SVC scheme, in which the instructions make the device determine first statistics associated with a first type of video block, determine second statistics associated with a second type of video unit, select a first VLC table from a plurality of VLC tables to be used in encoding the first type of video block based on the first statistics, select a second VLC table from the plurality of VLC tables to be used in encoding the second type of video unit based on the second statistics, encode video blocks of the first type based on the first VLC table and encode video blocks of the second type based on the second VLC table.
In some cases, the computer-readable medium may form part of a computer program product, which can be sold to manufacturers and / or used in a video encoding device. The computer program product may include the computer-readable medium, and in some cases, may also include packaging materials.
In other cases, this disclosure can be directed to a circuit, such as an integrated circuit, group of integrated circuits (chipset), application-specific integrated circuit (ASIC), field programmable port arrangement (FPGA), logic or various combinations
<td>configured from</td><td>same</td><td>for</td><td>perform a</td><td>or</td><td>more of</td>
<td>techniques described</td><td>on here.</td><td></td><td></td><td></td><td></td>
<td colspan="2">The details of</td><td>one or</td><td>more aspects</td><td>gives</td><td>revelation</td>
<td>are determined</td><td colspan="2">in the drawings</td><td>In attached is</td><td>at</td><td>description</td>
below. Other characteristics, objects and advantages of
8/40 techniques described in this disclosure will be apparent from the description and drawings, and from the claims.
BRIEF DESCRIPTION OF THE DRAWINGS
FIG. 1 is an exemplary block diagram illustrating a video encoding and decoding system.
FIG. 2 is a conceptual diagram illustrating video frames of a base layer and a scalable video bit stream enhancement layer.
FIG. 3 is a block diagram illustrating an example of a video encoder consistent with this disclosure.
FIG. 4 is a block diagram illustrating an example of a video decoder consistent with this disclosure.
FIG. 5 is an exemplary block diagram of a variable length encoding (VLC) encoding unit.
FIG. 6 is an exemplary block diagram of a VLC decoding unit.
FIG. 7 is a flow diagram illustrating a VLC technique for variable length coding consistent with this disclosure.
DETAILED DESCRIPTION
This disclosure describes techniques for encoding an enhancement layer in a scalable video encoding scheme (SVC). The techniques provide the selection of variable length coding tables (VLC) in an encoder and in a decoder. That is, the VLC table selection techniques that are reciprocal in that VLC table selection are performed on the encoder to encode information and on the decoder to decode the information. The techniques can be used in coding coefficients of
9/40 transformed, and are particularly useful in variable length encoding of refinement coefficients of an improvement layer of an SVC scheme. The refinement coefficients refer to coefficients of an improvement layer for which the corresponding coefficients of a preceding layer in the SVC scheme had values other than zero. On the contrary, the significant coefficients refer to coefficients of an improvement layer for which the corresponding coefficients of a previous layer in the SVC scheme had values of zero. The variable length coding of refinement coefficients can be performed separately from the variable length coding of significant coefficients.
According to the techniques of this disclosure, the VLC tables are selected for different types of video blocks, for example, intra and inter blocks. Tables can be selected once per coded unit, for example, once per frame, once per video information slice, once per FGS layer of a frame. The VLC tables for different types of video blocks can be selected based on statistics associated with the previously encoded blocks. For example, a VLC table for intra blocks can be selected based on the statistics. associated with the previously coded intra blocks, and a VLC table for inter blocks can be selected based on the statistics associated with the previously coded inter blocks.
In one example, the statistics for each type of video block can comprise a relationship of the number of refinement coefficients in the previously coded blocks of that type of block that had the same value
10/40 of signal relative to the number of refinement coefficients in the previously coded blocks of that type of block that had an inverted signal value. Based on the relationship for each type of block (intra and inter), a first VLC table can be selected to code the refinement coefficients associated with the intra blocks of a given frame, and a second VLC table can be selected to code the refinement coefficients associated with the inter blocks of the given frame. When the next frame is found, the ratios can be recalculated to facilitate VLC table selections for that frame.
FIG. 1 is a block diagram illustrating a video encoding and decoding system 10. As shown in FIG. 1, the system 10 includes a source device 2 which transmits the encoded video to a receiving device 6 via a communication channel 15. The source device 2 may include a video source 11, the video encoder 12 and a modulator / transmitter 14. The receiving device 6 can include a receiver / demodulator 16, a video decoder 18 and a display device 20. System 10 can be configured to apply techniques for VLC of video information associated with an enhancement layer in an SVC scheme .
SVC refers to video encoding in which a base layer and one or more scalable enhancement layers are used. For SVC, a base layer typically loads video data with a low level of quality. One or more enhancement layers carry additional video data to support higher spatial, temporal and / or signal-to-noise (SNR) levels. Improvement layers can be defined with respect to
11/40 previously coded layer. The improvement layers define at least two different types of coefficients, referred to as significant coefficients and refinement coefficients. The refinement coefficients can define values relative to the corresponding values of the previously encoded layer. Enhancement layer frames sometimes include only a portion of the total number of video blocks in the base layer or in the previous enhancement layer, for example, only those blocks for which the improvement is performed.
The significant coefficients refer to coefficients for which the corresponding coefficients in the preceding layer had values of zero. The refinement coefficients refer to coefficients for which the corresponding coefficients in the preceding layer had non-zero values in the preceding layer. Variable length coding of enhancement layers typically involves an approximation of two passages. A first pass is made to encode significant coefficients in variable length, and another pass is made to encode refinement coefficients in variable length. The techniques of this disclosure are particularly useful for encoding variable length coefficients of disclosure other than limited respect.
refinement, although it is necessarily to this
According to the techniques of this development, different VLC tables are selected for different types of video blocks. For example, a first VLC table can be selected for intra-block coding refinement coefficients and a second VLC table can be selected for coefficients of
12/40 inter block coding refinement. Intra blocks refer to blocks that are encoded based on blocks within that given encoded unit. Inter blocks refer to blocks that are encoded based on blocks from another encoded unit.
The VLC table selections can be based on the statistics associated with the previously coded blocks. For example, a VLC table for intra blocks can be selected based on the statistics associated with the previously coded intra blocks, and a VLC table for inter blocks can be selected based on the statistics associated with the previously coded inter blocks. VLC tables can be selected once per encoded unit, such as, once per video information frame, once per video information slice or once per FGS layer. FGS stands for Fine-Grain Signal-to-Noise Scalability, and is explained in more detail below.
The statistics for each type of video block can comprise a relationship of the number of refinement coefficients in the previously coded blocks that had the same signal value relative to the number of refinement coefficients in the previously coded blocks that had an inverted signal value. Based on a relationship for each type of block, that is, the relationship for intra blocks and the relationship for inter blocks, the first and second VLC tables can be selected to encode the refinement coefficients associated with the intra blocks and the inter blocks of the given frame (or other encoded unit), respectively. When the next frame (or another encoded unit) is found, the ratios can be calculated again to facilitate updated selections from the VLC table.
13/40
In the example of FIG. 1, the communication channel 15 may comprise any wireless or wired means of communication, such as a radio frequency (RF) spectrum or one or more physical transmission lines or any combination of wireless and wired means. Communication channel 15 can be part of a packet-based network, such as a local area network, a wide area network, or a global network, such as the Internet. The communication channel 15 generally represents any appropriate communication medium or collection of different communication means, for transmitting video data from the source device 2 to the receiving device 6.
The source device 2 generates encoded video data for transmission to the receiving device 6. In some cases, however, devices 2, 6 may operate in a substantially symmetrical manner. For example, each of the devices - 2, 6 can include encoding and video decoding components. In this way, system 10 can support unidirectional or bidirectional video transmission between video devices 2, 6, for example, for streaming video, broadcasting video or video telephony.
The video source 11 of the source device 2 can include a video capture device, such as a video camera, a video file containing previously captured video, or a video feed from a video content provider. As an additional alternative, video source 11 can generate data based on computer graphics as the source video or a combination of live and computer generated live video. In some cases, if the video source 11 is a video camera, the source device 2 and the receiving device 6 can form so-called camera phones or telephones
14/40 with video. In each case, the captured, pre-captured or computer generated video can be encoded by the video encoder 12 for transmission from the video source device 2 to the video decoder 18 of the video receiving device 6 through the modulator / transmitter 14, communication channel 15 and receiver / demodulator 16. The video encoding and decoding processes can implement the VLC table selection techniques described here to improve the processes. 0 display device 20 displays the decoded video data to a user, and can comprise some of a variety of display devices, such as a cathode ray tube, a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display or another type of display device.
Video encoder 12 and video decoder 18 can be configured to support SVC for spatial, temporal and / or noise ratio (SNR) scalability. In some respects, video encoder 12 and video decoder 18 can be configured to support fine-grained SNR scaling (FGS) encoding for SVC. Encoder 12 and decoder 18 can support varying degrees of scalability by supporting the encoding, transmission and decoding of a base layer and one or more scalable enhancement layers. Again, for scalable video encoding, a base layer loads the video data with a quality baseline level. One or more enhancement layers carry additional data to support higher spatial, temporal and / or SNR levels. The base layer can be transmitted in a way that is more reliable than the
15/40 transmission of improvement layers. For example, the most reliable portions of a modulated signal can be used to transmit the base layer, while the less reliable portions of the modulated signal can be used to transmit the enhancement layers.
In order to support SVC, video encoder 12 can include a base layer encoder 22 and one or more enhancement layer encoders 24 to encode a base layer and one or more enhancement layers, respectively . The techniques of this disclosure, which involve the selection of VLC table, are applicable to the encoding of video blocks of enhancement layers in SVC. More specifically, the techniques of this disclosure are applicable to VLC coefficients of refinement of video blocks of enhancement layers, although this disclosure is not necessarily limited in this regard. '
The video decoder 18 may include a combined base / enhancement decoder that decodes the video blocks associated with the base and enhancement layers. The video decoder 18 can decode the video blocks associated with the base and enhancement layers, and combines the decoded video to reconstruct the frames of a video sequence. The display device 20 receives the decoded video stream, and presents the video stream to a user.
Video encoder 12 and video decoder 18 can operate according to a video compression standard, such as, MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG- 4, Part 10, Advanced Video Encoding (AVC). Although not shown in FIG. 2, in some aspects, the video encoder 12 and the video decoder 18
16/40 can each be integrated with an encoder and an audio decoder, and can include appropriate MUX-DEMUX units, or other hardware and software, to handle both audio and video encoding in a common data stream or in separate data streams. If applicable, MUX-DEMUX units can conform to the ITU H.223 multiplexer protocol or to other protocols, such as the user datagram protocol (UDP).
The H.264 / MPEG-4 (AVC) standard was formulated by VCEG (Group of experts in video coding) ITU-T together with MPEG (Group of cinematographic experts) ISO / IEC as the product of a collective partnership known as JVT (Common video team). In some ways, the techniques described in this disclosure can be applied to devices that generally conform to the H.264 standard. The H.264 standard is described in the H.264 ITU-T recommendation, Advanced Video Encoding for generic audiovisual services, by the ITU-T Study Group, and dated March 2005, which can be referred to here as the H standard .264 or H.264 specification or H.264 / AVC standard or specification.
The JVT continues to work on an extension of SVC to H.264 / MPEG-4 AVC. The specification of the SVC extension under development is in the form of a Joint Draft (JD). The JSVM (Common Scalable Video Model) created by the JVT implements tools for use in scalable video, which can be used within System 10 for the various encoding tasks described in this disclosure. Detailed information regarding fine-grained SNR Scaling (FGS) coding can be found in the Joint Draft documents, and particularly in Joint Draft 6 (SVC JD6), Thomas Wiegand, Gary Sullivan, Julien Reichel,
17/40
Heiko Schwarz and Mathias Wien, Joint Draft 6: Scalable Video Coding, JVT-S 201, April 2006, Geneva, and in Joint Draft 9 (SVC JD9), Thomas Wiegand, Gary Sullivan, Julien Reichel, Heiko Schwarz and Mathias Wien, Joint Draft 9 of SVC Amendment, JVT-V 201, January 2007, Marrakesh, Morocco.
In some respects, for video broadcasting, the techniques described in this disclosure can be applied to enhanced H.264 video encoding to deliver real-time video services over multicast land mobile multimedia (TM3) systems using the Air Interface Specification of Direct Link Only (FLO), Forward Link Only Air Interface Specification for Terrestrial Mobile Multimedia Multicast to be published as Technical Standard TIA-1099 (the FLO Specification). That is to say, that the communication channel 15 may comprise a wireless information channel used to broadcast the wireless video information according to the FLO Specification or the like. The FLO specification includes examples that define the bitstream syntax and semantics and decoding processes appropriate for the FLO Aerial Interface. Alternatively, the video can be transmitted according to other standards such as DVB-H (digital video broadcast - portable device), ISDB-T (integrated services digital broadcast - terrestrial) or DMB (digital media broadcast). Thus, the source device 2 can be a mobile wireless terminal, a video streaming server or a video broadcast server. However, the techniques described in this disclosure are not limited to any particular type of broadcast, multicast or point-to-point system. In the case of broadcast, source device 2 can broadcast multiple channels of video data to multiple
18/40 reception, each of which may be similar to the receiving device 6 shown in FIG. 1.
The video encoder 12 and video decoder 18 each can be implemented as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable port arrangements (FPGAs), logic discrete, software, hardware, firmware or any combination of these. Each of the video encoders
12 and decoder 18 can be included in one or more encoders or decoders, both of which can be integrated as part of a combined encoder / decoder (CODEC) on a respective mobile device, subscriber device, broadcast device, server or similar. In addition, source device 2 and receiving device 6 may each include appropriate modulation, demodulation, frequency conversion, filtering and amplifier components for transmitting and receiving encoded video, as applicable, including wireless components and antennas enough radio frequency (RF) to support wireless communication. For ease of illustration, however, such components are summarized as being modulator / transmitter 14 of the source device 2 and receiver / demodulator 16 of the receiving device 6 in FIG. 1.
A video sequence includes a series of video frames. The video encoder 12 operates in pixel blocks (or transformed coefficient blocks) within the individual video frames in order to encode the video data. Video blocks can be fixed or variable in size, and can differ in size according to a specified encoding standard. In some
In 19/40 cases, each video frame is a coded unit, while in other cases, each video frame can be divided into a series of slices that form the coded units. Each slice can include a series of macroblocks, which can be arranged in sub-blocks (partitions and subpartitions). As an example, the ITU-T H.264 standard supports intra prediction in various block sizes, such as 16 by 16, 8 by 8 or 4 by 4 for luma components, and 8x8 for chroma components, as well as inter prediction in various block sizes, such as 16 by 16, 16 by 8, 8 by 16, 8 by 8, 8 by 4, 4 by 8 and 4 by 4 for luma components and corresponding sized sizes for chroma components.
Smaller video blocks can provide better resolution, and can be used for positions in a video frame that include higher levels of detail. Generally, macroblocks (MBs) and the various sub-blocks can be considered to be blocks of video. In addition, a slice can be considered to be a series of video blocks, such as MBs and / or sub-blocks. As noted, each slice can be an independent decodable unit of a video frame.
After intra or inter based predictive coding, additional coding techniques can be applied to the transmitted bit stream. These additional encoding techniques may include transformation techniques (such as a 4x4 or 8x8 integer transform used in H.264 / AVC or a discrete cosine transform (DCT)) and variable length encoding. The blocks of transformation coefficients can be referred to as video blocks. That is, the term video block refers to a video data block independent of the information domain. Thus, the video blocks can be in
20/40 a pixel domain or a transformed coefficient domain. The application of VLC table selection and VLC coding will be described generally in this disclosure with respect to the transform coefficient blocks.
This disclosure provides techniques for variable length coding of refinement coefficients. In addition, the refinement coefficients refer to coefficients that had values other than zero in the preceding layer, while the significant coefficients refer to coefficients that had values of zero in the preceding layer. According to this disclosure, encoder 12 and decoder 18 select different VLC tables for different types of video blocks. For example, encoder 12 and decoder 18 can select a first VLC for intra block encoding refinement coefficients and can select a second VLC table for inter block encoding refinement coefficients. The VLC table selections by encoder 12 and decoder 18 can be based on the statistics associated with the previously encoded blocks. For example, a VLC table for intra blocks can be selected based on the statistics associated with the previously coded intra blocks, and a VLC table for inter blocks can be selected based on the statistics associated with the previously coded inter blocks. In this way, encoder 12 and decoder 18 can perform reciprocal methods that encode an enhancement layer in an SVC scheme. As used here, the term encoding generally refers to at least a portion of the encoding or decoding processes. The video encoder 12 encodes the
21/40 data, while video decoder 18 decodes the data.
VLC tables can assign code words to different sets of transform coefficients. Sets of zero value coefficients can be represented by run lengths of zero, and tables can assign more likely careers to shorten VLC codes. Similarly, VLC tables can assign less likely careers to lengthen VLC codes. Therefore, selecting codes from VLC tables can improve coding efficiency. Alternatively, different coefficient patterns, such as coded block patterns, can be assigned with different variable length code words, with the most likely patterns being assigned with the shortest code words and the least likely patterns being assigned with the longest code words.
The formation of VLC tables could also be based on previous coding statistics, but in most cases, static VLC tables are used. In the case of static VLC tables, encoder 12 and decoder 18 simply teach an appropriate VLC table from a set of possible tables to encode intra-block refinement coefficients and select another appropriate VLC table from of the set of possible tables to code the inter-block refinement coefficients. Regardless of whether the VLC tables are static or dynamically formed, updates to the VLC tables could be made, as desired.
FIG. 2 is a diagram illustrating the video frames within a base layer 17 and enhancement layer 18 of a scalable video bit stream. As noted
22/40 above, the techniques of this disclosure are applicable to the coding of improvement layer data. The base layer 17 may comprise a bit stream that contains the encoded video data that represents the first level of spatial, temporal or SNR scalability. The enhancement layer 18 may comprise a bit stream that contains the encoded video data representing a second level of spatial, temporal and / or SNR scalability. Although a single enhancement layer is shown, several enhancement layers can be used in some cases. The bitstream of the enhancement layer can be decodable only in conjunction with the base layer (or the preceding enhancement layer, if multiple enhancement layers exist). The enhancement layer 18 contains references to video data decoded in the base layer 17. Such references can be used in the transform domain or in the pixel domain to generate the final decoded video data.
The base layer 17 and the enhancement layer 18 may contain intra (I), inter (P) and bidirectional (B) frames. Intra frames can include all intra coded video blocks. Frames I and P can include at least some inter-encoded video blocks (inter-blocks), but they can also include some intra-encoded blocks (intra-blocks). The different frames of the enhancement layer 18 do not need to include all the video blocks in the base layer 17. P frames in enhancement layer 18 depend on references to P frames in base layer 17. By decoding frames in enhancement layer 18 and base layer 17, a video decoder can increase the video quality of the decoded video. For example, the base layer 17 can include video encoded at a minimum frame rate of, for example,
For example, 15 frames per second, while enhancement layer 18 can include video encoded at a higher frame rate of, for example, 30 frames per second. To support encoding at different levels of quality, the base layer 17 and the enhancement layer 18 can be encoded with a higher quantization parameter (QP) and lower QP, respectively. In addition, the base layer 17 can be transmitted in a way that is more reliable than the transmission of the enhancement layer 18. As an example, the most reliable portions of a modulated signal can be used to transmit the base layer 17, while the less reliable portions of the modulated signal can be used to transmit the enhancement layer 18. The illustration of FIG. 2 it is merely exemplary, because the base and improvement layers could be defined in many different ways.
FIG. 3 is a block diagram illustrating an example of a video encoder 50 that includes a VLC unit 46 for encoding data consistent with this disclosure. The video encoder 50 of FIG. 3 can correspond to the enhancement layer encoder 24 of the source device 2 in FIG. 1. That is, the base layer encoding components are not illustrated in FIG. 3 for simplicity. Consequently, the video encoder 50 can be considered an enhancement layer encoder. Alternatively, the illustrated components of the video encoder 50 could also be implemented in combination with the base layer encoding modules or units, for example, in a pyramid encoder design that supports the scalable video encoding of the base layer and the improvement layer.
24/40
The video encoder 50 can perform intra and inter block encoding within the video frames. Intra coding depends on spatial prediction to reduce or remove spatial redundancy in the video within a given video frame. Inter-coding depends on temporal prediction to reduce or remove temporal redundancy in the video within the adjacent frames of a video sequence. For inter-coding, video encoder 50 performs motion estimation to follow the motion of matching video blocks between two or more adjacent frames. For intra coding, spatial prediction is used to identify other blocks within a frame that match well with the block being encoded. Spatial prediction components of intra coding, are not illustrated in FIG. 3.
As shown in FIG. 3, the video encoder 50 receives a current video block 31 (e.g., a video block from the enhancement layer) within a video frame to be encoded. In the example of FIG. 3, the video encoder 50 includes motion estimation unit 33, frame reference store 35, motion compensation unit 37, block transform unit 39, quantization unit 41, reverse quantization unit 42, transform unit reverse 44 and VLC unit 46. An unlock filter (not shown) can also be included to filter block edges to remove block artifacts. The video encoder 50 also includes adder 48 and adder 51. FIG. 3 illustrates the time prediction components of the video encoder 50 for the inter-encoding of the video blocks. Although not shown in FIG. 3 For ease of illustration, video encoder 50 may also include prediction components
25/40 space for the intra coding of some video blocks. The components of spatial prediction, however, are
<td>used</td><td>usually only for</td><td>The</td><td>coding</td><td colspan="2">layer</td>
<td>base.</td><td></td><td></td><td></td><td></td><td></td>
<td></td><td>The pet unit</td><td>in</td><td>movement 33</td><td colspan="2">compares the</td>
<td>block</td><td>video 31 with the blocks</td><td>in</td><td>one or more</td><td>frames</td><td>in</td>
<td>video</td><td>adjacent to generate</td><td>one</td><td colspan="2">or more vectors</td><td>in</td>
<td colspan="3">movement. 0 frame or frames</td><td>adj acentes</td><td>can</td><td>to be</td>
retrieved from the reference frame store 35, which can comprise any type of data storage device or memory to store the reconstructed video blocks from the previously encoded blocks. Motion estimation can be performed for blocks of varying sizes, for example, 16x16, 16x8, 8x16, 8x8 or smaller block sizes. The motion estimation unit 33 identifies a block in an adjacent frame that best matches the current video block 31, for example, based on a rate distortion model, and determines an offset between the blocks. On this basis, the motion estimation unit 33 produces a motion vector (MV) (or multiple MVs in the case of bidirectional prediction) that indicates the magnitude and trajectory of the displacement between the current video block 31 and a predictive block used for encode the current video block 31.
Motion vectors can have half pixel or quarter pixel precision, or even finer precision, allowing the video encoder 50 to follow the movement with higher precision than full pixel positions and obtain a prediction block best. When motion vectors with fractional pixel values are used, interpolation operations are performed on the motion compensation unit 37. The motion estimation unit 33 can identify the best vector
26/40 motion for a video block using a rate distortion model. Using the resulting motion vector, the motion compensation unit 37 forms a prediction video block by motion compensation.
The video encoder 50 forms a residual video block by subtracting the prediction video block produced by the motion compensation unit 37 from the original current video block 31 in adder 48. The block transform unit 39 applies a transform, such as a discrete cosine transform (DCT), to the residual block, producing the residual transform block coefficients. The quantization unit 41 quantizes the residual transform block coefficients to further reduce the bit rate. Adder 49A receives the base layer coefficient information, for example, from a base layer encoder (not shown) and is positioned between the block transform unit 39 and the quantization unit 41 to provide this information of base layer coefficient in the improvement layer coding. In particular, adder 49A subtracts base layer coefficient information from the output of block transform unit 39. In a similar way, adder 49B, which is positioned between the reverse transform unit 44 and the quantization unit inverse 42, also receives the base layer coefficient information from the base layer encoder (not shown). Adder 49B adds the base layer coefficient information back to the output of the inverse quantization unit 42.
The spatial prediction coding operates very similar to the temporal prediction coding. However, while temporal prediction coding depends on adjacent frame blocks (or other units
27/40 coded) to perform the coding, the spatial prediction depends on blocks within a common frame (the other coded unit) to perform the coding. Spatial prediction coding codes for intra blocks, while temporal prediction coding codes for inter blocks. Again, the spatial prediction components are not shown in FIG. 3 for simplicity.
The VLC unit 46 encodes the quantized transform coefficients according to a variable length encoding methodology to further reduce the bit rate of transmitted information. In particular, the VLC unit 46 applies techniques of this disclosure to encode the refinement coefficients of an improvement layer. VLC unit 46 may include VLC tables that map sets of coefficients into variable length code words.
The selection of the VLC table by the VLC unit 46 is performed based on the information collected for previously encoded frames. In addition, VLC tables are selected for different types of video blocks, for example, intra blocks and inter blocks. The VLC unit 46 can select VLC tables once per encoded unit, for example, once per frame, once per video information slice or once per FGS layer of a frame. The VLC tables for different types of video blocks can be selected based on the statistics associated with the previously encoded blocks. For example, the VLC 4 6 unit can select a VLC table for the intra blocks based on the statistics associated with the previously coded intra blocks, and the VLC 4 6 unit can select a VLC table for the inter blocks based on in the statistics associated with the inter-coded blocks
28/40 previously. In this case, the statistics associated with the previously encoded blocks can comprise the average number of non-zero coefficients with such previously encoded blocks.
After variable length encoding, the encoded video can be transmitted to another device. In addition, the inverse quantization unit 42 and the inverse transform unit 44 apply inverse quantization and inverse transformation, respectively, to reconstruct the residual block. Adder 51 adds the reconstructed residual block to the compensated motion prediction block produced by the motion compensation unit 37 to produce a reconstructed video block for storage in the frame reference store 35. The reconstructed video block is used by the unit motion estimator 33 and motion compensation unit 37 to encode a block in a subsequent video frame.
FIG. 4 is a block diagram illustrating an example of a video decoder 60, which can correspond to the video decoder 18 of FIG. 1 or a decoder from another device. The video decoder 60 includes a VLC unit 52A for enhancement layer information, which performs the reciprocal function of the VLC unit 46 of FIG. 3. That is to say, like the VLC unit 46, the VLC unit 52A encodes the refinement coefficients of an enhancement layer.
The video decoder 60 can also include another VLC 52B unit for base layer information. The intra prediction unit 55 can optionally perform spatial decoding of the base layer video blocks, and the output of the intra prediction unit 55 can be supplied to adder 53. The layer path of
The improvement can include the reverse quantization unit 58A, and the base layer path can include the reverse quantization unit 56B. The information in the base layer and improvement layer paths can be combined by the adder 57.
The video decoder 60 can perform the intra and inter decoding of the blocks within the video frames. In the example of FIG. 4, video decoder 60 includes VLC units 52A and 52B (mentioned above), motion compensation unit 54, reverse quantization units 56A and 56B, reverse transform unit 58 and frame reference store 62. Video decoder 60 also includes adder 64. Optionally, video decoder 60 can also include an unblocking filter (not shown) that filters the output of adder 64. In addition, adder 57 combines information in the base layer and enhancement layer paths, and the intra prediction 55 and adder 53 facilitate any spatial decoding of base layer video blocks.
According to this disclosure, the VLC unit 52A receives the encoded video bit stream and applies the VLC techniques described in this disclosure. In particular, for refinement coefficients, the VLC unit 5.2.A can select VLC tables for the different types of video block based on the information collected for previously encoded frames. The VLC 52A unit can select VLC tables once per encoded unit, for example, once per frame, once per video information slice, once per FGS layer of a frame. VLC tables for different types of video blocks can be selected based on the statistics associated with the blocks previously
30/40 coded. For example, the VLC 52A unit can select a VLC table for the intra blocks based on the statistics associated with the previously coded intra blocks, and the VLC 52A unit can select a VLC table for the inter blocks based on the statistics. associated with the previously coded inter blocks.
After decoding performed by the VLC unit 52A, the motion compensation unit 54 receives the motion vectors and one or more reconstructed reference frames from the reference frame store 62. The inverse quantization unit 56A quantizes in an inverse way , that is, de-quantize, the quantized block coefficients. After combining the enhancement layer and base information by the adder 57, the reverse transform unit 58 applies an inverse transform, for example, an inverse DCT, to the coefficients to produce residual blocks. The motion compensation unit 54 produces the motion compensated blocks which are added by the adder 64 to the residual blocks to form decoded blocks. If desired, an unlock filter can also be applied to filter the decoded blocks in order to remove the blocking artifacts. The filtered blocks are then placed in the reference frame store 62, which provides reference blocks from the motion compensation and also produces the decoded video to a drive display device (such as the device 20 of FIG. 1).
FIG. 5 is a block diagram illustrating an exemplary VLC unit 46, which may correspond to that shown in FIG. 3. The VLC unit 4 6 includes an encoding module 72, a statistics module 74, a table selection module 76 and the VLC tables 78. The tables
31/40 of VLC 78 generally refer to tables that can be stored in any position, for example, locally or outside the chip in a separate memory location. The VLC 78 tables can be updated, periodically, as desired.
encoding module 72 encodes refinement coefficients and significant coefficients in separate coding passages. The table selection by the VLC unit 46 for the encoding of coefficients associated with the different video blocks can be performed based on the information collected for previously encoded frames. For example, the statistics module 74 can perform statistical analysis of previously coded tables to facilitate table selection by the table selection module 76.
statistics module 74 determines the first statistics associated with a first type of video block (such as an intra block), and determines the second statistics associated with a second type of video block (such as an inter block). The table selection module 76 selects a first VLC table from a plurality of VLC tables 78 to be used in encoding the first type of video block based on the first statistics. In addition, the table selection module 76 selects a second VLC table from the plurality of VLC tables 78 to be used in encoding the second type of video block based on the second statistics. The encoding module 72 encodes the video blocks of the first type based on the first VLC table, and encodes the video blocks of the second type based on the second VLC table.
The techniques described here can be performed with respect to refinement coefficients, which
32/40 can be encoded in a separate coding pass for significant coefficients. The refinement coefficients can have the values restricted to -1, 0 and 1, which can be encoded by two bits of information. The first bit can indicate whether the coefficient is equal to 0 or not, and the second bit can indicate whether the sign (marked as s<sub>n</sub>) of the refinement coefficient is the same (coeff_ref_dir_flag = 0) or different (coeff_ref_dir_flag = 1) than the sign (marked as s<sub>n</sub>_i) the corresponding coefficient of the previous layer. The preceding layer is marked as s<sub>n</sub>-i. If the sign of the current coefficient is the same as that of the previous layer, then coeff_ref_dir_flag = 0, and if the sign of the current coefficient is different than that of the previous layer then coeff_ref_dir_flag = 1. The two refinement bits can be combined into an alphabet of three refinement symbols as follows in Table 1:
TABLE 1
<td>coeff ref flag</td><td>coeff ref_dir flag</td><td>ref_symbol</td>
<td> 0</td><td> -</td><td> 0</td>
<td> 1</td><td> 0</td><td> 1</td>
<td> 1</td><td> 1</td><td> 2</td>
Alternatively, another scheme could also be used to encode the refinement coefficients without departing from the techniques of this disclosure.
VLC tables 78 can comprise variable length code words that are mapped into different sets of coefficients, which can be defined by symbols, flags or other types of bits. The VLC 78 tables can be updated as desired. Any number of tables can be included in the VLC 88 tables. In some cases, two tables are used, although
33/40 more could be included. In any case, the encoding module 72 can access different tables of VLC tables for different types of video blocks. The statistics module 74 and the table selection module 76 determine which VLC table should be used for each type of video block being encoded.
Table 2 provides an example of a VLC table that could be used for coding refinement coefficients.
TABLE 2
<td>Group of ref symbol</td><td>Length of code</td><td>Codeword</td>
<td> {0,0,0}</td><td> 1</td><td> 1</td>
<td> {0,0,1}</td><td> 4</td><td> 0011</td>
<td> {0,0,2}</td><td> 5</td><td> 00101</td>
<td> {0,1,0}</td><td> 3</td><td> 011</td>
<td> {0,1,1}</td><td> 6</td><td> 000101</td>
<td> {0,1,2}</td><td> 8</td><td> 00000101</td>
<td> {0,2,0}</td><td> 5</td><td> 00100</td>
<td> {0,2,1}</td><td> 7</td><td> 0000101</td>
<td> {0,2,2}</td><td> 9</td><td> 000000101</td>
<td> {1,0,0}</td><td> 3</td><td> 010</td>
<td> {1,0,1}</td><td> 6</td><td> 000100</td>
<td> {1,0,2}</td><td> 8</td><td> 00000100</td>
<td> {1,1,0}</td><td> 6</td><td> 000011</td>
<td> {1,1,1}</td><td> 9</td><td> 000000100</td>
<td> {1,1,2}</td><td> 10</td><td> 0000000011</td>
<td> {1,2,0}</td><td> 7</td><td> 0000100</td>
<td> {1,2,1}</td><td> 10</td><td> 0000000010</td>
<td> {1,2,2}</td><td> 12</td><td> 000000000011</td>
<td> {2,0,0}</td><td> 5</td><td> 00011</td>
<td> {2,0,1}</td><td> 7</td><td> 0000011</td>
34/40
<td> {2,0,2}</td><td> 9</td><td> 000000011</td>
<td> {2,1,0}</td><td> 8</td><td> 00000011</td>
<td> {2,1,1}</td><td> 10</td><td> 0000000001</td>
<td> {2,1,2}</td><td> 12</td><td> 000000000010</td>
<td> {2,2,0}</td><td> 9</td><td> 000000010</td>
<td> {2,2,1}</td><td> 12</td><td> 000000000001</td>
<td> {2,2,2}</td><td> 12</td><td> 000000000000</td>
As shown in Table 2, different sets of refinement coefficients (as defined in Table 1) can be mapped to different code words of varying length. Table 2 also lists the respective bit lengths associated with the different code words. The code word mappings in different sets of refinement coefficients may differ in different VLC tables. Therefore, by selecting the appropriate table, coding efficiency can be achieved. According to this disclosure, for each coded unit (for example, each frame, slice or layer of FGS), the table selection module 7 6 of the VLC unit 4 6 selects a first VLC table for intra blocks and selects one second VLC table for inter blocks. The table selection can be based on the statistics associated with the previously coded intra blocks and the previously coded inter blocks. The encoding module 72 of the VLC unit 46 then uses the tables selected in the VLC process.
The statistics for each type of video block can be accumulated and analyzed by the statistics module 74. As an example, the statistics for each type of video block can comprise a relationship of the number of refinement coefficients in the previously coded blocks that had a same signal value
35/40 relative to the number of refinement coefficients in the previously coded blocks that had an inverted signal value. Based on this relationship for each type of video block, the table selection module 7 6 can select a VLC table to encode the refinement coefficients associated with that type of block in a given frame. When the next frame (or other encoded unit) is found, VLC unit 46 can recalculate the ratios for each type of block to facilitate VLC table selections for that frame (or the other encoded unit).
refinement symbol 1 in Table 1 corresponds to the scenario where the refinement symbol has the same value as the sign relative to the symbol in the previous layer (or in the base layer). The refinement symbol 2 in Table 1 corresponds to the scenario where the refinement symbol has an inverted value of the sign relative to the symbol in the preceding layer (or in the base layer). In other words, the symbol 1 means keep the same sign and the symbol 2 means invert the sign relative to the sign of the corresponding coefficient in the preceding layer.
The coding efficiency in SVC can be improved when the VLC table selection is based on the ratio of ref_symbols 1 and 2. Let s (l) and (2) denote respectively the number of refinement symbols 1 and 2 collected in the process of coding. The values for s (l) and s (2) could be set by sliding picture windows, or could accumulate over a full video sequence. In any case, a ratio r can be calculated in a number of ways. For example, a relationship r = (s<sub>max</sub>s<sub>m</sub>in) / sec<sub>max</sub>, can be calculated, where s<sub>max</sub> = max (s (1), s (2)) and Smin = min (s (1), s (2)). Alternatively, a ratio r = s (1) / (s (1) + s (2)) could be used. In each of these
36/40 cases, for each quantized value of r, an R value<sub>Q</sub> can be defined as being equal to floor (m * r), where m is some number greater than 1. Different VLC tables can be assigned depending on whether the ratio r is above or below the value R<sub>Q</sub>.
FIG. 6 is a block diagram illustrating an exemplary VLC unit 52A, which may correspond to that shown in FIG. 4. The VLC 52A unit performs reciprocal decoding functions relative to the encoding that is performed by the VLC 46 unit. Thus, while the VLC 46 unit receives quantized residual coefficients and generates a bit stream, the VLC 52A unit receives _ a stream of bits and generates quantized residual coefficients. The VLC 52A unit includes a decoding module 82, a statistics module 84, a table selection module 86 and a set of VLC tables 88. As in unit 46, the VLC tables 88 of unit 52A refer to usually tables that can be stored in any position, for example, locally or off the chip in a separate memory position. VLC 88 tables can be updated, periodically, as desired. Any number of tables can be included in the VLC 88 tables. In some cases two tables are used, although more could be included.
The VLC 82 decoding unit can perform separate decoding passes for significant coefficients and refinement coefficients. The techniques of this disclosure may be applicable to coding or refinement coefficients only, or could be used for both significant and refinement coefficients. The decoding performed by the VLC 52A unit is reciprocal to the encoding performed by the VLC unit 46.
37/40
The table selection by the VLC 52A unit for decoding the coefficients associated with the different video blocks can be performed based on the information collected for previously encoded blocks, for example, from previously encoded frames. For example, statistics module 84 can perform statistical analysis on previously decoded frame blocks to facilitate table selection by table selection module 86. In particular, statistics module 84 determines the first statistics associated with a first type of video block (such as, an intra block), and determines the second statistics associated with a second type of video block (such as, a block inter). The table selection module 86 selects a first VLC table from a plurality of VLC tables 88 to be used in encoding the first type of video block based on the first statistics. In addition, the table selection module 86 selects a second VLC table from the plurality of VLC tables 88 to be used in encoding the second type of video block based on the second statistics. The decoding module 82 decodes the video blocks of the first type based on the first VLC table and decodes the video blocks of the second type based on the second VLC table.
Table 2 above can also be seen as one of the VLC 88 tables. However, while the VLC 78 tables (FIG. 5) map sets of coefficients in variable length code words, the VLC 88 tables (FIG. 6) map the variable length code words back into the coefficient sets. In this way, the decoding performed by the VLC 52A unit
38/40 can be seen as being reciprocal to the coding performed by the VLC unit 46.
FIG. 7 is a flow chart illustrating a coding technique for variable length coding of the coefficients (e.g., typically refinement coefficients) of an improvement layer consistent with this disclosure. The coding process of FIG. 7 applies to encoding and decoding. As shown in FIG. 7, a statistics module 74, 84 determines statistics of the previously coded intra blocks (91). For example, statistics module 74, 84 can calculate for the pre-coded intra blocks, a relationship of refinement symbols that have the same signal value relative to refinement symbols having an inverted signal value. In addition, the statistics module 74, 84 can determine statistics for the previously encoded blocks (92), for example, by calculating previously encoded blocks, a list of refinement symbols that have the same sign value relative to refinement symbols having an inverted signal value.
table selection module 76, 86 selects a coding table for intra blocks based on the statistics of previously coded intra blocks (93). The coding table for intra blocks, for example, can be selected based on the value of the relationship associated with the intra blocks. In addition, the table selection module 76, 86 selects a coding table for the inter blocks based on the statistics of the previously coded inter blocks (94). The coding table for inter blocks can be selected based on the value of the relationship associated with the inter blocks.
39/40
Coding module 82, 84 intra codes blocks using the coding table selected for intra blocks (95), and inter codes blocks using the coding table selected for inter blocks (96). In particular, coding module 82, 84 performs table queries using the coding tables selected for the different types of block. The process can be repeated for each coded unit (97). The encoded units can be video frames, video frame slices, FGS layers or the like.
The techniques described here can be implemented in hardware, software, firmware or any combination of these. Any features described as modules or components can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be performed at least in part by a computer-readable medium that comprises instructions that, when executed, perform one or more of the methods described above. The computer-readable medium may be part of a computer program product, which may include packaging materials. The computer-readable medium can comprise random access memory (RAM) as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), read-only memory electrically erasable programmable (EEPROM), FLASH memory, magnetic or optical data storage and the like. The techniques additionally, or alternatively, can be performed at least in part by a computer-readable communication medium that loads or communicates the code in the form of
40/40 instructions or data structures that can be accessed, read and / or executed by a computer.
The code can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application-specific integrated circuits (ASICs), field programmable port arrangements (FPGAs) or another set equivalent circuits of integrated or discrete logic. Consequently, the term processor as used here can refer to some prior structure or any other structure suitable for implementing the techniques described here. In addition, in some respects, the functionality described here can be provided within the dedicated software modules or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
If implemented in hardware, this disclosure can be directed to a circuit, such as an integrated circuit, the chipset-specific application integrated circuit (ASIC), field programmable port arrangement (FPGA), logic or various configured combinations thereof to perform one or more of the techniques described here.
The various embodiments of the invention have been described. These and other modalities are within the scope of the following claims.
Contents8
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
11 priority claims, no other members on record
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 60883741 | United States of America | – | |
| 88374107 | United States of America | P | |
| 11868017 | United States of America | – | |
| 86801707 | United States of America | A | |
| 2008050261 | United States of America | W | |
| 11868017 | – | – | – |
| 2008050261 | – | – | – |
| 60883741 | – | – | – |
| US20070868017 | – | – | – |
| US20070883741P | – | – | – |
| WO2008US50261 | – | – | – |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Others concerning applications: alteration of classificationB15K | B15K | |
| Lapse as no evidence of payment of the annual fee has been furnished to inpi (acc. art. 87)LapsedB08K | B08K | |
| Application fees: dismissal - article 86 of industrial property lawB08F | B08F |
Numbers
- Publication
- PI0806304
- Publication, DOCDB
- PI0806304
- Publication, EPODOC
- BRPI0806304
- Application
- 6304
- Application, DOCDB
- PI0806304
- Application, EPODOC
- BR2008PI06304
Titles3
- Portuguese
- seleção de tabela de codificação de comprimento variável com base em estatìsticas de tipo de bloco para codificação de coeficiente de refinamento
- Portuguese
- SELEÇÃO DE TABELA DE CODIFICAÇÃO DE COMPRIMENTO VARIÁVEL COM BASE EM ESTATÍSTICAS DE TIPO DE BLOCO PARA CODIFICAÇÃO DE COEFICIENTE DE REFINAMENTO
- English
- SELECTION OF VARIABLE LENGTH CODING TABLE BASED ON BLOCK TYPE STATISTICS FOR REFINING COEFFICIENT CODING
Classification
- CPC, 7
- H04N19/34
- H04N19/13
- H04N19/159
- H04N19/176
- H04N19/187
- H04N19/30
- H04N19/91