Adaptive upsampling for scalable video coding
19 claims: 3 independent, 16 dependent
- 1CLAIMS REIVINDICAÇÕES 1. Method for encoding video data with spatial scalability, the method comprising:1. Método para codificar dados de vídeo com capacidade de escalonamento espacial, o método compreendendo: generating ascending sampled video data based on base layer video data, wherein the ascending sampled video data corresponds to a spatial resolution of the enhancement layer video data;and encode the enhancement layer video data based on the upwardly sampled video data, where generating the upwardly sampled video data includes interpolating values for one or more pixel locations of the upwardly sampled video data that correspond to locations between different base layer video blocks defined in the base layer video data. gerar dados de vídeo amostrados de maneira ascendente com base em dados de vídeo de camada base, em que os dados de vídeo amostrados de maneira ascendente correspondem a uma resolução espacial dos dados de vídeo de camada de aperfeiçoamento;e codificar os dados de vídeo de camada de aperfeiçoamento com base nos dados de vídeo amostrados de maneira ascendente, em que gerar os dados de vídeo amostrados de maneira ascendente inclui interpolar valores para uma ou mais localizações de pixel dos dados de vídeo amostrados de maneira ascendente que correspondem a localizações entre blocos de vídeo de camada base diferentes definidos nos dados de vídeo de camada base.
- 8Equipment that encodes video data with spatial scalability, which comprises:8. Equipamento que codifica dados de vídeo com capacidade de escalonamento espacial, o qual compreende: an ascending sampler configured to generate ascending sampled video data based on base layer video data, wherein the ascending sampled video data corresponds to a spatial resolution of the enhancement layer video data;and a video encoding unit configured to encode the enhancement layer video data based on the upwardly sampled video data, where the equipment interpolates values for one or more pixel locations of the upwardly sampled video data that correspond to locations between different base layer video blocks defined in the base layer video data. um amostrador ascendente configurado para gerar dados de vídeo amostrados de maneira ascendente com base em dados de vídeo de camada base, em que os dados de vídeo amostrados de maneira ascendente correspondem a uma resolução espacial dos dados de vídeo de camada de aperfeiçoamento;e uma unidade de codificação de vídeo configurada para codificar os dados de vídeo de camada de aperfeiçoamento com base nos dados de vídeo amostrados de maneira ascendente, em que o equipamento interpola valores para uma ou mais localizações de pixel dos dados de vídeo amostrados de maneira ascendente que correspondem a localizações entre blocos de vídeo de camada base diferentes definidos nos dados de vídeo de camada base.
- 19A computer-readable medium that includes instructions that, when executed on a processor, cause the processor to encode video data with spatial scalability, where the instructions cause the processor to:19. Meio passível de leitura por computador que compreende instruções que, quando da execução em um processador, fazem com que o processador codifique dados de vídeo com capacidade de escalonamento espacial, em que as instruções fazem com que o processador: gere dados de vídeo amostrados de maneira ascendente com base em dados de vídeo de camada base, em que os dados de vídeo amostrados de maneira ascendente manages sampled video data in an ascending manner based on base layer video data, where the sampled video data is ascending 8/54 8/54
Independent claims3
228 paragraphs in 7 sections, as filed
(54) Title: ASCENDENT SAMPLING (57) Summary:
ADAPTIVE FOR SCALABLE VIDEO CODING (30) Unionist Priority: 07/01/2008 us 11 / 970,413,
1/9/2007 US 60 / 884,099, 2/8/2007 US 60 / 888,912, 2/8/2007
US 60 / 888,912 (73) Holder (s): Qualcomm Incorporated (72) Inventor (s): Yan Ye, Yiliang Bao (74) Attorney (s): Montaury Pimenta, Machado &
Lioce (86) International Order: pct uS2008050546 de
08/01/2008 (87) International Publication: wo 2008 / 086377de
17/07/2008
<img file="BRPI0806521A2_D0001.tif" />
ADAPTIVE ASCENDING SAMPLING FOR SCALABLE VIDEO ENCODING
This claim claims the benefit of US provisional application No. 60 // 884 099, filed on January 9, 2007, and US provisional application No. 60/888 912, filed on February 8, 2007. All the content of both orders is incorporated here for reference.
TECHNICAL FIELD
This disclosure refers to digital video encoding and, more specifically, scalable video encoding (SVC) techniques that provide spatial scaling capabilities.
BACKGROUND
Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless communication devices, wireless broadcast systems, personal digital assistants (PDAs), laptop and desktop computers, cameras digital recording devices, video game devices, video game consoles, cell phones or satellite radio and the like. Digital video devices can implement block-based video compression techniques, such as those defined by the MPEG-2, MPEG-4, ITU-T H.261, H.263 or H.264 / MPEG-4, Part 10 standards , Coding
Advanced Video (AVC), to transmit and receive digital video more effectively. Video compression techniques perform spatial and temporal prediction to reduce or remove the redundancy inherent in video sequences.
Spatial prediction reduces redundancy between neighboring video blocks within a given video frame.
2/54
Time prediction, also known as motion estimation and compensation, reduces the time redundancy between video blocks in the past and / or future video frames of a video sequence. For time prediction, a video encoder performs motion estimation to track the movement of corresponding video blocks between two or more adjacent video frames. Motion vectors indicate the displacement of video blocks in relation to corresponding prediction video blocks in one or more reference frames. Motion compensation uses motion vectors to identify prediction video blocks from a frame of reference. A residual video block is formed by subtracting the prediction video block from the original video block to be encoded. The residual video block can be sent to a video decoder along with the motion vector, and the decoder can use this information to reconstruct the original video block or an approximation of the original video block. The video encoder can apply transform, quantization and entropy encoding processes to further reduce the bit rate associated with the residual block.
Some video encodings make use of scalable encoding techniques, which are particularly desirable for wireless video data communication. In general, scalable video encoding (SVC) refers to video encoding in which a video data is represented by a base layer and one or more layers of enhancement. For SVC, a base layer typically carries video data with a base spatial, temporal and / or signal-to-noise ratio (SNR) level. One or more layers of improvement carry data from
3/54 additional video to support spatial, temporal and / or SNR levels.
For spatial scalability, enhancement layers add spatial resolution to base layer frames. In SVC systems that support spatial scalability, interlayer prediction can be used to reduce the amount of data needed to transport the enhancement layer. In inter-camadã prediction, the enhancement layer video blocks can be encoded using prediction techniques that are similar to motion estimation and motion compensation. In particular, the enhancement layer residual video data blocks can be encoded using base layer reference blocks. However, the base and enhancement layers have different spatial resolutions. Therefore, the base layer video data can be sampled upwards in the spatial resolution of the enhancement layer video data, such as, for example, to form reference blocks for the generation of the residual enhancement layer data.
SUMMARY
In general, this disclosure describes adaptive techniques for ascending sampling of base layer video data in enhancement layer video data for spatial scalability. For example, base layer video data (residual base layer video blocks, for example) can be sampled upwards at a higher resolution, and data sampled upwards can be used to encode video data from layer of
4/54 improvement. As part of the superior sampling process, the techniques of this disclosure identify the conditions in which upward sampling by interpolation is preferable and other situations in which upward sampling by so-called closest neighbor copying techniques is preferable. Therefore, to encode enhancement layer data, the interpolation and closest neighbor copying techniques can be used in the upstream sampling of base layer data on an adaptive basis.
According to certain aspects of this disclosure, either interpolation or the nearest neighbor copy can be used to define sampled data in an upward fashion, which can be used as reference blocks in the encoding of enhancement layer video data. In particular, the decision on whether to perform interpolation or copy of nearest neighbor can be based on whether or not a pixel sampled upwardly corresponds to a border pixel location in the enhancement layer. Instead of considering only the base layer pixel locations in determining whether to interpolate or use the closest neighbor copy for upward sampling, the techniques described in this disclosure can consider upwardly sampled pixel locations with respect to block boundaries. in the improvement layer. Block boundaries in the enhancement layer may differ from block boundaries in the base layer.
In one example, this disclosure presents a method for encoding video data with spatial scalability. The method comprises generating upwardly sampled video data based on base layer video data, where the sampled video data
5/54 upwardly correspond to a spatial resolution of the enhancement layer video data, and encode the enhancement layer video data based on the upwardly sampled video data, where generating upwardly sampled video data includes interpolating values for one or more pixel locations of upwardly sampled video data that correspond to locations between different base layer video blocks defined in the base layer video data.
In another example, this disclosure features equipment that encodes video data with spatial scalability, the equipment being configured to generate video data sampled upwards based on base layer video data, in which the sampled video data upwardly correspond to a spatial resolution of enhancement layer video data, and to encode the enhancement layer video data based on the upwardly sampled video data, where the equipment interpolates values for one or more pixel locations of the upwardly sampled video data that correspond to locations between different blocks base layer video data defined in the base layer video data.
In another example, this disclosure features an apparatus for encoding video data capable of spatial scaling, the apparatus comprising a device for generating upwardly sampled video data based on base layer video data, in which the video data upwardly sampled correspond to a spatial resolution of enhancement layer video data, and a device for encoding the enhancement layer video data based on the
6/54 upwardly sampled video data, and a device for encoding the enhancement layer video data based on the upwardly sampled video data, wherein the device for generating the upwardly sampled video data includes a device for interpolating values for one or more pixel locations of the upwardly sampled video data that correspond to locations between different base layer video blocks defined in the data base layer video.
The techniques described in this disclosure can be implemented in hardware, software, firmware or any combination of them. If implemented in software, the software can run on a processor, such as a microprocessor, an application-specific integrated circuit (ASIC), a field programmable port arrangement (FPGA) or a digital signal processor (DSP). The software that performs the techniques can initially be stored on a medium that can be read by a computer and loaded and executed on the processor.
Therefore, this disclosure also includes a computer-readable medium that comprises instructions that, when executed on a processor, cause the processor to generate sampled video data in an upward manner based on base layer video data, in whereas the video data sampled upwardly corresponds to a spatial resolution of enhancement layer video data, and encode the enhancement layer video data based on the upwardly sampled video data, where generating the upwardly sampled video data includes interpolating values for one or more pixel locations of the upwardly sampled video data what
7/54 correspond to locations between different base layer video blocks defined in the base layer video data.
In other cases, this disclosure may refer to a circuit, such as an integrated circuit, a chip set, an application-specific integrated circuit (ASIC), a field programmable port arrangement (FPGA), a logic or several combinations of them, configured to perform one or more of the techniques described here.
Details of one or more aspects of the disclosure are presented in the accompanying drawings and in the following description. Other characteristics, objects and advantages of the techniques described in this disclosure will become evident with the description and drawings, and with the claims.
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1 is an exemplary block diagram showing a video encoding and decoding system that can implement the encoding techniques described here as part of the encoding and / or decoding process.
Figure 2 is a block diagram showing an example of a video encoder compatible with this disclosure.
Figure 3 is a block diagram showing an example of a video decoder compatible with this disclosure.
Figures 4 and 5 are conceptual diagrams showing upward sampling from a base layer to an improvement layer.
Figures 6-8 are conceptual diagrams showing ascending sampling techniques that can be used in accordance with this disclosure.
8/54
Figure 9 is a flow diagram showing a technique compatible with this disclosure.
DETAILED DESCRIPTION
This disclosure describes ascending sampling techniques useful in the encoding (i.e., encoding or decoding) of enhancement layer video blocks in a scalable video encoding (SVC) scheme. In SVC systems that support spatial scalability, base layer video data (residual video blocks from a base layer, for example) can be sampled upwards at a higher resolution, and the data sampled upwards from resolution higher can be used to encode the video data from the enhancement layer (such as residual video blocks from an improvement layer). In particular, the data sampled in an ascending manner are used as reference data in the encoding of video data of the enhancement layer with respect to the base layer. This means that the base layer video data is sampled upwards are sampled at the spatial resolution of the enhancement layer video data, and the resulting upwardly sampled data is used to encode the enhancement layer video data. .
As part of this bottom-up sampling process, the techniques of this disclosure identify the conditions for which bottom-up interpolation sampling is preferable, and the other conditions for which bottom-up sampling by so-called closest neighbor copying techniques is preferable. Interpolation can involve generating a weighted average
9/54 for an upwardly sampled value, where the weighted average is defined between two or more pixel values of the base layer. For nearest neighbor techniques, the value sampled in an ascending manner is defined as that of the pixel location in the base layer that is in closest spatial proximity to the pixel location sampled in an ascending manner. According to this disclosure, with the use of interpolation for some specific conditions of the upstream sampling and the closest neighbor copy for other conditions, the encoding of enhancement layer video blocks can be improved.
Either interpolation or the closest neighbor copy can be used to define sampled data upwardly to the spatial resolution of the enhancement layer. The data sampled in an ascending manner can be sampled in an ascending manner from the base layer data (such as, for example, from residual video blocks in the base layer). Upwardly sampled data can form blocks that can be used as references in the enhancement layer data encoding (in the encoding of residual video blocks from the enhancement layer, for example). The decision to perform interpolation or copy of nearest neighbor during the bottom sampling process can be based on whether the location of the pixel value sampled upward corresponds to an edge pixel location on the enhancement layer. This is in contrast to conventional bottom-up sampling techniques, which generally consider only the base layer pixel locations in determining whether to interpolate or use the nearest neighbor copy.
10/54
For example, conventional upward sampling techniques can interpolate a sampled pixel upward only when two base layer pixels used in the interpolation do not correspond to an edge of a base layer video block. In this disclosure, the term edge refers to pixel locations that correspond to an edge of a video block, and the inner term refers to pixel locations that do not correspond to an edge of a video block. A video block can refer to the block transform used in the video encoder-decoder (CODEC). As examples in H.264 / AVC, a video block can be 4x4 or 8x8 in size. The boundaries between blocks in the enhancement layer, however, may differ from the boundaries between blocks in the base layer. According to this disclosure, the decision of whether or not to perform interpolation may depend on whether the pixel value to be sampled upwardly corresponds or not to an edge pixel location in the enhancement layer.
If the two base layer pixels used in the interpolation correspond to the edges of two adjacent base layer video blocks, the upwardly sampled value may fall between the two base layer video blocks. In this case, conventional techniques use copy techniques of nearest neighbor for ascending sampling. Interpolation is conventionally avoided in this case because the different blocks of base layer video may have been encoded with different levels of quantization. For the closest neighbor copy, the value sampled in an ascending manner can be defined as that of the base layer pixel location in closest spatial proximity to the pixel location sampled in an ascending manner.
11/54
According to this disclosure, interpolation can be performed in various contexts where an upwardly sampled layer value falls between two base layer video blocks. If the upwardly sampled value was not itself associated with an edge pixel location in the enhancement layer, interpolation may be preferred over the nearest neighbor copy. In this case, unblocking filtering of the enhancement layer is unlikely to resolve massive artifacts in the reproduction of video frames. Therefore, interpolation may be preferred, although the different base layer video blocks may have been encoded at different levels of quantization. If the upwardly sampled value is itself associated with an edge pixel location on the enhancement layer and the upwardly sampled value falls between the two base layer video blocks, then the closest neighbor copy can be used according to this revelation. In this case, the enhancement layer unlock filtering will resolve the massive artifacts in the enhancement layer, and the risk of poor interpolation due to the fact that base layer video blocks have different levels of quantization can outweigh the potential benefits of interpolation in this context. In addition, the decision to interpolate or use a copy of the nearest neighbor may also depend on whether the two base layer video blocks were encoded using different encoding modes. In addition, optional adaptive low-pass filters can be applied before or after ascending sampling in order to further mitigate the problem of signal discontinuity across the boundaries between base layer coding blocks.
12/54
To simplify and facilitate the exemplification, in this disclosure the techniques of interpolation and copying of nearest neighbor are generically described in one dimension, although such techniques can typically be applied to both vertical and horizontal dimensions. For two-dimensional or two-dimensional closest neighbor interpolation techniques, the nearest neighbor interpolation or copy would first be applied to one dimension and then applied to the other dimension.
Figure 1 is a block diagram showing an encoding and decoding system 10. As shown in Figure 1, system 10 includes a source device 12 that transmits encoded video to a receiving device 16 via a communication channel. 15. The source device 12 may include a video source 20, a video encoder 22 and a modulator / transmitter 24. 0 The receiving apparatus 16 may include a receiver / demodulator 26, a video decoder 28 and a display apparatus 30. The system 10 may be configured to apply adaptive bottom sampling techniques, as described herein, during encoding and decoding information video layer enhancement of an SVC scheme. Encoding and decoding are more commonly referred to here as encoding.
In the example of Figure 1, the communication channel 15 can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines, or any combination of wireless and wired media. Communication channel 15 can be part of a packet-based network, such as a local area network, an extended area network, or a global network, such as the Internet. The channel
Communication 13/54 generally represents any means of communication or gathering different communication means suitable for transmitting video data from the source device 12 to the receiving device 16.
The source device 12 generates encoded video data for transmission to the receiving device 16. In some cases, however, the devices 12, 16 can function in a substantially symmetrical manner. For example, each of the devices 12, 16 can include video encoding and decoding components. Therefore, system 10 can support unidirectional or bidirectional video transmission between video devices 12, 16, such as, for example, video streaming, video broadcasting or video telephony.
The video source 20 of the video device 12 may include a video capture device, such as a video camera, a video file containing previously captured video, or a video feed from a video content provider. As another alternative, video source 20 can generate data based on computer graphics as the video source, or a combination of live video and computer generated video. In some cases, if the video source 20 is a video camera, the source device 12 and the receiving device 16 can form so-called camera phones or video phones. In each case, the captured, pre-captured or computer generated video can be encoded by the video encoder 22 for transmission from the video source device 12 to the video decoder 28 of the video receiving device 16 via the modulator / transmitter 22, communication channel 15 and receiver / demodulator 26.
Video encoding and decoding processes can implement bottom-up sampling techniques
14/54 adaptive using interpolation and copy of nearest neighbor, as described here, to improve the enhancement layer encoding process. The display device 30 displays the decoded video data for a user and can comprise any of several display devices, such as a cathode ray tube, a liquid crystal display (LCD), a plasma monitor, a organic light-emitting diode (OLED) or other type of display device. Spatial scaling capability allows a video decoder to reconstruct and display a higher spatial resolution video signal, such as CIF (Common intermediate format, 352 x 288 image resolution) as opposed to QCIF (Intermediate format quarter, 178 x 144 image resolution) by decoding the enhancement layer bit stream from an SVC bit stream.
Video encoder 22 and video decoder 28 can be configured to support scalable video encoding (SVC) for spatial scaling capability. In addition, the capacity for temporal spatial scaling and / or signal-to-noise ratio (SNR) can also be supported, although the techniques of this disclosure are not limited in this regard. In some ways, video encoder 22 and video decoder 28 can be configured to support encoding with fine-grained SNR scaling (FGS) to SVC. The encoder 22 and decoder 28 can support varying degrees of scalability by supporting the encoding, transmission and decoding of a base layer and one or more scalable enhancement layers. For scalable encoding, a base layer of video data carrier with a spatial, temporal or
15/54
Baseline SNR. One or more layers of enhancement carry additional data to support higher spatial, temporal or SNR levels. The base layer can be transmitted in a way that is more reliable than the transmission of enhancement layers. For example, the most secure parts of a modulated signal can be used to transmit the base layer, while the least secure parts of the modulated signal can be used to transmit the enhancement layers.
In order to support SVC, video encoder 22 may include a base layer encoder 32 and an enhancement layer encoder 34 to encode a base layer and the enhancement layer, respectively. In some cases, multiple enhancement layers can be supported, in which case multiple enhancement layer encoders can be presented to encode progressively detailed levels of video enhancement. The techniques of this disclosure, which involve ascending sampling of base layer data to spatial resolution of enhancement layer video data, so that data sampled upwards can be used to encode enhancement layer data, can be performed. by the enhancement layer encoder 34.
The video decoder 28 may comprise a base / combined enhancement decoder that decodes the video blocks associated with both the base layer and the enhancement layer and combines the decoded video to reconstruct the frames of a video sequence. On the decoding side, the techniques of this disclosure that involve ascending sampling of base layer data up to the spatial resolution of video data from
16/54 enhancement layer, so that the data sampled upwards can be used to encode enhancement layer data, can be performed by the video decoder 28. The display device 30 receives the decoded video sequence and displays the video stream to a user.
Video encoder 22 and video decoder 28 can operate according to a video compression standard, such as MPEG-1, MPEG-4, iTU-T H.263 or ITÜ-T H.264 / MPEG-4, Part 10, Advanced Video Encoding (AVC). Although not shown in Figure 1, in some ways the video encoder 22 and video decoder 28 can each be integrated with an audio encoder and decoder and may include appropriate MUX-DEMUX units, or other hardware and software, to process the encoding of both audio and video in a common data stream or in separate data streams. If applicable, MUXDEMUX units can conform to the ITU H.223 multiplexer protocol or other protocols, such as the user datagram protocol (UDP).
The H.264 / MPEG-4 (AVC) standard was formulated by the ITU-T Video Coding Specialist Group (VCEG) together with the ISSO / IEC Mobile Image Specialist Group (MPEG) as the product of a collective partnership known as the Joint Video Team (JVT). In some ways, the techniques described in this disclosure can be applied to devices that generally conform to the H.264 standard. 0 H.264 standard is described in ITU-T Recommendation H.264, Advanced Video Encoding for Generic Audiovisual Services, by the ITU-T Study Group, dated March 2005, which can be referred to here as the standard
17/54
Η.264 or Η.264 specification, or H.264 / AVC standard or specification.
The Joint Video Team (JVT) continues to work on scalable video encoding (SVC) extensions for H.264 / MPEG-4 AVC. For example, the Joint Scalable Video Model (JSVM) created by JVT implements tools for use in scalable video, which can be used within system 10 for the various encoding tasks described in this disclosure. Detailed information regarding fine-grained SNR Scaling Capacity (FGS) coding can be found in the Joint Draft documents and particularly in the Joint Draft 8 (JD8) of the SVC Amendment (revision 2), by Thomas Wiegand, Gary Sullivan, Julien Reichel, Heiko Schwarz and Mathias Wien, Draft SVC Amendment 8 (revision 2), JVTU201, October 2006, Hangzhou, China. In addition, additional details of an implementation of the techniques described here can be found in the proposal document JVT-W117 submitted to the ISSO / IEC MPEG & ITU-T VCEG Joint Video (JVT) Team (ISSO / IEC JTC1 / SC29 / WG11 and ITU-T SG16 Q.6), by Yan Ye and Yiliang Bao, in April 2007 and at the 23rd Meeting in San Jose, California, USA, and in the proposal document JVT-V116 submitted to the JVT by Yan Ye and Yiliang Bao in January 2007 at the 22nd Meeting in Marrakesh, Morocco.
In some ways, for video broadcast execution, the techniques described in this disclosure can be applied to enhanced H.264 video encoding to deliver real-time video services on terrestrial mobile multimedia (TM3) systems using the Specification Exclusive Direct Link Aerial Interface (FLO), Exclusive Direct Link Aerial Interface Specification for Mobile Multimedia Multicast
18/54
Land, to be published as Technical Standard TIA-1099 (the FLO Specification). This means that the communication channel 15 may comprise a wireless information channel used to broadcast wireless video information in accordance with the FLO Specification or the like. The FLO Specification includes examples that define the syntax and semantics of bit streams and decoding processes suitable for the FLO Aerial Interface. Alternatively, the video can be broadcast according to other standards, such as DVB-H (hand-held digital video broadcast), ISDB-T (terrestrial integrated services digital broadcast) or DMB (digital media broadcast) ).
Therefore, the source device 12 can be a mobile wireless terminal, a streaming video server or a video broadcast server. However, the techniques described in this disclosure are not limited to any specific type of broadcast, multicast or point-to-point system. In the case of broadcasting, the source device 12 can broadcast multiple channels of video data to several receiving devices, each of which may be similar to the receiving device 16 of Figure 1. As an example, the receiving device 16 it may comprise a wireless communication device, such as a mobile telephone device commonly referred to as a radio cellular telephone.
Video encoder 22 and video decoder 28 can each be implemented as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable port arrangements (FPGAs) , discrete logic, software, hardware, firmware, or any combination of them. Each video encoder 22 and
19/54 video decoder 28 can be included in one or more encoders or decoders, or one or the other of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective mobile device, subscriber handset, handset broadcast, server or similar. In addition, the source device 12 and receiver 16 may each include modulation, demodulation, frequency conversion, filtering and amplifier components suitable for transmitting and receiving encoded video, as applicable, including components and antennas. radio frequency (RF) wireless enough to support wireless communication. To facilitate the illustration, however, such components are summarized as being modulator / transmitter 24 of the original apparatus 12 and receiver / demodulator 26 of the receiving apparatus 16 of Figure 1.
A video sequence includes a series of video frames. The video encoder 22 operates in pixel blocks within individual video frames for encoding the video data. Video blocks can have fixed or variable sizes and may differ in size according to a specified encoding standard. Each video frame can be divided into a series of slices. Each slice can include a series of macroblocks, which can be arranged in sub-blocks. As an example, the ITU-T H.264 standard supports intra-prediction in several block sizes, such as 16 by 16, 8 by 8, 4 by 4 for luma components, and 8 x 8 for chroma components, as well as inter-prediction in several block sizes, such as 16 x 16, 16 by 8, 8 by 16, 8 by 8, 8 by 4, 4 by 8 and 4 by 4 for luma components and
20/54 corresponding step sizes for chroma components.
Smaller video blocks can have better resolution and can be used in locations in a video frame that include higher levels of detail. In general, macro-blocks (MBs) and the various sub-blocks can generally be referred to as video blocks. In addition, a slice can be considered as a series of video blocks, such as MBs and / or sub-blocks. Each slice can be an independently decodable unit. After prediction, a transform can be performed on the residual block of 8x8 or on the residual block of 4x4, and an additional transform can be applied to the DC coefficients of the 4x4 blocks for chroma components or light components if the intra-prediction mode_of 16x16 is used.
Following intra or interpretive coding, additional coding techniques can be applied to the transmitted bit stream. These additional coding techniques may include transformation techniques, such as the 4x4 or 8x8 integer transform used in the H-264 / AVC or a discrete cosine transform and entropy coding DCT, such as variable length coding (VLC), Huffman coding and / or execution length coding.
In accordance with the techniques of this disclosure, bottom-up sampling techniques are used to produce bottom-sampled video data for use in encoding (i.e., encoding or decoding) enhancement layer video data. Base layer video data can be sampled upwards to spatial resolution
21/54 of corresponding enhancement layer video blocks, and upwardly sampled data can be encoded as references in the enhancement layer video data. As part of this bottom-up sampling process, situations in which techniques are preferred, and other situations in which so-called closest neighbor copying techniques are preferred. Again, interpolation may involve generating a weighted average for an upwardly sampled value, where the weighted average is defined between two or more values of the base layer. For the closest neighbor copy, the layer value sampled upwardly is defined as the location of the base layer pixel in the closest spatial proximity to the pixel location sampled upwardly. With the use of interpolation in specific routes of the upward sampling and the copy of the nearest neighbor in other routes, the encoding of enhancement layer video blocks can be improved.
Upward sampling can change the boundaries of the blocks. For example, if the base layer and the enhancement layer each define 4 x 4 pixel video blocks, ascending sampling of the base layer to define more pixels according to the spatial resolution of the enhancement layer results in the The boundaries of the base layer blocks are different from those of the data sampled upwards. This observation can be explored so that decisions regarding interpolation and near neighbor techniques can be based on whether the values sampled upwards correspond or not to the pixel locations of the revelation identify of interpolation are
22/54 enhancement layer (that is, boundaries between blocks in the enhancement layer) and whether or not such locations also correspond to locations between borders between blocks in the base layer.
Encoder 22 and decoder 28 can perform reciprocal methods that each perform the upward sampling techniques described herein. The encoder 22 can use upward sampling to encode the enhancement layer information, and the decoder 28 can use the same upward sampling process to decode the enhancement layer information. The term encoding refers in general to either encoding or decoding.
Figure 2 is a block diagram showing an example of a video encoder 50 that includes an upward sampler 45 for upwardly sampling base layer video data at a spatial resolution associated with enhancement layer video data. The data sampled in an ascending manner is then used to encode video data from the enhancement layer. The video encoder 50 can correspond to the enhancement layer encoder 34 of the source device 12 of Figure 1. This means that the base layer encoding components are not shown in Figure 2 for simplicity. Therefore, video encoder 50 can be considered an enhancement layer encoder. In some cases, the components shown of the video encoder 50 can also be implemented in combination with modules or base layer encoding units, such as, for example, in a pyramid-shaped encoding design that supports scalable layer video encoding base and the improvement layer.
23/54
The video encoder 50 can perform intra and inter-coding of blocks within video frames. Intra-coding uses spatial prediction to reduce or remove spatial redundancy in video within a given video frame. Inter-coding uses time prediction to reduce or remove time redundancy in video within adjacent frames of a video sequence. For inter-encoding, video encoder 50 performs motion estimation in order to track the movement of corresponding video blocks between two or more adjacent frames. For intra-coding, spatial prediction that uses pixels from neighboring blocks within the same frame is applied to form a predictive block for the block that is encoded. The components of spatial prediction used in intra-coding are not shown in Figure 2.
As shown in Figure 2, video encoder 50 receives a current video block 31 (an enhancement layer video block, for example) within a video frame to be encoded. In the example in Figure 2, the video encoder 50 includes a motion estimation unit 33, a reference frame storage 35, a motion compensation unit 37, a block transform unit 39, a quantization unit 41, an inverse quantization unit 42, an inverse transform unit 44 and an entropy coding unit 46. An unlocking filter 32 can also be included to filter the boundaries between blocks in order to remove massive artifacts. The video encoder 50 also includes an adder 48, adder 49A and 49B and an adder 51. Figure 2 shows the time prediction components of the video encoder 50 for intercoding video blocks. Although not shown in
24/54
2 to facilitate illustration, the video encoder 50 may also include spatial prediction components for intra-encoding some video blocks.
The motion estimation unit 33 compares video block 31 with blocks in one or more adjacent video frames to generate one or more motion vectors. The adjacent frame or frames can be retrieved from reference frame storage 35, which can comprise any type of memory or data storage device for storing video blocks reconstructed from previously encoded blocks. Motion estimation can be performed for blocks of varying sizes, such as 16x16, 16x8, 8x16, 8x8 or smaller block sizes. The motion estimation unit 33 identifies a block in an adjacent frame that most closely corresponds to the video block 31, such as, for example, based on a rate distortion model, and determines an offset between the blocks. On this basis, the motion estimation unit 33 produces a motion vector (MV) (or several MVs in the case of bidirectional prediction) that indicates the magnitude and trajectory of the displacement between the current video block 31 and a predictive block used for encode the current video block 31.
Motion vectors can have half or quarter pixel precision, or even greater precision, allowing video encoder 50 to track motion more accurately than integer pixel locations and obtain a better prediction block. When motion vectors with fractional pixel values are used, interpolation operations are performed on motion compensation unit 37. Motion estimation unit 33 can identify the best vector
25/54 motion for a video block using a rate distortion model. Using the resulting motion vector, the motion compensation unit 37 forms a motion compensation prediction video block.
video encoder 50 forms a residual video block by subtracting the prediction video block produced by the motion compensation unit 37 from the original current video block 31 in adder 48. block transform unit 39 applies a transform, such as' a discrete cosine transform (DCT), to the residual block, producing residual transform block coefficients. At this point, additional compaction is applied by subtracting residual information from the base layer from the residual information from the improvement layer using the 49A adder. The bottom sampler 45 receives residual base layer information (from a base layer encoder, for example) and samples the base layer residual information upwards in order to generate sampled information in an upward manner. This upwardly sampled information is then subtracted (via adder 49A) from the residual enhancement layer information that is encoded.
As described in more detail below, the upward sampler 45 can identify situations in which interpolation is preferred, and other situations in which so-called closest neighbor copying techniques are preferred. Interpolation involves generating a weighted average for a sampled value in an ascending manner, where the weighted average is defined between two values of the base layer. For the closest neighbor copy, the value sampled upwards is defined as the base layer pixel location in
26/54 more intimate spatial proximity to the pixel location sampled in an ascending manner. According to this disclosure, the ascending sampler 45 uses interpolation in specific routes of the ascending sampling, and the nearest neighbor techniques in other routes. In particular, the upward sampler 45's decision to perform interpolation or nearest neighbor techniques can be based on whether the upwardly sampled value corresponds to a border pixel location in the enhancement layer or not. This is in contrast to conventional bottom-up sampling techniques, which generally only consider baseline pixel locations in determining whether to interpolate or use nearest neighbor techniques. The examples of interpolation and closest neighbor in this disclosure are described in one dimension, for simplicity, but such interpolation or nearest neighbor techniques would typically be applied sequentially in both the horizontal and vertical dimensions. An additional filter 47 can also be included to filter block edges of the base layer information prior to upward sampling by the upward sampler 45. Although shown in Figure 2 as being located before the rising sampler 45, the additional filter 47 can also be placed after the rising sampler 45 to filter the pixel locations in the upwardly sampled video data that are interpolated from two base layer pixels that correspond to two different base layer coding blocks. In any case, this additional filtering by filter 47 is optional and is resolved in more detail in this disclosure.
The quantization unit 41 quantizes the residual transform block coefficients for
27/54 further reduce the bit rate. Adder 49A receives the upstream sampled information from the upstream sampler 45 and is positioned between adder 48 and block transform unit 39. In particular, adder 49A subtracts a sampled data block upwardly from the unit output of block transform 39. Similarly, adder 49B, which is positioned between the reverse transform unit 44 and adder 51, also receives the upwardly sampled information from the upward sampler 45. Adder 49B adds the sampled data block upwardly back at the output of the reverse transform unit 44.
Spatial prediction coding works in a similar way to temporal prediction coding. However, while time prediction encoding uses blocks of adjacent frames (or other encoded units) to perform encoding, spatial prediction uses blocks within a common frame (another encoded unit) to perform encoding. Spatial prediction coding encodes intra-coded blocks, while temporal prediction coding encodes inter-coded blocks. Again, the spatial prediction components are not shown in Figure 2 for simplicity.
The entropy unit 46 encodes the transform coefficients quantized according to an entropy coding technique, such as variable length coding, binary arithmetic coding (CABAC), Huffman coding, execution length coding, coded block pattern (CBP) or similar coding, in order to further reduce the bit rate of the transmitted information. Entropy unit 46 can select a VLC table for
28/54 promote coding efficiency. After entropy encoding, the encoded video can be transmitted to another device. In addition, the reverse quantization unit 42 and the reverse transform unit 44 apply reverse quantization and reverse transformation, respectively, to reconstruct the residual block. Adder 49B adds back upwardly sampled data from upward sampler 45 (representing an upwardly sampled version of the base layer residual block), and adder 51 adds the reconstructed residual block to the produced motion compensated prediction block by the motion compensation unit 37 in order to produce a reconstructed video block for storage in the reference frame storage 35. The release filter 32 can perform release filtering before storing the frame of reference. Unblocking filtering may be optional in some examples.
According to this disclosure, the rising sampling interpolates values for one or more pixel locations of the rising sampled video blocks that correspond to a location between two different edges of two different base layer video blocks.
In one example, the ascending sampler 45 interpolates first values for the video data sampled upwardly based on the base layer video data for: (i) pixel locations of the ascended sampled video data that correspond to locations of internal pixels of enhancement layer video blocks defined in the enhancement layer video data, where at least some of the internal pixel locations of the enhancement layer video blocks
29/54 enhancement corresponds to locations between different base layer video blocks and (ii) pixel locations of the upwardly sampled video data that correspond to edge pixel locations of the enhancement layer video blocks and are not located between the different base layer video blocks. In this case, the ascending sampler 45 can define second values for the sampled video data in an ascending manner based on values of closest neighbors in the base layer video data for: (iii) pixel locations of the video data
<td>sampled</td><td colspan="2">in an upward way that</td><td colspan="3">match locations</td>
<td>of pixel</td><td>internal</td><td>of blocks</td><td>video</td><td>in</td><td>layer of</td>
<td colspan="2">improvement and</td><td>are located</td><td>in between</td><td>the</td><td>many different</td>
<td>blocks of</td><td>video of</td><td colspan="2">base layer when</td><td>two</td><td>many different</td>
<td>blocks of</td><td>layer</td><td>base define</td><td colspan="2">many different</td><td>modes of</td>
codification. The different coding modes can comprise an intra-coding mode and an inter-coding mode. In this case, the ascending sampler 45 considers not only the locations associated with the values sampled in an ascending manner (that is, if the values sampled in an ascending manner correspond to boundaries between blocks in the improvement layer and if the values fall between the boundaries between the blocks. in the base layer), but also if the two blocks of base layer video are encoded using different encoding modes (inter and intra coding modes, for example).
Figure 3 is a block diagram showing an example of a video decoder 60, which can correspond to the video decoder 28 of Figure 1 or a decoder of another device. The video decoder 60 includes a rising sampler 59, which performs functions similar to those of the rising sampler 45 of
30/54
Figure 2. This means that, like the rising sampler 45, the rising sampler 59 interpolates values for the sampled blocks upwardly to one or more pixel locations of the enhancement layer video blocks that correspond to a location between two different edges of two different base layer video blocks. In addition, like the rising sampler 45, the rising sampler 59 can select between interpolation and copy of nearest neighbor in the manner described herein. An upward motion sampler 61 can also be used to upwardly sample the motion vectors associated with the base layer. An optional filter 65 can also be used to filter block boundaries of the base layer data before the rising sampler performed by the rising sampler 59. Although not shown in Figure 3, filter 65 can also be placed after the upstream sampler 59 and used to filter the pixel locations in the upwardly sampled video data that are interpolated from two base layer pixels that correspond to two different base layer encoding blocks.
The video decoder 60 may include an entropy unit 52A for entropy decoding of base layer information. The intrapredictive components are not shown in Figure 3, but could be used if the video decoder 60 supports intra- and inter-prediction encoding. The enhancement layer path can include an inverse quantization unit 56A and an inverse transform unit 58B. The information in the base layer and enhancement layer paths can be combined by the adder 57. Before such a combination, however, the base layer information is sampled
31/54 by means of the ascending ascendant by the ascending sampler 59 according to the techniques described herein.
Video decoder 60 can perform block inter-encoding within video frames. In the example in Figure 3, video decoder 60 includes entropy units 52A and 52B, a motion compensation unit 54, reverse quantization units 56A and 56B, reverse transform units 58A and 58B and a reference frame storage 62 Video decoder 60 also includes adder 64. Video decoder 60 may also include an unlock filter 53 that filters the output of adder 64. Again, adder 57 combines information in the base layer and enhancement layer paths that follow upward sampling of the base layer path. ascending sampler 59. The motion sampler 61 can upwardly sample the motion vectors associated with the base layer, so that such motion vectors correspond to the spatial resolution of the enhancement layer video data.
For enhancement layer video blocks, the entropy unit 52A receives the encoded bit stream and applies an entropy decoding technique in order to decode the information. This can produce quantized residual coefficients, macro-block and sub-block coding mode and motion information, which can include motion vectors and block partitions. After decoding performed by the entropy unit 52A, the motion compensation unit 54 receives the motion vectors and one or more reconstructed reference frames from the reference frame storage 62. The reverse quantization unit 56A performs reverse quantization in, that is, yeah, it disquantifies them,
Massive 32/54. The quantized block coefficient blocks and the inverse transform unit 58A apply an inverse transform, i.e., an inverse DCT, to the coefficients in order to produce residual blocks. The output of the reverse transform unit 58A is combined with the base layer information sampled upwardly like the output of the upstream sampler 59. Adder 57 facilitates this combination. The motion compensation unit 54 produces moving compensated blocks that are added by the adder 64 to the residual blocks so as to form decoded blocks. 0 unlock filter 53 filters the decoded blocks in order to remove the filtered artifacts are then placed in the reference frame storage 62, which generates reference blocks from the motion compensation and also produces decoded video for a drive display device ( such as the apparatus 30 of Figure 1).
SVC can support several interlayer prediction techniques to improve coding performance. For example, when coding an enhancement layer macro-block, the macroblock mode, movement information and residual signals from the base or previous layer can be used. In particular, certain residual blocks in the base or previous layer can be correlated with the corresponding improvement layer residual blocks. For these application, residual prediction can reduce residual layer enhancement and improve coding performance.
In SVC, whether residual prediction is used or cannot be indicated using a PredRes bit indicator associated with the macro block, which can be encoded as a syntax element at the macro block level. If blocks, the energy
33/54
PredRes = 1, then the enhancement layer residue is encoded after subtracting the residual base layer block from it. When the enhancement layer bitstream represents a video signal with a higher spatial resolution, the residual base layer signal is sampled upwards to the resolution of the enhancement layer before being used in interlayer prediction. This is the function of the ascending samplers 45 and 49 of Figures 2 and 3, that is, the generation of the sampled video blocks in an ascending manner. In the SVC Joint Draft 8 (JD8), a bi-linear filter is proposed for the sampler in order to sample the residual base layer signal, with some exceptions on the boundaries between base layer blocks.
The JD8 SVC supports both dyadic scaling capabilities and extended scaling capabilities (ESS). In the dyadic spatial scaling capability, the enhancement layer video frame is twice the size of the base layer video in each dimension, and harvesting, if any, occurs at the borders of the macro blocks. In ESS, scaling ratios and arbitrary harvest parameters between base layer and enhancement video signals are allowed. When using ESS, the alignment of pixels between the base and enhancement layers can be arbitrary.
Figures 4 and 5 show examples of the relative positioning of pixels in the base layer and in the enhancement layer for scaling ratios of 2: 1 and 5: 3, respectively. In Figures 4 and 5, labels B refer to pixel locations associated with a base layer, and labels E refer to locations associated with an upwardly sampled data (which corresponds to the enhancement layer). The video block (s)
34/54 sampled (s) in an ascending manner is (are) used (s) as a reference in coding the improvement layer information. The BE label in the central pixel location of Figure 5 means that the same pixel location is superimposed on the base layer and on the upwardly sampled data that have the enhancement layer resolution.
As shown in Figures 4 and 5, upward sampling from a lower resolution base layer to the highest resolution of the enhancement layer occurs in two dimensions. This two-dimensional bottom-up sampling, however, can easily be done through successive one-dimensional bottom-up sampling processes for each pixel location. In Figure 5, for example, the 3x3 pixel array of base layer video data (labeled as B) is first sampled upwards in the horizontal direction to become a 5x3 array of intermediate values (labeled as X) . Then, the 5x3 X pixel array is sampled upwards in the vertical direction to become the final 5x5 pixel array of upwardly sampled video data that corresponds to the spatial resolution of the enhancement layer (labeled E ). Note that in Figure 5 the location of the pixel sampled in an ascending manner and the location of the pixel before ascending sampling can be co-located; for example, the central pixel is labeled three times as B, E and X. During the horizontal and vertical upward sampling processes, the decision as to whether to use interpolation or copy of nearest neighbor can be applied according to this disclosure.
In the following discussion, interpolation and closest neighbor copying techniques are discussed
35/54 in the horizontal dimension. It should be understood, however, that successive linear techniques can be applied in the vertical and horizontal directions for two-dimensional upward sampling.
When the spatial scaling capability is used, the residual signal from the base layer is sampled upwards. The upwardly sampled data generated has the same spatial dimension as the enhancement layer video information and is used as prediction data for the enhancement layer residual signal. Once again, the JD8 SVC proposes the use of a bi-linear filter in this upward sampling. For the dyadic spatial scaling capacity, the distances between the pixels used to derive weights in the bi-linear upward sampling in the horizontal direction are shown in Figure 6. Upward sampling in the vertical dimension is done in the same way as in the horizontal direction.
Figure 6 shows a row of pixels in a base layer block and a row of pixels in an upwardly sampled block that corresponds to the spatial resolution of an enhancement layer. As shown in Figure 6, the residual values sampled upwardly p (e0) and p (el) at the pixel locations eO and el are derived according to equations (1) and (2), where bO and bl are the locations nearest integer pixel in the base layer.
p (e0) = (1 -1 / 4) * /? (Z> 0) +1 / 4 * /? (M) (1)
Xel) = (1-3 / 4) * ρ (ά0) + 3/4 * p (bV) (2)
For ESS with a 5: 3 scaling ratio, the weights used in the bi-linear upward sampling in the
36/54 horizontal direction are shown in Figure 7. Again, bi-linear upward sampling in the vertical dimension is done in the same way as in the horizontal direction.
In Figure 7, the residual values sampled in an ascending manner p (eO) to p (e4) at pixel locations eO to e4 are derived as in equations (3) to (7), where bO to b4 are pixel locations of number base layer integers used in the interpolation and eO to e4 are pixel locations in the sampled layer in an ascending manner, which corresponds to the spatial resolution of the enhancement layer.
p (eO) = (1-4 / 5) * p (òO) + 4/5 * p (M) (3)
Xel) <l-2/5) * Xàl) + 2/5 * Xú2) (4) /> (e2) = X & 2) (5)
Xe3) = (l-3/5) * Xò2) + 3/5 * p (à3) (6) p (e4) = (1-1 / 5) * p (ò3) +1/5 * p (M ) (7)
There is discontinuity through the edges of blocks in the residual signal reconstructed in the base layer. Where discontinuity exists depends on the size of the transform used to encode the base layer video. When a 4x4 block transform is used, there is discontinuity at the boundaries between 4x4 blocks. When an 8x8 block transform is used, there is discontinuity in the boundaries between the 8x8 blocks. In JD8, if the two base layer pixels used in bi-linear interpolation (bO and bl in Figure 6) belong to two different blocks, then bi-linear interpolation is disabled. Instead, the values sampled upwards are derived by copying the nearest neighboring pixel in the base layer.
37/54
In the example of dyadic spatial scaling capability in Figure 6, if the pixels at location bO and bl belong to two blocks, then the pixels in eO and el are derived using equations (8) and (9) instead of ( 1 and 2);
p (e0) = p (bO) p (el) = p (bl)
In the ESS, the (8) (9) blocks between those of the parent that enhances encoding boundaries in the base layer are not aligned with the boundaries between the encoding blocks in the enhancement layer. Therefore, a block sampled in a way that matches the resolution of the layer may contain interpolated pixels from different base layer encoding blocks. For the ESS with a spatial ratio of 5: 3 (Figure 7), an example of block alignment can be found as in Figure 8. In Figure 8, pixels B0-B3 belong to a base layer coding block and pixels B4-B5 belong to a different base layer coding block (assuming a 4x4 transform is used). In this case, there may be a discontinuity of signal between pixels B3 and B4. In the upwardly sampled layer, what spatial resolution the E0-E4 pixels are copied from the base layer B3-B5 pixels. Assuming that the upwardly sampled pixels labeled eO to e7 belong to a coding block shown in an upward 8x8 way, the signal discontinuity of the base layer will be carried up to the 8x8 upwardly sampled layer. In particular, since the base layer pixels in b3 and b4 belong to two layer blocks according to conventional techniques, it corresponds to the improvement, of the interpolated layer of or base, of the pixel sampled upwards at the e5 location will be copied from the
38/54 base layer pixel at location b4 instead of being interpolated from pixels in b3 and b4. This forced copy (by conventional techniques defined in JD8) can aggravate the problem of signal discontinuity within the 8x8 enhancement layer block. This, in turn, can translate into a less accurate residual prediction.
In this disclosure, an adaptive residual upward sampling scheme is outlined. The adaptive residual upward sampling scheme can attenuate signal discontinuity within an upwardly sampled data block taking into account the relative alignment between blocks between the base layer and the enhancement layer. To further improve the residual signal quality sampled upwardly, adaptive low-pass filtering can be applied to the residual signal in the base layer (before upward sampling) and / or to the residual signal in the enhancement layer (after upward sampling) ). For example, filters 47 and 65 of Figures 2 and 3, respectively, may comprise low-pass filters that facilitate this filtering of the enhancement layer after upward sampling. Although in Figures 2 and 3 low-pass filtering is shown to be applied to the residual signal in the base layer before upward sampling, this low-pass filter (elements 47 and 65 in Figures 2 and 3) can be located after the upward sampler. 45 and after rising sampler 59 and before adder 49A and adder 57.
To summarize, in the residual up sampling process specified in the JD8 SVC, for each pixel location in the upwardly sampled data, the corresponding base layer pixel locations are first determined. If the base layer pixels belong to the
39/54 same base layer coding block in a given direction (horizontal or vertical), so the bilinear interpolation in that direction is called to obtain the value in the sampled data in an upward way. Otherwise (the base layer pixels belong to different base layer coding blocks in a given direction), the value in the data sampled upwards is determined by copying from the nearest neighbor pixel in the base layer.
The residual upward sampling currently specified in the JD8 SVC can aggravate the problem of signal discontinuity within enhancement layer blocks, especially in the case of extended spatial scaling capability, where the alignment between base and enhancement layer blocks can be arbitrary . Such signal discontinuity can distort the sampled signal upwardly and reduce its accuracy as the prediction signal when used in residual prediction.
According to this disclosure, the decision as to whether to call the interpolation or copy from the nearest neighboring pixel can be determined depending on the alignment between the base layer blocks and the enhancement layer. Take Figure 8 as an example. The enhancement layer pixel that corresponds to the e5 pixel location is arranged within an 8x8 coding block at the enhancement layer resolution. In this case, instead of copying from the base layer pixel at location b4, or p (e5) = p (b4), the interpolation between the base layer pixels at locations b3 and b4 can be called to attenuate the signal discontinuity and perfect precision accuracy. That is, the pixel value sampled upwards in e5 can be derived using equation (10);
40/54 p (e5) = (1-3 / 5) * p (Z> 3) + 3/5 * p (b4) (10)
The copy of the nearest neighboring pixel can be called only when both of the following conditions are true:
Cl. The upwardly sampled pixel to be interpolated is arranged at the boundary between the enhancement layer coding blocks; and
C2. The base layer pixels involved in the interpolation process belong to different base layer coding blocks.
A video encoding system can support more than one block transform. In H.264 / AVC, for example, both 8x8 integer block transforms are supported for the luma block, while only the 4x4 transform is applied to the chroma block. For the chroma component, there is an additional DC transform in the DC coefficients of the 4x4 blocks. Since this does not alter the fact that massive artifacts occur on boundaries between blocks of 4x4, this transform is not considered in the discussion of this revelation.
The block transform that is applied to the residual signal is encoded in the video bit stream as a syntax element at the macro-block level. In the context of SVC, the type of block transform applied to the base layer coding block is known to both the encoder and the decoder when upward sampling is performed. For the enhancement layer, the decoder can know the type of block transform used to encode the residue from the enhancement layer. However, on the encoder side, when the residual signal from the base layer is being sampled upwards, the real block transform that will be
41/54 used to encode the residue of the improvement layer is not yet known. One solution is to sample the base layer residue upwards based on the rules defined above, differently for the different types of block transforms attempted in the decision process on mode (s).
Two alternative methods can be used to mitigate this problem and provide common rules for both the encoder and the decoder to decide the enhancement layer encoding block size and, therefore, to identify a boundary between encoding blocks.
[Coding block rule A.] It can be assumed that the coding block size of the enhancement layer is 8x8 for the luma component and 4x4 for the chroma component; or [Coding block rule B.] It can be assumed that the coding block size of the enhancement layer is 8x8 for both the luma and chroma components.
Since the coding blocks in the base layer and in the enhancement layer are decided using one or the other of the above rules, if the pixel (s) involved in the interpolation process is (are) arranged at the borders between the coding blocks can be decided as follows:
The. A base layer pixel to be used in the interpolation of an upwardly sampled pixel can be considered to be arranged at the boundary between base layer coding blocks in a given direction (horizontal or vertical)
42/54 if it is either the last pixel within a base layer coding block or if it is the first pixel within a base layer coding block.
B. An upwardly sampled pixel to be interpolated can be considered to be arranged at the border between the enhancement layer blocks in a given direction (horizontal or vertical) if it is the last pixel within an enhancement layer coding block or if it is the first pixel within an enhancement layer coding block.
These rules consider the boundary between the coding blocks to be 1 pixel wide on each side. However, the boundary between the coding blocks can be considered to have widths other than 1 pixel on each side. In addition, the base layer and the enhancement layer may have different definitions of the boundaries of the coding blocks. For example, the base layer can define the boundary between the coding blocks as being one pixel wide on each side, while the enhancement layer can define the boundary between the coding blocks as being wider than that of a pixel. on each side, or vice versa.
It is noteworthy that the scope of this disclosure is not limited by the use of bi-linear interpolation. The bottom-up sampling decision based on the alignment of the blocks between the base layer and the enhancement layer can be applied to any interpolation scheme. The spatial ratios of 2: 1 and 5: 3, as well as the corresponding block alignments for these ratios and the corresponding weights shown in the interpolation equations,
43/54 are presented above as examples, but are not intended to limit the scope of this disclosure. In addition, the revealed scheme can be applied to residual upward sampling in systems and / or video encoding standards in which the encoding block size other than 4x4 and 8x8 can be used. Interpolation can also use weighted averages of several pixels located on one and the other side of the pixel to be interpolated.
To further alleviate the problem, the problem of signal discontinuity across the boundaries between the blocks in the base layer residue that may appear as internal pixels within the data sampled upwards, before the upward sampling, low-pass filtering on the residual signal of the base layer 'can reduce this discontinuity and improve the quality of the sampled waste in an upward way. For example, the operation on equations (11) and (12) can be performed on the pixel values at locations b3 and b4 on the base layer before equation (13) is applied to obtain the pixel values at e5 in Figure 8. Again, this can be implemented through filters 47 and 65 (low-pass filters, for example) located before rising sampler 45 (Figure 2) or rising sampler 59 (Figure 3).
ρ (ά3) = 1/4 * / 7 (ό2> 1/2 * ρ (ά3) + 1/4 * ΧΜ) (11) p (M) = l / 4 * p (à3> l / 2 * p (M) + l / 4 * p (W) (12) p (e5) = (l-3/5) * p (W) + 3/5 * p (M) (13)
In equations (11) and (12), the [1,2,1] smoothing filter is used as an example. Alternatively, a modified low-pass filter with less smoothing effect, for example, with derivation coefficients [1,6,1] instead of [1,2,1] can be used in (11) and
44/54 (12). In addition, an adaptive low-pass filter that adjusts the filtration intensity according to the nature and magnitude of the base layer discontinuity can be applied to the residual signal of the base layer before upward sampling is performed. The low-pass filter can be applied only to the pixels on the borders between the base layer blocks (b3 and b4 in Figure 8) or, alternatively, it can also be applied to pixels near the borders between the blocks (b2 and b5 in Figure 8) , for example).
The decision to apply low-pass filter to the signal of the base layer before the upward sampling can be based on the pixel locations of the enhancement layer and the base layer involved. For example, the additional low-pass filter can be applied to the pixels of the base layer if both of the following conditions are true:
1. The base layer pixels to be used in the interpolation process belong to different base layer coding blocks;
2. The upwardly sampled pixel to be interpolated corresponds to an internal pixel within an enhancement layer coding block. The upwardly sampled coding block can be determined using either the coding block rule A or the coding block rule B explained above.
Another way to reduce the discontinuity of the signal in the internal pixels to an encoding block sampled in an upward manner is to apply a low-pass filter to the sampled signal in an upward manner after
45/54 interpolation. This can be achieved by reordering the filters 47 and 65 in Figure 2 and Figure 3 to be applied after the upward samplers 45 and 59. Using the pixel at the e5 location in Figure 8 as an example, after p (e5) is obtained with using equations (10) and (11), the following can be applied:
p (e5) = 1/4 * p (e4) +1 / 2 * p (e5) +1 / .4 * p (e6) (14)
In equation (14), p (e4) and p (e6) are the pixel values sampled upwards at the e4 and e6 locations. Again, the [1,2,1] smoothing filter is used as an example. Alternative low-pass filtering can also be applied. For example, a modified smoothing filter with derivation coefficients [1,6,1] can be applied. Alternatively, adaptive low-pass filtering based on the nature and adaptive of the signal discontinuity can be applied.
The decision to apply an additional low-pass filter can be based on the locations of the pixels sampled upwardly and the base layer pixels involved. For example, the additional low-pass filter can be applied if both of the following conditions are true.
1. The upwardly sampled pixel corresponds to an internal pixel within an enhancement layer coding block. The enhancement layer coding block can be determined using either coding block rule A or coding block rule B (shown above) or any other coding block rules; and
46/54
2. The base layer pixels used in the interpolation process belong to different base layer encoding blocks.
SVC supports residual prediction to improve coding performance at the enhancement layer. In the spatial scaling capability, the residual signal from the base layer is sampled upwards before being used in the residual prediction. The discontinuity of locations where the signal quality can appear through base layer coding blocks can exist in the base layer residue, and this signal discontinuity can be carried up to the sampled signal in an upward manner. In the case of ESS, the inherited discontinuity can be arbitrary. To improve sampling in an ascending manner, disclosure is discussed an adaptive residual sampling scheme for scalable video encoding.
During residual up sampling, the scheme in the JD8 SVC prevents interpolation between pixels from different base layer coding blocks. In contrast, with the proposed adaptive bottom-up sampling scheme of this disclosure, the relative alignment of the blocks between the base layer and the improvement layer can be considered in the bottom-up sampling process. In particular, when the pixel to be sampled upwardly corresponds to an internal pixel in an enhancement layer coding block, interpolation and not the copy of the nearest neighbor pixel can be used to reduce signal discontinuity within a enhancement layer coding block.
In addition, as noted above, additional low-pass filtering can be applied to the signal
47/54 residual both before and after interpolation in order to further reduce the signal discontinuity that may exist within an enhancement layer coding block. In particular, if an upwardly sampled pixel corresponds to an internal pixel of the corresponding enhancement layer coding block and the base layer pixels involved are arranged at or near the boundaries of the phase layer coding blocks, then it can the following apply:
1. Prior to interpolation, an anti-aliasing filter or adaptive low-pass filter can be applied to the base layer pixels that are arranged at or near the boundaries between the base layer coding blocks;
2. After interpolation, the smoothing filter or low-pass filter can be applied to the enhancement layer pixel which is an internal pixel in an enhancement layer coding block.
Another factor to be considered when deciding whether to call interpolation or the copy of the nearest neighbor during residual upward sampling are the coding modes (inter-coded versus intrododed) of the base layer blocks involved. In SVC, upward sampling is only applied to the residual signal of inter-coded blocks. In JD8, when a base layer video block is intra-encoded, its residual signal is reset to zero in the residual image of the base layer. This results in an intense signal discontinuity through the residual blocks of the base layer when the encoding mode is changed. Therefore, it may be beneficial not to apply interpolation between two base layer pixels if they belong to blocks with encoding modes
48/54 different. In other words, the following decision can be used:
1. If two base layer pixels involved belong to two base layer coding blocks, one of which is intra-coded and the other inter-coded, then the nearest neighbor copy is used.
The above rule can be combined with the adaptive interpolation decision rules based on the alignment of the blocks shown in this disclosure. When combined, the adaptive interpolation decision can become the following:
The copy of the nearest neighboring pixel can be called when either of the following conditions 1 and 2 is true:
1. If the base layer pixels involved in the interpolation process belong to different base layer encoding blocks and the base encoding blocks have different encoding modes, or
2. If both of the following conditions are true:
The. The upwardly sampled pixel to be interpolated is arranged at the boundary between enhancement layer coding blocks.
B. The base layer pixels involved in the interpolation process belong to different base layer encoding blocks.
In the pseudo-code below, the underlined conditions can be added to the logic presented in JD8 during residual sampling and prediction in order to implement techniques compatible with this disclosure.
// Let bO and bl be the two layer pixels
49/54 base used in interpolation // Let w be the weight parameter used in bi-linear interpolation // Let e be the upwardly sampled pixel to be derived If (bO or bl belongs to a base layer coding block intra-coded OU (bO and bl belong to two base layer coding blocks E and are arranged on the border between blocks of the enhancement layer)) {
p (e) = (w> Vi)? p (bl): p (bO)}
else {
p (e) = (lw) * p (bO) + w * p (bl)}
The base layer block size is the actual block size used in base layer encoding. An FRext syntax element can be defined as part of the block header. If FRext is off, then the enhancement layer block size is also known as 4x4. However, if both the 4x4 and 8x8 transforms are allowed (that is, if FRext is turned on) for the enhancement layer, on the encoder side, the block size (that is, the transform size) in the enhancement layer is not it is even known when residual upward sampling is performed. In this situation (FRext linked to the enhancement layer), it can be assumed that the block size
50/54 improvement layer is 8x8 for luma and 4x4 for chroma. That is, for luma, the pixel sampled upwards in the location and is considered to be a border pixel between blocks if e = 8 * * ml or e = 8 * m (m is an integer). For chroma, the pixel sampled upwards in place and is considered to be a border pixel between blocks if e = 4 * ml or e = 4 * m.
For decoding, no syntax element is required. The JD8 decoding process (subparagraph G.8.10.3) can be modified as follows, with the modifications shown underlined.
For bi-linear interpolation for residual prediction, the Inputs are:
• One variable size (size = 8 or 4 for luma and size = 4 for chroma) • a TipoDeBlctrans arrangement [x, y] with x = O..mb 1 and y = 0..nb -1
The output of this process is an Interpres arrangement [x, y] with x = 0..m - 1 and y = ys..ye - 1.
• That the variable teml be derived as follows
- If IdxBlctrans [xl, y] is equal to IdxBlctrans [x2, yl] or 0 <(x%
SizeBlc) <(SizeBlbl-1) and TypeBlctrans [xl, yl] is equal to
TypeBlctrans [x2, y2]
Templ = r [xl, yl] * (16- (posX [x]% 16)) + r [x2, yl] * (posX [x]% 16) (G-544) • That the variable temp2 is derived from following way
- If IdxBlctrans [xl, y2] is equal to
51/54
IdxBlctrans [x2, y2] or 0 <(x%
SizeBlc) <(SizeBlobl-1) and TypeBlctrans [xl, y2] is equal to
TypeBlctrans [x2, y2]
Temp2 = r [xl, y2] * (16- (posX [x]% 16)) + r [x2, y2] * (posX [x]% 16) (G-547) • May Interpres be derived as follows - If IdxBlctrans [xl, yl] is equal to
IdxBlctrans [xl, y2] or 0 <(x%
SizeBlc) <(SizeBll-1) and TypeBlctrans [xl, yl] is equal to TypeBlctrans [xl, y2]
Interprs [x, y] = (templ ★ (16 - (posY [y]% 16)) + temp2 * (posY [y]% 16) + (128)) »8) (G-550)
Simulations were performed according to the verification conditions of basic experiment 2 (CE2) specified in JVT-V302, and the simulations showed improvements in the peak signal-to-noise ratio in video quality when the techniques of this revelation are used with respect to conventional techniques presented in JD8. In the case of ESS, the small change proposed in residual upward sampling provides a very simple and effective way to reduce blocking artifacts in the upwardly sampled residue within an enhancement layer coding block. The results of the simulations show that the proposed change intensifies the coding performance compared to JSVM_7_13 for all CE2 verification conditions. In addition, the proposed scheme considerably improves visual quality by suppressing annoying blocking artifacts in the reconstructed enhancement layer video.
52/54
Figure 9 is a flow diagram showing a technique compatible with this disclosure. Figure 9 will be described from the perspective of the encoder, although a similar process can be performed by the decoder. As shown in Figure 9, the base layer encoder 32 encodes base layer information (200) and the enhancement layer encoder 34 encodes improvement layer information (202). As part of the enhancement layer encoding process, the enhancement layer encoder 34 (the video encoder 50 shown in Figure 2, for example) receives base layer information (204). The filter 47 can perform optional border filtering between blocks of the base layer information (206). The upward sampler 45 upwardly samples the base layer information to generate upwardly sampled video blocks using the various techniques and rules defined here in order to select between interpolation and the nearest neighbor copy (208). The optional filter 47 can also be applied after the rising sampler 45 for the enhancement layer video pixels that are interpolated from the border pixels between base layer modes (210) (the e5 pixel in Figure 8 and in equation 14), and the video encoder 50 uses the sampled data upwardly to encode enhancement layer information (212).
The techniques described in this disclosure can be implemented on one or more processors, such as, for example, a general purpose microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable port arrangement on the field (FPGA) or other equivalent logical devices.
53/54
The techniques described here can be implemented in hardware, software, firmware or any combination of them. Any features described as modules or components can be implemented together in an integrated logic device or separately as discrete but inter-actionable logic devices. If implemented in software, the techniques can be performed, at least in part, by means of a computer-readable medium that comprises instructions that, when executed, would be one of the methods described above. The computer-readable data storage medium may be part of a computer program product, which may include packaging materials. 0 computer-readable media can comprise random access memory (RAM), such as synchronous random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory electrically erasable (EEPROM), FLASH memory, magnetic or optical data storage media and the like. In addition, or alternatively, the techniques can be performed, at least in part, by means of a computer-readable communication medium that carries or communicates code in the form of instructions or data structures and that can be accessed, read and / or run by a computer.
The program code can be executed by one or more processors, DSPs, general purpose microprocessors, ASICs, FPGAs or other equivalent integrated or discrete logic circuitry. Accordingly, the term processor as used herein may refer to any one of the preceding structure or any other structure suitable for implementing the techniques described herein. In this disclosure, the term processor intends to
54/54 cover any combination of one or more microprocessors, DSPs, ASICs, FPGAs or logic. In addition, in some respects, the functionality described here can be presented within dedicated software modules or hardware modules configured for encoding and decoding or incorporated into a combined video encoder-decoder (CODEC).
If implemented in hardware, this disclosure can refer to a circuit, such as an integrated circuit, chip set, ASIC, FPGA, logic or various combinations of them configured to perform one or more of the techniques described here.
Various embodiments of the invention have been described. These and other modalities are within the scope of the following claims.
Contents7
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
15 priority claims, no other members on record
Priority claims15
| Document | Office | Kind | Date |
|---|---|---|---|
| 60884099 | United States of America | – | |
| 88409907 | United States of America | P | |
| 60888912 | United States of America | – | |
| 88891207 | United States of America | P | |
| 11970413 | United States of America | – | |
| 97041308 | United States of America | A | |
| 2008050546 | United States of America | W | |
| 11970413 | – | – | – |
| 2008050546 | – | – | – |
| 60884099 | – | – | – |
| 60888912 | – | – | – |
| US20070884099P | – | – | – |
| US20070888912P | – | – | – |
| US20080970413 | – | – | – |
| WO2008US50546 | – | – | – |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Dismissal acc. art. 36, par 1 of ipl - no reply within 90 days to fullfil the necessary requirementsB11B | B11B | |
| Application fees: dismissal - article 86 of industrial property lawB08F | B08F | |
| Preliminary requirement: requests with searches performed by other patent offices: suspension of the patent application procedureB06U | B06U | |
| Objections, documents and/or translations needed after an examination request according art. 34 industrial property lawB06F | B06F | |
| Others concerning applications: alteration of classificationB15K | B15K |
Numbers
- Publication
- PI0806521
- Publication, DOCDB
- PI0806521
- Publication, EPODOC
- BRPI0806521
- Application
- 6521
- Application, DOCDB
- PI0806521
- Application, EPODOC
- BR2008PI06521
Titles2
- Portuguese
- AMOSTRAGEM ASCENDENTE ADAPTATIVA PARA CODIFICAÇÃO ESCALONÁVEL DE VÍDEO
- English
- ADAPTIVE ASCENDING SAMPLING FOR SCALABLE VIDEO ENCODING
Classification
- CPC, 7
- H04N19/59
- H04N19/105
- H04N19/117
- H04N19/159
- H04N19/187
- H04N19/30
- H04N19/33
