Tiling in video encoding and decoding
Abstract
Implementations are provided that refer, for example, to viewing tiles in video encoding and decoding. A specific method includes accessing a video image that includes multiple images in the accessed video image of at least one of the multiple images (824,826) and providing the information accessed and the decoded video image as an output (824,826). Some other implementations format or process information that indicates how multiple images included in a single video image are combined into a single video image, and format or process a coded representation of the multiple images combined.

Term
1.5 yearsleft in the term
Expires 11 April 2028.
- Priority
- Filed
- Granted
- Today
- Expires
28 claims: 4 independent, 24 dependent
- 1REIVINDICAÇÕES 1. Método, CARACTERIZADO pelo fato de que compreende:gerar informação indicando como múltiplas imagens incluídas em uma imagem de vídeo são combinadas em uma imagem de vídeo codificada, em que a informação indica se uma ou mais das múltiplas imagens são invertidas em relação à sua orientação pretendida;codificar a imagem de vídeo para prover uma representação codificada das múltiplas imagens combinadas;e prover a informação gerada e a imagem de vídeo codificado como saída.
- 2Método, de acordo com a reivindicação 1, CARACTERIZADO pelo fato de que as múltiplas imagens incluem uma primeira e uma segunda imagens, cada qual tendo uma orientação respectiva;em que a informação gerada inclui uma indicação de inversão para indicar que nenhuma das primeira e segunda imagens são invertidas em relação à sua orientação respectiva;e em que a codificação compreende ainda dispor a primeira e a segunda imagens na imagem de vídeo codificado em suas orientações respectivas.
- 3Método, de acordo com a reivindicação 1, CARACTERIZADO pelo fato de que as múltiplas imagens incluem uma primeira e uma segunda imagens, cada qual tendo uma orientação respectiva;em que a informação gerada inclui uma indicação de inversão para indicar que a primeira imagem está invertida horizontalmente com relação à sua respectiva orientação;e em que a codificação compreende ainda dispor a primeira e a segunda imagens na imagem de vídeo codificado de tal modo que a primeira imagem é invertida em uma direção horizontal com relação à sua orientação.
- 4Método, de acordo com a reivindicação 1, CARACTERIZADO pelo fato de que as múltiplas imagens incluem uma primeira e uma segunda imagens, cada qual tendo uma orientação respectiva;em que a informação gerada inclui uma indicação de inversão para indicar que a primeira imagem está invertida verticalmente com relação à sua respectiva orientação;e em que a codificação compreende ainda dispor a primeira e a segunda imagens na imagem de vídeo codificado, de tal modo que a primeira imagem é invertida em uma direção vertical com relação à sua orientação.
- 5Método, de acordo com a reivindicação 1, CARACTERIZADO pelo fato de que as múltiplas imagens incluem uma primeira e uma segunda imagens, em que a codificação compreende ainda dispor a primeira e a segunda imagens na imagem de vídeo codificado, e em que a primeira imagem é disposta ao lado da segunda imagem na imagem de vídeo codificado.
- 6Método, de acordo com a reivindicação 1, CARACTERIZADO pelo fato de que as múltiplas imagens incluem uma primeira e uma segunda imagens, em que a codificação compreende ainda dispor a primeira e a segunda imagens na imagem de vídeo codificado, e em que a primeira imagem é disposta sobre a segunda imagem na imagem de vídeo codifi2 cado.
- 7Método, de acordo com a reivindicação 1, CARACTERIZADO pelo fato de que as múltiplas imagens incluem uma primeira e uma segunda imagens, em que a codificação compreende ainda dispor a um nível de pixel a primeira e a segunda imagens na imagem de vídeo codificado, e em que os pixels da primeira e da segunda imagens são entrelaçados alternatívamente.
- 8Método, de acordo com a reivindicação 1, CARACTERIZADO pelo fato de que a informação gerada inclui uma indicação de inversão, em que a etapa de geração compreende formar uma mensagem de acordo com uma sintaxe de alto nível, em que a mensagem inclui a indicação de inversão.
- 9Método, de acordo com a reivindicação 8, CARACTERIZADO pelo fato de que a sintaxe de alto nível é selecionada de pelo menos o grupo de sintaxes de alto nível consistindo de um cabeçalho de seção, conjunto de parâmetros de sequência, conjunto de parâmetros de imagem, conjunto de parâmetros de vista, cabeçalho de unidade de camada de abstração de rede e uma mensagem de informação de otimização suplementar.
- 10Método CARACTERIZADO pelo fato de que compreende:acessar uma imagem de vídeo que inclui múltiplas imagens combinadas em uma única imagem, a imagem de vídeo sendo parte de uma corrente de vídeo recebida;acessar informação indicando como as múltiplas imagens na imagem de vídeo acessado são combinadas, em que a informação acessada indica se pelo menos uma das múltiplas imagens é invertida com relação à sua orientação pretendida;decodificar a imagem de vídeo para prover uma representação decodificada de pelo menos uma das múltiplas imagens;e prover pelo menos uma das informações acessadas e a representação decodificada como saída.
- 11Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que as múltiplas imagens incluem uma primeira e uma segunda imagens, em que a informação acessada inclui uma indicação de inversão para indicar que nem a primeira nem a segunda imagens são invertidas com relação à sua orientação pretendida, e em que a decodificação compreende ainda desmontar a primeira e a segunda imagens, ambas provenientes da imagem de vídeo codificado sem modificar as suas orientações pretendidas.
- 12Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que as múltiplas imagens incluem uma primeira e uma segunda imagens;em que a informação acessada inclui uma indicação de inversão para indicar que a primeira figura é invertida horizontalmente com relação à sua orientação pretendida, e em que a decodificação compreende ainda desmontar a primeira e a segunda imagens, ambas provenientes da imagem de vídeo codificado;e método compreendendo ainda, responsivo à indicação de inversão, inverter a primeira imagem em uma direção horizontal com relação à sua orientação pretendida.
- 13Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que as múltiplas imagens incluem uma primeira e uma segunda imagens, em que a informação gerada inclui uma indicação de inversão para indicar que a primeira figura é invertida verticalmente com relação à sua orientação pretendida, e em que a decodificação compreende ainda desmontar a primeira e a segunda imagem, ambas provenientes da imagem de vídeo codificado;e método compreendendo ainda, responsivo à indicação de inversão, inverter a primeira imagem em uma direção vertical com relação à sua orientação.
- 14Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que a decodificação compreende ainda desmontar a primeira e a segunda imagens a partir da imagem de vídeo codificado, e em que a primeira imagem é disposta ao lado da segunda imagem na imagem de vídeo codificado.
- 15Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que compreende ainda desmontar a primeira e a segunda imagens a partir da imagem de vídeo codificado, e em que a primeira imagem é disposta sobre a segunda imagem na imagem de vídeo codificado.
- 16Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que compreende ainda desmontar, a um nível de pixel, a primeira e a segunda figuras a partir da imagem de vídeo codificada, e em que pixels da primeira e segunda imagens são entrelaçados alternativamente.
- 17Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que as informações de acesso compreendem extrair uma indicação de inversão de uma mensagem formatada de acordo com uma sintaxe de alto nível, em que a mensagem inclui uma indicação de inversão.
- 18Método, de acordo com a reivindicação 17, CARACTERIZADO pelo fato de que a sintaxe de alto nível é selecionada de um grupo de sintaxes de alto nível consistindo de um cabeçalho de seção, um conjunto de parâmetros de sequência, um conjunto de parâmetros de imagem, um conjunto de parâmetros de vista, um cabeçalho de unidade de camada de abstração de rede e uma mensagem de informação de otimização suplementar.
- 19Método, de acordo com a reivindicação 10, CARACTERIZADO por:acessar a imagem de vídeo compreende acessar uma imagem de vídeo provida de acordo com um padrão de vídeo de vista única que trata todas as imagens como sendo de uma vista única (824);e acessar a informação compreende acessar informação provida de acordo com o padrão de vídeo de vista única (804), tal que o fornecimento da imagem de vídeo decodificado e a informação acessada possibilita suporte de múltiplas vistas pelo padrão de vídeo de vista única (826).
- 20Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que a informação acessada indica pelo menos uma de um local e uma orientação de ao menos uma das múltiplas imagens dentro da imagem de vídeo (822).
- 21Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que o acesso à imagem de vídeo, acesso à informação, decodificação da imagem de vídeo, e fornecimento da informação acessada e imagem de vídeo decodificada são realizados por um decodificador (1650).
- 22Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que a informação é compreendida em pelo menos um de um cabeçalho de seção, conjunto de parâmetros de sequência, conjunto de parâmetros de imagem, conjunto de parâmetros de vista, cabeçalho de unidade de camada de abstração de rede e uma mensagem de informação de otimização suplementar (804).
- 23Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que:pelo menos uma das múltiplas imagens é invertida na direção horizontal na imagem única, e a informação acessada indica que a inversão é na direção horizontal.
- 24Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que:pelo menos uma das múltiplas imagens é invertida na direção vertical na imagem única, e a informação acessada indica que a inversão é na direção vertical.
- 25Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que a informação acessada indica que uma segunda imagem das múltiplas imagens é invertida com relação à sua orientação pretendida.
- 26Método, de acordo com a reivindicação 10, CARACTERIZADO pelo fato de que compreende:pelo menos uma das múltiplas imagens inclui uma primeira imagem de uma primeira vista;e outra das múltiplas imagens incluem informação de profundidade para a primeira imagem.
- 27Aparelho, CARACTERIZADO pelo fato de ser configurado para realizar um ou mais dos métodos do tipo definido em qualquer uma das reivindicações 1 a 26.
- 28Meio legível por processador, CARACTERIZADO pelo fato de possuir armaze5 nado no mesmo uma estrutura de sinal de vídeo, a estrutura de sinal de vídeo compreendendo:uma seção de imagem codificada incluindo uma codificação de uma imagem de vídeo codificada, a imagem de vídeo codificada incluindo uma primeira e uma segunda figura 5 dispostas na imagem de vídeo codificado;e uma seção de sinalização incluindo uma codificação de uma indicação que indica se pelo menos uma das primeira e segunda figuras são invertidas com relação à suas respectivas orientações, a indicação permitindo decodificar as imagens de vídeo codificados em versões decodificadas da primeira e da segunda imagens. 10 29. Meio legível por máquina, CARACTERIZADO pelo fato de ter armazenado no mesmo instruções executáveis de máquina que, quando executadas, implementam um ou mais dos métodos do tipo definido em qualquer uma das reivindicações 1 a 26. 100
Independent claims28
510 paragraphs in 12 sections, as filed
(54) Title: TILE IN VIDEO ENCODING AND DECODING (30) Unionist Priority: 12/04/2007 us 60 / 923.014, 20/04/2007 US 60 / 925,400 (73) Holder (s): Thomson Licensing (72) Inventor (s): Dong Tian, Peng Yin, Purvin Bibhas Pandit (74) Attorney (s): NELLIE ANNE DAIEL-SHORES (86) International Order: pct US2008004747 of 11/04/2008 (87) International Publication: wo 2008 / 127676of 23/10/2008 (57) Abstract: Tiling in encoding and VIDEO DECODING. Implementations are provided that refer, for example, to viewing tiles in video encoding and decoding. A specific method includes accessing a video image that includes multiple images in the accessed video image of at least one of the multiple images (824,826) and providing the information accessed and the decoded video image as an output (824,826). Some other implementations format or process information that indicates how multiple images included in a single video image are combined into a single video image, and format or process a coded representation of the multiple images combined.
<img file="BRPI0809510A2_D0001.tif" />
"TILE IN VIDEO ENCODING AND DECODING"
REMISSIVE REFERENCE TO RELATED ORDERS
This claim claims the benefit of each of: (1) United States Provisional Order 60 / 923,014, filed on April 12, 2007, and entitled “Multiview Information” (Lawyer Dossier No. PU070078), and (2) United States Provisional Order 60 / 925,400, filed on April 20, 2007 and entitled “View Tiling in MVC Coding” (Lawyer Dossier No. PU070103). Each of these two orders is incorporated herein in full by reference.
TECHNICAL FIELD
These principles generally refer to video encoding and / or decoding.
BACKGROUND
Video manufacturers can use a tiling or tiling architecture of different lists in a single frame. The views can then be extracted from their respective locations and rendered.
SUMMARY
In general, a video image is accessed that includes multiple images combined into a single image. Information is accessed indicating how the multiple images, in the accessed video image, are combined. The video image is decoded to provide a decoded representation of the multiple images combined. The information accessed and the decoded video image are provided as an output.
According to another general aspect, information is generated indicating how multiple images included in a video image are combined into a single image. The video image is encoded to provide an encoded representation of the multiple images combined. The information generated and the encoded video image are provided as an output.
According to another general aspect, a signal or signal structure includes information indicating how multiple images included in a single video image are combined into a single video image. The signal or signal structure also includes a coded representation of the multiple images combined.
According to another general aspect, a video image is accessed which includes multiple images combined into a single image. Information is accessed which indicates how the multiple images in the accessed video image are combined. The video image is decoded to provide a decoded representation of at least one of the multiple images. The information accessed and the decoded representation are provided as an output.
According to another general aspect, a video image is accessed which includes multiple images combined into a single image. Information is accessed which indicates how the multiple images in the accessed video image are combined. The video image is decoded to provide a decoded representation of the multiple images combined. User input is received which selects at least one of the multiple images to display. A decoded output of at least one selected image is provided, the decoded output being provided based on the information accessed, the decoded representation, and the user input.
Details of one or more implementations are presented in the attached drawings and in the description below. Even if described in a specific way, it must be evident that implementations can be configured or incorporated in different ways. For example, an implementation can be performed as a method, or incorporated as a device configured to perform a set of operations, or incorporated as a device storing instructions to perform a set of operations, or incorporated into a signal. Other aspects and characteristics will become evident from the following detailed description, considered in conjunction with the attached drawings and the claims.
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1 is a diagram showing an example of four tiled views in a single frame;
Figure 2 is a diagram showing an example of four views turned and tiled in a single frame;
Figure 3 shows a block diagram for a video encoder to which the present principles can be applied, according to a modality of the present principles;
Figure 4 shows a block diagram for a video decoder to which the present principles can be applied, according to a modality of the present principles;
Figure 5 is a flow chart for a method for encoding images for a plurality of views using the MPEG-4 AVC Standard, according to one embodiment of the present principles;
Figure 6 is a flow chart for a method for encoding images for a plurality of views using the MPEG-4 AVC Standard, according to one embodiment of the present principles;
Figure 7 is a flow chart for a method for encoding images for a plurality of views and depths using the MPEG-4 AVC Standard, according to a modality of the present principles;
Figure 8 is a flow chart for a method for decoding a plurality of views and depths using the MPEG-4 AVC Standard, according to a modality of the present principles;
Figure 9 is a diagram showing an example of a depth signal, according to an embodiment of the present principles;
Figure 10 is a diagram showing an example of a depth signal added with a tile, according to an embodiment of the present principles;
Figure 11 is a diagram showing an example of five tiled views in a single frame, according to a modality of the present principles.
Figure 12 is a block diagram for an exemplary Multiple View Video Encoding (MVC) encoder to which the present principles can be applied, according to a modality of the present principles;
Figure 13 is a block diagram for an exemplary Multiple View Video Encoding (MVC) decoder to which the present principles can be applied, according to a modality of the present principles;
Figure 14 is a flow chart for a method for processing images for a plurality of views in preparation for encoding the images using the MPEG-4 AVC Standard multi-view video encoding (MVC) extension, according to one embodiment of the present Principles;
Figure 15 is a flowchart for a method for encoding images for a plurality of views using the MPEG-4 AVC Standard Multi-View Video Encoding (MVC) extension, according to one embodiment of the present principles;
Figure 16 is a flow chart for a method for processing images for a plurality of views in preparation for decoding the images using the MPEG-4 AVC Standard multi-view video encoding (MVC) extension, in accordance with a modality of present principles;
Figure 17 is a flowchart for a method for decoding images for a plurality of views using the MPEG-4 AVC Standard Multi-View Video Encoding (MVC) extension, according to a modality of the present principles;
Figure 18 is a flowchart for a method for processing images for a plurality of views and depths in preparation for encoding images using the MPEG-4 AVC Standard Multi-View Video Encoding (MVC) extension, according to one modality of the present principles;
Figure 19 is a flowchart for a method for encoding images for a plurality of views and depths using the MPEG-4 AVC Standard Multi-View Video Encoding (MVC) extension, according to one embodiment of the present principles;
Figure 20 is a flowchart for a method for processing images for a plurality of views and depths in preparation for decoding the images using the MPEG-4 AVC Standard Multi-View Video Encoding (MVC) extension, according to a modality of the present principles;
Figure 21 is a flowchart for a method for decoding images for a plurality of views and depths using the MPEG-4 AVC Standard Multi-View Video Encoding (MVC) extension, according to a modality of the present principles;
Figure 22 is a diagram showing examples of pixel-level tiling, according to a modality of the present principles; and
Figure 23 shows a block diagram for a video processing device, to which the present principles can be applied, according to a modality of the present principles.
DETAILED DESCRIPTION
Several implementations are directed at the methods and apparatus for viewing tiles in video encoding and decoding. Thus it will be considered that those versed in the technique will be able to conceive several arrangements that, although not described here or shown explicitly, incorporate the principles and are included within its spirit and scope.
All examples and conditional language, quoted here, are intended to have pedagogical purposes to assist the reader in understanding the present principles and concepts contributed by the inventor (s) to favor the technique, and should be considered to be without limitation to such examples and conditions specifically cited.
In addition, all statements made here citing principles, aspects and modalities of the present principles, as well as their specific examples, are intended to cover their equivalents, not only structural but also functional. Additionally, it is intended that such equivalents include the equivalents currently known as well as the equivalents developed in the future, that is, any elements developed that perform the same function, regardless of the structure.
Thus, for example, it will be considered by those skilled in the art that the 30 grams of blocks presented here represent conceptual views of sets of illustrative circuits incorporating the present principles. Similarly, it will be considered that any flowcharts, diagrams, state transition diagrams, pseudocode, and the like, represent various processes that can be substantially represented in computer-readable media and thus executed by a computer or processor, whether that computer or processor is shown or not explicitly.
The functions of the various elements shown in the figures can be provided through the use of dedicated hardware as well as hardware capable of running software in association with appropriate software. When provided by a processor, functions can be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which can be shared. In addition, the explicit use of the term “processor” or “controller” should not be considered as referring exclusively to hardware capable of running software, and may implicitly include, without limitation, digital signal processor (“DSP”) hardware, read memory (“ROM”) to store software, random access memory (“RAM”), and non-volatile storage medium.
Other hardware, conventional and / or special, can also be included. Similarly, any switches shown in the figures are only conceptual. Their function can be performed through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the specific technique can be selectable by the implementer as understood more specifically from of the context.
In the claims any elements expressed as a means to perform a specified function are intended to cover any form of performing that function including, for example, a) a combination of circuit elements that perform that function or b) software in any form, including, therefore, firmware, microcode or similar, combined with appropriate circuitry to run this software to perform the function. The present principles as defined by such claims reside in the fact that the functionalities provided by the various means cited are combined and united in the form in which demanded by the claims. It is therefore considered that any means that can provide these functionalities are equivalent to those shown here.
Reference in the specification to “a modality” or (“an implementation”) or “a modality” (or “an implementation”) of these principles means that a specific aspect, structure, characteristic, and so on described in connection with the modality is included in at least one modality of the present principles. Thus, the appearance of the phrase "in one modality" or "in some modality" appearing in various places throughout the specification does not all refer to the same modality.
It should be considered that the use of the terms "and / or" and "at least one of", for example, in the cases of "A and / or B" and "at least one of A and B", intends to cover the selection of first related option (A) only, or selecting the second related option (B) only, or selecting both options (A and B). As an additional example, in the cases of “A, B, and / or C” and “at least one of A, B and C”, this sentence is intended to cover the selection of only the first related option (A), or the selection only the second related option (B), or selecting only the third related option (C), or selecting only the first and second related option (A and B), or selecting only the first and third related option (A) and C), or selecting only the second and third related option (B and C), or selecting all three options (A and B and C). This can be extended, as is readily evident to those of ordinary knowledge in these techniques and in the techniques related, therefore, to the many related items.
In addition, it should be considered that although one or more modalities of these principles are described here with respect to the MPEG-4 AVC standard, these principles are not limited to just that standard and thus can be used with respect to other standards, recommendations, and extensions thereof, particularly video encoding standards, recommendations, and extensions thereof, including extensions to the MPEG-4 AVC standard, while maintaining the spirit of these principles.
In addition, it should be considered that although one or more different modalities of the present principles are described here with respect to the multi-view video encoding extension of the MPEG-4 AVC standard, the present principles are not limited only to that extension and / or this standard and thus can be used in relation to other video encoding standards, recommendations and extensions thereof related to multi-view video encoding, while maintaining the spirit of the present principles. Multi-view video encoding (MVC) is the compression structure for encoding multi-view streams. A Multiple View Video Encoding (MVC) sequence is a set of two or more video sequences that capture the same scene from a different point of view.
In addition, it must be considered that although one or more different modalities of the present principles are described here which use in-depth information regarding the video content, the present principles are not limited to such modalities and, thus, other modalities can be implemented which do not use in-depth information, while maintaining the spirit of the present principles.
Additionally, as used here, “high-level syntax” refers to the syntax present in the bit stream that resides hierarchically above the macroblock layer. For example, high-level syntax, as used here, can refer to, but is not limited to, section header level syntax, Supplemental Optimization Information (SEI) level, Image Parameter Set (PPS) level, Sequence Parameter Set (SPS) level, View Parameter Set (VPS), and Network Abstraction Layer (NAL) unit header level.
In this implementation of multi-video encoding (MVC) based on recommendation H.264 of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) of the International Telecommunication Union / Advanced Video Encoding Standard (AVC) Part 10 of the Group 4 by Cinematographic Specialists (MPEG-4) (hereinafter the “MPEG-4 AVC Standard”), the reference software obtains prediction of multiple views by coding each view with a single coder considering the cross-view references. Each view is encoded as a bit stream, separated by the encoder in its original resolution and subsequently all bit streams are combined to form a single bit stream which is then decoded. Each view produces a separate YUV decoded output.
Another approach to predicting multiple views involves grouping a set of views into pseudovists. In an example of this approach, we can tile the images from each N views out of a total of M views (sampled at the same time) in a larger frame or in a superframe with possible downward sampling or other operations. According to Figure 1, an example of the four tiled views in a single frame is usually indicated by the reference numeral 100. All four views are in their normal orientation.
According to Figure 2, an example of four overturned and tiled views in a single frame is usually indicated by the reference numeral 200. The top left view is in its normal orientation. The top right view is horizontally turned. The bottom left view is turned vertically. The bottom right view is turned horizontally as well as vertically. So, if there are four views, then an image from each view is arranged in a superframe like a tile. This results in a single non-coded input string with a high resolution.
Alternatively, we can downwardly sample the image to produce a lower resolution. Thus, we create multiple sequences each of which includes different views that are tiled together. Each such sequence then forms a pseudovista, where each pseudovista includes N different tiled views. Figure 1 shows a pseudovista, and Figure 2 shows another pseudovista. These pseudovistas can then be encoded using existing video encoding standards such as the ISO / IEC MPEG-2 Standard and the MPEG-4 AVC Standard.
Yet another approach to predicting multiple views simply involves coding the different views independently using a new standard and, after decoding, tiling the views as required by the reproduction apparatus.
Additionally, in another approach, views can also be tiled in the form of pixels. For example, in a super view that is made up of four views, pixel (x, y) can be from view 0, while pixel (x + 1, y) can be from view 1, pixel (x, y + 1) can be from view 2, and pixel (x + 1, y + 1) can be from view 3.
Many video manufacturers use such an arrangement or tiling structure for different views in a single frame and then extracting the views from their respective locations and rendering them. In such cases, there is no standard way to determine whether the bit stream has such a property. Thus, if a system uses the method of tiling images from different views in a large frame, then the method of extracting the different views is recorded.
However, there is no standard way to determine whether the bit stream has such a property. We propose a high level syntax to facilitate the rendering or reproduction apparatus to extract such information to assist in the display or other further processing. It is also possible that the subimages have different resolutions and some upward sampling may be necessary to eventually render the view. The user may wish to have the bottom-up sampling method also indicated in the high-level syntax. In addition, parameters for changing the depth focus can also be transmitted.
In one modality, we propose a new Supplementary Optimization Information (SEI) message to signal information from multiple views in an MPEG-4 AVC Standard compatible with the bit stream where each image includes sub-images that belong to a different view. The modality is intended, for example, for the easy and convenient display of video streams from multiple views on three-dimensional (3D) monitors which can use such a structure. The concept can be extended to other video coding standards and recommendations by signaling such information using high-level syntax.
In addition, in one modality, we propose a method of signaling how to arrange the views before they are sent to the multi-view video encoder and / or decoder. Advantageously, the modality can lead to a simplified implementation of multi-view coding, and can benefit from the coding efficiency. Certain views can be put together and form a pseudo view or super view and then the tiled view is treated as a normal view by a multi-view video encoder and / or decoder, common, for example, according to the implementation based on the MPEG Standard -4 current multi-view video encoding stroke. A new flag is proposed in the Sequence Parameter Set (SPS) extension for multi-view video encoding to signal the use of the pseudovista technique. The modality is intended for the easy and convenient display of video streams from multiple views on 3D monitors that use such a structure.
Encoding / decoding using a single view video encoding / decoding standard / recommendation
In this implementation of multivideo encoding (MVC), based on recommendation H.264 of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) of the International Telecommunication Union / Advanced Video Encoding Standard (AVC) Part 10 of Group 4 of Cinematographic Specialists (MPEG-4), telecommunications sector (ITU-T) (hereinafter the “MPEG-4 AVC Standard”), the reference software obtains multiview prediction by coding each view with a single encoder and considering cross-view references . Each view is encoded as a bit stream separated by the encoder at its original resolution and subsequently all bit streams are combined to form a single bit stream which is then decoded. Each view produces a separate YUV decoded output.
Another approach to multiview prediction involves tiling the images from each view (sampled at the same time) in a larger frame or in a superframe with a possible downward sampling operation. Turning now to Figure 1, an example of four tiled views in a single frame is usually indicated, usually indicated by reference numeral 100. Turning now to Figure 2, an example of four views turned and tiled in a single frame is usually indicated by the reference numeral 200. So, if there are four views, then an image for each view is arranged in a superframe as a tile. This results in a single non-coded input string with a high resolution. This signal can then be encoded using existing video encoding standards such as the ISO / IEC MPEG-2 standard and the MPEG-4 AVC standard.
Yet another approach to multiview prediction involves simply coding the different views independently using a new standard and, after decoding, tiling the views as required by the reproduction apparatus.
Many video manufacturers use such a layout or tiling structure for different views in a single frame and then extracting the views from their respective locations and rendering them. In such cases, there is no standard way of determining whether the bit stream has such a property. Thus, if a system uses the method of tiling images from different views in a large frame, then the method of extracting the different views is recorded.
Turning now to Figure 3, a video encoder capable of encoding video according to the MPEG-4 AVC standard is generally indicated by the reference numeral 300.
The video encoder 300 includes a frame ordering store 310 having an output in signal communication with a non-inverting input of a combiner 385. An output of combiner 385 is connected in signal communication with a first input of a transformer and quantizer 325. An output of the transformer and quantizer 325 is connected in signal communication with a first input of an entropy encoder 345 and a first input of an inversion transformer and inversion quantizer 350. An output of the entropy encoder 345 is connected in communication of signal with a first non-inverting input of a combiner 390. An output of the combiner 390 is connected in signal communication with a first input of an output store 335.
A first output of an encoder controller 305 is connected in signal communication with a second input of the frame ordering store 310, a second input of the inversion transformer and inversion quantizer 350, an input of a type decision module image 315, an entry for a macroblock type decision module (MB-type) 320, a second entry for an intraprediction module 360, a second entry for an unlock filter 365, a first entry of a motion compensator 370, a first entry of a motion estimator 375, and a second entry of a reference image store 380.
A second output of the encoder controller 305 is connected in signal communication with a first input of a Supplementary Optimization Information (SEI) insertion means 330, a second input of the transformer and quantizer 325, a second input of the entropy encoder 345 , a second input from the output store 335, and an input from the Sequence Parameter Set (SPS) and Image Parameter Set (PPS) 340 insertion medium.
A first output of the image type decision module 315 is connected in signal communication with a third input of a frame ordering store 310. A second output of the image type decision module 315 is connected in signal communication with a second entry of a macroblock type decision module 320.
An output of the Sequence Parameter Set (SPS) and Image Parameter Set (PPS) 340 inserts is connected in signal communication with a third non-inversion input of combiner 390. An output of the SEI insertion medium 330 is connected in signal communication with a second non-inverting input of combiner 390.
An output of the inversion quantizer and inversion transformer 350 is connected in signal communication with a first non-inversion input of a combiner 319. An output of the combiner 319 is connected in signal communication with a first input of the intraprediction module 360 and a first input of the unlock filter 365. An output of the unlock filter 365 is connected in signal communication with a first input of a reference image store 380. An output of the reference image store 380 is connected in signal communication with a second input of the motion estimator 375 and with a first input of a motion compensator 370. A first output of the motion estimator 375 is connected in signal communication. with a second input of the motion compensator 370. A second output of the motion estimator 375 is connected in signal communication with a third input of the entropy encoder 345.
An output of the motion compensator 370 is connected in signal communication with a first input of a switch 397. An output of the intraprediction module 360 is connected in signal communication with a second input of the switch 397. An output of the decision module of macroblock type 320 is connected in signal communication with a third input of switch 397 to provide a control input for switch 397. The third input of switch 397 determines whether or not the switch's “data” input (in comparison to the control input, that is, the third input) should be provided by the motion compensator 370 or the intraprediction module 360. A switch output 397 is connected in signal communication with a second non-inverting input of combiner 319 and with an inverting input of combiner 385.
The inputs of frame ordering store 310 and encoder controller 105 are available as inputs of encoder 300 to receive an input image 301. In addition, an input from the Supplementary Optimization Information (SEI) insert means 330 is available as an input from encoder 300, to receive metadata. An output from output store 335 is available as an output from encoder 300 to output a bit stream.
Turning now to Figure 4, a video decoder capable of performing video decoding in accordance with the MPEG-4 AVC Standard is generally indicated by reference numeral 400.
The video decoder 400 includes an input store 410 having an output connected in signal communication with a first input of the entropy decoder 445. A first output of the entropy decoder 445 is connected in signal communication with a first input of a transformer inversion and inversion quantizer 450. An output of the reversing transformer and the reversing quantizer 450 is connected in signal communication with a second non-inverting input of a combiner 425. An output of the combiner 425 is connected in signal communication with a second input of a combiner filter. unlocking 465 and a first entry of an intraprediction module 460. A second output of the unlocking filter 465 is connected in signal communication with a first input of a reference image store 480. One output of the reference image store 480 is connected in signal communication with a second input of a reference compensator. movement 470.
A second output of the entropy decoder 445 is connected in signal communication with a third input of the motion compensator 470 and a first input of the unlock filter 465. A third output of the entropy decoder 445 is connected in signal communication with an input of a decoder controller 405. A first output of the decoder controller 405 is connected in signal communication with a second input of the entropy decoder 445. A second output of the decoder controller 405 is connected in signal communication with a second input of the inversion transformer and inversion quantizer 450. A third output of the decoder controller 405 is connected in signal communication with a third input of the unlock filter 465. A fourth output of the decoder controller 405 is connected in signal communication with a second input of the intraprediction module 460, with a first input of the motion compensator 470, and with a second input of the reference image store 480.
A motion compensator output 470 is connected in signal communication to a first input of a switch 497. An output of the intrapredictive module 460 is connected in signal communication to a second input of switch 497. An output of switch 497 is connected in signal communication with a first non-inverting input of combiner 425.
An input of input store 410 is available as an input of decoder 400, to receive an input bit stream. A first output from the deblocking filter 465 is available as an output from the decoder 400, to output an image.
Turning to Figure 5, an exemplary method for encoding images for a plurality of views using the MPEG-4 AVC Standard is generally indicated by reference numeral 500. Method 500 includes an initial block 502 that passes control to a block function 504. Function block 504 arranges each view in a specific time instance as a tile-like subimage, and passes control to a function block 506. Function block 506 establishes a num_codec_viéws_minus1 syntax element, and passes control to a function block 508. Function block 508 establishes the org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1 syntax elements, and passes control to a 510 function block. function 510 sets a variable i equal to zero, and passes control to decision block 512. Decision block 512 determines whether or not variable i is less than the number of views. If so, then the control is passed to function block 514. Otherwise, control is passed to function block 524.
Function block 514 establishes a syntax element viewjdp], and passes control to a function block 516. Function block 516 establishes a syntax element num_parts [view_id [i], and passes control to a function block 518. Function block 518 sets a variable j equal to zero, and passes control to decision block 520. Decision block 520 determines whether or not the current value of variable j is less than the current value of the num_parts syntax element [view_id [i]. If so, then the control is passed to a 522 function block. otherwise, control is passed to function block 528.
Function block 522 establishes the following syntax elements, increments variable j, and then returns control to decision block 520:
depth_flag [view_id [i]] [j]; flip_dir [view_id [i]] [j]; Ioc_left_offset [view_id [i]] [j];
loc_top_offset [view_id [i]] [j]; frame_crop_left_offset [view_id [i]] [j];
frame_crop_rig ht_offset [view_id [ij] 0]; frame_crop_top_offset [view_id [i]] [j]; and frame_crop_bottom_offset [view_id [i]] de function block 528 establishes a syntax element upsample_view_flag [view_id [i]], and passes control to a decision block 530. Decision block 530 determines whether the current value of the upsample_view_flag [view_id [i]] syntax is equal to one or not. If so, then control is passed to function block 532. Otherwise, control is passed to decision block 534.
Function block 532 establishes an upsample_filter [view_id [i]] syntax element, and passes control to decision block 534.
Decision block 534 determines whether the current value of the upsample_filter [view_id [i]] syntax element is equal to three. If so, then control is passed to function block 536. Otherwise, control is passed to function block 540.
Function block 536 establishes the following syntax elements and passes control to a function block 538: vert_dim [view_id [i]]; hor_dim [view_id [i]]; and quantizer [view_id [i]].
Function block 538 sets the filter coefficients for each YUV component, and passes control to function block 540.
Function block 540 increments variable i, and returns control to decision block 512.
Function block 524 writes these syntax elements to at least one of: Sequence Parameter Set (SPS), Image Parameter Set (PPS), Supplemental Optimization Information message (SEI), Layer unit header Network Abstraction (NAL), and section header, and passes control to function block 526. Function block 526 encodes each image using the MPEG-4 AVC Standard or other single view codec, and passes control to a final block 599.
Turning to Figure 6, an exemplary method for decoding images for a plurality of views using the MPEG-4 AVC Standard is generally indicated by the reference numeral 600.
Method 600 includes an initial block 602 that passes control to a function block 604. Function block 604 analyzes the following syntax elements from at least one of the Sequence Parameter Set (SPS), Sequence Parameter Set Image (PPS), Supplementary Optimization Information (SEI) message, Network Abstraction Layer (NAL) unit header, and section header, and passes control to a 606 function block. Function block 606 parses a num_codec_views_minus1 syntax element, and passes control to a function block 608. Function block 608 parses the org_pic_width_in_mbs_minus1 and org_pic_heigth_in_mbs_minus1 syntax elements and passes control to a 610 function block. function 610 sets a variable i equal to zero, and passes control to decision block 612. Decision block 612 determines whether or not variable i is less than the number of views. If so, then control is passed to function block 614. Otherwise, control is passed to function block 624.
Function block 614 parses a view_id [i] syntax element, and passes control to a function block 616. Function block 616 parses a num_parts_minus1 [view_id [i]] syntax element, and passes control to a function block 618. Function block 618 sets a variable j equal to zero, and passes control to decision block 620. Decision block 629 determines whether or not the current value of variable j is less than the value current of the num_parts [view_id [i]] syntax element. If so, then control is passed to function block 622. Otherwise, control is passed to function block 628.
Function block 622 analyzes the following syntax elements, increments variable j, and then returns control to decision block 620:
depth_flag [view_id [i]] | j]; flip_dir [view_id [i]] [j]; loc_left_offset [view_id [i]] [j];
loc_top_offset [view__id [i]] [j]; frame_crop_left_offset [view_id [i]] [j]:
frame_crop_rig ht_offset [view_id [i]] [j]; frame_crop_top_offset [view_id [i]] [j]; and frame_crop_bottom_offset [view_id [i]] [j].
Function block 628 parses an upsample_view_flag [view_id [i]] syntax element, and passes control to a 630 decision block. Decision block 630 determines whether the current value of the upsample_view_flag syntax element [view_id [i] l is equal or not to one. If so, then the control is passed to function block 632. Otherwise, the control is passed to decision block 634.
Function block 632 parses an upsample_filter [view_id [i]] syntax element, and passes control to decision block 634.
Decision block 634 determines whether the current value of the upsample_filter [view_id [ij] syntax element is equal to three. If so, then control is passed to function block 636. Otherwise, control is passed to function block 640.
Function block 636 analyzes the following syntax elements and passes control to a function block 638: vert_dim [view_id [i]]; hor_dim [view_id [i]]; and quantizer [view_id [i] J.
Function block 638 analyzes the filter coefficients for each YUV component, and passes control to function block 640.
Function block 640 increments variable i, and returns control to decision block 612.
Function block 624 decodes each image using the MPEG-4 AVC standard or another single view codec, and passes control to a function block 626. Function block 626 separates each view from the image using the high syntax level, and passes control to a final 699 block.
Turning to Figure 7, an exemplary method for encoding images for a plurality of views and depths using the MPEG-4 AVC Standard is generally indicated by the reference numeral 700.
Method 700 includes an initial block 702 that passes control to a function block 704. Function block 704 arranges each corresponding view and depth in a specific time instance as a tile-shaped subimage, and passes control to a block function block 706. Function block 706 establishes a num_coded_views_minus1 syntax element, and passes control to a 708 function block. Function block 708 establishes the syntax elements org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1, and passes control to function block 710. Function block 710 establishes a variable i equal to zero, and passes control to a decision block 712. O decision block 712 determines whether or not variable i is less than the number of views. If so, then control is passed to function block 714. Otherwise, control is passed to function block 724.
Function block 714 establishes a syntax element view_id [i], and passes control to a function block 716. Function block 716 establishes a syntax element num_parts [view_id [i]], and passes control to a function block 718. Function block 718 sets a variable j equal to zero, and passes control to a decision block 720. Decision block 720 determines whether or not the current value of variable j is less than the value current of the num_parts [view_id [i]] syntax element. If so, then control is passed to function block 722. Otherwise, control is passed to function block 728.
Function block 722 establishes the following syntax elements, increments variable j, and then returns control to decision block 720:
depth_flag [view_id [i] j [j]; flip_dir [view_id [i]] [j]; loc_left_offset [view_id [i]] D];
loc_top_offset [view_id [i]] [j]; frame_crop_left_offset [view_id [i]] [j];
frame_crop_right_offset [view_id [i]] [j]; frame_crop_top_offset [view_id [i]] [j]; and frame_crop_bottom_offset [view_id [i]] [j].
Function block 728 establishes an upsample_view_flag [view_id [i] j syntax element, and passes control to a decision block 730. Decision block 730 determines whether the current value of the upsample_view_flag [view_id [ij] syntax element is or not equal to one. If so, then control is passed to function block 732. Otherwise, control is passed to decision block 734.
Function block 732 establishes an upsample_filter [view_id [i]] syntax element, and passes control to decision block 734.
Decision block 734 determines whether the current value of the upsample_filter [view_id [i]] syntax element is three or not. If so, then control is passed to function block 736. Otherwise, control is passed to function block 740.
Function block 736 establishes the following syntax elements and passes control to function block 738: vert_dim [view_id [i]]; hor_dim [view_id [i]]; and quantizer [view_id [i]].
Function block 738 establishes the filter coefficients for each YUV component, and passes control to function block 740.
Function block 740 increments variable i, and returns control to decision block 712.
Function block 724 records these syntax elements for at least one of: Sequence Parameter Set (SPS), Image Parameter Set (PPS), Supplemental Optimization Information (SEI) message, Header Layer unit Network Abstraction (NAL), and section header, and passes control to function block 726. Function block 726 encodes each image using the MPEG-4 AVC Standard or other single view codec, and passes control to a final block 799.
Turning to Figure 8, an exemplary method for decoding images for a plurality of views and depths using the MPEG-4 AVC Standard is generally indicated by reference numeral 800.
Method 800 includes a starting block 802 that passes control to a function block 804. Function block 804 analyzes the following syntax elements from at least one of the Sequence Parameter Set (SPS), Parameter Set of Image (PPS), Supplemental Optimization Information (SEI) message, Network Abstraction Layer (NAL) unit header, and section header, and passes control to an 806 function block. Function block 806 parses a num_codec_views_minus1 syntax element and passes control to a function block 808. Function block 808 parses the org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1 syntax elements, and passes control to a function block 810. function 810 sets a variable i equal to zero, and passes control to a decision block 812. Decision block 812 determines whether or not variable i is less than the number of views. If so, then control is passed to function block 814. Otherwise, control is passed to function block 824.
Function block 814 parses a view_id [i] syntax element, and passes the counter17 le to a function block 816. Function block 816 parses a num_parts_minus1 [view_id [i]] syntax element, and passes control to a function block 818. Function block 818 sets a variable j equal to zero, and passes control to a decision block 820. Decision block 820 determines whether or not the current value of variable j is less than the current value of the num_parts [view_id [i]] syntax element. If so, then control is passed to function block 822. Otherwise, control is passed to function block 828.
Function block 822 analyzes the following syntax elements, increments variable j, and then returns control to decision block 820:
depth_flag [view_id [i]] [j]: flip_dir [view_id [i]] [j]; loc_left_offset [view_id [i]] [j];
loc_top_offset [view_id [i]] [j]; frame_crop_left_offset [view_id [i]] [j];
frame_crop_rig ht_offset [view_id [i]] [j]; fram e_crop_top_offset [view_id [i]] [j]; and frame_crop_bottom_offset [view_id [i]] [j].
Function block 828 parses an upsample_view_flag [view_id [i]] syntax element, and passes control to an 830 decision block. Decision block 830 determines whether the current value of the upsample_view_flag syntax element [view_id [i]] is equal or not to one. If so, then control is passed to function block 832. Otherwise, control is passed to decision block 834.
Function block 832 parses an upsample_filter [view_id [i]] syntax element, and passes control to decision block 834.
Decision block 834 determines whether the current value of the upsample_filter [view_id [i]] syntax element is three or not. If so, then control is passed to function block 836. Otherwise, control is passed to function block 840.
Function block 836 analyzes the following syntax elements and passes control to a function block 838: vert_dim [view_id [i]]; hor_dim [view_id [i]]; and quantizer [view_id [i]].
Function block 838 analyzes the filter coefficients for each YUV component, and passes control to function block 840.
Function block 840 increments variable i, and returns control to decision block 812.
Function block 824 decodes each image using the MPEG-4 AVC Standard or other single view codec, and passes control to function block 826. Function block 826 separates each view and the corresponding depth from the image using the high level syntax, and passes control to a function block 827. Function block 827 potentially performs the synthesis of views using the extracted view and depth signals, and passes control to a final block 899.
With respect to the depth used in Figures 7 and 8, Figure 9 shows an example of a depth signal 900, where the depth is provided as a pixel value for each corresponding location in an image (not shown). Additionally, Figure 10 shows an example of two depth signals included in a 1000 tile. The upper right portion of tile 1000 is a depth signal having depth values corresponding to the image on the upper left portion of tile 1000. The lower right portion of tile 100 is a depth signal having depth values corresponding to the image on the lower left portion. of tile 1000.
Turning to Figure 11, an example of the five tiled views in a single frame is usually indicated by the reference numeral 1100. The top four views are in a normal orientation. The fifth view is also in a normal orientation, but is divided into two portions along the bottom of tile 1100. A left portion of the fifth view shows the “top” of the fifth view, and a right portion of the fifth view shows the “ bottom ”of the fifth view.
Encoding / decoding using a multiview video encoding / decoding standard / recommendation
Turning to Figure 12, an exemplary Multivista Video Encoding (MVC) encoder is generally indicated by reference numeral 1200. Encoder 1200 includes a combiner 1205 having an output connected in signal communication with an input from a transformer 1210. An output of transformer 1210 is connected in signal communication with an input of quantizer 1215. An output of the quantizer 1215 is connected in signal communication with an input of an entropy encoder 1220 and an input of a reverse quantizer 1225. An output of the reverse quantizer 1225 is connected in signal communication with an input of an inversion transformer 1230 An output of the 1230 reversing transformer is connected in signal communication to a first non-reversing input of a 1235 combiner. An output of combiner 1235 is connected in signal communication with an input of an intra predictor 1245 and an input of an unlock filter 1250. An output of the unlock filter 1250 is connected in signal communication with an input of a storage medium reference image 1255 (for view i). An output of the reference image storage medium 1255 is connected in signal communication with a first input of a motion compensator 1275 and a first input of a motion estimator 1280. An output of the motion estimator 1280 is connected in communication of signal with a second entry of the 1275 motion compensator.
An output of a reference image storage medium 1260 (for other views) is connected in signal communication with a first input of a disparity estimator 1270 and a first input of a disparity compensator 1265. An output of the disparity estimator 1270 is connected in signal communication with a second input of disparity compensator 1265.
An entropy decoder output 1220 is available as an output of encoder 1200. A non-inverting input of combiner 1205 is available as an input of encoder 1200, and is connected in signal communication with a second input of disparity estimator 1270 , and a second input of the motion estimator 1280. An output of a switch 1285 is connected in signal communication with a second non-inverting input of combiner 1235 and with an inverting input of combiner 1205. Switch 1285 includes a first input connected in signal communication with an output of motion compensator 1275, a second input connected in signal communication with an output of disparity compensator 1265, and a third input connected in signal communication with an output of the intra predictor 1245.
A 1240 mode decision module has an output connected to switch 1285 to control which input is selected by switch 1285.
Turning to Figure 13, an exemplary Multivista Video Encoding (MVC) decoder is generally indicated by reference number 1300. Decoder 1300 includes an entropy decoder 1305 having an output connected in signal communication with an input of an inversion quantizer 1310. An inversion quantizer output is connected in signal communication with an input of a reverse transformer 1315. An output of the reverse transformer 1315 is connected in signal communication with a first non-inverting input of a combiner 1320. An output of the combiner 1320 is connected in signal communication with an input of an unlocking filter 1325 and an input of a intra predictor 1330. An output of the unlocking filter 1325 is connected in signal communication with an input of a reference image storage medium 1340 (for view i). An output of the reference image storage medium 1340 is connected in signal communication with a first input of a 1335 motion compensator.
An output of a 1345 reference image storage medium (for other views) is connected in signal communication with a first input of a 1350 disparity compensator.
An entropy encoder 1305 input is available as an input to decoder 1300 to receive a waste bit stream. In addition, an input from a 1360 mode module is also available as an input to the 1300 decoder, to receive control syntax to control which input is selected by the 1355 switch. In addition, a second input of the motion compensator 1355 is available as an input of the decoder 1300, to receive motion vectors. In addition, a second input from the 1350 disparity compensator is available as an input to the 1300 decoder, to receive disparity vectors.
An output of a switch 1355 is connected in signal communication with a second non-inverting input of the combiner 1320. A first input of the switch 1355 is connected in signal communication with an output of the disparity compensator 1350. A second input of the switch 1355 is connected in signal communication with a 1335 motion compensator output. A third input of switch 1355 is connected in signal communication with an output of the intra predictor 1330. An output of the 1360 mode module is connected in signal communication with the switch 1355 to control which input is selected by the switch 1355. An output of the unlocking filter 1325 is available as an output of the decoder 1300.
Turning to Figure 14, an exemplary method for processing images for a plurality of views in preparation for encoding the images using the MPEG-4 AVC Standard multiview video encoding (MVC) extension is generally indicated by the reference numeral 1400.
Method 1400 includes an initial block 1405 that passes control to a function block 1410. Function block 1410 arranges each of the n views, out of a total of M views, in a specific time instance as a tile super image , and passes control to function block 1415. Function block 1415 establishes a num_coded_views_minus1 syntax element, and passes control to function block 1420. Function block 1420 establishes a view_id [i] syntax element for all views (num_coded_views_minus1 + 1), and passes control to function block 1425. Function block 1425 establishes reference dependency information between views for anchor images, and passes control to a 1430 function block. Function block 1430 establishes reference dependency information between views for non-anchor images, and passes control to a 1435 function block. Function block 1435 establishes a syntax element pseudo_view_present_fl39, and passes control to a decision block 1440. Decision block 1440 determines whether or not the current value of the syntax element pseudo_view_present_flag is true. If so, then the control is passed to a function block 1445. Otherwise, the control is passed to an end block 1499.
Function block 1445 establishes the following syntax elements, and passes control to a function block 1450: tiling_mode; org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus1. Function block 1450 requires a pseudo_view_info (view_id) syntax element for each coded view, and passes control to the final block 1499.
Turning to Figure 15, an exemplary method for encoding images for a plurality of views using the MPEG-4 AVC Standard multiview video encoding (MVC) extension is generally indicated by the reference numeral 1500.
Method 1500 includes an initial block 1502 that has an input parameter pseudo_view_id and passes control to a function block 1504. Function block 1504 establishes a num_sub_views_minus1 syntax element, and passes control to a function block 1506. O function block 1506 sets a variable i equal to zero, and passes control to a decision block 1508. Decision block 1508 determines whether or not variable i is less than the number of subviews. If so, then control is passed to function block 1510. Otherwise, control is passed to function block 1520.
Function block 1520 establishes a sub_view_id [i] syntax element, and passes control to function block 1512. Function block 1512 establishes a num_parts_minus1 [sub_view_ [id]] syntax element, and passes control to a function block 1514. Function block 1514 sets a variable j equal to zero, and passes control to a decision block 1516. Decision block 1516 determines whether or not variable j is less than the num_parts_minus1 syntax element [sub_view_ [id]]. If so, then control is passed to function block 1518. Otherwise, control is passed to decision block 1522.
Function block 1518 establishes the following syntax elements, increments variable j, and returns control to decision block 1516:
loçjeft_offset [sub_view_id [i]] [j]; loc_top_offset [sub_view_id [i]] [j];
frame_crop_left_offset [sub_view_id [i]] [j]; frame_crop_right_offset [sub_view_id [i]] [j];
frame_crop_top_offset [sub_view_id [i]] [j]; and frame_crop_bottom_offset [sub_view_id [i] [j].
Function block 1520 encodes the current image into the current view using multiview video encoding (MVC), and passes control to a final 1599 block.
Decision block 1522 determines whether or not an element of syntax tiling_mode is equal to zero. If so, then control is passed to function block 1524. Otherwise, control is passed to function block 1538.
Function block 1524 establishes a flip_dir [sub_view_id [i]] syntax element and an upsample_view_flag [sub_view_id [i]] syntax element, and passes control to a 1526 decision block. Decision block 1526 determines whether the value syntax element's current value upsample_view_flag [sub_view_id [i]] is equal to one or not. If so, then control is passed to function block 1528. Otherwise, control is passed to decision block 1530.
Function block 1528 establishes an upsample_flag [sub_view_id [i]] syntax element, and passes control to decision block 1530. Decision block 1530 determines whether a value of the upsample_fl39 [sub_view_id [i]] syntax element is or not equal to three. If so, control is passed to function block 1532. Otherwise, control is passed to function block 1536.
Function block 1532 establishes the following syntax elements, and passes control to a function block 1534: vert_dim [sub_view_id [i]]; hor_dim [sub_view_id [i]]; and quantizer [sub_view_id [i]]. Function block 1534 sets the filter coefficients for each YUV component, and passes control to function block 1536.
Function block 1536 increments variable i, and returns control to decision block 1508. Function block 1538 establishes a syntax element pixel_dist_x [sub_view_id [i]] and the syntax element flip_dist_y [sub_view_id [i]] , and passes control to function block 1540. Function block 1540 sets variable j equal to zero, and passes control to decision block 1542. Decision block 1542 determines whether or not the current value of variable j is less than the current value of the syntax element num_parts [sub_view_id [i]]. If so, then control is passed to function block 1544. Otherwise, control is passed to function block 1536.
Function block 1544 establishes a num_pixel_tiling_filter_coeffs_minus1 syntax element [sub_view_id [i]], and passes control to function block 1546. Function block 1546 establishes coefficients for all pixel tile filters, and passes control for function block 1536.
Turning to Figure 16, an exemplary method for processing images for a plurality of views in preparation for decoding the images using the MPEG-4 AVC Standard multiview video encoding (MVC) extension is generally indicated by the reference numeral 1600.
Method 1600 includes an initial block 1605 that passes control to a 1615 function block. The 1615 function block parses a num_coded_views_minus1 syntax element, and passes the control to a 1620 function block. The 1620 function block parses an element syntax view_id [i] for all views (num_coded_views_minus1 + 1), and passes control to a 1625 function block. Function block 1625 analyzes reference dependency information between views for anchor images, and passes control to a 1630 function block. Function block 1630 analyzes reference dependency information between views for non-anchor images, and passes control to a 1635 function block. Function 1635 parses a syntax element pseudo_view_present_flag, and passes control to a 1640 decision block. Decision block 1640 determines whether or not the current value of the pseudo_view_present_flag syntax element is true. If so, then the control is passed to a 1645 function block. Otherwise, the control is passed to a 1699 end block.
Function block 1645 analyzes the following syntax elements, and passes the control to a function block 1650: tiling_mode: org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus1. Function block 1650 requires a pseudo_view_info (view_id) syntax element for each coded view, and passes control to the final 1699 block.
Turning now to Figure 17, an exemplary method for decoding an image for a plurality of views using the MPEG-4 AVC Standard multiview video encoding (MVC) extension is generally indicated by the reference numeral 1700.
Method 1700 includes a starting block 1702 that starts with the input parameter pseudo_view_id and passes control to a 1704 function block. Function block 1704 parses a num_sub_views_minus1 syntax element, and passes control to a 1706 function block. Function block 1706 sets a variable i equal to zero, and passes control to a decision block 1708. Decision block 1708 determines whether or not variable i is less than the number of subviews. If so, then the control is passed to a 1710 function block. Otherwise, the control is passed to a 1720 function block.
Function block 1710 parses a sub_view_id [i] syntax element, and passes control to a 1712 function block. Function block 1712 parses a num_parts_minus1 [sub_view_id [i]] syntax element, and passes control to a function block 1714. Function block 1714 sets a variable j equal to zero, and passes control to a decision block 1716. Decision block 1716 determines whether or not the variable j is less than the num_parts_minus1 syntax element [sub_view_id [i]]. If so, then the control is passed to a 1718 function block. Otherwise, the control is passed to a 1722 decision block.
Function block 1718 establishes the following syntax elements, increments variable j, and returns control to decision block 1716:
loc_left_offset [sub_view_id [i]] 0]; loc_top_offset [su b_view_id [i]] [j];
frame_crop_left_offset [sub_view_id [i]] 0]; frame_crop_right_offset [sub_view_id [i]] j];
frame_crop_top_offset [sub_view_id [i]] | j]; and frame_crop_bottom_offset [sub_view_id [i] [j] ·
Function block 1720 decodes the current image into the current view using multiview video encoding (MVC), and passes control to a function block 1721. Function block 1721 separates each view from the image using high level syntax , and passes control to a final 1799 block.
The separation of each view from the decoded image is done using the high level syntax indicated in the bit stream. This high-level syntax can indicate the exact location and possible orientation of the views (and possible corresponding depth) present in the image.
Decision block 1722 determines whether a tiling_mode syntax element is equal to zero. If so, then the control is passed to a 1724 function block. Otherwise, the control is passed to a 1738 function block.
Function block 1724 parses an element of syntax flip_dir [sub_view_id [i]]; and an upsample_view_flag [sub_view_id [i]] syntax element, and passes control to a 1726 decision block. The 1726 decision block determines whether the current value of the upsample_view_flag [sub_view_id [i]] syntax element is equal to or not one. If so, then the control is passed to a 1728 function block. Otherwise, the control is passed to a 1730 decision block.
Function block 1728 parses an upsample_filter [sub_view_id [í]] syntax element, and passes control to decision block 1730. Decision block 1730 determines whether the value of the upsample_filter [sub_view_id [i]] syntax element is or not equal to three. If so, the control is passed to a 1732 function block. Otherwise, the control is passed to a 1736 function block.
Function block 1732 analyzes the following syntax elements, and passes control to a function block 1734: vert_dim [sub_view_id [i]]; hor_dim [sub_view_id [i]]; and quantizer [sub_view_id [i]]. Function block 1734 analyzes the filter coefficients for each YUV component and passes control to a 1736 function block.
Function block 1736 increments variable i, and returns control to decision block 1708.
Function block 1738 parses a pixel_dist_x [sub_view_id [i]] syntax element and flip_dist_y [sub_view_id [i]] syntax element, and passes control to a 1740 function block. Function block 1740 sets variable i equals zero, and passes control to a 1742 function block. Function block 1742 determines whether or not the current value of variable j is less than the current value of the num_parts [sub_view_id [i]] syntax element. If so, then the control is passed to a 1744 function block. Otherwise, the control is passed to a 1736 function block.
Function block 1744 parses a num_pixel_tiling_filter_coeffs_minus1 syntax element [sub_view_id [i]], and passes control to function block 1746. Function block 1776 analyzes coefficients for all pixel tile filters, and passes control for function block 1736.
Turning now to Figure 18, an exemplary method for processing images for a plurality of views and depths in preparation for encoding the images using the MPEG-4 AVC Standard multiview video encoding (MVC) extension is generally indicated by reference numeral 1800.
Method 1800 includes an initial block 1805 that passes control to a function block 1810. Function block 1810 arranges all N views and depth maps, out of a total of M views and depth maps, in a temporal instance and specific as a super image in the tile format, and passes control to an 1815 function block. Function block 1815 establishes a num_coded_views_minus1 syntax element, and passes control to an 1820 function block. Function block 1820 establishes a syntax element view_id [i] for all depths (num_coded_views_minus1 +1) corresponding to view_id [i], and passes control to function block 1825. Function block 1825 establishes information for reference dependency between views for anchor depth images, and passes control to an 1830 function block. Function block 1830 establishes reference dependency information between views for non-anchor depth images, and passes control to function block 1835. Function block 1835 establishes a pseudo_view_present_flag syntax element. θ passes control to an 1840 decision block. The 1840 decision block determines whether or not the current value of the pseudo_view_present_flag syntax element is true. If so, then the control is passed to an 1845 function block. Otherwise, the control is passed to an end 1899 block.
Function block 1845 establishes the following syntax elements, and passes control to function block 1850: tiling_mode; org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus1. Function block 1850 requires a pseudo_view_info (view_id) syntax element for each coded view, and passes control to the final block
1899.
Turning to Figure 19, an exemplary method for encoding images for a plurality of views and depths using the MPEG-4 AVC Standard multiview video encoding (MVC) extension is generally indicated by the reference numeral
1900.
Method 1900 includes a starting block 1902 that passes control to a 1904 function block. Function block 1904 establishes a num_sub_views_minus1 syntax element, and passes control to a 1906 function block. Function block 1906 sets a variable i equals zero, and passes control to a 1908 decision block. The 1908 decision block determines whether or not variable i is less than the number of subviews. If so, then the control is passed to a 1910 function block. Otherwise, the control is passed to a 1920 function block.
Function block 1910 establishes a sub_view_id [i] syntax element, and passes control to a 1912 function block. Function block 1912 establishes a num_parts_minus1 [sub_view_id [i]] syntax element, and passes control to a function block 1914. Function block 1914 sets a variable j equal to zero, and passes control to a decision block 1916. Decision block 1916 determines whether or not the variable j is less than the num_parts_minus1 [sub_view_id [i]] syntax element. If so, then the control is passed to a 1918 function block. Otherwise, the control is passed to a 1922 function block.
Function block 1918 establishes the following syntax elements, increments variable j, and returns control to decision block 1916:
loc_left_offset [sub_view_id [i]] [j]; loc_top_offset [sub_view_id [i]] [j];
frame_crop_left_offset [sub_view_id [i]] [j]; frame_crop_right_offset [sub_view_id [i]] D];
frame_crop_top_offset [sub_view_id [i]] [j]; θ frame_crop_bottom_offset [sub_view_id [i] [j].
Function block 1920 encodes the current depth of the current view using multiview video encoding (MVC), and passes control to a final 1999 block. The depth signal can be encoded in a manner similar to the way your video signal corresponding code is encoded. For example, the depth signal for a view can be included in a tile that includes only other depth signals, or just video signals, or both depth and video signals. The tile (pseudo-view) is then treated as a single view for MVC, and there are also presumably other tiles that are treated like other views for MVC.
Decision block 1922 determines whether or not a tiling_mode syntax element is equal to zero. If so, then the control is passed to a 1924 function block. Otherwise, the control is passed to a 1938 function block.
Function block 1924 establishes a flip_dir [sub_view_id [i]] syntax element and an upsample_view_flag [sub_view_id [i]] syntax element, and passes control to a 1926 decision block. The 1926 decision block determines whether the value syntax element's current value upsample_view_flag [sub_view_id [i]] is equal to one or not. If so, then the control is passed to a 1928 function block. Otherwise, the control is passed to a 1930 decision block.
Function block 1928 establishes an upsample_filter [sub_view_id [i]] syntax element, and passes control to decision block 1930. Decision block 1930 determines whether a value of the upsample_filter [sub_viewjd [i]] syntax element is or not equal to three. If so, the control is passed to a 1932 function block. Otherwise, the control is passed to a 1936 function block.
Function block 1932 establishes the following syntax elements, and passes control to a function block 1934: vert_dim [sub_view_id [i]]; hor_dim [sub_view_id [i]]; and quantizer [sub_view_id [i]]. Function block 1934 establishes the filter coefficients for each YUV component, and passes control to function block 1936.
Function block 1936 increments variable i, and returns control to decision block 1908.
Function block 1938 establishes a pixel_dist_x [sub_view_id [i]] syntax element and flip_dist_y [sub_view_id [i]] syntax element, and passes control to a 1940 function block. Function block 1940 establishes a variable j equal to zero, and passes control to a 1942 decision block. The 1942 decision block determines whether or not the current value of variable j is less than the current value of the num__parts syntax element [sub_view_id [i]]. If so, then control is passed to function block 1944. Otherwise, control is passed to function block 1936.
Function block 1944 establishes a num_pixel_tiling_filter_coeffs_minus1 syntax element [sub_view_id [i]], and passes control to a 1946 function block. Function block 1946 establishes the coefficients for all tiling and pixel filters, and passes control for function block 1936.
Turning to Figure 20, an exemplary method for processing images for a plurality of views and depths in preparation for decoding the images using the MPEG-4 AVC Standard multiview video encoding (MVC) extension is generally indicated by the numeral reference period 2000.
Method 2000 includes a starting block 2005 that passes control to a 2015 function block. Function 2015 parses a num_coded_views_minus1 syntax element, and passes control to a 2020 function block. Function 2020 parses an element syntax view_id [i] for all depths (num_coded_views_minus1 +1) corresponding to view_idp], and passes control to a 2025 function block. Function block 2025 analyzes reference dependency information between views for anchor depth images, and passes control to a 2030 function block. Function block 2030 analyzes reference dependency information between views for depth images non-anchor, and passes control to a 2035 function block. The 2035 function block parses a pseudo_view_present_flag syntax element, and passes the control to a 2040 decision block. Decision block 2040 determines whether or not the current value of the syntax element pseudo_view__present_flag is equal to true. If so, then the control is passed to a 2045 function block. Otherwise, the control is passed to a final 2099 block.
Function block 2045 analyzes the following syntax elements, and passes control to a function block 2050 tiling_mode; org_pic_width_in_mbs_minus1; and org_pic__h®ight_in_mbs_minus1. Function block 2050 requires a pseudo_view_info (view_id) syntax element for each coded view, and passes control to the final block
2099.
Turning to Figure 21, an exemplary method for decoding an image for a plurality of views and depths using the MPEG-4 AVC Standard multiview video encoding (MVC) extension is usually indicated by the reference numeral
2100.
Method 2100 includes an initial block 2102 that starts with the input parameter pseudo_view_id, and passes control to a 2104 function block. The function block
2104 parses a num_sub_views_minus1 syntax element, and passes control to function block 2106. Function block 2106 sets a variable i equal to zero, and passes control to decision block 2108. Decision block 2108 determines whether the variable i is less or less than the number of subviews. If so, then the control is passed to a 2110 function block. Otherwise, the control is passed to a 2120 function block.
Function block 2110 parses a sub_view_id [i] syntax element, and passes control to a 2112 function block. Function block 2112 parses a num_parts_minus1 [sub_view_id [i]] syntax element, and passes control to a function block 2114. Function block 2114 sets a variable j equal to zero, and passes control to a decision block 2116. Decision block 2116 determines whether or not variable j is less than the num_parts_minus1 [sub_view_id [i]] syntax element. If so, then control is passed to function block 2118. Otherwise, control is passed to decision block 2122.
Function block 2118 establishes the following syntax elements, increments variable j, and returns control to decision block 2116:
loc_left_offset [su b_view_id [i]] [j]; loc_top_offset [su b_view_id [i]] [j];
frame_crop_left_offset [sub_view_id [i]] [j]; frame_crop_right_offset [sub_view_id [i]] [j];
frame_crop_top_offset [sub_view_id [i]] [j]; and frame_crop_bottom_offset [sub_view_id [i] D].
Function block 2120 decodes the current image using multiview video encoding (MVC), and passes control to a function block 2121. Function block 2121 separates each view of the image using high level syntax, and passes the control for a final block 2199. The separation of each view using high level syntax is as previously described.
Decision block 2122 determines whether a tiling_mode syntax element is equal to zero. If so, then control is passed to function block 2124. Otherwise, control is passed to function block 2138.
Function block 2124 parses a flip_dir [sub_view_id [i]] syntax element and an upsample_view_flag [sub_view_id [i]] syntax element, and passes control to a 2126 decision block. Decision block 2126 determines whether the value syntax element's current value upsample_view_f! ag [sub_view_id [i]] is equal to one or not. If so, then the control is passed to a 2128 function block. Otherwise, the control is passed to a 2130 decision block.
Function block 2128 parses an upsample_filter [sub_view_id [i]] syntax element, and passes control to decision block 2130. Decision block 2130 determines whether a value of the upsample_filter [sub_view_id [i]] syntax element is or not equal to three. If so, the control is passed to a 2132 function block. Otherwise, the control is passed to a 2136 function block.
Function block 2132 analyzes the following syntax elements, and passes control to a function block 2134: vert_dim [sub_view_id [i]]; hor_dim [sub_view_id [i]]; and quantizer [sub_view_id [i]]. Function block 2134 analyzes the filter coefficients for each YUV component, and passes control to function block 2136.
Function block 2136 increments variable i, and returns control to decision block 2108.
Function block 2138 parses a pixel_dist_x [sub_view_id [i]] syntax element and the flip_dist_y [sub_view_id [i]] syntax element, and passes control to a 2140 function block. Function block 2140 sets the variable j equal to zero, and passes control to decision block 2142. Decision block 2142 determines whether or not the current value of variable j is less than the current value of the syntax element num_parts [sub_view_id [i]]. If so, then control is passed to function block 2144. Otherwise, control is passed to function block 2136.
Function block 2144 parses a num_pixel_tiling_filter_coeffs_minus1 syntax element [sub_view_id [i]], and passes control to function block 2146. Function block 2146 analyzes coefficients for all pixel tile filters, and passes control for function block 2136.
Turning to Figure 22, examples of pixel-level tiling are generally indicated by reference numeral 2200. Figure 22 is further described below.
VIEW TILING USING MPEG-4 AVC OR MVC
A multiview video encoding application is free-standing TV (or FTV). This application requires the user to be able to move freely between two or more views. To do this, the “virtual” views between two views need to be interpolated or synthesized. There are several methods for performing view interpolation. One method uses depth for view interpolation / synthesis.
Each view can have an associated depth signal. Thus, the depth can be considered as another form of video signal. Figure 9 shows an example of a depth signal 900. To enable applications such as FTV, the depth signal is transmitted together with the video signal. In the proposed tiling structure, the depth signal can also be added as one of the tiles. Figure 10 shows an example of depth signs added as tiles. The depth signs / tiles are shown on the right side of Figure 10.
When the depth is encoded as a tile of the entire frame, the high-level syntax should indicate which tile is the depth signal so that the renderer can properly use the depth signal.
In the case when the input sequence (such as that shown in Figure 1) is encoded using an MPEG-4 AVC Standard encoder (or an encoder corresponding to a different video encoding standard and / or recommendation), the high syntax The proposed level may be present, for example, in the Sequence Parameter Set (SPS), in the Image Parameter Set (PPS), in a section header, and / or in a Supplemental Optimization Information (SEI) message. A modality of the proposed method is shown in Table 1 where the syntax is present in a Supplemental Optimization Information (SEI) message.
In the case when the input streams of pseudovistas (such as those shown in Figure 1) are encoded using the multiview video encoding extension (MVC) of the MPEG-4 AVC Standard encoder (or an encoder corresponding to the encoding standard of multiview video against a different video encoding standard and / or recommendation), the proposed high-level syntax may be present in SPS, PPS, section header, SEI message, or in a specified profile. A modality of the proposed method is shown in Table 1. Table 1 shows elements of syntax present in the Set of Sequence Parameters (SPS) structure, including elements of syntax proposed according to a modality of the present principles
TABLE 1
<td>seq parameter set mvc extension () {</td><td>ç</td><td>Descriptor</td>
<td>na views min us 1</td><td></td><td>eu (v)</td>
<td>for (i = 0; i <= num views minus 1; i ++)</td><td></td><td></td>
<td>view id [i]</td><td></td><td>eu (v)</td>
<td>for (i = 0; i <= num views minus 1; i ++) {</td><td></td><td></td>
<td>in an anchor refs IO [i]</td><td></td><td>eu (v)</td>
<td>for (j = 0; j <in an anchor refs IO [í]: j<sup>++</sup> )</td><td></td><td></td>
<td>anchor ref IO [i] [j]</td><td></td><td>eu (v)</td>
<td>anchor refs H [i]</td><td></td><td>eu (v)</td>
<td>for (j = 0; j <in an anchor refs H [i]; j ++)</td><td></td><td></td>
<td>anchor ref H [i] [j]</td><td></td><td>eu (v)</td>
<td> }</td><td></td><td></td>
<td>for (i = 0; i <= num views minus 1; i ++) {</td><td></td><td></td>
<td>num non anchor refs IO [i]</td><td></td><td>eu (v)</td>
<td>for (j = 0; j <num non anchor refs IO [i]; j ++)</td><td></td><td></td>
<td>η ona ncho r ref l 0 [i] [j]</td><td></td><td>eu (v)</td>
<td>num non anchor refs M [i]</td><td></td><td>eu (v)</td>
<td>for (j = 0; j <num non anchor refs H [i]; j ++)</td><td></td><td></td>
<td>non anchor refJ1 [i] [j]</td><td></td><td>eu (v)</td>
<td> }</td><td></td><td></td>
<td>pseudo view present flag</td><td></td><td>u (l)</td>
<td>if (pseudo view present flag) {</td><td></td><td></td>
<td>tiling mode</td><td></td><td></td>
<td>org pic width in mbs minus1</td><td></td><td></td>
<td>org pic height in mbs minus1</td><td></td><td></td>
<td>for (i = 0; i <num views minus 1; i ++)</td><td></td><td></td>
<td>pseudo view info (i);</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
Table 2 shows the syntax elements for the pseu do_view_info syntax element in TABLE 1, according to a modality of the present principles. TABLE 2
<td>pseudo view info (pseudo view id) {</td><td>Ç</td><td>Descriptor</td>
<td>num sub views minus 1 [pseudo view id]</td><td> 5</td><td>eu (v)</td>
<td>if (in a sub views minus 1! = 0) {</td><td></td><td></td>
<td>for (i = 0; i <in a sub views minus 1 [pseudo view id]; i ++) {</td><td></td><td></td>
<td>sub view id [i]</td><td> 5</td><td>eu (v)</td>
<td>num parts minus1 [sub view id [i]]</td><td> 5</td><td>eu (v)</td>
<td>for (j = 0; j <= num parts minus1 [sub view id [i]]; j ++) {</td><td></td><td></td>
<td>loc left offset [sub view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td>loc top offset [sub view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td>cropJeft offset frame [sub view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td>frame crop right offset [sub view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td>crop top offset frame [sub view id [í]] [j]</td><td> 5</td><td>eu (v)</td>
<td>frame crop bottom offset [sub view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td> }</td><td></td><td></td>
<td>if (tiling mode == 0) {</td><td></td><td></td>
<td>flip dir [sub view id [i] [j]</td><td> 5</td><td>u (2)</td>
<td>upsample view flag [sub viewjd [i]]</td><td> 5</td><td>u (l)</td>
<td>if (upsample view flag [sub view id [i]])</td><td></td><td></td>
<td>upsample filter [sub view id [i]]</td><td> 5</td><td>u (2)</td>
<td>if (upsample fiter [sub view id [i]] == 3) {</td><td></td><td></td>
<td>vert dim [sub view id [i]]</td><td> 5</td><td>eu (v)</td>
<td>hor di m [sub view id [i]]</td><td> 5</td><td>eu (v)</td>
<td>quantizer [sub view id [i]]</td><td> 5</td><td>eu (v)</td>
<td>for (yuv = 0; yuv <3; yuv ++) {</td><td></td><td></td>
<td>for (y = 0; y <vert dim [sub view id [i]] -1; y ++) {</td><td></td><td></td>
<td>for (x = 0; x <hor dim [sub view id [i]] -1; x ++)</td><td></td><td></td>
<td>fi Ite r coeffs [sub view id [i]] [yuv] [y] [x]</td><td> 5</td><td>if (v)</td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td>} // se (tiling mode == 0)</td><td></td><td></td>
<td>or if (tiling mode == 1) {</td><td></td><td></td>
<td>pixel dist x [sub view id [i]]</td><td></td><td></td>
<td>pixel dist y [sub view id [i]]</td><td></td><td></td>
<td>for (j = 0; j <= num parts [sub view id [i]]; j ++) {</td><td></td><td></td>
<td>in a pixel tiling filter coeffs minus1 [sub view id [i]] [j]</td><td></td><td></td>
<td>for (coeff_dx = 0; coeff_dx <=</td><td></td><td></td>
<td>pixel tiling filter coeffs [sub view id [i]] [j]</td><td></td><td></td>
<td>} // for (j = 0; j <= num parts [sub view id [i]]; j ++)</td><td></td><td></td>
<td>} // or if (tiling mode == 1)</td><td></td><td></td>
<td>} // for (i = 0; i <in a sub views minus 1; i ++)</td><td></td><td></td>
<td>} // if (in a sub views minus 1! = 0)</td><td></td><td></td>
<td> }</td><td></td><td></td>
Semantics of the syntax elements presented in TABLE 1 and TABLE 2 pseudo_view_present_flag equal to true indicates that a certain view is a superview of multiple subviews.
tiling_mode equal to zero indicates that subviews are tiled at the image level. A value of 1 indicates that the tiling is done at the pixel level.
The new SEI message could use a value for the SEI payload type that was not used in the MPEG-4 AVC Standard or an extension of the MPEG-4 AVC Standard. The new SEI message includes several syntax elements with the following semantics.
num_coded_views_minus1 plus 1 indicates the number of coded views supported by the bit stream. The value of num_coded_views_minus1 is in the range 0 to 1023, inclusive.
org_pic_width_in_mbs_minus1 plus 1 specifies the width of an image in each view in units of macroblocks.
The variable for image width in units of macroblocks is derived as follows:
PicWidthlnMbs = org_pic_width_in_mbs_minus1 + 1
The image width variable for the luma component is derived as follows:
PicWidthlnSamplesL = PicWidthlnMbs * 16
The image width variable for the chroma components is derived as follows:
PicWidthlnSamplesC = PicWidthlnMbs * MbWidthC org_pic_height_in_mbs_minus1 plus 1 specifies the height of an image in each view in units of macroblocks.
The variable for image height in units of macroblocks is derived as follows:
PicHeightlnMbs = org_pic_height_in_mbs_minus1 + 1
The image height variable for the luma component is derived as follows:
PicHeightlnSamplesL = PicHeightlnMbs * 16
The image height variable for the chroma components is derived as follows:
PicHeightlnSamplesC = PicHeightlnMbs * MbHeightC num_sub_views_minus1 plus 1 indicates the number of coded views included in the current view. The value of num_coded_views_minus1 is in the range 0 to 1023, inclusive.
sub_view_id [i] specifies the sub_view_id of the subview with the decoding order indicated by i.
num_parts [sub_view_id [i]] specifies the number of parts into which the image of sub_view_id [i] is divided.
Ioc_left_offset [sub_view_id [i]] [j] and loc_top_offset [sub_view_id [i]] [j] specify the locations in left and top pixel offsets, respectively, where the current part is already located in the final reconstructed image of the view with equal sub_view_id the sub_view_id [i].
view_id [i] specifies the view_id of the view with the encoding order indicated by i. frame_crop_left_offset [view_id [i]] [j], frame_crop_right_offset [view_id [i]] [j], frame_crop_top_offset [view_id [i]] j], ® frame_crop_bottom_offset [view_id [i]] [j] specifies the image samples in the sequence of encoded video that form part of num_part je view_id i, in terms of a rectangular region specified in the frame coordinates for emission.
The CropUnitX and CropUnitY variables are derived as follows:
- If chroma_format_idc is equal to 0, CropUnitX and CropUnitY are derived as follows:
CropUnitX = 1
CropUnitY = 2 -frame_mbs_only_flag - Otherwise (chroma_format_idc is equal to 1, 2, or 3), CropUnitX and CropUnitY are derived as follows:
CropUnitX = SubWidthC
CropUnitY = SubHeightC * (2 - frame_mbs_only_flag)
The frame clipping rectangle includes the luma samples with horizontal frame coordinates starting from the following:
CropUnitX * frame_crop_left_offset for PicWidthlnSamplesL - (CropUnitX * frame_crop_right_offset + 1) and coordinates of vertical frames from CropUnitY * frame_crop_top_offset to (16 * FrameHeightlnMbs) - (CropUnitY * frame_crop_bottom, even_set_cropott). The frame_crop_left_offset value must be in the range of 0 to (PicWidthlnSamplesL / CropUnitX) - (frame_crop_right_offset + 1), inclusive; and the frame_crop_top_offset value must be in the range 0 to (16 * FrameHeightlnMbs / CropUnitY) - (frame_crop_bottom_offset +1), inclusive.
When chroma_format_idc is not equal to 0, the corresponding specified samples of the two chroma systems are the samples having SubWidthC, y / SubHeightC) frame coordinates, where (x, y) are the frame coordinates of the specified luma samples.
For decoded fields, the specified samples of the decoded field are samples that are comprised within the rectangle specified in the frame coordinates.
num_parts [view_id [i]] specifies the number of parts into which the image of view_id [i] is divided.
depth_flag [view_id [i]] specifies whether or not the current part is a depth signal. If depth_flag is equal to 0, then the current part is not a depth signal. If depth_flag is equal to 1, then the current part is a depth signal associated with the view identified by view_id [i].
flip_dir [sub_viewjd [i]] [j] specifies the turning direction of the current part. flip_dir equal to 0 indicates no turning, flip_dir equal to 1 indicates turning in a horizontal direction, flip_dir equal to 2 indicates turning in a vertical direction, and flip_dir equal to 3 indicates turning in horizontal and vertical directions.
flip_dir [view_id [i]] [j] specifies the turning direction for the current part. flip_dir equal to 0 indicates no turning, flip_dir equal to 1 indicates turning in a horizontal direction, flip_dir equal to 2 indicates turning in vertical direction, and flip_dir equal to 3 indicates turning in horizontal and vertical directions.
Ioc_left_offset [view_id [i]] [j], loc_top_offset [view_id [i]] [j] specifies the location in the pixel offsets, where the current part j is located in the final reconstructed image of the view with viewjd equal to viewjdp].
upsample_view_flag [view_id [i]] indicates whether the image belonging to the view specified by view_id [i] needs to be sampled upwards. upsample_view_flag [view_id [i]] equal to 0 specifies that the image with viewjd equal to viewjdp] will not be sampled upwards. upsample_view_flag [view_id [i]] equal to 1 specifies that the image with viewjd equal to viewjd [i] will be sampled upwards.
upsample Jilter [viewJd [i]] indicates the type of filter that should be used for ascending sampling. upsamplej<sup>:</sup>ilter [viewjd [i]] equal to 0 indicates that the 6-lead AVC filter should be used, upsample_filter [viewjd [i]] equal to 1 indicates that the 4-lead SVC filter should be used, upsample_filter [viewjd [i] ] equal to 2 indicates that the bilinear filter should be used, upsamplej<sup>r</sup>ilter [viewjd [i]] equal to 3 indicates that the special filter coefficients are transmitted. When upsample_filter [viewjd [i]] is not present, it is set to 0. In this mode, we use the custom 2D filter. It can be easily extended to 1 D filter, and some other non-linear filter.
vert_dim [view_id [i]] specifies the vertical dimension of the special 2D filter. hor_dim [view_id [i]] specifies the horizontal dimension of the special 2D filter. quantizer [view_id [i]] specifies the quantization factor for each coefficient of the filter.
filter_coeffs [view_id [i]] [yuv] [y] [x] specifies the quantized filter coefficients. yuv signals the component to which the filter coefficients apply, yuv equal to 0 specifies component Y, yuv equal to 1 specifies component U, and yuv equal to 2 specifies component V.
pixel_dist_x [sub_view_id [i]] and pixel_dist_y [sub_view_id [i]] respectively specify the distance in the horizontal direction and the vertical direction in the final reconstructed pseudo view between adjacent pixels in the view with sub_view_id equal to sub_view_id [i].
num_pixel_tiling_filter_coeffs_minus1 [sub_view_id [i] [j] plus one indicates the number of filter coefficients when the tiling mode is set to 1.
pixei_tiiing_filter_coeffs [sub_view_id [i] 0] signals the filter coefficients that are required to represent a filter that can be used to filter the tiled image. Examples of pixel-level tiling
Turning to Figure 22, two examples showing the composition of a pseudo-viewer using pixel tiling from four views are indicated respectively by reference numerals 2210 and 2220, respectively. The four views are indicated collectively by reference numeral 2250. The syntax values for the first example in Figure 22 are provided in TABLE 3 below.
TABLE 3
<td>pseudo view info (pseudo view id) {</td><td>Value</td>
<td>num sub views minus 1 [pseudo view id]</td><td> 3</td>
<td>sub view id [0]</td><td> 0</td>
<td>num parts minus1 [0]</td><td> 0</td>
<td>loc l eft offset [0] [0]</td><td> 0</td>
<td>loc top offset [0] [0]</td><td> 0</td>
<td>pixel dist x [0] [0]</td><td> 0</td>
<td>pixel dist y [0] [0]</td><td> 0</td>
<td>sub view id [1]</td><td> 0</td>
<td>num parts minus1 [1]</td><td> 0</td>
<td>loc left offset [1] [0]</td><td> 1</td>
<td>loc top offset [1] [0]</td><td> 0</td>
<td>pixel dist x [1] [0]</td><td> 0</td>
<td>pixel dist y [1] [0]</td><td> 0</td>
<td>sub view id [2]</td><td> 0</td>
<td>num parts minus1 [2]</td><td> 0</td>
<td>loc left offset [2] [0]</td><td> 0</td>
<td>loc top offset [2] [0]</td><td> 1</td>
<td>pixel dist x [2] [0]</td><td> 0</td>
<td>pixel dist y [2] [0]</td><td> 0</td>
<td>sub view id [3]</td><td> 0</td>
<td>num parts minus1 [3]</td><td> 0</td>
<td>loc left offset [3] [0]</td><td> 1</td>
<td>loc top offset [3] [0]</td><td> 1</td>
<td>pixel dist x [3] [0]</td><td> 0</td>
<td>pixel dist y [3] [0]</td><td> 0</td>
The syntax values for the second example in 22 are all identical except for the following two syntax elements: loc_left_offset [3] [0] equal to 5 and loc_top_offset [3] [0] equal to 3.
The offset indicates that the pixels corresponding to a view must start at a certain offset location. This is shown in Figure 22 (2220). This can be done, for example, when two views produce images in which common objects appear to be displaced from one view to the other. For example, if first and second cameras (representing first and second views) capture images of an object, the object may appear to be shifted five pixels to the right in the second view compared to the first view. This means that the pixel (i-5, j) in the first view corresponds to the pixel (i, j) in the second view. If the pixels in both views are simply tiled pixel by pixel, then there may not be much correlation between the adjacent pixels in the tile, and the gains in spatial coding may be small. Conversely, by shifting the tile so that the pixel (i-5, j) from a view is placed close to the pixel (i, j) from view two, spatial correlation can be increased and the spatial coding gain can also be increased. This is because, for example, the corresponding pixels for the object in the first and second views are being tiled next to each other.
Thus, the presence of loc_left_offset and loc_top_offset can benefit from the coding efficiency. The displacement information can be obtained through external means. For example, camera position information or global disparity vectors between views can be used to determine such displacement information.
As a result of the shift, some pixels in the pseudo view are not assigned pixel values from any view. Continuing the above example, when tiling the pixel (i-5, j) from view one along the pixel (i, j) from view two, for values of i = 0 ... 4 there is no pixel ( i-5, j) from view one to tile, so that these pixels are empty in the tile. For those pixels in the pseudo view (tile) that are not assigned pixel values from any view, at least one implementation uses an interpolation procedure similar to the subpixel interpolation procedure in motion compensation in stroke. That is, the empty tile pixels can be interpolated from adjacent pixels. Such interpolation can result in greater spatial correlation in the tile and more coding gain for the tile.
In video encoding, we can choose a different encoding type for each image, such as I, P and B images. For multiview video encoding, in addition, we define anchor images and non-anchor images. In one modality, we propose that the grouping decision can be made based on the type of image. This clustering information is signaled in the high-level syntax.
Turning to Figure 11, an example of five tiled views in a single frame is usually indicated by the reference numeral 1100. In particular, the sequence “ballroom” is shown with 5 tiled views in a single frame. Additionally, it can be seen that the fifth view is divided into two parts so that it can be arranged in a rectangular frame. Here, each view is of QVGA size so that the total frame dimension is 640x600. Since 600 is not a multiple of 16 it must be extended to 608.
For this example, the possible SEI message could be as shown in TABLE 4.
TABLE 4
<td>multiview display info (payloadSize) {</td><td>Value</td>
<td>num coded views minus 1</td><td> 5</td>
<td>org pic width in mbs minus1</td><td> 40</td>
<td>org piç height in mbs minus1</td><td> 30</td>
<td></td><td></td>
<td>view id [0]</td><td> 0</td>
<td>num parts [view id [0]]</td><td> 1</td>
<td></td><td></td>
<td>depth flag [view id [0]] [0]</td><td> 0</td>
<td>flip dir [view id [0]] [0]</td><td> 0</td>
<td>loc left offset [view id [0]] [0]</td><td> 0</td>
<td>loc top offset [view id [0]] [0]</td><td> 0</td>
<td>frame crop left offset [view id [0]] [0]</td><td> 0</td>
<td>frame crop right offset [view id [0]] [0]</td><td> 320</td>
<td>crop top offset frame [view id [0]] [0]</td><td> 0</td>
<td>frame crop bottom offset [view id [0]] [0]</td><td> 240</td>
<td></td><td></td>
<td>upsample view flag [view jd [0]]</td><td> 1</td>
<td>if (upsample view flag [view id [0]]) {</td><td></td>
<td>vert d im [view id [0]]</td><td> 6</td>
<td>hor dim [view id [0]]</td><td> 6</td>
<td>quantizer [view id [0]]</td><td> 32</td>
<td>for (yuv = 0; yuv <3; yuv ++) {</td><td></td>
<td>for (y = 0; y <vert dim [view id [i]] -1; y ++) {</td><td></td>
<td>for (x = 0; x <hor dim [view id [i]] -1; x ++)</td><td></td>
<td>filter coeffs [view id [i]] [yuv] [y] [x]</td><td>XX</td>
<td></td><td></td>
<td></td><td></td>
<td>view id [1]</td><td> 1</td>
<td>num parts [view id [1]]</td><td> 1</td>
<td></td><td></td>
<td>depth flag [view id [0]] [0]</td><td> 0</td>
<td>flip dir [view id [1]] [0]</td><td> 0</td>
<td>loc left offset [view id [1]] [0]</td><td> 0</td>
<td>loc top offset [view id [1]] [0]</td><td> 0</td>
<td>frame crop left offset [view id [1]] [0]</td><td> 320</td>
<td>frame crop right offset [view id [1]] [0]</td><td> 640</td>
<td>crop top offset frame [view id [1]] [0]</td><td> 0</td>
<td>frame crop bottom offset [view id [1]] [0]</td><td> 320</td>
<td></td><td></td>
<td>upsample view flag [view id [1]]</td><td> 1</td>
<td>if (upsample view flag [viewjd [1]]) {</td><td></td>
<td>vert dim [view id [1]]</td><td> 6</td>
<td>hor dim [view id [1]]</td><td> 6</td>
<td>quantizer [view id [1]]</td><td> 32</td>
<td>for (yuv = 0; yuv <3; yuv ++) {</td><td></td>
<td>for (y = 0; y <vert dim [view id [i]] -1; y ++) {</td><td></td>
<td>for (x = 0; x <hor dim [view id [i]] -1; x ++)</td><td></td>
<td>filter coeffs [view id [i]] [yuv] [y] [x]</td><td>XX</td>
<td></td><td></td>
<td></td><td></td>
<td></td><td></td>
<td>.......... (similarly for view 2.3)</td><td></td>
<td></td><td></td>
<td>view id [4]</td><td> 4</td>
<td>num parts [view id [4]]</td><td> 2</td>
<td></td><td></td>
<td>depth flag [view id [0]] [0]</td><td> 0</td>
<td>flip dir [view id [4]] [0]</td><td> 0</td>
<td>loc left offset [view id [4]] [0]</td><td> 0</td>
<td>loc top offset [view id [4]] [0]</td><td> 0</td>
<td>frame crop left offset [view id [4]] [0]</td><td> 0</td>
<td>frame crop right offset [view id [4]] [0]</td><td> 320</td>
<td>crop top offset frame [view id [4]] [0]</td><td> 480</td>
<td>frame crop bottom offset [view id [4]] [0]</td><td> 600</td>
<td></td><td></td>
<td>flip dir [view id [4]] [1]</td><td> 0</td>
<td>loc left offset [view id [4]] [1]</td><td> 0</td>
<td>loc top offset [view id [4]] [1]</td><td> 120</td>
<td>frame crop left pffset [view id [4]] [1]</td><td> 320</td>
<td>frame crop right offset [view id [4]] [1]</td><td> 640</td>
<td>frame crop top offset [view id [4]] [1]</td><td> 480</td>
<td>frame crop bottom offset [view id [4]] [1]</td><td> 600</td>
<td></td><td></td>
<td></td><td></td>
<td>upsample view fl3g [view id [4]]</td><td> 1</td>
<td>sefuDsamDle view flaííview idf 411) {</td><td></td>
<td>vert dim [view id [4]]</td><td> 6</td>
<td>hor dim [view id [4]]</td><td> 6</td>
<td>qua ntize r [view id [4]]</td><td> 32</td>
<td>for (yuv = 0; yuv <3; yuv ++) {</td><td></td>
<td>for (y = 0; y <vert dim [view id [i]] -1; y ++) {</td><td></td>
<td>for (x = 0; x <hor dim [view id [i]] -1; x ++)</td><td></td>
<td>fi I have coeffs [view id [i]] [yuv] [y] [x]</td><td>XX</td>
<td></td><td></td>
TABLE 5 shows the general syntax structure for transmitting multiview information for the example shown in Table 4.
TABLE 5
<td>multiview display info (payloadSize) {</td><td>Ç</td><td>Descriptor</td>
<td>num coded views minus1</td><td> 5</td><td>eu (v)</td>
<td>org pic width in mbs minus1</td><td> 5</td><td>eu (v)</td>
<td>o rg p iç he ig ht inm bs minus 1</td><td> 5</td><td>eu (v)</td>
<td>for (i = 0; i <= num coded views minus1; i ++) {</td><td></td><td></td>
<td>view id [i]</td><td> 5</td><td>eu (v)</td>
<td>num parts [view id [i]]</td><td> 5</td><td>eu (v)</td>
<td>for (j = 0; j <= num parts [i]; j ++) {</td><td></td><td></td>
<td>depth flag [view id [i]] [j]</td><td></td><td></td>
<td>flip dir [view id [i]] [j]</td><td> 5</td><td>u (2)</td>
<td>loc left offset [view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td>loc top offset [view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td>frame crop left offset [view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td>frame crop right offset [view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td>crop top offset frame [view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td>frame crop bottom offset [view id [i]] [j]</td><td> 5</td><td>eu (v)</td>
<td> }</td><td></td><td></td>
<td>upsample view flag [view id [i]]</td><td> 5</td><td>u (1)</td>
<td>if (upsample view flag [view id [i]])</td><td></td><td></td>
<td>upsample filter [view id [i]]</td><td> 5</td><td>u (2)</td>
<td>if (upsample_fiter [view_id [i]] == 3) {</td><td></td><td></td>
<td>vert dim [view id [i]]</td><td> 5</td><td>eu (v)</td>
<td>hor dim [view id [i]]</td><td> 5</td><td>eu (v)</td>
<td>quantizer [view id [i]]</td><td> 5</td><td>eu (v)</td>
<td>for (yuv = 0; yuv <3; yuv ++) {</td><td></td><td></td>
<td>for (y = 0; y <vert dim [view id [i]] -1; y ++) {</td><td></td><td></td>
<td>for (x = 0; x <hor dim [view id [i]] -1; x ++)</td><td></td><td></td>
<td>filter coeffs [view id [i]] [yuv] [y] [x]</td><td> 5</td><td>if (v)</td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
Referring to Figure 23, a 2300 video processing device is shown. The 2300 video processing device can be, for example, a signal converter or other device that receives encoded video and provides, for example, decoded video for display. for a user or for storage. Thus, the 2300 device can provide its output to a television, computer monitor, or a computer or other processing device.
The device 2300 includes a decoder 2310 that receives a data signal 2320. The data signal 2320 may include, for example, an AVC compatible stream or an MVC compatible stream. The decoder 2310 decodes all or part of the received signal 2320 and outputs a decoded video signal 2330 and tiling information 2340. The decoded video 2330 and tiling information 2340 are provided to a selector 2350. The 2300 device also includes a 2360 user interface that receives a 2370 user input. The 2360 user interface provides a 2380 image selection signal, based on the 2370 user input, for the 2350 selector. image 2380 and user entry 2370 indicate which of the multiple images a user wants to have displayed. The 2350 selector provides the selected image (s) as a 2390 output. The 2350 selector uses the image selection information 2380 to select which of the images in the 2330 encoded video to provide as the 2390 output. The 2350 selector uses the 2340 tiling information to locate the selected image (s) in the 2330 decoded video.
In many implementations, the 2350 selector includes the 2360 user interface, and in other implementations no 2360 user interface is required because the 2350 selector receives the 2370 user input directly without a separate interface function being performed. Selector 2350 can be implemented in software or as an integrated circuit, for example. The 2350 selector can also incorporate the 2310 decoder.
More generally, decoders of various implementations described in that application can provide a decoded output that includes an entire tile. Additionally or alternatively, decoders can provide a decoded output that includes only one or more selected images (images or depth signals, for example) from the tile.
As noted above, high-level syntax can be used to perform signaling in accordance with one or more modalities of these principles. High-level syntax can be used, for example, but is not limited to signaling any of the following: the number of coded views present in the largest frame, the original width and height of all views; for each coded view, the view identifier corresponding to the view; for each coded view, the number of parts by which the frame of a view is divided; for each part of the view, the turning direction (which can be, for example, no turning, only horizontal turning, only vertical turning or horizontal and vertical turning); for each part of the view, the left position in pixels or the number of macroblocks where the current part belongs in the final frame for the view; for each part of the view, the top position of the part in pixels or number of macroblocks where the current part belongs in the final frame for the view; for each part of the view, the left position, in the current large decoded / coded frame, of the clipping window in pixels or number of macroblocks; for each part of the view, the position on the right, in the current large decoded / coded frame, of the clipping window in pixels or number of macroblocks; for each part of the view, the top position, in the current large decoded / coded frame, of the clipping window in pixels or number of macroblocks; and, for each part of the view, the bottom position, in the current large decoded / coded frame, of the pixel clipping window or the number of macroblocks; for each coded view if the view needs to be sampled upwardly before output (where if upward sampling needs to be performed, a high-level syntax can be used to indicate the method for upward sampling (including, but not limited to, filter 6 AVC leads, 4-lead SVC filter, bilinear filter or a special 1D filter, 2D linear or non-linear).
It should be noted that the terms, "encoder" and "decoder" connote general structures and are not limited to any specific functions or features. For example, a decoder can receive a modulated carrier that carries an encoded bit stream, and demodulates the encoded bit stream, just as it decodes the bit stream.
Several methods have been described. Many of these methods are detailed to provide broad disclosure. However, it is observed that variations are considered to vary one or many of the specific characteristics described for these methods. In addition, many of the features that are cited are known in the art and, consequently, are not described in great detail.
Additionally, reference was made to the use of high-level syntax to send certain information in various implementations. However, it must be understood that other implementations use lower-level syntax, or in fact other mechanisms in general (such as, for example, sending information as part of the encoded data) to provide the same information (or variations of that information).
Several implementations provide tiling and appropriate signage to allow multiple views (images, more generally) to be tiled in a single image, encoded as a single image, and sent as a single image. Signaling information can allow a post-processor to separate views / images. In addition, the multiple images that are tiled could be seen, but at least one of the images could be depth information. These implementations can provide one or more advantages. For example, users may wish to display multiple views in a tiled way, and these various implementations provide an efficient way to code and transmit or store such views by tiling them before encoding and transmitting / storing them in a tiled way.
Implementations that tile multiple views in the context of stroke and / or MVC also provide additional benefits. Stroke is only used ostensibly for a single view, so no additional views are expected. However, such AVC-based implementations can provide multiple views in an AVC environment because the tiled views can be arranged so that, for example, a decoder knows that the tiled images belong to the different views (for example, the upper left image in the pseudovista is view 1, top right image is view 2, etc.).
Additionally, MVC already includes multiple views, so that multiple views should not be included in a single pseudo view. Additionally, MVC has a limit on the number of views that can be supported, and such MVC-based implementations effectively increase the number of views that can be supported by allowing (as in AVC-based implementations) that additional views are tiled. For example, each pseudovista can correspond to one of the supported views of MVC, and the decoder may be aware that each “supported view” actually includes four views in a pre-arranged tiled order. Thus, in such an implementation, the number of possible views is four times the number of “supported views”.
The implementations described here can be implemented, for example, in a method or process, in an apparatus, or in a software program. Even if discussed only in the context of a single form of implementation (for example, discussed only as a method), the implementation of discussed features can be implemented in other ways (for example, an apparatus or program). A device can be implemented, for example, in appropriate hardware, software and firmware. The methods can be implemented, for example, on an apparatus such as, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a device programmable logic. Processing devices also include a communication device, such as, for example, computers, cell phones, personal / portable digital assistants ('PDAs'), and other devices that facilitate the communication of information between end users.
Implementations of the various processes and features described here can be incorporated into a variety of different equipment or applications, particularly, for example, equipment or applications associated with data encoding and decoding. Examples of equipment include video encoders, video decoders, video codecs, network servers, signal converters, laptops, personal computers, cell phones, PDAs, and other communication devices. As must be evident, the equipment can be mobile and even installed in a mobile vehicle.
In addition, the methods can be implemented upon instructions being carried out by a processor, and such instructions can be stored in a processor-readable medium such as, for example, an integrated circuit, a software carrier or other storage device such as, for example, example, a hard disk, a floppy disk, a random access memory (“RAM”), or a read memory (“ROM”). The instructions can form a tangible embedded application program in a processor-readable medium. Of course, a processor may include a processor-readable medium having, for example, instructions for carrying out a process. Such application programs can be transferred to, and executed by, a machine comprising any suitable architecture. Preferably, the machine is implemented on a computer platform having hardware such as one or more central processing units (“CPU”), random access memory (“RAM”), and input / output interfaces (“L / O ”). The computer platform may also include an operating system and microinstruction code. The various processes and functions described here can be part of the microinstruction code or part of the application program, or any combination of them, which can be performed by a CPU. In addition, several other peripheral units can be connected to the computer platform such as an additional data storage unit and a printing unit.
As should be evident to those skilled in the art, implementations can also produce a signal formatted to carry information that can, for example, be stored or transmitted. The information may include, for example, instructions for carrying out a method, or data produced by one of the described implementations. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a portion of radio frequency spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream, producing syntax, and modulating a carrier with the encoded data stream and syntax. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known.
It should be further understood that, due to some of the system components, constituents and methods illustrated in the attached drawings are preferably implemented in software, the actual connections between the system components or the process function blocks may differ depending on the way in which the present principles are programmed. Given the present teachings, those of ordinary knowledge in the relevant technique will be able to consider these and similar implementations or configurations of the present principles.
Some implementations have been described. However, it will be understood that several modifications can be made. For example, elements from different implementations can be combined, supplemented, modified, or removed to produce other implementations. In addition, those of ordinary skill in the art will understand that other structures or processes may be substitutes for those disclosed and the resulting implementations will perform at least substantially the same function (s) in at least substantially the same form (s) to obtain at least substantially the same. same result (s) as the deployments revealed. In particular, although illustrative modalities are described here with reference to the attached drawings, it should be understood that the present principles are not limited to those precise modalities, and that several changes and modifications can be made to them by those skilled in the relevant technique without departing from the relevant technique. spirit or scope of these principles. Consequently, these and other implementations are considered by this request and are within the scope of the following claims.
Contents12
29 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29
164 members in 21 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 60923014 | United States of America | – | |
| 92301407 | United States of America | P | |
| 92301407 | United States of America | P | |
| 60925400 | United States of America | – | |
| 92540007 | United States of America | P | |
| 92540007 | United States of America | P | |
| 2008004747 | United States of America | W | |
| 2008004747 | United States of America | W | |
| 2008004747 | – | – | – |
| 60923014 | – | – | – |
| 60925400 | – | – | – |
| US20070923014P | – | – | – |
| US20070925400P | – | – | – |
| WO2008US04747 | – | – | – |
Members164
| Document | Office | Kind | |
|---|---|---|---|
| AU2008239653A1 | Australia | A1 | |
| WO2008127676A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008127676A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2008127676A9 | World Intellectual Property Organization (WIPO) | A9 | |
| MX2009010973A | Mexico | A | |
| EP2137975A2 | European Patent Office (EPO) | A2 | |
| KR20100016212A | Republic of Korea | A | |
| CN101658037A | China | A | |
| US2010046635A1 | United States of America | A1 | |
| JP2010524398A | Japan | A | |
| RU2009141712A | Russian Federation | A | |
| ZA201006649B | South Africa | B | |
| AU2008239653B2 | Australia | B2 | |
| EP2512135A1 | European Patent Office (EPO) | A1 | |
| EP2512136A1 | European Patent Office (EPO) | A1 | |
| ZA201201942B | South Africa | B | |
| BRPI0809510A2This record | Brazil | A2 | |
| AU2012278382A1 | Australia | A1 | |
| CN101658037B | China | B | |
| JP5324563B2 | Japan | B2 | |
| BRPI0823512A2 | Brazil | A2 | |
| JP2013258716A | Japan | A | |
| RU2521618C2 | Russian Federation | C2 | |
| US8780998B2 | United States of America | B2 | |
| KR20140098825A | Republic of Korea | A | |
| US2014301479A1 | United States of America | A1 | |
| KR101467601B1 | Republic of Korea | B1 | |
| AU2012278382B2 | Australia | B2 | |
| JP5674873B2 | Japan | B2 | |
| EP2512135B1 | European Patent Office (EPO) | B1 | |
| KR20150046385A | Republic of Korea | A | |
| AU2008239653C1 | Australia | C1 | |
| JP2015092715A | Japan | A | |
| AU2015202314A1 | Australia | A1 | |
| EP2887671A1 | European Patent Office (EPO) | A1 | |
| US2015281736A1 | United States of America | A1 | |
| RU2014116612A | Russian Federation | A | |
| US9185384B2 | United States of America | B2 | |
| US2015341665A1 | United States of America | A1 | |
| US9219923B2 | United States of America | B2 | |
| US9232235B2 | United States of America | B2 | |
| US2016080757A1 | United States of America | A1 | |
| EP2512136B1 | European Patent Office (EPO) | B1 | |
| KR101646089B1 | Republic of Korea | B1 | |
| PT2512136T | Portugal | T | |
| DK2512136T3 | Denmark | T3 | |
| US9445116B2 | United States of America | B2 | |
| AU2015202314B2 | Australia | B2 | |
| ES2586406T3 | Spain | T3 | |
| KR20160121604A | Republic of Korea | A | |
| PL2512136T3 | Poland | T3 | |
| US2016360218A1 | United States of America | A1 | |
| HUE029776T2 | Hungary | T2 | |
| US9706217B2 | United States of America | B2 | |
| JP2017135756A | Japan | A | |
| US2017257638A1 | United States of America | A1 | |
| KR20170106987A | Republic of Korea | A | |
| KR101766479B1 | Republic of Korea | B1 | |
| US9838705B2 | United States of America | B2 | |
| US2018048904A1 | United States of America | A1 | |
| RU2651227C2 | Russian Federation | C2 | |
| US9973771B2 | United States of America | B2 | |
| EP2887671B1 | European Patent Office (EPO) | B1 | |
| US9986254B1 | United States of America | B1 | |
| US2018152719A1 | United States of America | A1 | |
| ES2675164T3 | Spain | T3 | |
| PT2887671T | Portugal | T | |
| DK2887671T3 | Denmark | T3 | |
| TR201809177T4 | Türkiye | T4 | |
| LT2887671T | Lithuania | T | |
| US2018213246A1 | United States of America | A1 | |
| KR101885790B1 | Republic of Korea | B1 | |
| KR20180089560A | Republic of Korea | A | |
| PL2887671T3 | Poland | T3 | |
| HUE038192T2 | Hungary | T2 | |
| SI2887671T1 | Slovenia | T1 | |
| EP3399756A1 | European Patent Office (EPO) | A1 | |
| US10129557B2 | United States of America | B2 | |
| US2019037230A1 | United States of America | A1 | |
| RU2684184C1 | Russian Federation | C1 | |
| KR101965781B1 | Republic of Korea | B1 | |
| KR20190038680A | Republic of Korea | A | |
| US10298948B2 | United States of America | B2 | |
| US2019253727A1 | United States of America | A1 | |
| HK1255617A1 | Hong Kong, China | A1 | |
| US10432958B2 | United States of America | B2 | |
| BRPI0809510B1 | Brazil | B1 | |
| BR122018004903B1 | Brazil | B1 | |
| BR122018004904B1 | Brazil | B1 | |
| BR122018004906B1 | Brazil | B1 | |
| KR102044130B1 | Republic of Korea | B1 | |
| KR20190127999A | Republic of Korea | A | |
| JP2019201435A | Japan | A | |
| JP2019201436A | Japan | A | |
| US2019379897A1 | United States of America | A1 | |
| RU2709671C1 | Russian Federation | C1 | |
| RU2721941C1 | Russian Federation | C1 | |
| KR102123772B1 | Republic of Korea | B1 | |
| KR20200069389A | Republic of Korea | A | |
| US10764596B2 | United States of America | B2 |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Requested change of headquarter approvedB25G | B25G | |
| Patent or certificate of addition of invention granted [chapter 16.1 patent gazette]GrantedPRAZO DE VALIDADE: 10 (DEZ) ANOS CONTADOS A PARTIR DE 08/10/2019, OBSERVADAS AS CONDICOES LEGAIS. (CO) 10 (DEZ) ANOS CONTADOS A PARTIR DE 08/10/2019, OBSERVADAS AS CONDICOES LEGAISB16A | B16A | |
| Patent application procedure suspended [chapter 6.1 patent gazette]B06A | B06A | |
| Objections, documents and/or translations needed after an examination request according [chapter 6.6 patent gazette]B06F | B06F | |
| Others concerning applications: alteration of classificationB15K | B15K | |
| Others concerning applications: alteration of classificationAS CLASSIFICACOES ANTERIORES ERAM: H04N 7/26 , H04N 7/50B15K | B15K | |
| Requested transfer of rights approvedB25A | B25A |
Numbers
- Publication
- PI0809510
- Publication, DOCDB
- PI0809510
- Publication, EPODOC
- BRPI0809510
- Application
- 9510
- Application, DOCDB
- PI0809510
- Application, EPODOC
- BR2008PI09510
Titles2
- Portuguese
- LADRILHAMENTO EM CODIFICAÇÃO E DECODIFICAÇÃO DE VÍDEO
- English
- TILE IN VIDEO ENCODING AND DECODING
Classification
- CPC, 11
- H04N19/597
- H04N19/46
- H04N19/70
- H04N19/61
- H04N19/172
- H04N19/182
- H04N2213/003
- H04N13/161
- H04N13/194
- H04N19/174
- H04N13/111
- IPC, 2
- H04N7 26
- H04N7 50