Efficient encoding of multiple views
Abstract
A new method of encoding multiple view image information into an image signal (200) is described, comprising: adding to the image signal (200) a first image (220) of pixel values representing one or more objects (110, 112) captured by a first camera (101); - adding to the image signal (200) a map (222) comprising for respective sets of pixels of the first image (220) respective values, representing a three-dimensional position in space of a region of the one or more objects (110, 112) represented by the respective set of pixels; and adding to the image signal (200) a partial representation (223) of a second image (224) of pixel values representing one or more objects (110, 112) captured by a second camera (102), the partial representation (223) comprising at least information of the majority of the pixels representing regions of the one or more objects (110, 112) not visible to the first camera (101). Advantages are less required information for higher precision and increased usability.
Term
0.5 yearsto projected expiry
Projected expiry 23 March 2027, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
18 claims: 10 independent, 8 dependent
- 1Claims Zastrzeżenia patentowe 1. Sposób kodowania informacji obrazu wielu widoków w jednym sygnale obrazu (200) obejmujący:A method for coding image information of multiple views in one image signal (200) comprising: - adding to the image signal (200) an image of the first (220) pixel values representing one or more objects (110, 112) captured by the first camera (101);- dodawanie do sygnału obrazu (200) obrazu pierwszego (220) wartości pikseli reprezentujących jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę pierwszą (101);- adding to the image signal (200) a map (222) containing for the respective sets of pixels of the first image (220) corresponding values representing a three-dimensional position in the space of one or more objects (110, 112) represented by a respective set of pixels;- dodawanie do sygnału obrazu (200) mapy (222) zawierającej dla odpowiednich zbiorów pikseli obrazu pierwszego (220) odpowiednie wartości reprezentujące pozycję trójwymiarową w przestrzeni obszaru jednego lub większej liczby obiektów (110, 112) reprezentowanych przez odpowiedni zbiór pikseli;- dostarczanie obrazu drugiego (224) wartości pikseli reprezentujących wymieniony jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę drugą (102) w innej lokalizacji niż kamera pierwsza;- providing a second image (224) of pixel values representing said one or more objects (110, 112) captured by the second camera (102) at a location other than the first camera;- determining which areas are present in the second image and not in the first image (220) in response to the discrepancy between the first image (220) and the second image (224);wherein the method is characterized by comprising: - ustalanie, które obszary są obecne w obrazie drugim i nie w obrazie pierwszym (220) w odpowiedzi na rozbieżności między obrazem pierwszym (220) i obrazem drugim (224);przy czym sposób jest znamienny obejmowaniem: - adding to the image signal (200) a partial representation (223) of the second image (224), the partial representation (223) comprising at least the majority of the pixels of the designated areas and determining the regions of the second image (224) that do not have to be coded, - dodawania do sygnału obrazu (200) reprezentacji częściowej (223) obrazu drugiego (224), przy czym reprezentacja częściowa (223) zawiera co najmniej większość pikseli wyznaczonych obszarów i określanie obszarów obrazu drugiego (224), które nie muszą być kodowane,
- 5Sposób według któregokolwiek z powyższych zastrzeżeń, w którym wartości mapy (222) są dokładnie dostrajane przez osobę przed dodawaniem do sygnału obrazu (200). The method of any of the preceding claims, wherein the map values (222) are accurately tuned by the person before adding to the image signal (200).
- 6Sposób według któregokolwiek z powyższych zastrzeżeń, w którym reprezentacja częściowa (223) jest dokładnie dostrajana przez osobę przed dodawaniem do sygnału obrazu (200). A method according to any one of the preceding claims, wherein the partial representation (223) is precisely tuned by the person before adding to the image signal (200).
- 8Sposób według któregokolwiek z powyższych zastrzeżeń, w którym etap dodawania do sygnału obrazu (200) reprezentacji częściowej (223) obrazu drugiego (224) obejmuje określanie i dodawanie do sygnału obrazu reprezentacji otaczającego kształtu, otaczającego obszar pikseli w reprezentacji częściowej (223) obrazu drugiego (224). A method according to any one of the preceding claims, wherein the step of adding to the image signal (200) a partial representation (223) of the second image (224) comprises determining and adding to the image signal a representation of the surrounding shape surrounding the pixel area of the partial representation (223) of the image the second (224).
- 9A method according to any one of the preceding claims, wherein the image analysis, such as, for example, morphological analysis, is performed on the areas included in the partial representation (223) of the second image (224) and the modification is performed on the partial representation (223) prior to the addition of the representation. partial (223) to the image signal (200), wherein the morphological analysis includes, for example, determining the largest width of the respective areas. 9. Sposób według któregokolwiek z powyższych zastrzeżeń, w którym analiza obrazu, taka jak np. analiza morfologiczna, jest realizowana na obszarach zawartych w reprezentacji częściowej (223) obrazu drugiego (224) a modyfikacja jest realizowana na reprezentacji częściowej (223) przed dodawaniem reprezentacji częściowej (223) do sygnału obrazu (200), przy czym analiza morfologiczna obejmuje na przykład ustalanie największej szerokości odpowiednich obszarów.
- 10An apparatus (310) for generating coding in an image signal (200) of image information of a plurality of views, including:10. Urządzenie (310) do generowania kodowania w sygnale obrazu (200) informacji obrazu wielu widoków, zawierające: - means (340) adapted to add to the image signal (200) an image of the first (220) pixel values representing one or more objects (110, 112) captured by the first camera (101);- środki (340) przystosowane do dodawania do sygnału obrazu (200) obrazu pierwszego (220) wartości pikseli reprezentujących jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę pierwszą (101);- means (341) adapted to add to the image signal (200) a map (222) containing for the respective sets of pixels of the first image (220) corresponding values representing a three-dimensional position in the space of one or more objects (110, 112) represented by a corresponding set pixels;and - środki (341) przystosowane do dodawania do sygnału obrazu (200) mapy (222) zawierającej dla odpowiednich zbiorów pikseli obrazu pierwszego (220) odpowiednie wartości reprezentujące trójwymiarową pozycję w przestrzeni obszaru jednego lub większej liczby obiektów (110, 112) reprezentowanych przez odpowiedni zbiór pikseli;i - means adapted to provide a second image (224) of pixel values representing one or more objects (110, 112) captured by the second camera (102) at a location other than the first camera;- środki przystosowane do dostarczania obrazu drugiego (224) wartości pikseli reprezentujących jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę drugą (102) w innej lokalizacji niż kamera pierwsza;- means adapted to determine at least which areas are present in the second image and not in the first picture (220) in response to discrepancies between the first image (220) and the second image (224);and the device is characterized by the inclusion: - środki przystosowane do ustalania co najmniej, które obszary są obecne w obrazie drugim i nie w obrazie pierwszym (220) w odpowiedzi na rozbieżności między obrazem pierwszym (220) i obrazem drugim (224);i urządzenie jest znamienne zawieraniem: - means (342) adapted to add to the image signal (200) a partial representation (223) of the second image (224), the partial representation (223) comprising at least the information of the majority of pixels of the regions, and defining regions in the second image (224), which do not have to be coded. - środków (342) przystosowanych do dodawania do sygnału obrazu (200) reprezentacji częściowej (223) obrazu drugiego (224), przy czym reprezentacja częściowa (223) zawiera co najmniej informację większości pikseli obszarów, i określające obszary w obrazie drugim (224), które nie muszą być kodowane.
- 13An image signal receiver (400), comprising:13. Odbiornik (400) sygnału obrazu, zawierający: - means (402) adapted to extract from the image signal (200) the image of the first (220) pixel values representing one or more objects (110, 112) captured by the first camera (101);- środki (402) przystosowane do ekstrakcji z sygnału obrazu (200) obrazu pierwszego (220) wartości pikseli reprezentujących jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę pierwszą (101);- means (404) adapted for extraction from the image signal (200) of the map (222) containing for the respective sets of pixels of the first image (220) corresponding values representing a three-dimensional position in the space of one or more objects (110, 112) represented by a corresponding set pixels;- środki (404) przystosowane do ekstrakcji z sygnału obrazu (200) mapy (222) zawierającej dla odpowiednich zbiorów pikseli obrazu pierwszego (220) odpowiednie wartości reprezentujące trójwymiarową pozycję w przestrzeni obszaru jednego lub większej liczby obiektów (110, 112) reprezentowanych przez odpowiedni zbiór pikseli;- means (406) adapted to extract from the image signal (200) a partial representation (223) of the second image (224) the pixel values representing one or more objects (110, 112) captured by the second camera (102) at a different location than the first camera;and further comprising: a partial representation (223) comprising at least the information of a majority of pixels representing regions of one or more objects (110, 112) present in the second image and not in the first image (220), and the image signal (200) includes for the second image (224) only the part (223) of the second image. - środki (406) przystosowane do ekstrakcji z sygnału obrazu (200) reprezentacji częściowej (223) obrazu drugiego (224) wartości pikseli reprezentujących jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę drugą (102) w innej lokalizacji niż kamera pierwsza, i znamienny zawieraniem ponadto: reprezentacji częściowej (223) zawierającej co najmniej informację większości pikseli reprezentujących obszary jednego lub większej liczby obiektów (110, 112) obecnych w obrazie drugim i nie w obrazie pierwszym (220), a sygnał obrazu (200) zawiera dla obrazu drugiego (224) tylko część (223) obrazu drugiego.
- 14A display (415) capable of generating at least two image views, including:14. Wyświetlacz (415) zdolny do generowania co najmniej dwóch widoków obrazu zawierający: - an image signal receiver (400) according to claim 13;- odbiornik (400) sygnału obrazu według zastrzeżenia 13;- image regenerator (410), adapted to generate from two images of image signal data received by the receiver (400) of the image signal;and - regenerator (410) obrazu, przystosowany do generowania z dwóch obrazów z danych sygnału obrazu odebranych przez odbiornik (400) sygnału obrazu;i - an image rendering unit (412) adapted to generate from successive images two images in a format suitable for the display. - zespół (412) renderowania obrazu, przystosowany do generowania z dwóch obrazów kolejnych obrazów w formacie odpowiednim dla wyświetlacza.
- 15A method of extracting image information of a plurality of views from an image signal (200) comprising:15. Sposób ekstrakcji informacji obrazu wielu widoków z sygnału obrazu (200) obejmujący: - ekstrakcję z sygnału obrazu (200) obrazu pierwszego (220) wartości pikseli reprezentujących jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę pierwszą (101);- extracting from the image signal (200) the image of the first (220) pixel value representing one or more objects (110, 112) captured by the first camera (101);- ekstrakcję z sygnału obrazu (200) mapy (222) zawierającej dla odpowiednich zbiorów pikseli obrazu pierwszego (220) odpowiednie wartości reprezentujące trójwymiarową pozycję w przestrzeni obszaru jednego lub większej liczby obiektów (110,112) reprezentowanych przez odpowiedni zbiór pikseli;- extraction from the image signal (200) of the map (222) containing for the respective sets of pixels of the first image (220) corresponding values representing a three-dimensional position in the space of one or more objects (110,112) represented by a respective set of pixels;- ekstrakcję z sygnału obrazu (200) reprezentacji częściowej (223) obrazu drugiego (224) wartości pikseli reprezentujących jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę drugą (102) w innej lokalizacji niż kamera pierwsza, i sposób jest znamienny zawieraniem ponadto: reprezentacji częściowej (223) zawierającej co najmniej informację większości pikseli reprezentujących obszary jednego lub większej liczby obiektów (110, 112) obecnych w obrazie drugim i nie w obrazie pierwszym (220), a sygnał obrazu (200) zawiera dla obrazu drugiego (224) tylko część (223) obrazu drugiego. - extracting from the image signal (200) a partial representation (223) of the second image (224) of pixel values representing one or more objects (110, 112) captured by the second camera (102) at a location other than the first camera, and the method is characterized by including further: a partial representation (223) comprising at least the information of a majority of pixels representing regions of one or more objects (110, 112) present in the second image and not in the first image (220), and the image signal (200) includes for the second image (224) ) only part (223) of the second image.
- 16An image signal (200) containing:16. Sygnał obrazu (200) zawierający: - a first image (220) of pixel values representing one or more objects (110, 112) captured by the first camera (101);- obraz pierwszy (220) wartości pikseli reprezentujących jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę pierwszą (101);- a map (222) containing for the respective sets of pixels of the first image (220) corresponding values representing a three-dimensional position in the space of one or more objects (110, 112) represented by a respective set of pixels, and characterized by: - mapę (222) zawierającą dla odpowiednich zbiorów pikseli obrazu pierwszego (220) odpowiednie wartości reprezentujące trójwymiarową pozycję w przestrzeni obszaru jednego lub większej liczby obiektów (110, 112) reprezentowanych przez odpowiedni zbiór pikseli, i znamienny zawieraniem: - partial representation (223) of only part of the second image (224) of pixel values representing one or more objects (110, 112) captured by the second camera (102) at a location other than the first camera, the partial representation (223) comprising at least information of the majority of pixels representing the areas of one or more objects (110, 112) present in the second image and not in the first image (220). - reprezentacji częściowej (223) tylko części obrazu drugiego (224) wartości pikseli reprezentujących jeden lub większą liczbę obiektów (110, 112) przechwytywanych przez kamerę drugą (102) w innej lokalizacji niż kamera pierwsza, przy czym reprezentacja częściowa (223) zawiera co najmniej informację większości pikseli reprezentujących obszary jednego lub większej liczby obiektów (110, 112) obecnych w obrazie drugim i nie w obrazie pierwszym (220).
Independent claims10
64 paragraphs, as filed
The invention relates to a method for coding a plurality of image views in an image signal, such as, for example, a compressed television signal according to one of the MPEG standards.
The invention also relates to: an apparatus for generating such a signal, a receiver for receiving such a signal, a method for acquiring coded information from a signal, so that it can be used to generate multiple views, and the efficiently encoded signal itself.
[0003] Currently, work is underway to normalize the coding of 3D image information. For example, in Redert A et al .: "ATTEST: advanced threedimensional television television technologies" 3D Data Processing Visualization And Transmission, 2002. Proceedings. First intemational Symposium, June 19-21 2002, Piscataway, NJ, USA, IEEE, June 19, 2002 (2002-06-19), pages 313-319, ΧΡ010596672 ISBN: 0-7695-1521-4 elements of the three-dimensional television system for emission environment. The article provides a brief overview of some of the tools defined in the MPEG standard and which are relevant for 3D video applications.
[0004] There are several ways to represent three-dimensional objects, for example as a set of voxels (popular for example in displaying medical data or industrial inspection of components), or as a number of image views captured from different directions and intended to be viewed from different directions, e.g. two the eyes of one observer or by many observers, or a moving observer, etc.
[0005] A popular format is the left-right format in which the image is captured by the camera on the left and the image is captured by the camera on the right. These images can be displayed on different displays, for example the left image can be displayed during the first set of time moments, the right image at the time of the interlaced second set of time moments, the left and right observer's eye being locked synchronously with the shutter glasses. A projector with polarizing means is another example of a display capable of generating three-dimensional experience, at least rendering part of the three-dimensional scene information, namely, what looks approximately like a certain direction (namely stereo).
[0006] Different scene zooming qualities may be used, e.g. a 3D scene may be represented as a set of flat layers, one after the other. But these different qualities can be coded using existing formats.
[0007] Another popular display is the autostereoscopic display. This display is formed, for example, by placing the LCD behind a collection of lenses, so that a group of pixels is imaged in the area in space by respective lenses. In this way, a number of cones in space are generated, which two to two contain left and right pictures for the left and right eyes, so that without glasses the user can be positioned in a number of areas in space and perceive 3D. However, the data for these pixel groups must be generated from left and right pictures. Another option is that the user can see the object from a number of intermediate directions, between the left and right views of the stereo coding, where intermediate views can be generated by calculating the discrepancy field between the left and right image and subsequent interpolation. WO 02/097733 describes the representation of 3D images at various angles through a normal image, a depth image and additional images corresponding to different viewing points.
[0008] The disadvantage of the left-right coding of the prior art is that significant amounts of data are needed to obtain intermediate views and that the obtained results may be somewhat disappointing. It is difficult to calculate exactly the appropriate discrepancy field, which leads to interpolation artifacts, such as the part of the background adhering to the object of the foreground. The wish leading to the technological embodiments described below was to have a coding method that can lead to relatively accurate results when converting to various formats, such as a collection of views with intermediate views, but which do not contain an excessive amount of data.
[0009] Such requirements are at least partially met by a method of coding image information of multiple views in an image signal (200) comprising:
adding to the image signal (200) the first image (220) having pixel values representing one or more objects (110, 112) captured by the first camera (101);
adding to the image signal (200) a map (222) comprising, for respective sets of pixels of the first image (220), corresponding values representing a three-dimensional position in the space of one or more objects (110, 112) represented by a respective set of pixels;
- providing a second image (224) having pixel values representing said one or more objects (110, 112) captured by the second camera (102) at a location other than the first camera;
- determining which areas are present in the second image and not in the first image (220) in response to the discrepancy between the first image (220) and the second image (224); and adding to the image signal (200) a partial representation (223) of the second image (224), the partial representation (223) comprising at least the majority of the pixels of the designated areas and determining the regions of the second image (224) that need not be coded, and the signal obtained by means of a method or device enabling the implementation of the method.
[0010] The inventors have found that if it is understood that for the sake of quality it is best to add to the images a left and right map containing information about a three-dimensional scene structure representing at least that part of the three-dimensional scene information that is required to allow the given application (with the desired quality), an interesting coding format can be developed. For interpolation of the view, the map can be, for example, a precisely segmented map of discrepancy, whose discrepancy vectors will lead to good interpolation of intermediate views. It should be noted that this map can be optimally tuned on the creation / transmission side according to its use on the receiving side, i.e. for example according to how the three-dimensional environment will be simulated on the display, which means that it will usually have different properties,
[0011] The map can e.g. be fine-tuned or even created by an operator who can visually check on his side how a number of planned pixels behaves when receiving a signal. Today, and even more so in the future, some of the content is already being prepared using a computer, such as a three-dimensional dinosaur model or graphic overlay, which implies that it is not too problematic to create at least areas containing accurate discrepancies maps, depth maps, or similar pixel maps created man-made objects.
[0012] This is certainly true in gaming applications where the user may, for example, move slightly from the scene and may wish to see the scene differently, but in the near future the invention may also be valid for 3D television captured by two cameras or even generated based on e.g. traffic parallax. The growing number of studies (eg BBC) uses, for example, a virtual environment in the news.
[0013] The map may be encoded with a small data overhead, e.g. as an image of gray values, compressed according to the MPEG-2 standard and attached to a left-right image (or images for several times of mobile video) already in the signal .
[0014] However, having this map, as the inventors have found, enables further reduction of the amount of data, as part of the scene is imaged by both cameras. Although pixel information may be useful for bidirectional interpolation (e.g., mirror reflections towards one of the cameras may be reduced), in fact, information that is not very important may be in the double-coded parts. Therefore, having a map, you can determine which parts of the second image (eg the right image) must be coded (and transmitted) and which parts are less important for the given application. On the receiving side, good quality reconstruction of the missing data can be carried out.
[0015] For example, approximately a straight scene (capture) when the object with a substantially flat surface towards the cameras (which can be located parallel or at a small angle to the scene) and not too close, the missing part of the first (left) image that is captured in the second (right) image, it consists of pixels of the background object (e.g., scene elements in the distance of infinity.
[0016] An interesting embodiment includes encoding a partial second map of divergence or depth or the like. This partial, e.g. depth map, will generally contain the depth values of the area that could not be visualized by the camera first. From this data, depth can then be inferred on the receiving side, which unveiled part belongs to the object of the front plan, the first depth (marked 130 in Fig. 1), and which part belongs to the background (132). This may allow for better interpolation strategies, e.g. the amount of stretching and filling holes may be fine tuned, pseudo perspective ear rendering can be rendered in the intermediate image instead of just the background pixels, etc. Another example is that the trapezoidal distortion of angle cameras can be coded in this second map for compensation on the receiving side.
[0017] In the case of trapezoidal distortions from camera captures (typically slightly convergent), a vertical divergence will generally occur in addition to the horizontal. The vertical component can be coded vector or in the second map, as already foreseen eg in the proposal for "representation of additional data" in the subgroup MPEG-4 Video-3DAV (eg ISO / IEC JTC1 / SC29 / WG11, doc. MPEG2005 / 12603, 12602, 12600, 12595). The discrepancy components may be mapped in the luminance and / or additional image chrominance, e.g. the horizontal discrepancy may be mapped with a high luminance discrepancy, and vertical discrepancies may be mapped using a scheme in one or two chrominance components (so that part of the data is in U , and through mathematical division, the same amount of additional data in V).
[0018] The advantages of the partial left + right + "depth" format relative to the coding for middle view + "depth" + bilateral occlusion data are as follows: Converting the occlusion data to the center view instead of storing it in the original view of the capture camera leads to inaccurate processing (in particular if depth map (map) depth is automatically acquired and of lower quality / consistency, including temporal and spatial imperfections), and thus low coding efficiency, also in the calculation of the intermediate view additional inaccuracies will appear.
10019] These and other aspects of the method and apparatus of the invention will become apparent on the basis and explained with reference to the embodiments and embodiments described below and with reference to the accompanying drawings which serve only as non-limiting specific illustrations exemplifying a more general concept and wherein the hyphens are used for indicating that the component is optional, where components without a hyphen are not necessarily key.
[0020] In the drawings:
Fig. 1 schematically illustrates scene capture using at least two cameras;
Fig. 2 schematically shows several options for encoding the required data in the image signal;
Fig. 3 schematically illustrates an illustrative device for generating an image signal; and
Fig. 4 schematically shows an illustrative receiving device capable of using a signal.
[0021] Fig. 1 shows the first camera 101 catching the first scene image comprising a nearer object 110 and a distal object 112. Its field of view is delimited by lines 103 and 104. Its background view is obscured by the proximal object, namely the region 132 on the left side of the tangent 120 it is not visible. The second camera 102, however, is capable of capturing portions of this region 132 in a second image which for simplicity can be considered and called a right image (but should not be interpreted less than if captured a bit more to the right from the second image), the second camera is also capable of capturing a further portion 130 of a proximal object 110.
[0022] Fig. 2 shows symbolically how these captured images will look like a set of pixels. The image signal 200 may, for example, have a designated JPEG encoding format and include a coded photograph of the scene, or it may be encoded in the MPEG-4 format of the movie shot. In the latter case, 3D data 210 contains the required information to reconstruct the scene at a given point in time.
[0023] The image 220 is a left image captured by the first camera including the object 110 and the background 112.
[0024] Map 222 is a map including any information regarding how objects are located in three-dimensional space, including at least the information required to render a number of required views (statically or dynamically, e.g. in interaction with a moving user in the game) on display. Several such representations are possible, e.g. it may be a depth map containing, for example, an orthogonal approximate (e.g., the average for all areas of the object) the distance of the object in the background from the center of the camera, in their two-dimensional positions as seen by the first camera, or it may be a discrepancy or parallax, or only the horizontal component of the discrepancy.
[0025] Depth and parallax etc. may be mathematically related to each other.
[0026] This depth map may e.g. be an exact pixel, or it may have one value for each block of 8x8 pixels and can be coded e.g. as an image.
[0027] Additional information may be added to the depth map (which may include scalars or multiples into a set of pixels, the set may potentially contain only one pixel), such as e.g. accuracy data (describing how reliable some part of the depth map is) designated based on the matching algorithm for its acquisition.
The partial data structure 223 (part of the right image 224) includes background pixel information (e.g., only luminance, or color, or any other common representation, such as a texture model capable of generating pixels in the area) that can only be seen through the camera the second one (with the nearer object 225 moved by the parallax). This coded partial area or at least the data required to acquire pixel values in a portion of the larger coded shape of the area according to the image patch algorithm may be slightly smaller than the actual exposed area captured by the right image, when used on the receiving side, it may tolerate some missing pixels , for example by generating them using simple extrapolation, stretching, etc.
The coded area can also be larger (e.g., up to twice the width and a similar buffer size in the vertical direction). This may be of interest for example in the case of uncertainty as to the accuracy of the shape in the case of automatic acquisition, or in the case where bi-directional interpolation is desired for some reasons.
[0030] It can also be for coding reasons. It can be cheaper to encode whole blocks and you can benefit from additional coded pixels, while encoding a compound shape can be expensive. In addition, on the transmitting side, manual or semi-automatic analysis may be performed on the right image data, which is proposed as the result of the previous acquisition stage, which may be useful in complementing the data in the left image. For example, you can look at the properties of a pixel to identify specular reflection and decide to encode the pixel area containing the reflection in both images.
[0031] Also, the shape of the areas of differences can be analyzed by means of morphological analysis, in particular the size or width of the area can be determined. Small areas may contain significant coding overhead, but can often be approximated on the receiving side in the absence or with little information. Hence, small areas can be omitted in the partial second image. This can be under the control of the operator who checks the effect of each removal.
[0032] (The surrounding or exact) shape of the area may e.g. be encoded with a polygonal approximation or a limiting frame, and the internal pixel values (texture) may be directly coded, or with the coefficients of linear transformation representation by shape or another mathematical model. Also inversely, parts that do not have to be encoded / transmitted may be indicated.
The partial representation may be mapped (e.g., a simple shift in blanking lines, morphing or sub-blocking, which are re-arranged according to a pre-determined order) on the user image or data (e.g., regeneration model) not used in the image first.
[0034] If the first image with a depth map accompanying it is a middle image, for each page there may be partial second images, i.e. a certain angular distance (baseline) between which it can be interpolated.
[0035] The first camera may image the background, and the second camera may image the background with, e.g. a message presenter obscuring its part, e.g. from the same point of view at a different time. Ie. cameras do not have to be actual cameras present at the same time at a certain time, but rather, for example, one of the views can be eg taken from the image store.
[0036] Optionally, at least for the portion around the imaged exposed areas of the object in the second image, the second depth map 239 (part of a depth map 240), or a similar representation, can be added to the signal. This depth map can contain the boundary between a near and far object. With this information, the receiving side can attach different pixels to the appropriate object / depth layers during interpolation.
[0037] Further data 230 may also be added to the signal - e.g., in proprietary fields such as information about the separation or generally three-dimensional composition of the objects in the scene. The indication can be as simple as the line following the image of the scene abroad (if, for example, a depth map is not sufficient or sufficiently accurate for self-delimiting objects) or even something as complex as the mesh (eg local depth structure in exposed parts), or information obtained from them.
[0038] Also included may be camera position information and stage range information enabling the receiving side to perform more advanced reconstructions of many (at least two) views.
[0039] Fig. 3 shows a device 310 for generating an image signal. Usually it will be a chip or part of an integrated circuit, or a processor with the appropriate software. The device may be contained in a larger device, such as a dedicated authoring device in the studio, and may be attached to a computer or may be included in a computer. In the exemplary embodiment, the first camera 301 and the second camera 302 are connected to the input of the device 310. Each camera has a distractor (308, 309, respectively) that can use, for example, a laser beam or a throwing grid, etc.
[0040] The apparatus has a discrepancy estimation unit 312 that is adapted to determine the discrepancy between the at least two images, at least by considering the geometry of the object (by using the information in the depth map). Various techniques for estimating discrepancies are known in the prior art, e.g. by the sum of absolute differences in pixel values in related blocks ..
[0041] It is adapted to determine at least which areas are present on only one of the images and which are present on both, but additionally may comprise assemblies that are capable of applying matching criteria for areas and pixels.
[0042] There may be a depth mapmap 314 capable of generating and / or analyzing and / or refining depth maps (or a similar representation as a discrepancy map) or determined by a discrepancy estimation unit 312 or extracted from an inputted camera signal including, for example, data distances. Optionally, a rendering unit 316 may be included that may generate intermediate views, for example, such that the studio artist may check the effect of any modifications and / or more efficient coding. This is accomplished by a user interface unit 318 that may, e.g., allow the user to change the value in the partial representation 223, or change its shape (e.g., by enlarging or reducing it). The user can also modify the map 222. The 335 display and the user input means can be attached to this. 8
The apparatus is capable of transmitting the ultimately composed image signal to the network 330 by means of signal transmission and composition means 339 that the skilled person can find for the respective network (e.g., conversion to a television signal includes upward conversion to transmission frequency, internet transmission includes packetizing, furthermore there are defect protection teams, etc.).
[0143] The movie network should not be interpreted in a limiting manner and should also include i.e. transmission to a memory unit or storage medium by, for example, an internal network of the device, such as a bus.
[0044] Fig. 4 shows an illustrative receiver 400, which can again be e.g. an integrated circuit (part) and which contains means for extracting relevant information from an image signal that can be received from the network 330, at least:
means (402) adapted to extract the image of the first (220) pixel values representing one or more objects (110, 112) captured by the first camera (101);
means (404) adapted to be extracted from the image signal (200) of the map, e.g. a depth map corresponding to the positions of the first picture objects; and
- means (406) adapted to extract the partial representation (223) of the second image (224) of the pixel values representing one or more objects (110, 112) captured by the second camera (102).
[0045] Of course, further means may be present, because the receiver (and the extraction method) may reflect any of the potential embodiments for generating, so that for example, there may be means for extracting further data, such as an indication of the boundary between two objects.
[0046] This acquired information is sent to the image regenerator, which may generate e.g. a full left and right image. The image rendering unit 412 can generate e.g. an intermediate view (e.g., by one or two-way interpolation, or any other known algorithm), or signals required for two views (stereo) on the autostereoscopic display. Depending on the type of 3D display and how 3D is actually represented, these two units can be implemented in different combinations.
The receiver can usually be combined with or included in the 3D display 415, which can render at least two views, or the regenerated signal can be stored in a memory device 420, e.g. a disk write module 422, or in semiconductor memory, etc.
[0048] The algorithmic components disclosed in this text can practically be (in whole or in part) implemented in hardware (e.g., in parts of a specialized integrated circuit) or in the form of software operating on a special digital signal processor, e.g. a generic processor, etc.
[0049] A computer program product should be understood as the physical implementation of a set of instructions that enables a processor - generic or special purpose after a series of loading steps (which may include intermediate conversion steps, such as switching to an intermediate language and final processing language). implementation of any of the characteristic functions of the invention. In particular, a computer program product can be implemented as data on a medium such as a disk or tape, data present in memory, data moved over a network link - wired or wireless - or program code on paper. In addition to the program code, the specific data required by the program can also be implemented as a computer program product.
[0050] Some of the steps required for the operation of the method may already be present in the processor's functionality, instead of those described in the computer program product, such as data entry or output steps.
[0051] It should be noted that the above-mentioned embodiments illustrate and do not limit the invention. In addition to the combination of elements of the invention joined in the claims, other combinations of elements are possible. Any combination of elements can be implemented in one dedicated element.
[0052] No reference in the brackets in the claim claim is intended to limit the claim. The word "comprises" does not exclude the presence of elements or aspects not specified in the claim. The applied singular in relation to an element does not exclude the presence of many such elements.
22 members in 10 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 06112096 | European Patent Office (EPO) | A | |
| 06112096 | – | – | – |
| EP20060112096 | – | – | – |
Members22
| Document | Office | Kind | |
|---|---|---|---|
| WO2007113725A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2007113725A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP2005757A2 | European Patent Office (EPO) | A2 | |
| KR20090007384A | Republic of Korea | A | |
| CN101416520A | China | A | |
| JP2009531927A | Japan | A | |
| RU2008143205A | Russian Federation | A | |
| US2010231689A1 | United States of America | A1 | |
| RU2431938C2 | Russian Federation | C2 | |
| CN101416520B | China | B | |
| JP5317955B2 | Japan | B2 | |
| KR101340911B1 | Republic of Korea | B1 | |
| EP2005757B1 | European Patent Office (EPO) | B1 | |
| EP3104603A1 | European Patent Office (EPO) | A1 | |
| ES2599858T3 | Spain | T3 | |
| PL2005757T3This record | Poland | T3 | |
| EP3104603B1 | European Patent Office (EPO) | B1 | |
| US9986258B2 | United States of America | B2 | |
| ES2676055T3 | Spain | T3 | |
| TR201810139T4 | Türkiye | T4 | |
| EP3104603B2 | European Patent Office (EPO) | B2 | |
| ES2676055T5 | Spain | T5 |
Numbers
- Publication
- 2005757
- Publication, DOCDB
- 2005757
- Publication, EPODOC
- PL2005757T
- Application
- 7735242
- Application, DOCDB
- 07735242
- Application, EPODOC
- PL07735242T
Titles2
- English
- EFFICIENT ENCODING OF MULTIPLE VIEWS
- Polish
- WYDAJNE KODOWANIE WIELU WIDOKÓW
Classification
- CPC, 6
- H04N19/597
- H04N2213/005
- H04N19/00
- H04N13/161
- H04N13/128
- H04N13/00
- IPC, 2
- H04N19 597
- H04N13 00