Method and system for randomly accessing multiview videos
9 claims: 2 independent, 7 dependent
- 1マルチビュービデオにランダムにアクセスする方法であって、 複数のマルチビュービデオについて参照ピクチャバッファ内に参照ピクチャを保持するステップであって、該参照ピクチャバッファ に保持する参照ピクチャ は、前記複数のマルチビュービデオの時間参照ピクチャ及び空間参照ピクチャからなるステップと、 Vフレームは、前記時間参照ピクチャを用いることなく前記空間参照ピクチャのみを用いて予測されるフレームであって、前記マルチビュービデオのうちの特定の1つについて、前記時間参照ピクチャを用いることなく、前記空間参照ピクチャのみを用いて前記Vフレームを予測するステップと、 前記Vフレームを処理する際に、前記マルチビュービデオのうちの前記特定の1つに関連する全ての時間参照ピクチャを前記参照ピクチャバッファから削除するステップとを含む、マルチビュービデオにランダムにアクセスする方法。
- 2複数の符号化されたマルチビュービデオに対応するビットストリームを受け取ることをさらに含む、請求項1に記載の方法。
- 3前記Vフレームは、デコーダにおける前記マルチビュービデオへのランダムアクセスを可能にする、請求項1に記載の方法。
- 4前記マルチビュービデオにおけるVフレームの配置は、エンコーダによって決められる、請求項1に記載の方法。
- 5前記参照ピクチャバッファは合成参照ピクチャをインデックス付けし、前記Vフレームは該合成参照ピクチャから予測される、請求項1に記載の方法。
- 6前記Vフレームは低域フレームから予測される、請求項1に記載の方法。
- 7前記マルチビュービデオのうちの選択された1つの中の特定のフレームへの前記ランダムアクセスは、該マルチビュービデオのうちの該選択された1つの中の該特定のフレームに先行する前記Vフレームを復号化し、次に、該マルチビュービデオのうちの該選択された1つの中の以後のフレームを復号化することによって提供される、請求項3に記載の方法。
- 8近傍マルチビュービデオからの参照ピクチャが、前記マルチビュービデオのうちの前記特定の1つの中の前記以後のフレームの予測のために復号化される、請求項7に記載の方法。
- 9マルチビュービデオにランダムにアクセスするシステムであって、 複数のマルチビュービデオについて参照ピクチャバッファ内に参照ピクチャを保持する手段であって、該参照ピクチャバッファ に保持する参照ピクチャ は、前記複数のマルチビュービデオの時間参照ピクチャ及び空間参照ピクチャからなる手段と、 Vフレームは、前記時間参照ピクチャを用いることなく前記空間参照ピクチャのみを用いて予測されるフレームであって、前記マルチビュービデオのうちの特定の1つについて、前記時間参照ピクチャを用いることなく、前記空間参照ピクチャのみを用いて前記Vフレームを予測する手段と、 前記Vフレームを処理する際に、前記マルチビュービデオのうちの前記特定の1つに関連する全ての時間参照ピクチャを前記参照ピクチャバッファから削除する手段とを備える、マルチビュービデオにランダムにアクセスするシステム。
Independent claims9
103 paragraphs, as filed
The present invention relates comprehensively to coding and decoding of multi-view video, and particularly to randomly accessing the multi-view video.
Coding and decoding of multi-view video is essential for applications such as 3D television (3DTV), free-viewpoint television (FTV), and surveillance with multiple cameras. Coding and decoding of multi-view video is also known as dynamic lightfield compression.
FIG. 1 shows a prior art "simulcast" system 100 for encoding multi-view video. Cameras 1 to 4 acquire the frame sequence of scene 5, that is, videos 101 to 104. Each camera has a different view of the scene. Each video is individually encoded 111-114 to become the corresponding encoded video 121-124. This system uses traditional 2D video coding techniques. Therefore, the system does not correlate different videos obtained from different viewpoints by the camera when predicting the frame of the coded video. Individual coding reduces compression efficiency and thus increases network bandwidth and storage.
FIG. 2 shows a prior art parallax compensation prediction system 200 that uses correlation between views. The videos 201 to 204 are encoded 211 to 214 to become the encoded videos 231 to 234. The videos 201 and 204 are individually encoded using a standard video encoder such as MPEG-2, or H.264, also known as MPEG-4 Part 10. These individually encoded videos are "reference" videos. The remaining videos 202 and 203 are encoded using time predictions and interview predictions based on the reconstructed reference videos 251 and 252 obtained from decoders 221 and 222. This prediction is usually adaptively determined block by block (SC Chan et al., "The data compression of simplified dynamic light fields" (Proc. IEEE Int. Acoustics, Speech, and Signal Processing Conf., April, 2003)). ..
Figure 3 shows the "lifting-based" wavelet decomposition of the prior art (W. Sweldens, "The data compression of simplified dynamic light fields" (J. Appl. Comp. Harm. Anal., Vol. 3, no. 2). , Pp. 186-200, 1996)). Wavelet decomposition is an effective technique for static lightfield compression. The input sample 301 is divided into an odd sample 302 and an even sample 303. Odd samples are predicted 320 from even samples. The prediction error forms the high frequency sample 304. This high frequency sample is used to update the even sample 330 to form the low frequency sample 305. Since this decomposition is reversible, linear or non-linear operations can be incorporated into predictive and update steps.
The lifting scheme enables motion compensation time conversion, or motion compensation time filtering (MCTF), which, in the case of video, filters substantially along the trajectory of the motion in time. A review of MCTF for video coding in Ohm et al., "Interframe wavelet codingmotion picture representation for universal scalability" (Signal Processing: Image Communication, vol. 19, no. 9, pp. 877-908, October 2004). Has been done. The lifting scheme can be based on any wavelet nucleus such as Haar or 5/3 Dovesy and any motion model such as block-based translation or affine global motion without affecting the reconstruction.
For coding purposes, the MCTF breaks the video into high and low frames. These frames are then spatially transformed to reduce the remaining spatial correlation. The converted low-frequency frame and high-frequency frame are entropy-coded together with the related motion information to form an encoded bit stream. The MCTF can be implemented with temporally adjacent videos as input using the lifting method shown in FIG. The MCTF can also be applied iteratively to the output low frequency frame.
The compression efficiency of MCTF-based video is comparable to that of video compression standards such as H.264 / AVC. Video also has inherent time scalability. However, this method cannot be used for direct coding of multi-view video with correlations between videos obtained from multiple views. This is because there is no efficient view prediction method that considers the temporal correlation.
Lifting techniques have also been used to encode static light fields, i.e., a single multi-view image. Instead of motion-compensated time filtering, the encoder performs parallax-compensated view-to-view filtering (DCVF) between stationary views in the spatial region (Chang et al., "Inter-view wavelet compression of light fields with disparity compensated lifting" (SPIE Conf). on Visual Communications and Image Processing, See 2003)). For coding, DCVF decomposes the static light field into high-frequency and low-frequency images, and then spatially transforms these images to reduce the remaining spatial correlation. The transformed image is entropy-encoded with the relevant parallax information to form a coded bitstream. DCVF is usually performed using an image acquired from spatially adjacent camera views as input using a lifting-based wavelet transform method as shown in FIG. DCVF can also be applied iteratively to the output low frequency image. DCVF-based static write-field compression provides higher compression efficiency than encoding multiple frames individually. However, this method also fails to encode multi-view video that uses both temporal and spatial correlation between views. This is because there is no efficient view prediction method that considers the temporal correlation.
<p> A method and system for decomposing a multi-view video acquired for a scene by multiple cameras is presented.</p>
<p> Each multi-view video contains a frame sequence and each camera provides a different view of the scene.</p><p> One prediction mode is selected from the time prediction mode, the spatial prediction mode, the view composition prediction mode, and the intra prediction mode.</p><p> The multiview video is then decomposed into low-frequency frames, high-frequency frames, and side information according to the selected prediction mode.</p><p> New videos that reflect the composite view of the scene can also be generated from one or more of the multi-view videos.</p><p> In particular, one embodiment of the invention provides a method of randomly accessing a multiview video. Multi-view video is captured for a scene by corresponding cameras arranged in multiple orientations so that there is overlap of views between any pair of cameras. V-frames are generated from multi-view video. V-frames are coded using only spatial prediction. V-frames are then periodically inserted into the encoded bitstream to provide random time access to the multi-view video.</p>
One embodiment of the present invention provides a complex time / view processing method for encoding and decoding frames of a multi-view video. Multi-view video is video captured by multiple cameras with different postures for a scene. In the present invention, the camera orientation is defined as both a 3D (x, y, z) position and a 3D (θ, ρ, φ) orientation. Each posture corresponds to a "view" of the scene.
The method has a temporal correlation between frames in each video acquired for a particular camera orientation and a spatial correlation between synchronized frames in video acquired from multiple camera views. Also, as described below, "composite" frames can be correlated.
In one embodiment, the temporal correlation uses motion compensation time filtering (MCTF) and the spatial correlation uses parallax compensation view-to-view filtering (DCVF).
In another embodiment of the invention, spatial correlation uses the prediction of one view from a composite frame generated from a "neighborhood" frame. Neighboring frames are one or more frames that are temporally or spatially adjacent, eg, frames before and after the current frame in the time domain, or at the same time, but from cameras with different poses or scene views. It is a frame.
Each frame of each video contains a pixel macroblock. Therefore, the method for coding and decoding multi-view video according to one embodiment of the present invention is macroblock adaptive. Coding and decoding of the current macroblock within the current frame is performed using several possible prediction modes, including various forms of time prediction, spatial prediction, view synthesis prediction, and intra prediction. In order to determine the best prediction mode for each macroblock, one embodiment of the invention provides a method of selecting the prediction mode. This method can be used for any number of camera arrangements.
Describes how to manage the reference picture list for compatibility with existing single-view coding and decoding systems. Specifically, the present specification describes a method of inserting and deleting a reference picture from the picture buffer according to the reference picture list. Reference pictures include time reference pictures, spatial reference pictures, and composite reference pictures.
As used herein, a reference picture is defined as any frame used to "predict" the current frame during coding and decoding. Usually, the reference picture is spatially or temporally adjacent to, or "neighbored", to the current frame.
It is important to note that the same operation applies to both encoders and decoders, as the same set of reference pictures is used to encode and decode the current frame at any given time. is there.
One embodiment of the invention allows random access to frames of multi-view video during coding and decoding. This increases the coding efficiency.
MCTF / DCVF decomposition FIG. 4 shows the MCTF / DCVF decomposition 400 according to one embodiment of the present invention. Frames of input videos 401-404 are acquired by cameras 1-4 with different orientations for a scene 5. As shown in FIG. 8, some of the cameras 1a and 1b are in the same position, but may be in different orientations. It is assumed that there is some amount of view overlap between any pair of cameras. The camera orientation can change during the acquisition of multi-view video. Normally, the cameras are synchronized with each other. Each input video provides a different "view" of the scene. Input frames 401 to 404 are sent to the MCTF / DCVF decomposition 400. This decomposition produces a coded low-frequency frame 411, a coded high-frequency frame 412, and associated side information 413. The high-frequency frame uses the low-frequency frame as a reference picture to encode the prediction error. The decomposition is performed according to the selected prediction mode 410. Prediction modes include spatial prediction mode, time prediction mode, view composition prediction mode, and intra prediction mode. The prediction mode can be adaptively selected for each macroblock for each current frame. When using intra-prediction, the current macroblock is predicted from other macroblocks in the same frame.
FIG. 5 shows a preferable alternating grid pattern of the low frequency frame (L) 411 and the high frequency frame (H) 412 in the vicinity of the frame 510. These frames have a spatial (view) dimension 501 and a temporal dimension 502. In essence, this pattern is such that the low and high frames alternate in spatial dimensions at each time, and the low and high frames alternate in time for each video.
This grid pattern has several advantages. This pattern provides scalability in space and time when the decorator reconstructs only the low frequency frame by uniformly distributing the low frequency frame in both spatial and temporal dimensions. This pattern also aligns high-frequency frames with adjacent low-frequency frames in both spatial and temporal dimensions. This maximizes the correlation between the reference pictures for predicting the error in the current frame, as shown in FIG.
According to the lifting-based wavelet transform, the high frequency frame 412 is generated by predicting one sample set from the other sample set. This prediction can be achieved using several modes, including various forms of time prediction, various forms of spatial prediction, and view composition prediction according to embodiments of the present invention described below.
The means for predicting the high frequency frame 412 and the information necessary for making this prediction are called side information 413. When performing time prediction, the time mode is signaled as part of the side information along with the corresponding motion information. When performing spatial prediction, the spatial mode is signaled as part of the side information along with the corresponding parallax information. When performing view composition prediction, the view composition mode is signaled as part of the side information along with the corresponding parallax information, motion information and depth information.
As shown in FIG. 6, the prediction of each current frame 600 uses the neighborhood frame 510 in both the spatial dimension and the temporal dimension. The frame used to predict the current frame is called a reference picture. The reference picture is kept in a reference list that is part of the coded bitstream. The reference picture is stored in the decrypted picture buffer.
In one embodiment of the invention, the MCTF and DCVF are adaptively applied to each current macroblock for each frame of the input video, decomposed low-frequency frames, and high-frequency frames and related sides. Produce information. Thus, each macroblock is adaptively processed according to the "best" prediction mode. The optimal method for selecting the prediction mode will be described later.
In one embodiment of the invention, the MCTF is first applied individually to each video frame. The resulting frame is then further decomposed by DCVF. In addition to the final decomposed frame, the corresponding side information is also generated. If done on a macroblock basis, the choice of MCTF and DCVF prediction modes will be considered separately. As an advantage, this predictive mode selection essentially supports time scalability. In this way, the lower time rate of video can be easily accessed in the compressed bitstream.
In another embodiment, the DCVF is first applied to the frame of the input video. The resulting frame is then temporally decomposed by MCTF. In addition to the final decomposed frame, side information is also generated. If done on a macroblock basis, the choice of MCTF and DCVF prediction modes will be considered separately. As an advantage, this choice essentially supports spatial scalability. Thus, a reduced number of views in the compressed bitstream can be easily accessed.
The decomposition described above can be iteratively applied to the resulting set of low frequency frames from the previous decomposition step. As an advantage, the MCTF / DCVF resolution 400 of the present invention can effectively remove both temporal and spatial (between views) correlations and achieve very high compression efficiencies. The compression efficiency of the multi-view video encoder of the present invention is superior to conventional simulcast coding, which encodes each video of each view independently.
MCTF / DCVF decomposition coding As shown in FIG. 7, the outputs 411 and 412 of the decomposition 400 are supplied to the signal encoder 710, and the output 413 is supplied to the side information encoder 720. The signal encoder 710 performs conversion, quantization, and entropy coding to remove the correlation remaining in the decomposed low-frequency frame 411 and high-frequency frame 412. Such operations are known in the art (Netravali and Haskell, "Digital Pictures: Representation, Compression and Standards" (Second Edition, Plenum Press, 1995)).
The side information encoder 720 encodes the side information 413 generated by the decomposition 400. In addition to the prediction mode and the reference picture list, the side information 413 includes motion information corresponding to time prediction, parallax information corresponding to spatial prediction, and view composition information and depth information corresponding to view composition prediction.
Side information can be encoded in the MPEG-4 Visual standard ISO / IEC 14496-2 "Information technology --Coding of audio-visual objects --Part 2: Visual" (2nd edition, 2001), or more recent H.264. It can be achieved by known and established techniques such as the / AVC standard and the techniques used in ITU-T Recommendation H.264 "Advanced video coding for generic audiovisual services" (2004).
For example, the motion vector of a macroblock is usually encoded using a prediction method of obtaining a prediction vector from the vector in the macroblock in the reference picture. Next, an entropy coding process is applied to the difference between the prediction vector and the current vector. This process typically uses prediction error statistics. The parallax vector can be encoded using a similar procedure.
In addition, encoding the depth information of each macroblock using a predictive coding method that obtains the predicted value from the macroblock in the reference picture, or simply by using a fixed length code to directly represent the depth value. Can be done. If pixel-level depth precision is extracted and compressed, texture coding techniques that apply transformation techniques, quantization techniques and entropy coding techniques can be applied.
The encoded signals 711 to 713 from the signal encoder 710 and the side information encoder 720 can be multiplexed 730 to generate the encoded output bitstream 731.
Decoding MCTF / DCVF decomposition Bitstream 731 can be decrypted 740 to generate output multiview video 741 corresponding to input multiview videos 401-404. You can optionally generate a composite video as well. In general, the decoder performs the reverse operation of the encoder to reconstruct the multi-view video. Once all low and high frames have been decoded, a complete set of frames with coding quality is reconstructed and available in both the spatial (view) and temporal dimensions.
Depending on how many iteration levels of decomposition were applied in the encoder and what type of decomposition was applied, the reduced number of videos and / or reduced time rates were decoded as shown in Figure 7. can do.
View composition As shown in FIG. 8, view compositing is the process of generating compositing video frame 801 from one or more actual multiview video frames 803. In other words, view compositing provides a means of compositing the frame 801 corresponding to the new selected view 802 in scene 5. This new view 802 may correspond to a "virtual" camera 800 that does not exist at the time the input multiview videos 401-404 are acquired, or may correspond to the captured camera view, and thus. The composite view is used for its prediction and encoding / decoding as described below.
When using one video, the composition is based on extrapolation or warping, and when using multiple videos, the composition is based on interpolation.
Given the pixel values of frame 803 of one or more multi-view videos and the depth values of multiple points in the scene, the pixels in frame 801 of composite view 802 are combined from the corresponding pixel values in frame 803. can do.
View compositing is commonly used in computer graphics to render still images for multiple views (see Buehler et al., "Unstructured Lumigraph Rendering" (Proc. ACM SIGGRAPH, 2001)). This method requires external and internal parameters of the camera.
View compositing for compressing multi-view video is new. In one embodiment of the invention, a synthetic frame is generated that is used to predict the current frame. In one embodiment of the invention, the composite frame is generated for a designated high frequency frame. In another embodiment of the invention, the composite frame is generated for a particular view. The composite frame acts as a reference picture, and the current composite frame can be predicted from these reference pictures.
One problem with this technique is that you don't know the depth value for scene 5. Therefore, the present invention uses known techniques to estimate depth values, for example, based on the correspondence of features in multiview video.
Alternatively, for each composite video, the present invention generates a plurality of composite frames corresponding to the candidate depth values. For each macroblock of the current frame, the best matching macroblock from the set of composite frames is obtained. The composite frame in which this best match is found indicates the depth value of that macroblock of the current frame. This process is repeated for each macroblock in the current frame.
The difference between the current macroblock and the composite block is encoded and compressed by the signal encoder 710. The side information in this multi-view mode is encoded by the side information encoder 720. The side information is a signal showing the view composition prediction mode, the depth value of the macroblock, and the optional displacement vector to compensate for the misalignment between the macroblock in the current frame to be compensated and the best matching macroblock in the composition frame. Including.
Prediction mode selection In the macroblock adaptive MCTF / DCVF decomposition, the prediction mode m of each macroblock can be selected by adaptively minimizing the cost function for each macroblock.
<maths num="1"><img file="JP5106830B2_D0001.tif" /></maths>
Where J (m) = D (m) + λR (m), D is the strain, λ is the weight parameter, R is the rate, m is the set of candidate prediction modes, and m<sup>*</sup>Indicates the optimal prediction mode selected based on the minimum cost criterion.
Candidate mode m includes various time prediction modes, spatial prediction modes, view composition prediction modes, and intra prediction modes. The cost function J (m) depends on the rate and distortion that results from coding the macroblock with a particular prediction mode m.
Strain D measures the difference between the reconstructed macroblock and the original macroblock. The reconstructed macroblock is obtained by encoding and decoding the macroblock using a given prediction mode m. A common distortion measure is the sum of squares of differences. The rate R corresponds to the number of bits required to encode the macroblock, including prediction error and side information. The weight parameter λ controls the rate-distortion trade-off of macroblock coding and can be derived from the quantization step size.
Detailed aspects of the coding process and the decoding process will be described in more detail below. In particular, the various data structures used by the coding and decoding processes will be described. It should be understood that the data structures used in the encoder are the same as the corresponding data structures used in the decoder, as described herein. It should also be understood that the processing steps of the decoder essentially follow the same processing steps as the encoder, but in reverse order.
Reference picture management FIG. 9 shows reference picture management for a prior art single-view coding and decoding system. The time-referenced picture 901 is managed by the single-view reference picture list (RPL) manager 910, which determines the insertion 920 and deletion 930 of the time-referenced picture 901 into the decryption picture buffer (DPB) 940. The reference picture list 950 is also retained to indicate the frames stored in the DPB940. The RPL is used for reference picture management operations such as insert 920 and delete 930, as well as time prediction 960 in both encoders and decoders.
In a single-view encoder, the time-referenced picture 901 applies a set of normal coding operations, including prediction, transformation and quantization, and then performs operations including these inverse, inverse quantization, inverse transformation and motion compensation. Generated as a result of application. In addition, the time reference picture 901 is inserted into the DPB940 and added to the RPL950 only when the time picture is needed to predict the current frame in the encoder.
In a single-view decoder, the same time-referenced picture 901 is generated by applying a set of normal decoding operations, including inverse quantization, inverse transformation, and motion compensation, to the bitstream. Like the encoder, the time-referenced picture 901 is inserted into the DPB940 and added to the RPL950 only when necessary to predict the current frame in the decoder.
FIG. 10 shows reference picture management for multi-view coding and decoding. In addition to the time reference picture 1003, the multiview system also includes a spatial reference picture 1001 and a composite reference picture 1002. These reference pictures are collectively called the multi-view reference picture 1005. These multi-view reference pictures 1005 are managed by the multi-view RPL manager 1010, which determines the insertion 1020 and deletion 1030 of the multi-view reference picture 1005 into the multi-view DPB1040. For each video, a multi-view reference picture list (RPL) 1050 is also retained to indicate the frames stored in the DPB. That is, RPL is an index of DPB. The multi-view RPL is used for reference picture management operations such as insert 1020 and delete 1030, as well as for predicting the current frame 1060.
Note that the multi-view system prediction 1060 differs from the single-view system prediction 960 because it allows predictions from different types of multi-view reference pictures 1005. Further details regarding the multi-view reference picture management 1010 will be described later.
Multiview Reference Picture List Manager A set of multiview reference pictures 1005 can be indicated in the multiview RPL1050 before the encoder encodes the current frame or the decoder decodes the current frame. As defined conventionally and herein, a set may have nothing (empty set), but may have one or more elements. An identical copy of the RPL is maintained by both the encoder and decoder for each current frame.
All frames inserted into the multi-view RPL1050 are initialized and marked as predictable, using proper syntax. According to the H.264 / AVC standard and reference software, the "used_for_reference" flag is set to "1". In general, the reference picture is initialized so that the frame can be used for prediction in a video coding system. A picture order count (POC) is assigned to each reference picture for compatibility with traditional single-view video compression standards such as H.264 / AVC. Typically, for single-view coding and decoding systems, the POC corresponds to the temporal ordering of the pictures, eg frame numbers. For multi-view coding and decoding systems, chronological order alone is not sufficient to assign a POC to each reference picture. Therefore, in the present invention, a unique POC is obtained for all multi-view reference pictures according to a certain rule. One rule is to assign POCs to time-referenced pictures in chronological order, and then reserve a sequence of very high POC numbers, such as 10,000-10,100, for spatial and composite reference pictures. .. Other POC allocation rules, or simply "ordering" rules, are described in more detail below.
All frames used as multi-view reference pictures are held in the RPL and stored in the DPB so that they are treated as conventional reference pictures by the encoder 700 or decoder 740. As a result, the coding process and the decoding process can be performed as before. Further details regarding the storage of the multi-view reference picture will be described later. The RPL and DPB are updated correspondingly for each current frame to be predicted.
Definition of multi-view rules and signal transduction The process of managing the RPL is coordinated between the encoder 700 and the decoder 740. In particular, encoders and decoders keep the same copy of the multiview reference picture list when predicting a particular current frame.
Several rules are possible to manage the multiframe reference picture list. Therefore, the particular rules used are either inserted into bitstream 731 or provided as sequence-level side information, eg, configuration information transmitted to the decoder. In addition, this rule allows sequences to be synthesized using different predictive structures, such as 1D arrays, 2D arrays, arcs, crosses, and view interpolation or warping techniques.
For example, a composite frame is generated by warping the corresponding frame of one of the multiview videos captured by the camera. Alternatively, a traditional model of the scene can be used during synthesis. Other embodiments of the invention define some multi-view reference picture management rules that depend on the view type, insertion order, and camera characteristics.
The view type is whether the reference picture is a frame from a video other than the video in the current frame, or whether the reference picture is a composite from another frame, or whether the reference picture is in another reference picture. Indicates whether it depends. For example, the composite reference picture can be kept separately from the reference picture from the same video as the current frame or the reference picture from the spatially adjacent video.
Insertion order indicates how the referenced pictures are ordered within the RPL. As an example, a reference picture in the same video as the current frame can be given a lower order value than the reference picture in the video taken from the adjacent view. In this case, this reference picture is placed earlier in the multi-view RPL.
The camera characteristic indicates the characteristic of the camera used to acquire the reference picture or the virtual camera used to generate the composite reference picture. These properties include translation and rotation with respect to the fixed coordinate system, ie the "orientation" of the camera, internal parameters that describe how 3D points are projected onto the 2D image, lens distortion, color calibration information, illumination levels, etc. .. As an example, the proximity of a particular camera to an adjacent camera can be automatically determined based on the camera characteristics, and only the video captured by the adjacent camera is considered part of the particular RPL.
As shown in FIG. 11, one embodiment of the present invention reserves a portion 1101 of each reference picture list for the time reference picture 1003 and another portion 1102 for the composite reference picture 1002. Use a rule that reserves part 1103 for spatial reference picture 1001. This is an example of a rule that depends only on the view type. The number of frames included in each portion can vary based on the predictive dependence of the current frame during coding or decoding.
Specific control rules can be specified by standards, explicit or implicit rules, or as side information in a coded bitstream.
Storing pictures in DPB The multi-view RPL manager 1010 holds the RPL so that the order in which the multi-view reference pictures are stored in the DPB corresponds to the "usefulness" of the pictures in increasing the efficiency of coding and decoding. Specifically, the reference picture towards the beginning of the RPL can be predictably encoded with fewer bits than the reference picture towards the end of the RPL.
As shown in FIG. 12, the optimization of the order in which the multi-view reference pictures are kept in the RPL can have a great influence on the coding efficiency. For example, according to the POC assignments described above for initialization, a multi-view reference picture can be assigned a very large POC value. This is because multi-view reference pictures do not occur in the normal temporal ordering of video sequences. Therefore, the default ordering process for most video codecs may place such multi-view reference pictures earlier in the reference picture list.
Default ordering is not desirable because time-referenced pictures from the same sequence usually show stronger correlation than spatial-referenced pictures from other sequences. Thus, the multi-view reference pictures are explicitly sorted by the encoder, and the encoder then signals this sort to the decoder, or the encoder and decoder implicitly sort the multi-view reference pictures according to predetermined rules. ..
As shown in FIG. 13, ordering of reference pictures is facilitated by view mode 1300 for each reference picture. Note that view mode 1300 also affects the multi-view prediction process 1060. In one embodiment of the invention, three different types of view modes, described in more detail below, are used: I view, P view and B view.
Prior to discussing the detailed operation of multi-view reference picture management, FIG. 14 shows prior art reference picture management for a single video coding and decoding system. Only the time reference picture 901 is used for the time prediction 960. The time prediction dependence between the time reference pictures of the video in the acquisition order or the display order 1401 is shown. The reference pictures are sorted in coding order 1402, 1410, in which each reference picture is at time t in this coding order 1402.<sub>0</sub>~ t<sub>6</sub>Is encoded or decoded at. Block 1420 shows the ordering of reference pictures by time. Intraframe I<sub>0</sub>Time t when is encoded or decoded<sub>0</sub>In, the DBP / RPL is empty because the time reference picture is not used for time prediction. One-way interframe P<sub>1</sub>Time t when is encoded or decoded<sub>1</sub>Then, frame I<sub>0</sub>Is available as a time reference picture. Time t<sub>2</sub>And t<sub>3</sub>Then, frame I<sub>0</sub>And P<sub>1</sub>Both are interframe B<sub>1</sub>And B<sub>2</sub>It can be used as a reference frame for bidirectional time prediction. Time-referenced pictures and DBP / RPL are managed for future pictures as well.
To illustrate the case of multi-view according to one embodiment of the invention, consider the three different types of views described above and shown in FIG. 15, ie, I view, P view, and B view. Shows the predictive dependency of the multiview between the reference pictures of the video in display order 1501. As shown in FIG. 15, the reference pictures of the video are rearranged in the coding order 1502 for each view mode 1510, and in this coding order 1502, each reference picture is t.<sub>0</sub>~ t<sub>2</sub>It is encoded or decoded at the time indicated by. The order of the multi-view reference pictures is shown in block 1520 by time.
I-view is the simplest mode that allows for more complex modes. The I-view uses traditional coding and prediction modes that do not use spatial or synthetic prediction. For example, I-views can be encoded using traditional H.264 / AVC techniques without the use of multi-view extensions. When placing spatial reference pictures from an I-view sequence in a reference list in another view, these spatial reference pictures are usually placed after the time reference picture.
As shown in Figure 15, for I view, frame I<sub>0</sub>Is t<sub>0</sub>When encoded or decoded in, the multiview reference picture is not used for prediction. Therefore, DBP / RPL is empty. Frame P<sub>0</sub>Time t when is encoded or decoded<sub>1</sub>Then I<sub>0</sub>Is available as a time reference picture. Frame B<sub>0</sub>Time t when is encoded or decoded<sub>2</sub>Then, frame I<sub>0</sub>And P<sub>0</sub>Both are available as time reference pictures.
P-views are more complex than I-views in that they allow prediction from another view and utilize spatial correlation between views. Specifically, the sequence encoded using the P-view mode uses a multi-view reference picture from another I-view or P-view. A composite reference picture can also be used in the P view. When placing a multiview reference picture from an I view in a reference list of another view, the P view is placed after both the time reference picture and the multiview reference picture derived from the I view.
As shown in Figure 15, for P view, frame I<sub>2</sub>Is t<sub>0</sub>When encoded or decoded in, the composite reference picture S<sub>20</sub>And spatial reference picture I<sub>0</sub>Is available for prediction. Further details regarding the generation of the composite picture will be described later. P<sub>2</sub>Time t when is encoded or decoded<sub>1</sub>Then I<sub>2</sub>Is a composite reference picture S as a time reference picture<sub>21</sub>And the spatial reference picture P from the I view<sub>0</sub>Available with. Time t<sub>2</sub>Now, two time reference pictures I<sub>2</sub>And P<sub>2</sub>, And composite reference picture S<sub>22</sub>And spatial reference picture B<sub>0</sub>Exists, and predictions can be made from these reference pictures.
The B view is similar to the P view in that it uses a multi-view reference picture. One important difference between P-view and B-view is that P-view uses reference pictures from the view itself and one other view, while B-view can refer to pictures in multiple views. That is. When using a composite reference picture, the B view is placed before the spatial reference picture because the composite view usually has a stronger correlation than the spatial reference.
As shown in Figure 15, for B view, I<sub>1</sub>Is t<sub>0</sub>When encoded or decoded in, the composite reference picture S<sub>10</sub>And spatial reference picture I<sub>0</sub>And I<sub>2</sub>Is available for prediction. P<sub>1</sub>Time t when is encoded or decoded<sub>1</sub>Then I<sub>1</sub>Is a composite reference picture S as a time reference picture<sub>11</sub>, And the spatial reference picture P from the I and P views, respectively.<sub>0</sub>And P<sub>2</sub>Available with. Time t<sub>2</sub>Now, two time reference pictures I<sub>1</sub>And P<sub>1</sub>Exists and the composite reference picture S<sub>12</sub>And spatial reference picture B<sub>0</sub>And B<sub>2</sub>Exists, and predictions can be made from these reference pictures.
It should be emphasized that the example shown in FIG. 15 relates to only one embodiment of the present invention. Many different types of predictive dependencies are supported. As an example, spatial reference pictures are not limited to pictures in different views at the same time. Spatial reference pictures can also include reference pictures for different views at different times. Also, the number of bidirectional and unidirectional predictive intrapictures between intrapictures can vary. Similarly, the configurations of I-view, P-view, and B-view can change. In addition, several composite reference pictures, each generated using different picture sets or different depth maps or processes, may be available.
compatibility One important advantage of multi-view picture management according to embodiments of the present invention is compatibility with existing single-view video coding systems and designs. This multi-view picture management not only makes minimal changes to the existing single-view video coding standard, but also describes the software and hardware from the existing single-view video coding system as described herein. It also makes it possible to use it for video coding.
The reason for this is that most traditional video coding systems transmit the coding parameters to the decoder in a compressed bitstream. Therefore, the syntax for transmitting such parameters is specified by existing video coding standards such as the H.264 / AVC standard. For example, a video coding standard defines a prediction mode for a given macroblock in the current frame from other time-related reference pictures. The standard also defines the methods used to encode and decode the resulting prediction errors. Other parameters specify the type or size of the transformation, the quantization method, and the entropy coding method.
Therefore, the multi-view reference picture of the present invention is implemented with only a limited number of modifications to standard coding and decoding components such as reference picture lists, decoded picture buffers, and predictive structures of existing systems. can do. Note that the macroblock structure, transformation, quantization and entropy coding are unchanged.
View composition As described above for FIG. 8, view compositing is the process of generating frame 801 corresponding to composited view 802 of virtual camera 800 from frame 803 obtained from existing video. In other words, view compositing provides a means of compositing frames corresponding to a new selected view of a scene by a virtual camera that does not exist at the time the input video was acquired. Given the pixel values of one or more real video frames and the depth values of the points in the scene, the pixels in the frame of the composite video view can be extrapolated and / or interpolated.
Prediction from composite view FIG. 16 shows the process of generating a reconstructed macroblock using view compositing mode when depth 1901 information is contained in the coded multi-view bitstream 731. The depth of a given macroblock is decoded by the side information decoder 1910. View compositing 1920 is performed using the depth 1901 and the spatial reference picture 1902 to generate the compositing macroblock 1904. Next, the reconstructed macroblock 1903 is formed by adding the composite macroblock 1904 and the decoded residual macroblock 1905 to 1930.
Details of multi-view mode selection in encoder FIG. 17 shows the process of selecting a predictive mode during coding or decoding of the current frame. Motion estimation 2010 is performed for the current macroblock 2011 using the time reference picture 2020. Using the resulting motion vector 2021, the first coding cost cost using time prediction<sub>1</sub>2030 seeking 2031. The prediction mode associated with this process is m<sub>1</sub>Is.
Parallax estimation 2040 is performed for the current macroblock using the spatial reference picture 2041. Using the resulting parallax vector 2042, using spatial prediction, a second coding cost cost<sub>2</sub>2050 to find 2051. The prediction mode associated with this process is m<sub>2</sub>Indicated by.
Depth estimation 2060 is performed for the current macroblock based on the spatial reference picture 2041. View composition is performed based on the estimated depth. Third coding cost cost using view composite prediction with depth information 2061 and composite view 2062<sub>3</sub>2070 seeking 2071. The prediction mode associated with this process is m<sub>3</sub>Is.
A fourth coding cost cost using intra-prediction, using adjacent pixels 2082 of the current macroblock<sub>4</sub>2080 to find 2081. The prediction mode associated with this process is m<sub>4</sub>Is.
cost<sub>1</sub>, Cost<sub>2</sub>, Cost<sub>3</sub>And cost<sub>4</sub>Find the minimum cost in 2090, mode m<sub>1</sub>, M<sub>2</sub>, M<sub>3</sub>And m<sub>4</sub>The mode with the lowest cost is selected as the best prediction mode 2091 of the current macroblock 2011.
View composition using depth estimation View compositing mode 2091 can be used to estimate the depth information and displacement vectors of the compositing view from decoded frames of one or more multi-view videos. The depth information may be the depth of each pixel estimated from the stereo camera or the depth of each macroblock estimated from the macroblock matching, depending on the process to be applied.
The advantage of this approach is that as long as the encoder has access to the same depth and displacement information as the decoder, no depth and displacement vectors are needed in the bitstream, resulting in lower bandwidth. The encoder can achieve this as long as the decoder uses exactly the same depth and displacement estimation process as the encoder. Therefore, in this embodiment of the present invention, the difference between the current macroblock and the synthetic macroblock is encoded by the encoder.
The side information in this mode is encoded by the side information encoder 720. The side information includes a signal indicating the view composition mode and the reference view (s). The side information can also include depth and displacement correction information, which is the difference between the depth and displacement used by the encoder for view compositing and the value estimated by the decoder.
FIG. 18 shows the macroblock decoding process using view compositing mode when depth information is estimated or inferred by the decoder and not transmitted by the coded multi-view bitstream. Depth 2101 is estimated from 2110 from spatial reference picture 2102. Next, view compositing 2120 is performed using the estimated depth and spatial reference pictures to generate compositing macroblock 2121. The reconstructed macroblock 2103 is formed by the addition 2130 of the composite macroblock and the decoded residual macroblock 2104.
Spatial random access Intraframes, also known as I-frames, are typically spaced across the video to provide random access to frames in traditional video. This allows the decoder to access any frame in the decoding sequence, but reduces compression efficiency.
For the multi-view coding and decoding system of the present invention, a new type of frame referred to herein as a "V-frame" is provided to enable random access and improved compression efficiency. V-frames are similar to I-frames in the sense that they are coded without time prediction. However, V-frames also allow predictions from other cameras or from synthetic video. Specifically, a V-frame is a frame in a compressed bitstream predicted from a spatial reference picture or a composite reference picture. By periodically inserting a V-frame into the bitstream instead of an I-frame, the present invention enhances coding efficiency while providing time-random access as is possible with an I-frame. Therefore, V-frames do not use time-referenced frames. FIG. 19 shows the use of an I-frame for the first view and the use of a V-frame for subsequent views at the same time 1900. Note that for the grid configuration shown in Figure 5, V-frames do not occur at the same time for all views. A V frame can be assigned to any of the low frequency frames. In this case, the V frame is predicted from the low frequency frame of the neighborhood view.
The H.264 / AVC video coding standard suggests that an IDR frame similar to an MPEG-2 I-frame with a closed GOP removes all referenced pictures from the decoder picture buffer. As a result, the frame before the IDR frame cannot be used for predicting the frame after the IDR frame.
In the multi-view decoder described herein, it is suggested that V-frames can likewise remove all time-referenced pictures from the decoder picture buffer. However, the spatial reference picture can be left in the decoder picture buffer. As a result, the frame before the V frame in a given view cannot be used to predict the time of the frame after the V frame in the same view.
In order to access one particular frame in a multi-view video, the V-frame for that view must first be decoded. As mentioned above, this can be achieved by prediction from spatial or composite reference pictures without the use of time reference pictures.
After decoding the V frame of the selected view, the subsequent frames of that view are decoded. Since these subsequent frames are likely to have a predictive dependency on the reference picture from the neighborhood view, the reference picture in these neighborhood views is also decoded.
Although the present invention has been described as an example of preferred embodiments, it is understood that various other adaptations and modifications can be made within the spirit and scope of the invention. Therefore, the object of the appended claims is to cover all such modifications and modifications that fall within the true spirit and scope of the invention.
<figref num="1">It is a block diagram of the prior art system for encoding a multi-view video.</figref><figref num="2">It is a block diagram of the parallax compensation prediction system of the prior art for encoding a multi-view video.</figref><figref num="3">It is a flow chart of the wavelet decomposition process of the prior art.</figref><figref num="4">It is a block diagram of MCTF / DCVF decomposition by one Embodiment of this invention.</figref><figref num="5">It is a block diagram as a function of time and space of a low-pass frame and a high-pass frame after MCTF / DCVF decomposition according to one embodiment of the present invention.</figref><figref num="6">It is a block diagram of the prediction of the high-pass frame from the adjacent low-pass frame by one embodiment of the present invention.</figref><figref num="7">FIG. 6 is a block diagram of a multi-view coding system using macroblock adaptive MCTF / DCVF decomposition according to one embodiment of the present invention.</figref><figref num="8">It is the schematic of the video synthesis by one Embodiment of this invention.</figref><figref num="9">It is a block diagram of the reference picture management of the prior art.</figref><figref num="10">It is a block diagram of the multi-view reference picture management by one Embodiment of this invention.</figref><figref num="11">FIG. 3 is a block diagram of a multi-view reference picture in a decoded picture buffer according to one embodiment of the present invention.</figref><figref num="12">It is a graph which compares the coding efficiency of the ordering of different multi-view reference pictures.</figref><figref num="13">FIG. 6 is a block diagram of view mode dependencies on a multi-view reference picture list manager according to one embodiment of the present invention.</figref><figref num="14">FIG. 5 is a diagram of prior art reference picture management for a single-view coding system that uses predictions from time reference pictures.</figref><figref num="15">FIG. 5 is a reference picture management diagram for a multi-view coding and decoding system using prediction from a multi-view reference picture according to one embodiment of the present invention.</figref><figref num="16">FIG. 5 is a block diagram of view composition in a decoder using depth information encoded and received as side information according to one embodiment of the present invention.</figref><figref num="17">FIG. 6 is a block diagram of cost calculation for selecting a prediction mode according to one embodiment of the present invention.</figref><figref num="18">FIG. 3 is a block diagram of view composition in a decoder using depth information estimated by the decoder according to one embodiment of the present invention.</figref><figref num="19">FIG. 6 is a block diagram of a multi-view video that achieves spatial random access in a decoder using V-frames according to one embodiment of the present invention.</figref>
20 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP09261653A | Cites | Japan |
| JP2004187265A | Cites | Japan |
96 members in 5 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 11292168 | United States of America | – | |
| 29216805 | United States of America | A | |
| 29216805 | United States of America | A | |
| 2005292168 | – | – | – |
| US20050292168 | – | – | – |
Members96
| Document | Office | Kind | |
|---|---|---|---|
| US2004230706A1 | United States of America | A1 | |
| US2005204069A1 | United States of America | A1 | |
| US2005216617A1 | United States of America | A1 | |
| US7000036B2 | United States of America | B2 | |
| US2006075154A1 | United States of America | A1 | |
| US2006132610A1 | United States of America | A1 | |
| WO2006064710A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2006146138A1 | United States of America | A1 | |
| US2006146141A1 | United States of America | A1 | |
| US2006146143A1 | United States of America | A1 | |
| US7174274B2 | United States of America | B2 | |
| US2007030356A1 | United States of America | A1 | |
| US2007079022A1 | United States of America | A1 | |
| US2007109409A1 | United States of America | A1 | |
| US2007121722A1 | United States of America | A1 | |
| EP1793609A2 | European Patent Office (EPO) | A2 | |
| EP1793610A1 | European Patent Office (EPO) | A1 | |
| EP1793611A2 | European Patent Office (EPO) | A2 | |
| JP2007159111A | Japan | A | |
| JP2007159112A | Japan | A | |
| JP2007159113A | Japan | A | |
| EP1825690A1 | European Patent Office (EPO) | A1 | |
| CN101036397A | China | A | |
| JP2007259433A | Japan | A | |
| JP2008022549A | Japan | A | |
| US2008103754A1 | United States of America | A1 | |
| US2008103755A1 | United States of America | A1 | |
| US2008109580A1 | United States of America | A1 | |
| US7373435B2 | United States of America | B2 | |
| JP2008524873A | Japan | A | |
| JP2008172749A | Japan | A | |
| EP1978750A2 | European Patent Office (EPO) | A2 | |
| US7468745B2 | United States of America | B2 | |
| US7489342B2 | United States of America | B2 | |
| EP1793611A3 | European Patent Office (EPO) | A3 | |
| US7516248B2 | United States of America | B2 | |
| EP1793609A3 | European Patent Office (EPO) | A3 | |
| US7600053B2 | United States of America | B2 | |
| CN100562130C | China | C | |
| EP1978750A3 | European Patent Office (EPO) | A3 | |
| US7671894B2 | United States of America | B2 | |
| US7710462B2 | United States of America | B2 | |
| US7728877B2 | United States of America | B2 | |
| US7728878B2 | United States of America | B2 | |
| US2010322311A1 | United States of America | A1 | |
| US7886082B2 | United States of America | B2 | |
| US7903737B2 | United States of America | B2 | |
| US2011106521A1 | United States of America | A1 | |
| JP2011155683A | Japan | A | |
| US2011194452A1 | United States of America | A1 | |
| US2011200229A1 | United States of America | A1 | |
| JP4762936B2 | Japan | B2 | |
| JP4786534B2 | Japan | B2 | |
| JP2012016044A | Japan | A | |
| JP2012016045A | Japan | A | |
| JP2012016046A | Japan | A | |
| JP4890201B2 | Japan | B2 | |
| US2012062756A1 | United States of America | A1 | |
| US8145802B2 | United States of America | B2 | |
| JP2012114942A | Japan | A | |
| US2012159009A1 | United States of America | A1 | |
| JP4995330B2 | Japan | B2 | |
| JP5013993B2 | Japan | B2 | |
| JP2012230671A | Japan | A | |
| US2012314027A1 | United States of America | A1 | |
| JP5106830B2This record | Japan | B2 | |
| JP5116394B2 | Japan | B2 | |
| JP5154679B2 | Japan | B2 | |
| JP5154680B2 | Japan | B2 | |
| JP5154681B2 | Japan | B2 | |
| US8407373B2 | United States of America | B2 | |
| WO2013073282A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US8451895B2 | United States of America | B2 | |
| JP5274766B2 | Japan | B2 | |
| US2013227178A1 | United States of America | A1 | |
| WO2014010537A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US8639857B2 | United States of America | B2 | |
| US2014101344A1 | United States of America | A1 | |
| EP1825690B1 | European Patent Office (EPO) | B1 | |
| US8823821B2 | United States of America | B2 | |
| US8824548B2 | United States of America | B2 | |
| EP2781090A1 | European Patent Office (EPO) | A1 | |
| US8854486B2 | United States of America | B2 | |
| JP2015502057A | Japan | A | |
| CN104429079A | China | A | |
| US9026689B2 | United States of America | B2 | |
| EP2870766A1 | European Patent Office (EPO) | A1 | |
| JP5744333B2 | Japan | B2 | |
| JP2015519834A | Japan | A | |
| US2015199284A1 | United States of America | A1 | |
| EP1793609B1 | European Patent Office (EPO) | B1 | |
| JP5773935B2 | Japan | B2 | |
| US9229883B2 | United States of America | B2 | |
| CN104429079B | China | B | |
| EP2870766B1 | European Patent Office (EPO) | B1 | |
| EP2781090B1 | European Patent Office (EPO) | B1 |
24 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5106830
- Publication, DOCDB
- 5106830
- Publication, EPODOC
- JP5106830B
- Application
- 305320
- Application, DOCDB
- 2006305320
- Application, EPODOC
- JP20060305320
Titles2
- Japanese
- マルチビュービデオにランダムにアクセスする方法及びシステム
- English
- Random access methods and systems for multi-view video
Classification
- CPC, 13
- H04N19/615
- H04N19/597
- H04N19/105
- H04N19/159
- H04N19/176
- H04N19/147
- H04N19/172
- H04N19/46
- H04N19/13
- H04N19/63
- H04N19/61
- H04N19/103
- H04N19/573
- IPC, 1
- H04N7 32
