Method and apparatus of motion vector prediction for scalable video coding
Abstract
Inter-layer motion mapping information may be used to enable temporal motion vector prediction (TMVP) of an enhancement layer of a bitstream. For example, a reference picture and a motion vector (MV) of an inter-layer video block may be determined. The reference picture may be determined based on a collocated base layer video block. For example, the reference picture may be a collocated inter-layer reference picture of the reference picture of the collocated base layer video block. The MV may be determined based on a MV of the collocated base layer video block. For example, the MV may be determined by determining the MV of the collocated base layer video block and scaling the MV of the collocated base layer video block according to a spatial ratio between the base layer and the enhancement layer. TMVP may be performed on the enhancement layer picture using the MV of the inter-layer video block.

Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
24 claims: 12 independent, 12 dependent
- 1裝置,包括:經由一視訊解碼器接收一位元流,該位元流包括一基礎層以及一增強層;基於該位元流中的一指示,確定是否使用一層間參考圖像或一時間增強層圖像作為一並置圖像(ColPic)以用於一增強層圖像的時間移動向量預測(TMVP);在確定使用該層間參考圖像作為該並置圖像(ColPic)以用於該增強層圖像的TMVP時,經由該視訊解碼器將該層間參考圖像加入用於該增強層圖像的一參考圖像列表,其中該層間參考圖像包括從一基礎層圖像的一紋理確定的一紋理、從該基礎層圖像的一移動向量確定的一移動向量、以及從該基礎層圖像的一參考圖像索引確定的一參考圖像索引;以及使用該層間參考圖像作為該並置圖像(ColPic)以用於該增強層圖像的TMVP以經由該視訊解碼器解碼該增強層圖像。
- 2如申請專利範圍第1項所述的方法,更包括:經由該視訊解碼器上取樣該基礎層圖像的該紋理以確定該層間參考圖像的該紋理;以及經由該視訊解碼器縮放該基礎層圖像的該移動向量以確定該層間參考圖像的該移動向量。
- 3如申請專利範圍第1項所述的方法,其中該層間參考圖像是在一不同時間實例作為該增強層圖像。
- 4如申請專利範圍第1項所述的方法,其中該層間參考圖像是在一相同時間實例作為該增強層圖像。
- 5如申請專利範圍第1項所述的方法,更包括:標記該層間參考圖像的一區域以表明該區域在TMVP中不可用。
- 6如申請專利範圍第5項所述的方法,其中標記該層間參考圖像的該區域包括將與該區域對應的一參考圖像索引設定為-1的一值。
- 7如申請專利範圍第5項所述的方法,其中標記該層間參考圖像的該區域是回應於確定該基礎層圖像中的一對應區域是被內部編碼的而被執行。
- 8如申請專利範圍第1項所述的方法,更包括:經由該視訊解碼器確定該基礎層圖像的一壓縮移動向量;以及基於該基礎層圖像的該壓縮移動向量以經由該視訊解碼器確定該層間參考圖像的該移動向量。
- 9如申請專利範圍第8項所述的方法,其中確定該層間參考圖像的該移動向量包括根據該基礎層與該增強層之間的一空間比率來縮放該基礎層圖像的該壓縮移動向量、以及基於該縮放的移動向量確定該層間參考圖像的該移動向量。
- 10如申請專利範圍第1項所述的方法,更包括:藉由複製該基礎層圖像的一參考圖像列表以經由該視訊解碼器來確定用於該層間參考圖像的一參考圖像列表。
- 11如申請專利範圍第10項所述的方法,更包括:藉由複製該基礎層圖像的該參考圖像列表中的一對應參考圖像的一圖像順序計數(POC)以經由該視訊解碼器來確定該層間參考圖像的該參考圖像列表中的每一參考圖像的一POC。
- 12如申請專利範圍第1項所述的方法,其中解碼包括使用該層間參考圖像的一移動向量來預測該增強層圖像的一移動向量。
- 13一種解碼器,包括:一處理器,被配置為:接收一位元流,該位元流包括一基礎層以及一增強層; 基於該位元流中的一指示,確定是否使用一層間參考圖像或一時間增強層圖像作為一並置圖像(ColPic)以用於一增強層圖像的時間移動向量預測(TMVP);在確定使用該層間參考圖像作為該並置圖像(ColPic)以用於該增強層圖像的TMVP時,將該層間參考圖像加入用於該增強層圖像的一參考圖像列表,其中該層間參考圖像包括從一基礎層圖像的一紋理確定的一紋理、從該基礎層圖像的一移動向量確定的一移動向量、以及從該基礎層圖像的一參考圖像索引確定的一參考圖像索引;以及使用該層間參考圖像作為該並置圖像(ColPic)以用於該增強層圖像的TMVP而使用TMVP解碼該增強層圖像。
- 14如申請專利範圍第13項所述的解碼器,其中該紋理是藉由上取樣該基礎層圖像的該紋理而被確定,以及該移動向量是藉由縮放該基礎層圖像的該移動向量而被確定。
- 15如申請專利範圍第13項所述的解碼器,其中該層間參考圖像是在一不同時間實例作為該增強層圖像。
- 16如申請專利範圍第13項所述的解碼器,其中該層間參考圖像是在一相同時間實例作為該增強層圖像。
- 17如申請專利範圍第13項所述的解碼器,其中該層間參考圖像的一區域被標記以表明該區域在TMVP中不可用。
- 18如申請專利範圍第17項所述的解碼器,其中與該區域對應的一參考圖像索引被設定為-1的一值以表明該區域在TMVP中不可用。
- 19如申請專利範圍第17項所述的解碼器,其中該層間參考圖像的該區域被標記以回應於確定該基礎層圖像中的一對應區域是被內部編碼的。
- 20如申請專利範圍第13項所述的解碼器,其中該處理器更被配置為:確定該基礎層圖像的一壓縮移動向量;以及 基於該基礎層圖像的該壓縮移動向量以確定該層間參考圖像的該移動向量。
- 21如申請專利範圍第20項所述的解碼器,其中該處理器被配置為確定該層間參考圖像的該移動向量包括該處理器被配置為根據該基礎層與該增強層之間的一空間比率來縮放該基礎層圖像的該壓縮移動向量、以及基於該縮放的移動向量來確定該層間參考圖像的該移動向量。
- 22如申請專利範圍第13項所述的解碼器,其中該處理器更被配置為:藉由複製該基礎層圖像的一參考圖像列表以確定用於該層間參考圖像的一參考圖像列表。
- 23如申請專利範圍第22項所述的解碼器,其中該處理器更被配置為藉由複製該基礎層圖像的該參考圖像列表中的一對應參考圖像的一圖像順序計數(POC)以確定該層間參考圖像的該參考圖像列表中的每一參考圖像的一POC。
- 24如申請專利範圍第13項所述的解碼器,其中該處理器被配置為藉由使用該層間參考圖像的一移動向量預測該增強層圖像的一移動向量以使用TMVP來解碼該增強層圖像。
Independent claims24
179 paragraphs in 1 section, as filed
Method and device for predicting motion vector of slicing video coding
METHOD AND APPARATUS OF MOTION VECTOR PREDICTION FOR SCALABLE VIDEO CODING
Cross-references to related applications
This application requires the U.S. Provisional Patent Application No. 61/694,555 filed on August 29, 2012, the U.S. Provisional Patent Application No. 61/734,650 filed on December 7, 2012, and the U.S. Provisional Patent Application No. 61/734,650 filed on August 16, 2013. The rights and interests of U.S. Provisional Patent Application No. 61/866,822, the content of which is incorporated herein by reference.
In the past two decades, digital video compression technology has been greatly developed and standardized, making efficient digital video communication, distribution and consumption possible. Most of the widely used commercial standards were developed by ISO/IEC and ITU-T, such as MPEC-2 and H.263 (MPEG-4 Part 10). Due to the emergence and maturity of video compression technology, High Efficiency Video Coding (HEVC) was developed.
Compared with traditional digital video services via satellite, cable and terrestrial transmission channels, more and more video applications can be used on the client and network side in a heterogeneous environment, such as but not limited to video chat, mobile video and streaming video . Smart phones, tablets, and TVs dominate the client side, where video can be transmitted via the Internet, mobile networks, and/or a combination of the two. To improve user experience and video service quality, adjustable video coding (SVC) can be used. In SVC, once the signal is encoded at the highest resolution, it can be selected from a subset of the data stream according to the application The specific rate and resolution required and supported by the client device are decoded. The international video standards MPEC-2 video, H.263, MPEG4 visual and H.264 have tools and/or configuration files that support tunability modes.
The inter-layer motion mapping information is used to enable the temporal motion vector prediction (TMVP) of the enhancement layer of the bit stream. For example, the reference image of the enhancement layer video block may be determined based on the collocation of the base layer video block. The enhancement layer video block is associated with the enhancement layer of the bit stream, and the collocated base layer video block is associated with the base layer of the bit stream. For example, the enhancement layer video block is associated with the enhancement layer image, and the collocated base layer video block is associated with the base layer image. The juxtaposed base layer video block can be determined by selecting the video block of the juxtaposed base layer image characterized by the largest overlap area with the enhancement layer video block. The video block can be an operation unit at any level of the bit stream. The video block can be of any size (for example, block size (e.g., 16x16), PU, SPU, etc.).
The reference image of the enhancement layer video block can be determined by determining the reference image of the juxtaposed base layer video block. The reference image of the enhancement layer video block may be a collocated enhancement layer image of the reference image of the collocated base layer video block. You can determine the reference image of the juxtaposed base layer video block, use the reference image of the juxtaposed base layer video block to determine the reference image of the inter-layer video block, and use the reference image of the inter-layer video block to determine the reference of the enhancement layer video block Image to determine the reference image of the enhancement layer video block. The inter-layer video block may be collocated with the enhancement layer video block and/or the base layer video block.
The MV of the enhancement layer video block may be determined based on the motion vector (MV) of the collocated base layer video block. The MV of the enhancement layer video block can be determined by determining the MV of the collocated base layer video block and scaling the MV of the collocated base layer video block according to the spatial ratio between the base layer and the enhancement layer to determine the MV of the enhancement layer video block.
By determining the MV of juxtaposed base layer video blocks, according to the base layer and enhancement layer The MV of the base layer video block is scaled and juxtaposed to determine the MV of the inter-layer video block and the MV of the enhancement layer video block is predicted based on the MV of the inter-layer video block to determine the MV of the enhancement layer video block. For example, based on the MV of the inter-layer video block, the MV of the enhancement layer video block can be predicted by performing time scaling on the MV of the inter-layer video block. The inter-layer video block may be collocated with the enhancement layer video block and/or the base layer video block.
The MV and/or reference image of the inter-layer video block can be used on the enhancement layer video block to perform TMVP. The enhancement layer video block can be decoded according to the reference image and/or MV of the enhancement layer video block, and/or the reference image and/or MV of the inter-layer video block.
The method includes receiving a bit stream including a base layer and an enhancement layer, and using Temporal Motion Vector Prediction (TMVP) to decode the enhancement layer in the encoded bit stream. The inter-layer reference image can be used as a collocated reference image for the enhancement layer TMVP.
Using TMVP to decode the enhancement layer in the encoded bit stream may include using TMVP to decode the enhancement layer image. Decoding the enhancement layer image using TMVP may include determining a motion vector (MV) field of the inter-layer reference image and decoding the enhancement layer image based on the MV domain of the inter-layer reference image. The MV domain of the inter-layer reference image may be determined based on the MV domain of the collocated base layer image. The MV field includes the reference image index and MV of the video block of the inter-layer reference image. For example, the MV domain may include one or more indexes of one or more video blocks of the inter-layer reference image (for example, depending on whether it is a P slice or B slice) and MV. Determining the MV domain of the inter-layer reference image may include determining the compressed MV domain of the concatenated base layer image and determining the MV domain of the inter-layer reference image according to the compressed MV domain of the concatenated base layer image.
Determining the MV domain of the inter-layer reference image may include determining the MV and reference image of the video block of the inter-layer reference image. Determining the MV and reference image of the video block of the inter-layer reference image may include determining the reference image of the inter-layer video block based on the reference image of the juxtaposed base layer video block and determining the inter-layer video block based on the MV of the juxtaposed base layer video block. MV. The juxtaposition base can be determined by selecting the video block in the juxtaposed base layer image characterized by the largest overlap area with the inter-layer reference image Layer video block.
Determining the reference image of the inter-layer video block may include determining the reference image of the juxtaposed base layer video block and determining the reference image of the inter-layer video block. The reference image of the inter-layer video block may be a juxtaposed inter-layer reference image of the reference image of the base layer video block. Determining the MV of the inter-layer video block may include determining the MV of the collocated base layer video block and scaling the MV of the collocated base layer video block according to the spatial ratio between the base layer and the enhancement layer to determine the MV of the inter-layer video block.
The MV domain of the enhancement layer video block may be determined based on the MV domain of the inter-layer video block. The enhancement layer video block may be collocated with the middle layer video block and/or the base layer video block. For example, the reference image of the enhancement layer video block may be determined based on the reference image of the inter-layer video block (for example, it may be a collocated enhancement layer image). The MV of the enhancement layer video block may be determined based on the MV of the inter-layer video block. For example, the MV of the inter-layer video block may be scaled (eg, time scaled) to determine the MV of the enhancement layer video block. The enhancement layer video block can be decoded based on the MV domain of the enhancement layer video block.
The method includes receiving the bit stream including the base layer and the enhancement layer and the inter-layer motion mapping information, and performing the inter-layer motion prediction of the enhancement layer. It can be determined that the inter-layer motion prediction for the enhancement layer is enabled based on the inter-layer mapping information.
The inter-layer mapping information can be signaled at the sequence level of the bit stream. For example, the inter-layer mapping information can be a variable (such as a flag), which can be signaled at the sequence level of the bit stream. The inter-layer mapping information can be inferred at the sequence level of the bit stream. The inter-layer mapping information can be signaled via variables (such as flags) in the video parameter set (VPS) of the bit stream (for example, the inter-layer mapping information can be the flags in the VPS of the bit stream). For example, the inter-layer mapping information can be signaled via variables (flags) in the sequence parameter set (SPS) of the bit stream (for example, the inter-layer mapping information can be the flags in the SPS in the bit stream). For example, the inter-layer mapping information can be signaled through variables (flags) in the picture parameter set (PPS) of the bitstream (for example, the inter-layer mapping information may be the flags in the PPS of the bitstream).
<p>100, 200, 300, 400, 500, 600, 701, 702, 703, 800, 810, 900, 910, 1000, 1010, 1020, 1030, 1100, 111011120, 1200, 1210, 1220, 1300, 1310 Graphics</p><p>310NeighbRefPic</p><p>314, 412PU</p><p>320, 420CurrRefPic</p><p>322,422CurrRefPU</p><p>330,440CurrPic</p><p>332,442CurrPU</p><p>334NeighbPU</p><p>340, 350, 450, 460, 1302, 1304, 1312, 1314MV</p><p>410ColRefPic</p><p>430ColPic</p><p>432ColPU</p><p>802, 804, 812, 814, 1002, 1012, 1022, 1034, 1102, 1114, 1124Short-term MV</p><p>902, 904, 912, 914, 1004, 1014, 1024, 1032, 1202, 1214, 1224Long-term MV</p><p>1104, 111211122, 1204, 1212, 1222Interlayer MV</p><p>1400,1700Communication system</p><p>1402, 1402a, 1402b, 1402c, 1402dWTRU</p><p>1403, 1404, 1405RAN</p><p>1406, 1407, 1409Core network</p><p>1408PSTN</p><p>1410Internet</p><p>1412Other networks</p><p>1414a, 1414b, 1480a, 1480b, 1480cBase station</p><p>1415, 1416, 1417Air interface</p><p>1418Processor</p><p>1420Transceiver</p><p>1422Transmit/Receive Components</p><p>1424Speaker/Microphone</p><p>1426Keyboard</p><p>1428Display/Touchpad</p><p>1430 Non-removable memory</p><p>1432Removable memory</p><p>1434Power</p><p>1436GPS Chipset</p><p>1438 Peripheral devices</p><p>1440a, 1440b, 1440cNode B</p><p>1442a, 1442bRNC</p><p>1444MGW</p><p>1446MSC</p><p>1448SGSN</p><p>1450GGSN</p><p>1460a, 1460b, 1460ceNode B</p><p>1462MME</p><p>1464Service Gateway</p><p>1466PDN Gateway</p><p>1482ASN Gateway</p><p>1484MIP-HA</p><p>1486AAA server</p><p>1488Gateway</p><p>1502Video signal</p><p>1504Conversion</p><p>1506Quantification</p><p>1508Entropy coding</p><p>1510,1610Inverse quantization</p><p>1512,1612Inverse conversion</p><p>1520, 1602Video bitstream</p><p>1560, 1660Spatial prediction</p><p>1562, 1662Time forecast</p><p>1564,1664Reference image memory</p><p>1566, 1666Inner loop filter</p><p>1580Mode decision block</p><p>1608Entropy decoding</p><p>1702Encoder</p><p>1704Communication Network</p><p>1706Decoder</p><p>1708, 1710Connect</p><p>AAAAuthentication, authorization, accounting</p><p>ASNAccess service network</p><p>BL, pBLbase layer image</p><p>ColPicColocated Image</p><p>ColPUCo-located PU</p><p>ColRefPicCo-located reference image</p><p>CurrPicCurrent image</p><p>CurrPUCurrent PU</p><p>CurrRefPicCurrent reference image</p><p>CurrRefPUBest matching block</p><p>GGSNGateway GPRS Support Node</p><p>GPSGlobal Positioning System</p><p>IPInternet Protocol</p><p>Iub, IuCS, IuPS, iur, S1, X2Interface</p><p>MGWMedia Gateway</p><p>MIP-HAMobile IP local agent</p><p>MMEMobility Management Entity</p><p>MSCMobile Switching Center</p><p>MVMotion vector</p><p>NeighbPUNeighboring PU</p><p>NeighbRefPicNeighboring reference image</p><p>PDNPacket Data Network</p><p>PSTNPublic Switched Telephone Network</p><p>PUPrediction Unit</p><p>R1, R3, R6, R8Reference point</p><p>RANRadio access network</p><p>SGSNServing GPRS Support Node</p><p>SPUMinimum PU</p><p>TB,TDTime distance</p><p>WTRUWireless Transmission/Receiving Unit</p>
Figure 1 is a diagram showing an example of a tunable structure with additional inter-layer prediction for SVC spatial tunable coding.
Figure 2 is a diagram showing an example inter-layer prediction structure considered for HEVC tunable coding.
Fig. 3 is a diagram showing an example of spatial motion vector (MV) prediction (SMVP).
Fig. 4 is a diagram showing an example of temporal MV prediction (TMVP).
Figure 5 is a diagram showing an example of the prediction structure replication from the base layer to the up-sampled base layer.
Fig. 6 is a diagram showing an example relationship between the up-sampled SPU of the base layer and the SPU of the original base layer.
FIGS. 7A to 7C are diagrams showing an example relationship between a segment of a base layer image and a segment of a processed base layer image.
Fig. 8A is a diagram showing MV prediction between temporal and short-term MVs.
Figure 8B is a diagram showing the MV prediction of the short-term MV based on the mapped short-term MV.
Fig. 9A is a diagram showing an example of MV prediction between temporal and long-term MVs.
FIG. 9B is a diagram showing an example of MV prediction of a temporal long-term MV based on a mapped long-term MV.
FIG. 10A is a diagram showing an example of MV prediction based on a temporal long-term MV versus a temporal short-term MV.
FIG. 10B is a diagram showing an example of MV prediction of a temporal short-term MV based on a mapped long-term MV.
FIG. 10C is a diagram showing an example of MV prediction based on temporal short-term MV versus temporal long-term MV.
FIG. 10D is a diagram showing an example of MV prediction of a temporal long-term MV based on a mapped short-term MV.
FIG. 11A is a diagram showing an example of MV prediction based on the prohibition of inter-layer MV to temporal short-term MV.
FIG. 11B is a diagram showing an example of disabled MV prediction for inter-layer MV according to the temporal short-term MV.
FIG. 11C is a diagram showing an example of MV prediction for disabling the inter-layer MV according to the mapped short-term MV.
FIG. 12A is a diagram showing an example of MV prediction based on the prohibition of inter-layer MV to temporal long-term MV.
FIG. 12B is a diagram showing an example of disabled MV prediction of inter-layer MV according to time long-term MV.
FIG. 12C is an example diagram showing the forbidden MV prediction of the inter-layer MV according to the mapped long-term MV.
Fig. 13A is a diagram showing an example of MV prediction between two inter-layer MVs when Te=Tp.
FIG. 13B is an example diagram showing forbidden MV prediction between two inter-layer MVs when TeTp.
FIG. 14A is a system diagram of an example communication system in which one or more disclosed embodiments may be implemented.
FIG. 14B is a system diagram of an example wireless transmission/reception unit (WTRU) that can be used in the communication system shown in FIG. 14A.
FIG. 14C is a system diagram of an example radio access network and an example core network that can be used in the communication system shown in FIG. 14A.
FIG. 14D is a system diagram of another example radio access network and another example core network that can be used in the communication system shown in FIG. 14A.
FIG. 14E is a system diagram of another example radio access network and another example core network that can be used in the communication system shown in FIG. 14A.
Figure 15 is a block diagram showing an example of a block-based video codec.
Figure 16 is a block diagram showing an example of a block-based video decoder.
Figure 17 is a diagram showing an example communication system.
For example, with the tunability extension of H.264, the encoding and/or decoding (e.g., transmission and/or reception) of bit streams (e.g., partial bit streams) can be provided in order to reserve and part bit streams. The rate of the metastream is closely related to the reconstruction quality while providing video services with lower temporal resolution, lower spatial resolution, and/or reduced fidelity. Figure 1 is a diagram showing an example of a tunable structure with additional inter-layer prediction for SVC spatial tunability coding. FIG. 100 may show an example of a dual-layer SVC inter-layer prediction mechanism that can improve adjustable coding efficiency. A similar mechanism can be used on the multi-layer SVC coding structure. In the graph 100, the base layer and the enhancement layer may represent two adjacent spatially adjustable layers with different resolutions. Among the layers (e.g., base layer and/or enhancement layer), for example, an H.264 encoder may use motion compensation prediction and/or intra prediction. Inter-layer prediction can use base layer information (for example, spatial texture, motion vector, reference image index value, residual signal, etc.) to improve the coding efficiency of the enhancement layer. When decoding the enhancement layer, the SVC may be completely reconstructed without using the reference image from the lower layer (for example, the dependent layer of the current layer).
In a tunable coding system (for example, HEVC tunable coding extension), inter-layer prediction is used, for example, to determine the relationship between multiple layers and/or to improve the tunable coding efficiency. Figure 2 is a diagram showing an example inter-layer prediction structure considered for HEVC tunable coding. For example, the graph 200 may show an example of a tunable structure with additional inter-layer prediction for HEVC spatial tunable coding. It can be predicted by motion compensation from the reconstructed base layer signal (for example, after up-sampling if the spatial resolution between the two layers is different), the current time prediction in the enhancement layer, and/or the average value of the base layer reconstruction signal and the time prediction signal To produce the prediction of the enhancement layer. A complete reconstruction of the lower layer image can be performed. Similar implementations can be used for tunable coding systems with more than two layers (for example, HEVC tunable coding systems with more than two layers).
For example, by using pixels from the encoded video image to predict the pixels in the current video image, HEVC can use advanced motion compensation prediction technology to determine the inherent inter-layer image redundancy in the video signal. The displacement between the current prediction unit (PU) to be coded and one or more matching blocks in the reference image (for example, a neighboring PU) can be represented by a motion vector in motion compensation prediction. MV can include two components, MVx and MVy. MVx and MVy can represent the displacement in the horizontal and vertical directions, respectively. MVx and MVy can be directly encoded or not directly encoded.
Advanced Motion Vector Prediction (AMVP) can be used to predict the MV from one or more MVs of neighboring PUs. The difference between the real MV and the MV predictor (predictor) can be encoded. By encoding the difference of the MV (for example, encoding only), the number of bits used to encode the MV can be reduced. The MV for prediction can be obtained from the spatial and/or temporal neighborhood. The spatial neighborhood may refer to those spatial PUs around the current coded PU. The temporal neighborhood may refer to juxtaposed PUs in neighboring images. In HEVC, in order to obtain accurate MV predictors, prediction candidates from the spatial and/or temporal neighborhood can be put together to form a candidate list and the best predictor is selected to predict the MV of the current PU. For example, the best MV predictor can be selected based on Lagrangian rate-distortion (RD) consumption, etc. The MV difference can be encoded into a bit stream.
Fig. 3 is a diagram showing an example of spatial MV prediction (SMVP). The graph 300 may show examples of the neighboring reference image 310, the current reference image 320, and the current image 330. In the current picture (CurrPic 330) to be encoded, the hashed square may be the current PU (CurrPU 332). CurrPU 332 may have the best matching block (CurrRefPU 322) located in the reference image (CurrRefPic 320). Can predict CurrPU's MV (MV2 340). For example, in HEVC, the spatial neighborhood of the current PU may be the PU above, on the left side, on the upper left side, on the lower left side, or on the upper right side of the current PU 332. For example, the neighboring PU 334 shown is the upper neighbor of CurrPU 332. For example, the reference picture (NeighbRefPic 310), PU 314, and MV (MV1 350) of the neighboring PU (NeighbPU) are known because NeighbPU 334 is coded before CurrPU 332.
Fig. 4 is a diagram showing an example of temporal MV prediction (TMVP). The graphic 400 may include 4 images, such as a collocated reference image (ColRefPic) 410, a CurrRefPic 420, a collocated image (ColPic) 430, and a CurrPic 440. In the current image (CurrPic 440) to be encoded, the hashed square (CurrPU 442) may be the current PU. The hashed square (CurrPU 442) may have the best matching block (CurrRefPU 422) located in the reference image (CurrRefPic 420). Can predict CurrPU's MV (MV2 460). For example, in HEVC, the temporal neighborhood of the current PU may be a collocated PU (ColPU) 432, for example, ColPU 432 is a part of a neighboring image (ColPic) 430. For example, ColPU's reference image (ColRefPic 410), PU 412, and MV (MV1 450) are known because ColPic 430 is coded before CurrPic 440.
The movement between PUs is a uniform translation. The MV between the two PUs is proportional to the time distance between the moments when the two associated images are captured. Before predicting the MV of the current PU, a scalable motion vector predictor (for example, in AMVP). For example, the time distance between CurrrPic and CurrRefPic can be referred to as TB. For example, the time distance between CurrPic and NeighbRefPic (e.g., Figure 3) or ColPic and ColRefPic (e.g., Figure 4) may be referred to as TD. Given TB and TD, the scaled predictor of MV2 (for example, MV) can be equal to:<maths><img he="206" wi="1894" file="TW201804792A_D0001.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></maths>
Can support short-term reference images and long-term reference images. For example, reference pictures stored in the decoded picture buffer (DPB) may be marked as short-term reference pictures or long-term reference pictures. For example, in equation (1), if one or more of the reference images are long-term reference images, scaling of the movement vector may be disabled.
This article describes the use of MV prediction for multi-layer video coding. The examples described herein may use the HEVC standard as the underlying single-layer coding standard and a tunable system with two spatial layers (e.g., enhancement layer and base layer). The examples described in this article can be applied to Use other types of basic single-layer codecs, with more than two layers, and/or other tunable coding systems that support other types of tunability.
When starting to decode a video segment (e.g., P segment or B segment), one or more reference pictures in the DPB can be added to the reference picture list (e.g., list 0) of the P segment used for motion compensation prediction and/or Two reference image lists for segment B (for example, list 0 and list 1). The adjustable coding system can use the temporal reference image of the enhancement layer and/or the processed reference image from the base layer (for example, if the spatial resolutions of the two layers are different, the base layer image is upsampled). Carry out motion compensation prediction. When predicting the MV of the current image in the enhancement layer, the inter-layer MV pointing to the processed reference image from the base layer can be used to predict the time MV of the temporal reference image pointing to the enhancement layer. Temporal MV can also be used to predict inter-layer MV. Because these two types of MVs are almost irrelevant, it will result in a loss of efficiency for the MV prediction of the enhancement layer. The single-layer codec does not support the temporal MV between the base layer images to predict the temporal MV between the enhancement layer images, and the two are highly correlated and can be used to improve the MV prediction performance.
The MV prediction process can be simplified and/or the compression efficiency for multi-layer video coding can be improved. The MV prediction in the enhancement layer can be backward compatible with the MV prediction process of the single-layer encoder. For example, there may be MV prediction implementations that do not require any changes to block-level operations in the enhancement layer, so that the single-layer encoder and decoder logic can be reused in the enhancement layer. This can reduce the implementation complexity of the adjustable system. The MV prediction of the enhancement layer can distinguish a temporal MV pointing to a temporal reference picture in the enhancement layer and an inter-layer MV pointing to a processed (e.g., up-sampled) reference picture from the base layer. This can improve coding efficiency. The MV prediction in the enhancement layer can support the MV prediction between the temporal MV between the enhancement layer images and the temporal MV between the base layer images, which can improve coding efficiency. When the spatial resolution between the two layers is different, the time MV between the base layer images can be scaled according to the ratio of the spatial resolution between the two layers.
The implementation described herein is related to the inter-layer mobile information mapping algorithm used for the base layer MV, for example, so that the base layer MV mapped in the AMVP process can be used for prediction Enhancement layer MV (for example, TMVP mode in Figure 4). There may be no change in block-level operations. The single-layer encoder and decoder can be used for the MV prediction of the enhancement layer without change. This article will describe MV prediction tools that include block-level changes for the enhancement layer encoding and decoding process.
The interlayer may include a processed base layer and/or an upsampled base layer. For example, interlayer, processed base layer, and/or upsampled base layer may be used interchangeably. Inter-layer reference images, processed base layer reference images, and/or up-sampled base layer reference images can be used interchangeably. Inter-layer video blocks, processed base layer video blocks, and/or up-sampled base layer video blocks can be used interchangeably. There may be a timing relationship between the enhancement layer, between the layers, and between the base layer. For example, the video block and/or image of the enhancement layer may be temporally associated with the corresponding video block and/or image of the inter-layer and/or base layer.
The video block can be an operation unit at any level and/or any level of the bit stream. For example, a video block may be an operation unit at the image level, block level, segment level, and so on. The video block can be of any size. For example, a video block can refer to a video block of any size, such as a 4x4 video block, an 8x8 video block, a 16x16 video block, and so on. For example, a video block may refer to a prediction unit (PU), a minimum PU (SPU), and so on. The PU may be a video block unit used to carry information related to motion prediction, for example, including reference image index and MV. One PU may include one or more smallest PUs (SPUs). Although SPUs in the same PU refer to the same reference image with the same MV, in some implementations, storing mobile information in units of SPUs can make it easier to obtain mobile information. Mobile information (for example, MV domain) can be stored in units of video blocks (for example, PU and SPU, etc.). Although the examples described here are described with reference to images, video blocks, PUs and/or SPUs, and arbitrary operation units of any size (for example, images, video blocks, PUs, SPUs, etc.) can also be used.
The texture of the reconstructed base layer signal may be processed for inter-layer prediction of the enhancement layer. For example, when the spatial tunability between two layers is enabled, the inter-layer reference image processing may include upsampling of one or more base layer images. It may not be possible to correctly generate information related to the movement of the processed reference image from the base layer (for example, MV, reference image list, reference image Like index etc.). When the temporal MV predictor comes from the processed base layer reference image (for example, as shown in Figure 4), the missing motion information may affect the prediction of the enhancement layer MV (for example, via TMVP). For example, when the selected processed base layer reference image is used as a temporally adjacent image (ColPic) including a temporally collocated PU (ColPU), if the MV predictor for the processed base layer reference image is not correctly generated ( MV1) and reference image (ColRefPic), TMPV cannot work normally. In order to enable the TMVP for the enhancement layer MV prediction, the inter-layer movement information mapping can be used to implement, for example, as described herein. For example, the MV domain (including MV and reference image) for the processed base layer reference image can be generated.
One or more variables can be used to describe the reference image of the current video clip, for example, the reference image list ListX (for example, X is 0 or 1), the reference image index refIdx in ListX, and so on. Using the example of Figure 4, in order to obtain the reference image (ColRefPic) of the collocated PU (ColPU), the reference image of the PU (for example, each PU) (ColPU) in the processed reference image (ColPic) can be generated . This can be decomposed into a list of reference images for generating ColPic and/or a reference image index for ColPUs (e.g., each ColPU) in ColPic. Given the reference image list, the generation of the reference image index for the PU in the processed base layer reference image is described here. The implementation related to the composition of the reference image list of the processed base layer reference image is described here.
Because the base layer and the processed base layer are related, it can be assumed that the base layer and the processed base layer have the same or substantially the same prediction dependency relationship. The prediction dependency relationship of the base layer image can be copied to form a reference image list of the processed base layer image. For example, if the base layer image BL1 is a temporal reference image of another base layer image BL2 with a reference image index refIdx in the reference image list ListX (for example, X is 0 or 1), the processing of BL1 The base layer image pBL1 may be added to the same reference image list ListX with the same index refIdx of the processed base layer image pBL2 of BL2 (for example, X is 0 or 1). Fig. 5 is a diagram showing an example of copying the prediction structure from the base layer to the up-sampling base layer. The graph 500 shows an example of spatial tunability, which is the same as the movement information of the sampling base layer (indicated by the dotted line in the figure). The same B-level structure used for the motion prediction of the base layer is copied (indicated by the solid line in the figure).
The reference image of the processed base layer PU may be determined based on the collocated base layer prediction unit (PU). For example, the collocated base layer PU of the processed base layer PU may be determined. The juxtaposed base layer PU can be determined by selecting the PU in the juxtaposed base layer image characterized by the largest overlap area with the processed base layer PU, for example, as described herein. The reference image of the collocated base layer PU may be determined. The processed reference image of the base layer PU may be determined as a collocated base layer reference image of the reference image of the collocated base layer PU. The processed reference image of the base layer PU can be used for the TMVP of the enhancement layer and/or for decoding the enhancement layer (for example, collocated enhancement layer PU).
The processed base layer PU is associated with the processed base layer image. The MV domain of the processed base layer image may include a reference image of the processed base layer PU, for example, a TMVP for an enhancement layer image (for example, a collocated enhancement layer PU). The reference image list is associated with the processed base layer image. The reference image list of the processed base layer image may include one or more of the reference images of the processed base layer PU. The processed images in the base layer (for example, each image) can inherit the same image order count (POC) and/or short-term/long-term image tags from the corresponding images in the base layer.
The spatial tunability with 1.5 times the upsampling rate is used as an example. Fig. 6 is a diagram showing an example relationship between the up-sampled SPU of the base layer and the SPU of the original base layer. The graph 600 may show the up-sampled SPU of the base layer (e.g., marked as u<sub>i</sub>Block) and the SPU of the original base layer (for example, marked as b<sub>j</sub>The relationship between the blocks). For example, given various upsampling rates and coordinates in the image, the SPUs in the upsampled base layer image may correspond to various numbers and/or ratios of SPUs from the original base layer image. For example, SPU u<sub>4</sub>It can cover 4 SPU areas in the base layer (for example, b<sub>0</sub>, B<sub>1</sub>, B<sub>2</sub>, B<sub>3</sub>). SPU u<sub>1</sub>Can cover two base layer SPUs (for example, b<sub>0</sub>And b<sub>1</sub>,). SPU u<sub>0</sub>Can cover a single base layer SPU (for example, b<sub>0</sub>). The MV domain mapping implementation can be used to estimate the reference image index and MV of the SPU in the processed base layer image, for example, using the movement information of the corresponding SPU from the original base layer image.
The MV of the processed base layer PU may be determined according to the MV of the collocated base layer PU. For example, the collocated base layer PU of the processed base layer PU may be determined. The MV of the collocated base layer PU can be determined. The MV of the base layer PU can be scaled to determine the MV of the processed base layer PU. For example, the MV of the base layer PU may be scaled according to the spatial ratio between the base layer and the enhancement layer to determine the MV of the processed base layer PU. The processed MV of the base layer PU can be used for the TMVP of the enhancement layer (e.g., collocated enhancement layer PU) and/or used to decode the enhancement layer (e.g., collocated enhancement layer PU).
The processed base layer PU may be associated (e.g., temporally associated) with an enhancement layer image (e.g., PU of an enhancement layer image). The MV domain of the collocated enhancement layer image is based on the MV of the processed base layer PU, for example, the TMVP for the enhancement layer image (for example, the collocated enhancement layer PU). The MV of the enhancement layer PU (for example, the collocated enhancement layer PU) may be determined based on the MV of the processed base layer PU. For example, the MV of the processed base layer PU may be used to predict (e.g., spatial prediction) the MV of the enhancement layer PU (e.g., collocated enhancement layer PU).
The reference image of the SPU (for example, each SPU) in the processed base layer image may be selected based on the reference image index of the corresponding SPU in the base layer. For example, for the SPU in the processed base layer image, the main rule applied to determine the reference image index is the reference image index whose corresponding SPU from the base layer image uses the most frequently. For example, suppose an SPU u in the processed base layer image<sub>h</sub>Corresponding to K SPU b from the base layer<sub>i</sub>(i=0,1,...,K-1), there are M reference images with index values {0,1,...,M-1} in the reference image list of the processed base layer image. Assume that the secondary index is {r<sub>0</sub>,r<sub>1</sub>,..,r<sub>k-1</sub>} Set of reference images to predict K corresponding SPUs in the base layer, where for i=0,1,...,K-1, there are<i>r</i><sub><i>i</i></sub><img file="TW201804792A_D0002.tif" wi="43" he="58" img-format="tif" img-content="character" orientation="portrait" inline="no" />{0,1,...,<i>M</i>-1}, then u<sub>h</sub>The reference image index of can be determined by equation (2):<maths><img he="110" wi="1846" file="TW201804792A_D0003.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></maths>Where C(r<sub>i</sub>), i=0,1,...,K-1 is a counter used to indicate how many times the reference image r is used. For example, if the base layer image has 2 reference images (M=2) marked as {0,1}, and the given u in the processed base layer image<sub>h</sub>Corresponds to {0,1,1,1} (for example, {r<sub>0</sub>,r<sub>1</sub>,...,R<sub>3</sub>} Is equal to {0,1,1,1}) predicted 4 (K=4) base layer SPU, then according to equation (2), r(u<sub>h</sub>) Is set to 1. For example, because two images with a smaller time distance have a higher correlation, the reference image r with the smallest POC of the image processed so far is selected.<sub>i</sub>(For example, to break C(r<sub>i</sub>) Constraints).
The different SPUs in the processed base layer image may correspond to various numbers and/or ratios of SPUs from the original base layer (for example, as shown in FIG. 6). The reference image index of the base layer SPU with the largest coverage area can be selected to determine the reference image of the corresponding SPU in the processed base layer. For the given SPU u in the processed base layer<sub>h</sub>, The reference image index can be determined by equation (3):<maths><img he="113" wi="1816" file="TW201804792A_D0004.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></maths>Among them, S<sub>i</sub>Is the i-th corresponding SPU b in the base layer<sub>i</sub>Coverage area. For example, when the coverage area of two or more corresponding SPUs is the same, the reference image with the smallest POC distance to the currently processed image is selected.<sub>i</sub>, To break the equation (3)S<sub>i</sub>Constraints.
The internal mode can be used to control the corresponding base layer SPU b<sub>j</sub>Encode. Reference image index (for example, the corresponding base layer SPU b<sub>j</sub>The reference image index of) can be set to -1 and ignored when applying equation (2) and/or equation (3). If the corresponding base layer SPU b<sub>j</sub>Is internally coded, SPU u<sub>h</sub>The reference image index can be set to -1 or marked as invalid for TMVP.
For the given SPU u in the processed base layer<sub>h</sub>, Its corresponding SPU b<sub>i</sub>The area may be different. For example, the area-based implementation described here can be used to estimate where The MV of the SPU (for example, each SPU) in the processed base layer image.
To estimate an SPU u in the processed base layer image<sub>h</sub>MV, available in the base layer SPU candidate b<sub>i</sub>Selected and SPU u<sub>h</sub>The base layer SPU b with the largest coverage area (for example, the largest overlap area)<sub>1</sub>MV. For example, equation (4) can be used:<maths><img he="170" wi="1947" file="TW201804792A_D0005.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></maths>Where MV' represents the resulting SPU u<sub>h</sub>MV, MV<sub>i</sub>Represents the i-th corresponding SPU b in the base layer<sub>i</sub>MV, and N is the upsampling factor (for example, N can be equal to 2 or 1.5), depending on the spatial ratio (for example, spatial resolution) between the two layers (for example, the base layer and the enhancement layer). For example, an upsampling factor (eg, N) can be used to scale the resulting MV determined from the PU of the base layer to calculate the MV of the PU in the processed base layer image.
The weighted average can be used to determine the MV of the SPU in the processed base layer. For example, a weighted average may be used to determine the MV of the SPU in the processed base layer by using the MV associated with the corresponding SPU in the base layer. For example, using weighted average can improve the accuracy of the processed base layer MV. For SPU u in the processed base layer<sub>h</sub>, Can be determined for one or more (for example, each) and u<sub>h</sub>SPU b of overlapping base layer<sub>i</sub>To obtain the SPU u<sub>h</sub>MV. For example, it is shown by equation (5):<maths><img he="216" wi="1920" file="TW201804792A_D0006.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></maths>Where B is the reference image index in the base layer equal to r(u<sub>h</sub>) SPU b<sub>i</sub>The subset of is, for example, determined by equation (2) and/or equation (3).
One or more filters (for example, median filter, low-pass Gaussian filter, etc.) can be applied to the MV group represented by B in equation (5), for example, to obtain the mapped image represented by MV' MV. The average value of confidence can be used to improve the accuracy of the assessed MV, As shown in equation (6):<maths><img he="226" wi="1903" file="TW201804792A_D0007.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></maths>Where the parameter w<sub>i</sub>Is estimated SPU u<sub>h</sub>Base layer SPU b<sub>i</sub>(For example, each base layer SPU b<sub>i</sub>) Reliable measurement of MV. Different metrics can be used to derive w<sub>i</sub>Value. For example, w can be determined according to the predicted residual amount during the motion compensation prediction period<sub>i</sub>, Or according to MV<sub>i</sub>Determine the degree of coherence with its neighboring MV to determine w<sub>i</sub>。
The movement information of the processed base layer image can be mapped from the original movement domain of the base layer, for example, it can be used to perform time motion compensation prediction in the base layer. A motion compensation algorithm (for example, as supported in HEVC) can be applied to the motion field of the base layer, for example, to generate a compressed motion field of the base layer. The movement information of one or more of the processed base layer images can be mapped from the compressed movement domain of the base layer.
The lost movement information for the processed base layer image can be generated, for example, as described here. There is no need to make additional changes to block-level operations, and TMVP supported by a single-layer codec (for example, the HEVC codec) can be used for the enhancement layer.
When the corresponding base layer reference image includes one or more segments, the reference image list generation process and/or the MV mapping process can be used, for example, as shown here. If there are multiple segments in the base layer reference image, the segment cutting can correspond to the processed base layer image from the base layer image. For the processed fragments in the base layer, a reference image list generation step can be performed to derive a suitable fragment type and/or reference image list.
FIGS. 7A to 7C are diagrams showing an example relationship between a segment of a base layer image and a segment of a processed base layer image, for example, for 1.5 times the spatial adjustability. FIG. 7A is a graph 701 showing an example of segment cutting in the base layer. FIG. 7B is a graph 702 showing an example of segmentation of the mapped segment in the processed base layer. Figure 7C shows the base layer after processing The graphic 703 of the example of the adjusted segment cutting in.
The base layer image may include multiple segments, for example, two segments as shown in the graph 701. The segment cuts mapped in the processed base layer image may cross the boundaries of adjacent coding tree blocks (CTBs) in the enhancement layer, for example, when the base layer is upsampled (e.g., as shown in graph 702). This is because the spatial ratio between the base layer image and the enhancement layer image is different. Segment cuts (e.g., in HEVC) can be aligned with CTB boundaries. The segment cutting in the processed base layer can be adjusted to align the boundaries of the segments with the boundaries of the CTB, for example, as shown in graph 703.
The enhancement layer TMVP derivation process may include constraints. For example, if there is a segment in the corresponding base layer image, the processed base layer image can be used as a collocated image. When there is more than one segment in the corresponding base layer image, it is not necessary to perform inter-layer movement information mapping on the processed base layer reference image (for example, reference image list generation and/or MV mapping as described herein) . If there is more than one segment in the corresponding base layer image, the temporal reference image can be used as a concatenated image for the TMVP derivation process of the enhancement layer. The number of segments in the base layer image can be used to determine whether to use the inter-layer reference image and/or the temporal reference image as the concatenated image for the TMVP of the enhancement layer.
If there is a segment in the corresponding base layer image and/or if the segment information (for example, the reference image list of the segment in the corresponding base layer image, the segment type, etc.) is exactly the same, then the processed The base layer image can be used as a collocated image. When two or more segments in the corresponding base layer image have different segment information, no inter-layer movement information mapping is performed on the processed base layer reference image (for example, reference image list generation and/or as described here) Or MV mapping). If two or more segments in the corresponding base layer image have different segment information, the temporal reference image can be used as a collocated image for the TMVP export process of the enhancement layer.
Mobile information mapping enables various single-layer MV prediction techniques to be used in tunable coding systems. Block-level MV prediction operations can be applied to improve enhancement layer coding performance. The MV prediction of the enhancement layer is described here. There is no change in the MV prediction process of the basic layer.
Temporal MV refers to an MV that points to a reference image from the same enhancement layer. The inter-layer MV refers to an MV that points to another layer (for example, a processed base layer reference image). The mapped MV refers to the MV generated for the processed base layer image. The mapped MV may include a mapped time MV and/or a mapped inter-layer MV. The mapped time MV refers to the mapped MV derived from the time prediction of the last coding layer. The mapped inter-layer MV refers to the mapped MV generated from the inter-layer prediction of the last coding layer. For tunable coding systems with more than two layers, there may be mapped inter-layer MVs. The time MV and/or the mapped time MV may be a short-term MV or a long-term MV, for example, depending on the MV pointing to a short-term reference image or a long-term reference image. The temporal short-term MV and the mapped short-term MV refer to the temporal MV and the mapped temporal MV using the short-term temporal reference in the respective coding layers. The temporal long-term MV and the mapped long-term MV refer to the temporal MV and the mapped temporal MV using a long-term temporal reference in their respective coding layers. Time MV, mapped time MV, mapped inter-layer MV, and inter-layer MV can be regarded as different types of MVs.
The enhancement layer MV prediction may include one or more of the following. The MV prediction of the temporal MV from the inter-layer MV and/or the mapped inter-layer MV may be enabled or disabled. The MV prediction of the inter-layer MV from the time MV and/or the mapped time MV may be enabled or disabled. The MV prediction from the mapped time MV to the time MV can be enabled. The MV prediction of the inter-layer MV from the inter-layer MV and/or the mapped inter-layer MV may be enabled or disabled. For the long-term MV included in the MV prediction (for example, including the temporal long-term MV and the mapped long-term MV), MV prediction without using MV scaling may be adopted.
Short-term inter-MV prediction using MV scaling (for example, similar to single-layer MV prediction) can be enabled. Fig. 8A is a diagram showing MV prediction between temporal short-term MVs. FIG. 8B is a diagram showing the MV prediction of the temporal short-term MV based on the mapped short-term MV. In the graph 800, the temporal short-term MV 802 can be predicted based on the temporal short-term MV 804. In the graph 810, the temporal short-term MV 812 can be predicted based on the mapped short-term MV 814.
For example, due to the large POC spacing, it can provide long-term MV prediction without MV scaling. This is similar to MV prediction for single-layer encoding and decoding. Figure 9A shows the long-term Diagram of an example of MV prediction between MVs. FIG. 9B is a diagram showing an example of MV prediction from the mapped long-term MV to the temporal long-term MV. In the graph 900, the temporal long-term MV 902 can be predicted from the temporal long-term MV 904. In the graph 910, the temporal long-term MV 912 can be predicted from the mapped long-term MV 914.
For example, because the distance between two reference images is very long, it is possible to provide prediction without MV scaling between short-term MV and long-term MV. This is similar to MV prediction for single-layer encoding and decoding. Fig. 10A is a diagram showing an example of MV prediction from a temporal long-term MV to a temporal short-term MV. FIG. 10B is a diagram showing an example of MV prediction from the mapped long-term MV to the temporal short-term MV. Fig. 10C is a diagram showing an example of MV prediction from a temporal short-term MV to a temporal long-term MV. FIG. 10D is a diagram showing an example of MV prediction from the mapped short-term MV to the temporal long-term MV.
In the graph 1000, the temporal short-term MV 1002 can be predicted from the temporal long-term MV 1004. In FIG. 1010, the temporal short-term MV 1012 can be predicted from the mapped long-term MV 1014. In Fig. 1020, the temporal long-term MV 1024 can be predicted from the temporal short-term MV 1022. In FIG. 1030, the temporal long-term MV 1032 can be predicted from the mapped short-term MV 1034.
The prediction of temporal short-term MV from inter-layer MV and/or mapped inter-layer MV may be disabled. The prediction of the inter-layer MV from the temporal short-term MV and/or the mapped short-term MV may be disabled. FIG. 11A is a diagram showing an example of forbidden MV prediction from inter-layer MV to temporal short-term MV. FIG. 11B is a diagram showing an example of forbidden MV prediction from the temporal short-term MV to the inter-layer MV. FIG. 11C is a diagram showing an example of forbidden MV prediction from mapped short-term MV to inter-layer MV.
The graph 1100 shows an example of disabled MV prediction from the inter-layer MV 1104 to the temporal short-term MV 1102. For example, the temporal short-term MV 1102 may not be predicted from the inter-layer MV 1104. The graph 1110 shows an example of disabled MV prediction from the temporal short-term MV 1114 to the inter-layer MV 1112. For example, the inter-layer MV 1112 may not be predicted from the temporal short-term MV 1114. The graph 1120 shows an example of disabled MV prediction from the mapped short-term MV 1124 to the inter-layer MV 1122. example For example, the inter-layer MV 1122 may not be predicted from the mapped short-term MV 1124.
The prediction of temporal long-term MV from inter-layer MV and/or mapped inter-layer MV may be disabled. The prediction of the inter-layer MV from the temporal long-term MV and/or the mapped long-term MV may be disabled. Fig. 12A is a diagram showing an example of forbidden MV prediction from inter-layer MV to temporal long-term MV. Fig. 12B is a diagram showing an example of forbidden MV prediction from time-long-term MV to inter-layer MV. Fig. 12C is a diagram showing an example of forbidden MV prediction from the mapped long-term MV to the inter-layer MV.
Graph 1200 shows an example of disabled MV prediction from inter-layer MV 1204 to temporal long-term MV 1202. For example, the temporal long-term MV 1202 may not be predicted from the inter-layer MV 1204. The graph 1210 shows an example of forbidden MV prediction from the temporal long-term MV 1214 to the inter-layer MV 1212. For example, the inter-layer MV 1212 may not be predicted from the temporal long-term MV 1214. The graph 1220 shows the disabled MV mapping from the mapped long-term MV 1224 to the inter-layer MV 1222. For example, the inter-layer MV 1222 may not be predicted from the mapped long-term MV 1224.
For example, if two inter-layer MVs have the same time interval in the enhancement layer and the processed base layer, predicting the inter-layer MV from another inter-layer MV can be enabled. If the time interval of the MV between the two layers in the enhancement layer and the processed base layer is different, the prediction between the MVs between the two layers may be disabled. This is because the prediction cannot produce good coding performance due to the lack of clear MV correlation.
Fig. 13A is a diagram showing an example of MV prediction between two inter-layer MVs when Te=Tp. FIG. 13B is a diagram showing an example of MV prediction between two inter-layer MVs that is disabled when TeTp. TMVP can be used as an example (for example, as shown in Figures 13A to 13B). In the diagram 1300, the current inter-layer MV (e.g., MV2) 1302 can be predicted from another inter-layer MV (e.g., MV1) 1304. The time interval between the current image CurrPic and its temporally adjacent image ColPic (for example, including the juxtaposed PU ColPU) is denoted as T<sub>e</sub>. The time interval between their respective reference pictures (for example, CurrRefPic and ColRefPic) is denoted as T<sub>p</sub>. CurrPic and ColPic can be In the enhancement layer, CurrRefPic and ColRefPic may be in the processed base layer. If T<sub>e</sub>=T<sub>p</sub>, Then MV1 can be used to predict MV2.
For example, because POC-based MV scaling may fail, MV scaling for predicted MVs between two inter-layer MVs may be disabled. In FIG. 1310, for example, because the time interval between the current image CurrPic and its adjacent image ColPic (for example, T<sub>e</sub>) And their respective reference images are not equal (for example, T<sub>p</sub>), it may not be possible to predict the current inter-layer MV (for example, MV2) 1312 from another inter-layer MV (for example, MV1) 1314.
For example, if the inter-layer MV and the mapped inter-layer MV have the same time distance, the prediction of the inter-layer MV based on the mapped inter-layer MV without scaling can be enabled. If their time distances are different, the prediction of the inter-layer MV based on the mapped inter-layer MV may be disabled. Table 1 summarizes examples of different conditions for MV prediction for enhancement layer coding of SVC.
<tables><img he="1625" wi="2101" file="tw201804792a_d0008.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></tables><tables><img he="2198" wi="2037" file="tw201804792a_d0009.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></tables>
For the implementation of mobile information mapping between different coding layers, the MV mapping of the inter-layer MV may be disabled, for example, as described here. The mapped inter-layer MV is not available for MV prediction in the enhancement layer.
MV prediction including inter-layer MV may be disabled. For the enhancement layer, the time MV can be predicted based on other time MVs (for example, only the time MV). This is equivalent to MV prediction for a single-layer codec.
Devices (for example, processors, encoders, decoders, WTRUs, etc.) can be connected Receive a bit stream (e.g., a scalable bit stream). For example, the bitstream may include a base layer and one or more enhancement layers. The TMVP can be used to decode the base layer (e.g., base layer video block) and/or the enhancement layer (e.g., enhancement layer video block) in the bitstream. TMVP can be performed on the base layer and enhancement layer in the bit stream. For example, TMVP may be performed on the base layer (for example, the base layer video block) in the bit stream without any change, for example, refer to the description in FIG. 4. The inter-layer reference image may be used to perform TMVP on the enhancement layer (e.g., enhancement layer video block) in the bitstream, for example, as described herein. For example, the inter-layer reference image may be used as a collocated reference image for the TMVP of the enhancement layer (e.g., enhancement layer video block). For example, the compressed MV domain of the juxtaposed base layer image can be determined. The MV domain of the inter-layer reference image can be determined according to the compressed MV domain of the collocated base layer image. The MV domain of the inter-layer reference image can be used to perform TMVP on the enhancement layer (e.g., enhancement layer video block). For example, the MV domain of the inter-layer reference image can be used to predict the MV domain of an enhancement layer video block (for example, and an enhancement layer video block).
The MV domain of the inter-layer reference layer image can be determined. For example, the MV domain of the inter-layer reference layer image may be determined based on the MV domain of the collocated base layer image. The MV domain may include one or more MVs and/or reference image indexes. For example, the MV domain may include the reference image index and the MV of the PU in the inter-layer reference layer image (for example, for each PU in the inter-layer reference layer image). The enhancement layer image can be decoded based on the MV domain (e.g., collocated enhancement layer images). It can be based on the MV domain to perform TMVP on the enhancement layer image.
Syntax messaging (for example, advanced syntax messaging) for inter-layer motion prediction can be provided. At the sequence level, it is possible to enable or disable inter-layer movement information mapping and MV prediction. At the image/segment level, it is possible to enable or disable inter-layer movement information mapping and MV prediction. For example, the decision of whether to enable and/or disable certain inter-layer motion prediction technologies can be made based on the consideration of improving coding efficiency and/or reducing system complexity. For example, because the added grammar can be applied to the images of the sequence (for example, all images), the transmission at the sequence level is less expensive than the transmission at the image/segment level. For example, because the sequence of images (e.g., each image) can receive its own motion prediction implementation and/or MV prediction is implemented, and image/segment level communication can provide better flexibility.
Can provide serial-level messaging. Inter-layer movement information mapping and/or MV prediction can be signaled at the sequence level. If sequence-level signaling is used, the images in the sequence (for example, all images) can use the same motion information mapping and/or MV prediction. For example, the syntax shown in Table 2 may indicate whether to allow mobile information mapping and/or MV prediction between sequence-level layers. The syntax in Table 2 can be used on parameter sets, for example, such as video parameter set (VPS) (for example, in HEVC), sequence parameter set (SPS) (for example, in H.264 and HEVC), and Picture Parameter Set (PPS) (for example, in H.264 and HEVC) and so on, but not limited to this.
<tables><img he="872" wi="1870" file="tw201804792a_d0010.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></tables>
The inter-layer motion vector presence flag (inter_layer_mvp_present_flag) may indicate that the inter-layer motion prediction is used at the sequence level or at the image/slice level. For example, if the flag is set to 0, the communication is at the image/segment level. If the flag is set to 1, the motion mapping and/or MV prediction signaling is at the sequence level. The inter-layer motion mapping sequence enabling flag (inter_layer_motion_mapping_seq_enabled_flag) may indicate whether to use inter-layer motion mapping (for example, inter-layer motion prediction) at the sequence level. Adding a motion vector prediction sequence enabling flag (inter_layer_add_mvp_seq_enabled_flag) between layers can indicate whether block MV prediction (for example, additional block MV prediction) is adopted at the sequence level.
Provide image/segment level communication. Can be signaled between layers at the image/slice level Mobile information mapping. If image/segment level messaging is used, the images of the sequence (e.g., each image) can receive its own messaging. For example, images of the same sequence may use different mobile information mapping and/or MV prediction (for example, based on the received transmission). For example, the syntax in Table 3 can be used in the segment header to indicate whether the inter-layer movement information mapping and/or MV prediction are used for the current image/segment in the enhancement layer.
<tables><img he="1866" wi="1989" file="tw201804792a_d0011.tif" img-content="drawing" img-format="tif" orientation="portrait" inline="no" /></tables>
The inter-layer motion mapping segment enablement flag (inter_layer_motion_mapping_slice_enabled_flag) can be used to indicate whether to apply inter-layer motion mapping to the current segment. Adding the motion vector prediction slice enable flag (inter_layer_add_mvp_slice_enabled_flag) between layers can be used to indicate whether to apply to the current slice Additional block MV prediction.
It is proposed to use MV predictive coding in a multi-layer video coding system. The inter-layer mobile information mapping algorithm described here is used to generate mobile-related information for the processed base layer, for example, so that the TMVP process in the enhancement layer can utilize the correlation between the time MV of the base layer and the enhancement layer . Because the block-level operation can be unchanged, the single-layer encoder and decoder can be used for MV prediction in the enhancement layer without any change. The MV prediction may be based on the analysis of the characteristics of different types of MVs in the adjustable system (for example, for the purpose of improving the efficiency of MV prediction).
Although the two-layer SVC system with spatial tunability is described here, the content of this disclosure can be extended to SVC systems with more than two layers and other tunability modes.
Inter-layer motion prediction can be performed on the enhancement layer in the bitstream. The inter-layer movement prediction can be signaled, for example, as described here. Inter-layer motion prediction can be signaled at the sequence level of the bitstream (for example, using inter_layer_motion_mapping_seq_enabled_flag, etc.). For example, the inter-layer motion prediction can be signaled via variables in the video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), etc. in the bit stream.
A device (e.g., processor, encoder, decoder, WTRU, etc.) can perform any of the functions described herein. For example, the encoder may include a processor configured to receive a bit stream (e.g., an adjustable bit stream). The bit stream can include a base layer and an enhancement layer. The decoder may decode the enhancement layer in the bitstream using Temporal Motion Vector Prediction (TMVP) that uses the inter-layer reference picture as a collocated reference picture for the TMVP of the enhancement layer. The enhancement layer video blocks, the inter-layer video blocks, and/or the base layer video blocks may be collocated (for example, collocated in time).
The decoder can use TMVP to decode the enhancement layer image. For example, the decoder may determine the MV domain of the inter-layer reference image based on the MV domain of the collocated base layer image. The inter-layer reference image and the enhancement layer reference image are collocated. The MV field of the inter-layer reference image may include the reference image index and MV of the video block of the inter-layer reference image. The decoder may decode the enhancement layer image based on the MV domain of the inter-layer reference image. For example, the decoder may determine the MV domain of the enhancement layer image based on the MV domain of the inter-layer reference image and decode the enhancement layer image based on the MV domain of the enhancement layer image.
The MV domain of the inter-layer reference image may be determined based on the compressed MV domain. For example, the decoder may determine the compressed MV domain of the collocated base layer image and determine the MV domain of the inter-layer reference image based on the compressed MV domain of the collocated base layer image.
The decoder can determine the MV and reference image of the video block of the inter-layer reference image. For example, the decoder may determine the reference image of the inter-layer video block based on the reference image of the collocated base layer video block. The decoder may determine the MV of the inter-layer video block based on the MV of the collocated base layer video block. The decoder can determine the juxtaposed base layer video block by selecting the video block of the juxtaposed base layer image characterized by the largest overlap area of the inter-layer video block. The decoder may determine the MV and/or reference image of the video block of the enhancement layer (for example, the collocated video block of the enhancement layer image) based on the MV and/or reference image of the video block of the inter-layer reference image.
The decoder may determine the reference image of the juxtaposed base layer video block, and determine the reference image of the inter-layer video block based on the reference image of the juxtaposed base layer video block. For example, the reference image of the inter-layer video block may be a juxtaposed inter-layer reference image of the reference image of the base layer video block. The decoder may determine the reference image of the video block of the enhancement layer image based on the reference image of the inter-layer video block. For example, the reference image of the enhancement layer may be a collocated enhancement layer reference image of the reference image of the inter-layer video block. The enhancement layer video blocks, the inter-layer video blocks, and/or the base layer video blocks may be collocated (for example, collocated in time).
The decoder can determine the MV of the video block between layers. For example, the decoder may determine the MV of the collocated base layer video block, and scale the MV of the collocated base layer video block according to the spatial ratio between the base layer and the enhancement layer to determine the MV of the inter-layer video block. The decoder may determine the MV of the enhancement layer video block based on the MV of the inter-layer video block. For example, the decoder can use the MV of the inter-layer video block to predict the MV of the enhancement layer video block, for example, by time scaling the MV of the inter-layer video block.
The decoder may be configured to determine the reference image of the enhancement layer video block based on the collocated base layer video block, determine the MV of the enhancement layer video block based on the MV of the collocated base layer video block, and/or determine the reference image based on the enhancement layer video block The MV of the image and the enhancement layer video block is used to decode the enhancement layer video block. For example, the decoder can select the maximum overlap area with the enhancement layer video block as a special feature. The video blocks of the base-layer image can be collocated to determine the collocated base-layer video blocks.
The decoder can determine the reference image of the collocated base layer video block. The decoder may use the reference images of the base layer video blocks to be juxtaposed to determine the reference images of the inter-layer video blocks. The decoder can determine the reference image of the enhancement layer video block. For example, the reference image of the enhancement layer video block may be a collocated enhancement layer image of the reference image of the collocated base layer video block and the reference image of the collocated inter-layer video block. The enhancement layer video blocks, the inter-layer video blocks, and/or the base layer video blocks may be collocated (for example, collocated in time).
The decoder can determine the MV of the collocated base layer video block. The decoder may scale and concatenate the MV of the base layer video block according to the spatial ratio between the base layer and the enhancement layer to determine the MV of the inter-layer video block. The decoder can predict the MV of the enhancement layer video block based on the MV of the inter-layer video block, for example, by time scaling the MV of the inter-layer video block.
The decoder may include a processor that can receive a bit stream. The bit stream can include a base layer and an enhancement layer. The bit stream may include inter-layer movement mapping information. The decoder may determine that the inter-layer motion prediction for the enhancement layer can be enabled based on the inter-layer mapping information. The decoder can perform the inter-layer motion prediction of the enhancement layer based on the inter-layer mapping information. The inter-layer mapping information can be signaled at the sequence level of the bit stream. For example, the inter-layer mapping information can be signaled by variables (such as flags) in the VPS, SPS, and/or PPS in the bitstream.
Although described from the perspective of a decoder, the functions described here (for example, the reverse function of the functions described here) can be performed by other devices, such as an encoder.
Figure 14A is a system diagram of an example communication system 1400 in which one or more embodiments may be implemented. The communication system 1400 may be a multiple access system that provides content, such as voice, data, video, message sending, broadcast, etc., to multiple users. The communication system 1400 can allow multiple wireless users to access these contents through system resource sharing (including wireless bandwidth). For example, the communication system 1400 may use one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), Single carrier FMDA (SC-FDMA) etc.
As shown in Figure 14A, the communication system 1400 may include wireless transmission/reception units (WTRU) 1402a, 1402b, 1402c, and/or 1402d (often collectively referred to as WTRU 1402), and a radio access network (RAN) 1403/1404/ 1405, core network 1406/1407/1409, public switched telephone network (PSTN) 1408, Internet 1410 and other networks 1412. However, it will be understood that the disclosed implementation takes into account any number of WTRUs, base stations, networks, and/or network elements. Each of the WTRUs 1402a, 1402b, 1402c, and 1402d may be any type of device configured to operate and/or communicate in a wireless environment. As an example, the WTRU 1402a, 1402b, 1402c, 1402d may be configured to transmit and/or receive wireless signals, and may include user equipment (UE), base stations, fixed or mobile user units, pagers, mobile phones, personal Digital assistants (PDA), smart phones, notebook computers, portable easy-to-net machines, personal computers, wireless sensors, consumer electronics, etc.
The communication system 1400 may also include a base station 1414a and a base station 1414b. Each of the base stations 1414a, 1414b may be configured to wirelessly interface with at least one of the WTRUs 1402a, 1402b, 1402c, and 1402d to facilitate access to one or more communication networks, such as the core network 1406/1407, 1409, Internet 1410 and/or any device type of network 1412. As an example, the base stations 1414a, 1414b may be base station transceiver stations (BTS), node B), evolved node B (eNodeB), local node B, local eNB, website controller, access point (AP), Wireless routers and so on. Although each of the base stations 1414a, 1414b is described as a single element, it will be understood that the base stations 1414a, 1414b may include any number of interconnected base stations and/or network elements.
The base station 1414a may be a part of RAN 1403/1404/1405, and RAN 1403/1404/1405 may also include other base stations and/or network components (not shown), such as base station controller (BSC), radio network control Device (RNC), relay node, etc. The base station 1414a and/or the base station 1414b may be configured to transmit and/or receive wireless signals within a specific geographic area, which may be referred to as a cell (not shown). Cells can also be cut into cell sectors. For example, the cell associated with the base station 1414a can be cut into three sectors. Therefore, in a In an embodiment, the base station 1414a may include three transceivers, that is, one sector for each cell. In another embodiment, the base station 1414a can use multiple input multiple output (MIMO) technology, so multiple transceivers can be used for each sector of the cell.
The base stations 1414a, 1414b can communicate with one or more of the WTRUs 1402a, 1402b, 1402c, 1402d via the air interface 1415/1416/1417, and the air interface 1415/1416/1417 can be any suitable wireless communication link (For example, radio frequency (RF), microwave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish the air interface 1415/1416/1417.
More specifically, as described above, the communication system 1400 may be a multiple access system, and may use one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and so on. For example, the base station 1414a and WTRU 1402a, 1402b, 1402c in RAN 1403/1404/1405 may use radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may use wideband CDMA (WCDMA) to Create an air interface 1415/1416/1417. WCDMA may include communication protocols such as high-speed packet access (HSPA) and/or evolved HSPA (HSPA+). HSPA may include High Speed Downlink Packet Access (HSDPA) and/or High Speed Uplink Packet Access (HSUPA).
In another embodiment, the base station 1414a and the WTRUs 1402a, 1402b, 1402c may use radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may use Long Term Evolution (LTE) and/or LTE Advanced ( LTE-A) to establish the air interface 1415/1416/1417.
In other embodiments, the base station 1414a and WTRU 1402a, 1402b, 1402c may use, for example, IEEE 802.16 (ie, Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS -2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), GSM EDGE (GERAN), etc. Technology.
The base station 1414b in Figure 14A can be a wireless router, a local Node B, a local eNode B, or an access point, for example, and can use any appropriate RAT to facilitate local areas such as commercial places, residences, vehicles, campuses, etc. Wireless connection in the area. In one embodiment, the base station 1414b and the WTRUs 1402c, 1402d may implement radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, the base station 1414b and the WTRUs 1402c, 1402d may use radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In another embodiment, the base station 1414b and the WTRUs 1402c, 1402d may use a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, etc.) to establish a pico cell or a femto cell. As shown in FIG. 14A, the base station 1414b may have a direct connection to the Internet 1410. Therefore, the base station 1414b may not need to access the Internet 1410 through the core network 1406/1407/1409.
RAN 1403/1404/1405 may communicate with core network 1406/1407/1409, which may be configured to provide voice to one or more of WTRU 1402a, 1402b, 1402c, 1402d , Data, applications, and/or Internet Protocol-based Voice over Internet Protocol (VoIP) services, etc. of any type of network. For example, the core network 1406/1407/1409 can provide call control, billing services, mobile location-based services, prepaid calls, Internet connections, video distribution, etc. and/or perform high-level security functions, such as user authentication. Although not shown in Figure 14A, it will be understood that the RAN 1403/1404/1405 and/or the core network 1406/1407/1409 may be the same RAT as the RAN 1403/1404/1405 or other RATs that use different RATs. RAN communicates directly or indirectly. For example, in addition to connecting to the RAN 1403/1404/1405 that is using the E-UTRA radio technology, the core network 1406/1407/1409 can also communicate with another RAN (not shown) that uses the GSM radio technology.
The core network 1406/1407/1409 can also serve as a gateway for the WTRU 1402a, 1402b, 1402c, and 1402d to access the PSTN 1408, the Internet 1410, and/or other networks 1412. PSTN 1408 may include a circuit-switched telephone network that provides plain old telephone service (POTS). The Internet 1410 may include a global system of interconnecting computer networks and devices using public communication protocols. Such protocols include, for example, the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP) and the Internet Protocol (IP) in the TCP/IP Internet Protocol Group. The network 1412 may include a wired or wireless communication network owned and/or operated by other service providers. For example, the network 1412 may include another core network connected to one or more RANs, which may use the same RAT as the RAN 1403/1404/1405 or a different RAT.
Some or all of the WTRUs 1402a, 1402b, 1402c, and 1402d in the communication system 1400 may include multi-mode capabilities. Multiple transceivers. For example, the WTRU 1402c shown in Figure 14A may be configured to communicate with base station 1414a and base station 1414b, which may use cellular-based radio technology, and base station 1414b may use IEEE 802 radio technology.
Figure 14B is a system diagram of an example WTRU 1402. As shown in Figure 14B, the WTRU 1402 may include a processor 1418, a transceiver 1420, a transmission/reception element 1422, a speaker/microphone 1424, a keyboard 1426, a display/touchpad 1428, a non-removable memory 1430, a removable Memory 1432, power supply 1434, global positioning system (GPS) chipset 1436 and other peripheral devices 1438. It will be understood that the WTRU 1402 may include any sub-combination of the foregoing elements while remaining consistent with the implementation. Similarly, the base station 1414a and base station 1414b and/or the nodes represented by the base station 1414a and base station 1414b, such as but not limited to base station transceiver station (BTS), node B, website controller, access Point (AP), Local Node B, Evolved Local Node B (eNodeB), Local Evolved Node B (HeNB), Gateway and Proxy Node of Local Evolved Node B, which may include the parts described in Figure 14B and here Or all components.
The processor 1418 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DSP core, a controller, a micro Controller, dedicated integrated circuit (ASIC), field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), state machine, etc. The processor 1418 may perform signal encoding, data processing, power control, input/output processing, and/or enable the WTRU 1402 to operate wirelessly Any other functions operated in the environment. The processor 1418 may be coupled to a transceiver 1420, which may be coupled to a transmission/reception element 1422. Although FIG. 14B describes the processor 1418 and the transceiver 1420 as separate components, it will be understood that the processor 1418 and the transceiver 1420 may be integrated together in an electronic package or chip.
The transmission/reception element 1422 may be configured to transmit signals to or receive signals from a base station (e.g., base station 1414a) via the air interface 1415/1416/1417. For example, in one embodiment, the transmission/reception element 1422 may be an antenna configured to transmit and/or receive RF signals. In another embodiment, the transmission/reception element 1422 may be a transmitter/detector configured to transmit and/or receive, for example, IR, UV, or visible light signals. In another embodiment, the transmission/reception element 1422 may be configured to transmit and receive both RF and optical signals. It should be understood that the transmission/reception element 1422 may be configured to transmit and/or receive any combination of wireless signals.
In addition, although the transmission/reception element 1422 is described as a separate element in Figure 14B, the WTRU 1402 may include any number of transmission/reception elements 1422. More specifically, the WTRU 1402 may use, for example, MIMO technology. Therefore, in one embodiment, the WTRU 1402 may include two or more transmission/reception elements 1422 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 1415/1416/1417.
The transceiver 1420 may be configured to modulate the signal to be transmitted by the transmission/reception element 1422 and/or demodulate the signal received by the transmission/reception element 1422. As described above, the WTRU 1402 may have multi-mode capabilities. The transceiver 1420 may therefore include multiple transceivers that enable the WTRU 1402 to communicate via multiple RATs such as UTRA and IEEE 802.11.
The processor 1418 of the WTRU 1402 may be coupled to the following devices, and may receive user input data from the following devices: speaker/microphone 1424, keyboard 1426, and/or display/touch panel 1428 (e.g., liquid crystal display (LCD)) Display unit or organic light emitting diode (OLED) display unit). The processor 1418 can also output user data to the speaker/microphone 1424, the keyboard 1426, and/or the display/touchpad 1428. In addition, the processor 1418 can be from any class A type of appropriate memory accesses information, and can store data in any type of appropriate memory, such as non-removable memory 1430 and/or removable memory 1432. The non-removable memory 1430 may include random access memory (RAM), read-only memory (ROM), hard disk, or any other type of memory device. The removable memory 1432 may include a user identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and so on. In other embodiments, the processor 1418 may access information from a memory that is not physically located on the WTRU 1402, such as a server or a home computer (not shown), and may store data in the memory. middle.
The processor 1418 may receive power from the power supply 1434 and may be configured to distribute and/or control power to other components in the WTRU 1402. The power supply 1434 may be any suitable device for powering the WTRU 1402. For example, the power supply 1434 may include one or more dry batteries (for example, nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), solar cells, fuel cells, etc. Wait.
The processor 1418 may also be coupled to a GPS chipset 1436, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 1402. In addition, in addition to or as an alternative to the information from the GPS chipset 1436, the WTRU 1402 can receive location information from base stations (e.g., base stations 1414a, 1414b) via the air interface 1415/1416/1417 and/or based on two or The timing of the signals received by more neighboring base stations determines its location. It will be understood that while maintaining the consistency of the implementation, the WTRU 1402 may use any suitable location determination method to obtain location information.
The processor 1418 may be coupled to other peripheral devices 1438, which may include one or more software and/or hardware modules that provide additional features, functions, and/or wired or wireless connections. For example, peripheral devices 1438 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos or videos), universal serial diversion (USB) ports, vibration devices, TV transceivers, hands-free headsets, Bluetooth (BluetoothR) module, frequency modulation (FM) radio unit, digital music player, media player, video game console module, Internet browsing Device and so on.
Figure 14C is a structural diagram of the RAN 1403 and the core network 1406 according to an embodiment. As described above, for example, the RAN 1403 may use UTRA radio technology to communicate with the WTRUs 1402a, 1402b, and 1402c via the air interface 1415. The RAN 1403 can also communicate with the core network 1406. As shown in Figure 14C, the RAN 1403 may include Node Bs 1440a, 1440b, 1440c, and each of the Node Bs 1440a, 1440b, 1440c includes one or more for communicating with the WTRU 1402a, 1402b, 1402c via the air interface 1415 Transceiver. Each of the Node Bs 1440a, 1440b, 1440c may be associated with a specific cell (not shown) within the RAN 1403. The RAN 1403 may also include RNC 1442a, 1442b. It will be understood that the RAN 1403 may include any number of Node Bs and RNCs while maintaining the consistency of the implementation.
As shown in Figure 14C, Node Bs 1440a and 1440b can communicate with RNC 1442a. In addition, Node B 1440c can communicate with RNC 1442b. Node Bs 1440a, 1440b, 1440c can communicate with RNC 1442a, 1442b via the Iub interface, respectively. RNC 1442a, 1442b can communicate with each other via the Iur interface. Each of the RNC 1442a, 1442b can be configured to control the respective Node B 1440a, 1440b, 1440c to which it is connected. In addition, each of RNC 1442a, 1442b can be configured to perform or support other functions, such as outer loop power control, load control, admission control, packet scheduling, handover control, macro diversity, security functions, data encryption, and so on.
The core network 1406 shown in Figure 14C may include a media gateway (MGW) 1444, a mobile switching center (MSC) 1446, a serving GPRS support node (SGSN) 1448, and/or a gateway GPRS support node (GGSN) 1450 . Although each of the aforementioned elements is described as part of the core network 1406, it will be understood that any of these elements may be owned or operated by entities other than the core network operator.
The RNC 1442a in the RAN 1403 can be connected to the MSC 1446 in the core network 1406 via the IuCS interface. MSC 1446 can be connected to MGW 1444. MSC 1446 and MGW 1444 can provide WTRUs 1402a, 1402b, 1402c with access to circuit-switched networks such as PSTN 1408 to facilitate WTRUs 1402a, 1402b, 1402c and traditional landlines Communication between communication devices.
The RNC 1442a in the RAN 1403 can also be connected to the SGSN 1448 in the core network 1406 via the IuPS interface. SGSN 1448 can be connected to GGSN 1450. SGSN 1448 and GGSN 1450 may provide WTRUs 1402a, 1402b, 1402c with access to a packet-switched network such as the Internet 1410 to facilitate communication between WTRUs 1402a, 1402b, 1402c and IP-enabled devices.
As described above, the core network 1406 may also be connected to the network 1412, and the network 1412 may include other wired or wireless networks owned or operated by other service providers.
Figure 14D is a structural diagram of the RAN 1404 and the core network 1407 according to an embodiment. As described above, for example, the RAN 1404 may use E-UTRA radio technology to communicate with the WTRUs 1402a, 1402b, and 1402c via the air interface 1416. The RAN 1404 can also communicate with the core network 1407.
The RAN 1404 may include eNodeBs 1460a, 1460b, and 1460c, but it will be understood that the RAN 1404 may include any number of eNodeBs while maintaining consistency with various embodiments. Each of the eNode Bs 1460a, 1460b, and 1460c may include one or more transceivers for communicating with the WTRU 1402a, 1402b, 1402c via the air interface 1416. In an embodiment, eNodeBs 1460a, 1460b, and 1460c may implement MIMO technology. Thus, for example, the eNodeB 1460a may use multiple antennas to transmit wireless signals to and/or receive wireless signals from the WTRU 1402a.
Each of the eNodeBs 1460a, 1460b, and 1460c can be associated with a specific cell (not shown), and can be configured to handle wireless resource management decisions, handover decisions, user scheduling in uplink and/or downlink, etc. Wait. As shown in Figure 14D, eNodeBs 1460a, 1460b, and 1460c can communicate with each other via the X2 interface.
The core network 1407 shown in FIG. 14D may include a mobility management entity (MME) 1462, a service gateway 1464, and/or a packet data network (PDN) gateway 1466. Although each of the foregoing units is described as part of the core network 1407, it will be understood that these Any of the units can be owned and/or operated by entities other than the core network operator.
The MME 1462 can be connected to each of the eNodeBs 1460a, 1460b, and 1460c in the RAN 1404 via the S1 interface, and can act as a control node. For example, the MME 1462 may be responsible for user authentication of WTRUs 1402a, 1402b, 1402c, bearer activation/deactivation, selection of a specific service gateway during the initial connection of WTRUs 1402a, 1402b, 1402c, and so on. The MME 1462 may also provide a control plane function for switching between the RAN 1404 and other RANs (not shown) that use other radio technologies such as GSM or WCDMA.
The service gateway 1464 may be connected to each of the eNodeBs 1460a, 1460b, and 1460c in the RAN 104b via the S1 interface. The service gateway 1464 can generally route and forward user data packets to/from WTRUs 1402a, 1402b, and 1402c. The service gateway 1464 may also perform other functions, such as anchoring the user plane during inter-eNB handover, triggering paging when downlink data is available for WTRU 1402a, 1402b, 1402c, managing and storing the context of WTRU 1402a, 1402b, 1402c and many more.
Service gateway 1464 can also be connected to PDN gateway 1466. PDN gateway 1466 can provide WTRU 1402a, 1402b, 1402c with access to a packet switching network (such as Internet 1410) for the convenience of WTRU 1402a, 1402b, 1402c Communication with IP enabling devices.
The core network 1407 can facilitate communication with other networks. For example, the core network 1406 may provide WTRUs 1402a, 1402b, and 1402c with access to a circuit-switched network (such as PSTN 1408) to facilitate communication between the WTRUs 1402a, 1402b, and 1402c and traditional landline communication devices. For example, the core network 1407 may include or communicate with an IP gateway (for example, an IP Multimedia Subsystem (IMS) server), and the IP gateway serves as an interface between the core network 1407 and the PSTN 1408. In addition, the core network 1407 may provide WTRUs 1402a, 1402b, 1402c with access to the network 1412, which may include other wired or wireless networks owned and/or operated by other service providers.
Figure 14E is a structural diagram of the RAN 1405 and the core network 1409 according to an embodiment. The RAN 1405 may use IEEE 802.16 radio technology to communicate with each other via the air interface 1417 The Access Service Network (ASN) through which WTRUs 1402a, 1402b, and 1402c communicate. As discussed further below, the links between different functional entities of WTRU 1402a, 1402b, 1402c, RAN 1405 and core network 1409 may be defined as reference points.
As shown in Figure 14E, RAN 1405 may include base stations 1480a, 1480b, 1480c and ASN gateways 1482, but it will be understood that RAN 1405 may include any number of base stations and ASN gateways to be consistent with the implementation . Each of the base stations 1480a, 1480b, and 1480c may be associated with a specific cell (not shown) in the RAN 1405 and may include one or more transceivers that communicate with the WTRU 1402a, 1402b, 1402c via the air interface 11417. In an example, the base stations 1480a, 1480b, and 1480c may implement MIMO technology. Thus, for example, the base station 1480a may use multiple antennas to transmit wireless signals to or receive wireless signals from the WTRU 1402a. Base stations 1480a, 1480b, and 1480c can provide mobility management functions, such as call handoff triggering, tunnel establishment, radio resource management, traffic classification, service quality policy execution, and so on. The ASN gateway 1482 can act as a traffic aggregation point, and is responsible for paging, caching user profiles, routing to the core network 1409, and so on.
The air interface 1417 between the WTRU 1402a, 1402b, 1402c and the RAN 1405 may be defined as the R1 reference point for implementing the IEEE 802.16 specification. In addition, each of the WTRUs 1402a, 1402b, and 1402c may establish a logical interface with the core network 1409 (not shown). The logical interface between the WTRU 1402a, 1402b, 1402c and the core network 1409 can be defined as an R2 reference point, which can be used for authentication, authorization, IP host configuration management, and/or mobility management.
The communication link between each of the base stations 1480a, 1480b, and 1480c may be defined as an R8 reference point that includes an agreement to facilitate WTRU handover and data transfer between base stations. The communication link between the base stations 1480a, 1480b, 1480c and the ASN gateway 1482 can be defined as the R6 reference point. The R6 reference point may include agreements to facilitate mobility management based on mobility events associated with each of WTRUs 1402a, 1402b, 1402c.
As shown in Figure 14E, the RAN 1405 can be connected to the core network 1409. The communication link between the RAN 1405 and the core network 1409 can be defined as including, for example, facilitating data transfer and The R3 reference point for the agreement on mobility management capabilities. The core network 1409 may include a mobile IP home agent (MIP-HA) 1484, an authentication, authorization, and accounting (AAA) server 1486, and a gateway 1488. Although each of the aforementioned elements is described as part of the core network 1409, it will be understood that any of these elements may be owned or operated by entities other than the core network operator.
MIP-HA can be responsible for IP address management, and can enable WTRUs 1402a, 1402b, and 1402c to roam between different ASNs and/or different core networks. The MIP-HA 1484 can provide WTRUs 1402a, 1402b, and 1402c with access to a packet-switched network (such as the Internet 1410) to facilitate communication between the WTRUs 1402a, 1402b, and 1402c and IP-enabled devices. The AAA server 1486 may be responsible for user authentication and supporting user services. Gateway 1488 can facilitate intercommunication with other networks. For example, the gateway 1488 may provide WTRUs 1402a, 1402b, and 1402c with access to a circuit-switched network (such as PSTN 1408) to facilitate communication between the WTRUs 1402a, 1402b, and 1402c and traditional landline communication devices. In addition, the gateway 1488 may provide the WTRU 1402a, 1402b, 1402c with a network 1412, which may include other wired or wireless networks owned or operated by other service providers.
Although not shown in Figure 14E, it will be understood that the RAN 1405 can be connected to other ASNs, and the core network 1409 can be connected to other core networks. The communication link between the RAN 1405 and other ASNs may be defined as the R4 reference point, which may include an agreement to coordinate the mobility of WTRUs 1402a, 1402b, and 1402c between the RAN 1405 and other ASNs. The communication link between the core network 1409 and other core networks can be defined as an R5 reference, which can include an agreement to promote intercommunication between the local core network and the visited core network.
Figure 15 is a block diagram showing a block-based video codec (for example, a hybrid video codec). The input video signal 1502 can be processed block by block. The video block unit includes 16x16 pixels. Such a block unit may be called a macro block (MB). In High Efficiency Video Coding (HEVC), an extended block size (for example, referred to as "coding unit" or CU) can be used to effectively compress high-resolution (for example, greater than and equal to 1080p) video signals. In HEVC, the CU can be up to 64x64 pixels. CU can be cut into prediction units (PU). For prediction units Meta can use the method of individual prediction.
For input video blocks (for example, MB or CU), spatial prediction 1560 and/or temporal prediction 1562 can be performed. Spatial prediction (eg, "intra prediction") can use pixels in adjacent blocks that have been coded in the same video image/slice to predict the current video block. Spatial prediction can reduce the inherent spatial redundancy in the video signal. Temporal prediction (for example, "inter prediction" or "motion compensation prediction") may use pixels of a video image that has been coded (for example, it may be referred to as a "reference image") to predict the current video block. Time prediction can reduce the inherent time redundancy in the signal. The temporal prediction of the video block can be signaled via one or more motion vectors, and the motion vector can be used to indicate the amount and/or direction of the motion between its prediction block and the current block in the reference image. If multiple reference images are supported (for example, in the case of H.264/AVC and/or HEVC), for each video block, its reference image index will be additionally sent. The reference picture index may be used to identify which reference picture in the reference picture memory 1564 (for example, it may be referred to as a "decoded picture buffer" or DPB) that the temporal prediction signal comes from.
After spatial and/or temporal prediction, the mode decision block 1580 in the encoder can select a prediction mode. The prediction block is subtracted from the current video block 1516. The predicted residual can be converted 1504 and/or quantized by 1506. The quantized residual coefficient may be inversely quantized by 1510 and/or inversely transformed 1512 to form a reconstructed residual, and then the reconstructed residual is added back to the prediction block 1526 to form a reconstructed video block.
Before the reconstructed video block is put into the reference image storage 1564 and/or used to encode the subsequent video block, for example, but not limited to, deblocking filter and sample adaptive offset can be applied to the reconstructed video block And/or in-loop filtering 1566 such as adaptive loop filter. To form the output video bit stream 1520, the coding mode (for example, inter prediction mode or intra prediction mode), prediction mode information, motion information, and/or quantized residual coefficients will be sent to the entropy coding unit 1508 for compression And/or shrink to form a bit stream.
Fig. 16 is a diagram showing an example of a block-based video decoder. The video bitstream 1602 is disassembled and/or entropy decoded in the entropy decoding unit 1608. The coding mode and/or prediction information can be It is sent to the spatial prediction unit 1660 (for example, if it is intra coding) and/or the temporal prediction unit 1662 (for example, if it is inter coding) to form a prediction block. If it is inter coded, the prediction information may include the prediction block size, one or more motion vectors (for example, it may be used to indicate the direction and amount of motion), and/or one or more reference indexes (for example, It can be used to indicate from which reference image the prediction signal is obtained).
The temporal prediction unit 1662 may apply motion compensation prediction to form a temporal prediction block. The residual conversion coefficient may be sent to the inverse quantization unit 1610 and the inverse conversion unit 1612 to reconstruct the residual block. In 1626, the prediction block and the residual block are added. The reconstructed block is first passed through inner loop filtering before being stored in the reference image storage 1664. The reconstructed video in the reference image storage 1664 can be used to drive the display device and/or used to predict future video blocks.
A single-layer video codec can use a single video sequence input and generate a single compressed bit stream that is sent to the single-layer decoder. Video codecs can be designed for digital video services (for example, such as but not limited to sending TV signals via satellite, cable, and terrestrial transmission channels). With the development of video center applications in a heterogeneous environment, multi-layer video coding technology as an extension of the video standard can be developed to enable various applications. For example, tunable video coding technology is designed to handle more than one video layer, where each layer can be decoded to reconstruct a video with a specific spatial resolution, temporal resolution, fidelity, and/or view. Video signal. Although single-layer encoders and decoders have been described with reference to Figures 15 and 16, the concepts described here also utilize multi-layer encoders and decoders, for example, for multi-layer or tunable coding techniques. The encoder in Figure 15 and/or the decoder in Figure 16 can perform any of the functions described here. For example, the encoder of Figure 15 and/or the decoder of Figure 16 may use the MV of the enhancement layer PU to perform TMVP on the enhancement layer (for example, an enhancement layer image).
Fig. 17 is a diagram showing an example of a communication system. The communication system 1700 may include an encoder 1702, a communication network 1704, and a decoder 1706. The encoder 1702 can communicate with the communication network 1704 via the connection 1708. The connection 1708 may be a wired connection or a wireless connection. The encoder 1702 is similar to the block-based video codec in Figure 15. Encoder 1702 may include a single-layer codec (For example, as shown in Figure 15) or multi-layer codec.
The decoder 1706 can communicate with the communication network 1704 via the connection 1710. The connection 1710 may be a wired connection or a wireless connection. The decoder 1706 is similar to the block-based video decoder in Figure 16. The decoder 1706 may include a single-layer codec (for example, as shown in FIG. 16) or a multi-layer codec. The encoder 1702 and/or the decoder 1706 can be combined with any of a variety of wired communication devices and/or wireless transmission/reception units (WTRU), such as, but not limited to, digital televisions, wireless broadcasting systems, and network components / Terminal, server (for example, content or website server (for example, such as hypertext transfer protocol (HTTP) server), personal digital assistant (PDA), laptop or desktop computer, tablet computer, digital camera, Digital recording devices, video game devices, video game consoles, cellular or satellite wireless phones, digital media players, etc.
The communication network 1704 is suitable for communication systems. For example, the communication network 1704 may be a multiple access system that provides content (for example, voice, data, video, messaging, broadcast, etc.) to multiple wireless users. The communication network 1704 can enable multiple wireless users to access these contents through system resource sharing (including wireless bandwidth). For example, the communication network 1704 may use one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), Single carrier FMDA (SC-FDMA) etc.
The method described here can be implemented by a computer program, software, or firmware, which can be included in a computer-readable medium executed by a computer or a processor. Examples of computer-readable media include electronic signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (for example, internal hard drives) And removable disks), magneto-optical media and optical media, such as compact discs (CD) or digital versatile discs (DVD). The processor associated with the software is used to implement a radio frequency transceiver for WTRU, UE, terminal, base station, RNC or any host computer.
11 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11
37 members in 9 offices
Priority claims15
| Document | Office | Kind | Date |
|---|---|---|---|
| 201261694555 | United States of America | P | |
| 201261694555 | United States of America | P | |
| 61694555 | United States of America | – | |
| 201261734650 | United States of America | P | |
| 201261734650 | United States of America | P | |
| 61734650 | United States of America | – | |
| 201361866822 | United States of America | P | |
| 201361866822 | United States of America | P | |
| 61866822 | United States of America | – | |
| 201261694555P | – | – | – |
| 201261734650P | – | – | – |
| 201361866822P | – | – | – |
| US201261694555P | – | – | – |
| US201261734650P | – | – | – |
| US201361866822P | – | – | – |
Members37
| Document | Office | Kind | |
|---|---|---|---|
| US2014064374A1 | United States of America | A1 | |
| WO2014036259A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201424390A | Taiwan Province of China | A | |
| AU2013308719A1 | Australia | A1 | |
| KR20150046228A | Republic of Korea | A | |
| CN104604230A | China | A | |
| EP2891311A1 | European Patent Office (EPO) | A1 | |
| JP2015529420A | Japan | A | |
| MX2015002536A | Mexico | A | |
| AU2013308719B2 | Australia | B2 | |
| AU2016201903A1 | Australia | A1 | |
| JP5961761B2 | Japan | B2 | |
| MX341900B | Mexico | B | |
| JP2016213857A | Japan | A | |
| KR101754999B1 | Republic of Korea | B1 | |
| KR20170081741A | Republic of Korea | A | |
| JP6220013B2 | Japan | B2 | |
| AU2016201903B2 | Australia | B2 | |
| TW201804792AThis record | Taiwan Province of China | A | |
| US9900593B2 | United States of America | B2 | |
| JP2018029361A | Japan | A | |
| CN104604230B | China | B | |
| US2018131952A1 | United States of America | A1 | |
| CN108156463A | China | A | |
| TWI637625B | Taiwan Province of China | B | |
| JP6431966B2 | Japan | B2 | |
| TWI646822B | Taiwan Province of China | B | |
| KR101955700B1 | Republic of Korea | B1 | |
| KR20190025758A | Republic of Korea | A | |
| EP3588958A1 | European Patent Office (EPO) | A1 | |
| KR102062506B1 | Republic of Korea | B1 | |
| US10939130B2 | United States of America | B2 | |
| US2021120257A1 | United States of America | A1 | |
| US11343519B2 | United States of America | B2 | |
| CN108156463B | China | B | |
| CN115243046A | China | A | |
| EP3588958B1 | European Patent Office (EPO) | B1 |
Numbers
- Publication
- 201804792
- Publication, DOCDB
- 201804792
- Publication, EPODOC
- TW201804792
- Application
- 106126144
- Application, DOCDB
- 106126144
- Application, EPODOC
- TW20176126144
Titles4
- English
- METHOD AND APPARATUS OF MOTION VECTOR PREDICTION FOR SCALABLE VIDEO CODING
- Chinese
- 可條整視訊編碼移動向量預測的方法及裝置
- English
- METHOD AND APPARATUS OF MOTION VECTOR PREDICTION FOR SCALABLE VIDEO CODING
- English
- Method and device for predicting motion vector of slicing video coding
Classification
- CPC, 11
- H04N19/31
- H04N19/30
- H04N19/52
- H04N19/70
- H04N19/46
- H04N19/51
- H04N19/33
- H04N19/587
- H04N19/59
- H04N19/105
- H04N19/139
- IPC, 1
- H04N19 00