Video encoding of screen content data.
Abstract
Una entrada para un codificador recibe datos de imagen en movimiento que comprenden una secuencia de cuadros a ser codificados, cada cuadro comprende una pluralidad de bloques en dos dimensiones, cada bloque comprende una pluralidad de pixeles en esas dos dimensiones. Un módulo de predicción de movimiento lleva a cabo la codificación al por lo menos parte de cada uno de la pluralidad de cuadros, codificar cada bloque con relación a la porción de referencia respectiva de otro cuadro de la secuencia, la porción de referencia respectiva está desplazada desde el bloque por un vector de movimiento respectivo. De conformidad con la presente invención, los datos de imagen en movimiento de la pluralidad de cuadros comprende una corriente de captura de pantalla y el módulo de predicción de movimiento está configurado para restringir cada uno de los vectores de movimiento de la corriente de captura de pantalla en un número entero de pixeles en por lo menos una de las dimensiones.

Term
8.2 yearsleft in the term
Expires 19 December 2034.
- Priority
- Filed
- Granted
- Today
- Expires
21 claims: 5 independent, 16 dependent
- 1REIVINDICACIONES 1. Un sistema codificador para codificar datos de imagen en movimiento que comprende una secuencia de cuadros, cada cuadro comprende una pluralidad de bloques en dos dimensiones con cada bloque que comprende una pluralidad de valores píxeles en las dos dimensiones, y los datos de imagen en movimiento que comprenden contenido de captura de pantalla y/o contenido de vídeo de cámara, el sistema codificador comprende:un procesador;y un dispositivo de memoria legible por computadora acoplada a dicho procesador, que causa que dicho procesador configure el sistema codificador para codificar los datos de imagen en movimiento para producir datos codificados al llevar cabo funciones que incluyen: decidir si o no la precisión de un vector de movimiento (“MV”) para al menos uno de los cuadros está controlada en una base región por región;si la precisión de MV para al menos uno de los cuadros no está controlada en una base región por región, decidir si la precisión de MV para al menos uno de los cuadros es precisión de muestra de entero o precisión de muestra de cuarto;establecer un valor de un indicador en un encabezado que aplica para al menos uno de los cuadros de la secuencia de video, el indicador que índica sí o no la precisión de MV para al menos uno de los cuadros está controlada en una base región por región y, si la precisión de i jyi p i MV para el al menos uno de los cuadros no está coi . ^“indÜWrlÍ ^8181 11111 ¾......* región por región, además indica si la precisión de MV para al menos uno de los cuados es precisión de muestra de entero o precisión de muestra de cuarto: y si la precisión de MV para al menos uno de los cuadros está controlada en una base región por región, para cada región de la una o más regiones del al menos uno de los cuadros: decidir, con base al menos en parte en si el tipo de contenido para la región es contenido de captura de pantalla o contenido de video de cámara, si la precisión de MV para la reglón es precisión de muestra de entero o precisión de muestra de cuarto;y establecer un valor de una bandera en un encabezado para la región, la bandera que indica si la precisión de MV para la región es precisión de muestra de entero, o precisión de muestra de cuarto;dicho dispositivo de memoria legible por computadora adicionalmente causa dicho procesador configure el sistema codificar para emitir los datos codificados como parte de una corriente de bits, dicha corriente de bits que incluye el indicador y, si la precisión de MV para el al menos uno de los cuadros es controlada en una base región por región, una bandera para cada región de la una o más regiones dei al menos uno de los cuadros que Indica la precisión de MV para la región. un módulo de predicción de movimiento para usarse al codificar los datos de imagen en movimiento al, por lo menos en parte, de cada uno de la pluralidad de cuadros, codificar cada bloque con relación a una porción 3' de referencia respectiva de otro cuadro de la secuenc · » ' i referencia respectiva desplazada desde el bloque por un respectivo ;en donde los datos de imagen en movimiento de la pluralidad de cuadros comprende una corriente de captura de pantalla y el módulo de predicción de movimiento está configurado para restringir cada uno de los vectores de movimiento de la corriente de captura de pantalla en un número entero de pixeles en por lo menos una de las dimensiones.
- 2El sistema codificador de acuerdo con la reivindicación 1, en donde dicho procesador adicionalmente configura el sistema codificador para seleccionar la precisión de MV individualmente en cada una de las dos dimensiones.
- 3El sistema codificador de acuerdo con la reivindicación 1, en donde los bloque son bloques o macrobloques de un estándar de codificación de video H.26x.
- 4Una terminal de usuario que comprende un sistema codificador tal como el descrito en la reivindicación 1, y un transmisor configurado para transmitir los datos codificados sobre una red a una terminal remota.
- 5El sistema codificador de acuerdo con la reivindicación 1, en donde el encabezado que aplica para el al menos uno de los cuadros es un grupo de parámetros de secuencia o un grupo de parámetros de imágenes, en donde las regiones son rebanadas, y en donde el encabezado para la región es un encabezado de rebanada.
- 6El sistema codificador de acuerdo con la reivindicación 1, en donde las funciones adicionalmente comprenden:recibir, desde una aplicación o un sistema operativo, una indicación de si el tipo de contenido es contenido de captura de de video de cámara;medir una heurística de rendimiento que indica si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara;determinar datos estadísticos históricos que Indican si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara;o realizar un análisis de múltiples pases para determinar si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara.
- 7El sistema codificador de acuerdo con la reivindicación 1, en donde la precisión de MV es precisión de muestra de entero si el tipo de contenido es contenido de captura de pantalla, y en donde la precisión de MV es precisión de muestra de cuarto si el tipo de contenido es contenido de video de cámara.
- 8En un sistema de computadora que comprende una o más unidades de procesamiento y memoria, un método que comprende:codificar cuadros de una secuencia de video para producir datos codificados, cada uno de los cuadros incluye una o más reglones, en donde la codificación incluye: decidir si o no la precisión de un vector de movimiento (“MV”) para al menos uno de los cuadros está controlada en una base reglón por reglón;si la precisión de MV para al menos uno de los cuadros no INDUSTRIA está controlada en una base reglón por reglón, declc MV para al menos uno de los cuadros es precisión de muestra de entero o precisión de muestra de cuarto;establecer un valor de un indicador en un encabezado que aplica para al menos uno de los cuadros de la secuencia de video, el Indicador que índica sí o no la precisión de MV para al menos uno de los cuadros está controlada en una base reglón por región y, si la precisión de MV para el al menos uno de los cuadros no está controlada en una base región por reglón, además Indica si la precisión de MV para al menos uno de los cuados es precisión de muestra de entero o precisión de muestra de cuarto;y si la precisión de MV para al menos uno de los cuadros está controlada en una base región por región, para cada región de la una o más regiones del al menos uno de los cuadros: decidir, con base al menos en parte en si el tipo de contenido para la región es contenido de captura de pantalla o contenido de video de cámara, si la precisión de MV para la región es precisión de muestra de entero o precisión de muestra de cuarto;y establecer un valor de una bandera en un encabezado para la reglón, la bandera que Indica si la precisión de MV para la región es precisión de muestra de entero o precisión de muestra de cuarto;y emitir los datos codificados como parte de una corriente de bits, dicha corriente de bits que incluye el indicador y, si la precisión de MV para el al menos uno de los cuadros es controlada en una base reglón por región, una bandera para cada región de la una o menos uno de los cuadros que indica la precisión de MV para la región.
- 9El método de acuerdo con la reivindicación 8, en donde el encabezado que aplica para al menos uno de los cuadros es un grupo de parámetros de secuencia o un grupo de parámetros de imagen, en donde las regiones son rebanadas, y en donde el encabezado para la región es un encabezado de rebanada.
- 10El método de acuerdo con la reivindicación 8, adicionalmente comprende:recibir, desde una aplicación o un sistema operativo, una indicación de si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara.
- 11El método de acuerdo con la reivindicación 8, adlcionalmente comprende, durante la codificación:medir una heurística de rendimiento que indica si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara.
- 12El método de acuerdo con la reivindicación 8, adicionalmente comprende, durante la codificación:determinar datos estadísticos históricos que indican si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara.
- 13El método de acuerdo con la reivindicación 8, adicionalmente comprende, durante la codificación:realizar un análisis de múltiples pases para determinar s¡ el tipo de contenido es contenido de captura de pantalla o co cámara.
- 14El método de acuerdo con la reivindicación 8, en donde la precisión de MV es precisión de muestra de entero si el tipo de contenido es contenido de captura de pantalla, y en donde la precisión de MV es precisión de muestra de cuarto si el tipo de contenido es contenido de video de cámara.
- 15Un dispositivo de memoria legible por computadora que provoca que un procesador realice las funciones de un método que comprende:codificar cuadros de una secuencia de video para producir datos codificados, cada uno de los cuadros incluye una o más regiones, en donde la codificación Incluye: decidir si o no la precisión de un vector de movimiento (“MV”) para al menos uno de los cuadros está controlada en una base reglón por región;si la precisión de MV para al menos uno de los cuadros no está controlada en una base región por región, decidir si la precisión de MV para al menos uno de los cuadros es precisión de muestra de entero o precisión de muestra de cuarto;establecer un valor de un indicador en un encabezado que aplica para al menos uno de los cuadros de la secuencia de video, el indicador que índica sí o no la precisión de MV para al menos uno de los cuadros está controlada en una base reglón por región y, si la precisión de MV para el al menos uno de los cuadros no está controlada en una base región por región, además indica si la precisión de M de los cuados es precisión de muestra de entero o precisión de muestra de cuarto;y si la precisión de MV para al menos uno de los cuadros está controlada en una base región por región, para cada región de la una o más regiones del al menos uno de los cuadros: decidir, con base al menos en parte en si el tipo de contenido para la región es contenido de captura de pantalla o contenido de video de cámara, si la precisión de MV para la región es precisión de muestra de entero o precisión de muestra de cuarto;y establecer un valor de una bandera en un encabezado para la región, la bandera que indica si la precisión de MV para la región es precisión de muestra de entero o precisión de muestra de cuarto;y emitir los datos codificados como parte de una corriente de bits, dicha corriente de bits que incluye el indicador y, si la precisión de MV para el al menos uno de los cuadros es controlada en una base región por región, una bandera para cada región de la una o más regiones del al menos uno de los cuadros que indica la precisión de MV para la región.
- 16El dispositivo de memoria legible por computadora de acuerdo con la reivindicación 15, en donde el encabezado que aplica para al menos uno de los cuadros es un grupo de parámetros de secuencia o un grupo de parámetros de imagen, en donde las regiones son rebanadas, y en donde el encabezado para la región es un encabezado de rebanada. TT Tk ir TF-% τ ................. « I i, A 1J I 1 fVl r 1 ·®ι·* :’so
- 17El dispositivo de memoria legible por com con la reivindicación 15, en donde las funciones adicionalmente comprenden:recibir, desde una aplicación o un sistema operativo, una indicación 5 de si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara.
- 18El dispositivo de memoria legible por computadora de acuerdo con la reivindicación 15, en donde las funciones adicionalmente comprenden, durante la codificación:10 medir una heurística de rendimiento que indica si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara.
- 19El dispositivo de memoria legible por computadora de acuerdo con la reivindicación 15, en donde las funciones adicionalmente 15 comprenden, durante la codificación:determinar datos estadísticos históricos que indican si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara.
- 20El dispositivo de memoria legible por computadora de acuerdo 20 con la reivindicación 15, en donde las funciones adicionalmente comprenden, durante la codificación:realizar un análisis de múltiples pases para determinar si el tipo de contenido es contenido de captura de pantalla o contenido de video de cámara.
- 2125 21. El dispositivo de memoria legible por computadora de acuerdo TT Tk ir TF-% τ a................ « I i, Λ 1J I |ΙΕ^Β<^·μ C» 1 fVl r 1 ·®ι·* :’so con la reivindicación 15, en donde la precisión de muestra de entero si el tipo de contenido es contenido de captura de pantalla, y en donde la precisión de MV es precisión de muestra de cuarto si el tipo de contenido es contenido de video de cámara. T λ Λ D Τ ® 1 J V1 i 1 ΙΤΙΑ’’’ rjo
Independent claims21
225 paragraphs in 18 sections, as filed
(54) Title: VIDEO ENCODING IN DISPLAY CONTENT DATA. (54) Title: VIDEO ENCODING OF SCREEN CONTENT DATA.
(57) Summary
An input for an encoder receives moving image data comprising a sequence of frames to be encoded, each frame comprising a plurality of blocks in two dimensions, each block comprising a plurality of pixels in those two dimensions. A motion prediction module performs encoding at least part of each of the plurality of frames, encoding each block relative to the respective reference portion of another frame in the sequence, the respective reference portion being shifted from the block by a respective motion vector. In accordance with the present invention, the moving picture data of the plurality of frames comprises a screen capture stream and the motion prediction module is configured to constrain each of the motion vectors of the screen capture stream by an integer number of pixels in at least one of the dimensions.
(57) Abstract
An input of an encoder receives moving image data comprising a sequence of frames to be encoded, each frame comprising a plurality of blocks in two dimensions with each block comprising a plurality of pixels in those two dimensions. A motion prediction module performs encoding by, for at least part of each of a plurality of said frames, coding each block relative to a respective reference portion of another frame of the sequence, with the respective reference portion being offset from the block by a respective motion vector. According to the present disclosure, the moving image data of this plurality of frames comprises a screen capture stream, and the motion prediction module is configured to restrict each of the motion vectors of the screen capture stream to an integer number of pixels in at least one of said dimensions.
PATENT TITLE No. 360925
Headlines):
Home:
Denomination:
Classification:
Inventors):
MICROSOFT TECHNOLOGY LICENSING, LLC
One Microsoft Way, Redmond, Washington, 98052-6399, EU .A.
VIDEO ENCODING IN DISPLAY CONTENT DATA.
CIP: H04N19 / 51; H04N19 / 42; H04N19 / 109; H04N19 / 136; H04N19 / 174;
CPC:
H04 N19 / 51; H04N19 / 523 H04N19 / 42;
SERGEY SILKIN; SERGEY S, LEE
10/199; H04N19 / 136; H04N19 / 174; H04N19 / 523
YOU ZHOU; CHIH-LUNG UN; MING- CHIEH 'SABLIN: YOU ZHOU
I * ........ ...... = ................................. ................. ..........
<img file="MX360925B_D0001.tif" />
Number
MX / a / 2016/009023
International:
2014
Validity: Twenty years Expiration Date: Exp Date
The reference patent
Pursuant to the date of pn
Who subscribes to the present Ul ulo lo (Official Gazette! Of the Federation '
25/01/2006, 06/05/2009. 06/01/2010, and 12 · sections I and III of the Regulations (07/28/2004 and 07/09/2007); articles 1 «, 3 °, 4“, 5 “I Industrial Property (DOF 27 / 12Π999 refom, 'powers in the Deputy General Directors, Coi Coordinators Departmental and other subordinates of! Insti 07/29/2004, 04/08/2004 and 09/13/2007).
Lad Industry!
non-extendable, counted to echos.
the Industrial Property Law
999, 01/26/2004, 06/16/2005, s 1 °, 3 'fraction V subsection a), 4 ° B and 07/01/2002, 07/15/2004, of the Mexican Institute of the
3<sup>and</sup> and 5 'clause a) of the Agreement that delegates Regional offices, Divisional Deputy Directors, strial. (DOF 12/15/1999, amended on 02/04/2000,
This document is signed with an advanced electronic signature (FIEL), based on articles 7 BIS 2 of the Industrial Property Liability Law; 3rd of its Regulations, and 1 s ection III, 2 section V, 26 BIS and 26 TER of! Agreement that establishes the lineamierrtos for the use of the Portal of Payments and Electronic Services (PASÉ) of the Mexican Institute of Industrial Property, in the procedures indicated,
THE DIVISIONAL DIRECTOR OF PATENTS
NAHANNY CANAL REYES
<img file="MX360925B_D0002.tif" />
p »| Original string:
NAHANNY MARISOL CANAL REYES | 00001QQOQQ04G3252793 | Tributary Administration Service | 1695 | jMX / 2019/284 | MX / a / 20T6 / 009023 | T (patent tulO PCTj1027 | RGZ | Pág (s) | XCHW / HTCzPj + aw6k4
Seal Digital:
ulTThZHnbtcGtxompjR5UBhCxS6BakPK5t6CGU4IY4t¡dLcMwOxdlXjdatOr02tc + S5FM9sPqcjxkyq8jDs1vH8aa '
29bLkJ3wkgvG6Bl0PpmlN1zqCFPUObe30uYhbnws.hrcr / dUMQT60p / 9VMÜPio9GIKkxM ++ eTj4YToOewEaeL¡0hDJb3 hWW7zfEscgHh1uyOJ + dgEzNsYSRk1SpHqAbSOo4wNc | fUFTjffMGv736anK2VhlOt / dEmHzwJKEvh8iWYmoReuRmnM
US / a9ymFyuGqIWRTD2ghreuPJLMARm2x5iNhuEI + Izxxty3eFDCKZS78tPFbnkQIL / le2AaQ ==
Arenal No. SSO, Floor 1, Pueblo Santa wairfa Tepeipari, Xochimilco, Mexico City, (56) 53340700 www.gob.rrtx / imdi
<img file="MX360925B_D0003.tif" />
MX / 2019/284
IMPIC
VIDEO CODING IN C DATA
SCREEN
Field of the Invention
In modern communication systems a video signal can be sent from one terminal to another over a medium, such as a wired and / or wireless network, often a packet-based network, such as the Internet. For example, the video may be part of a VolP (Internet Protocol over voice) call conducted from a VolP client application executed in the user's terminal, such as a desktop or laptop computer, a Tablet or a Smartphone.
Typically, the video frames are encoded by an encoder at the transmission terminal in order to compress them for transmission over the network. Coding for a given frame may comprise inter-frame coding, whereby blocks are coded relative to other blocks in the same frame. In this case, a target block is coded in terms of the difference (the residual) between that block and the neighboring block. Alternatively, the coding of some frames may comprise interframe coding, whereby the blocks in the target frame are coded relative to the corresponding portions in a previous frame, typically based on motion prediction. In this case, a target block is encoded in terms of a motion vector, which identifies the displacement between the block and the corresponding portion from which to forecast and the differential (residual) between the block and the portion
<img file="MX360925B_D0004.tif" />
®>
...... J | € g corresponding from which the corresponding forecast is going to be predicted in the receiver decodes the frames of the received video signal based on the appropriate type of prediction, in order to decompress them to output them in a protocol on the decoder side.
When encoding (compressing) a video, motion vectors are used to generate the interframe prediction for the current frame, the encoder first looks for a similar block (the reference block) in a previous coded frame that best matches the block current (target block) and signals the offset between the reference block and the target block for the decoder as part of the encoded view stream. Displacement is typically represented as the horizontal and vertical "x" and "y" coordinates and is referred to as the motion vector.
The reference “block” is in fact not limited to being in an actual block position in the reference frame, ie it is not limited to the same grid as the target blocks, rather it is a corresponding size portion of the offset of the reference frame relative to the position of the target block by the motion vector. In accordance with the present standards, motion vectors are represented in fractional pixel resolution. For example, in the H.264 standard, each motion vector is represented with a resolution of X pixel. As an example, when a 16x16 block in the current frame is to be predicted from another 16x16 block in the previous frame, the left pixel 1 is the position of the target block,
ASI ® ”ft € 3
- í then the motion vector is (4.0). Otherwise it will be predicted from a reference block that is only, that is,% of pixel to the Left of the target block, the motion vector is (3.0). The reference block at the fractional pixel position does not actually exist per se, rather, it is generated by interpolation between pixels in the reference frame. Sub-pixel motion vectors can perform well in terms of compression efficiency.
Brief description of the invention
However, with the use of a fractional pixel resolution more bits are incurred to encode the motion vector than when the motion is calculated at the whole pixel resolution and more processing resources are also incurred to find the best reference of coincidence. For video encoding this may be convenient, for example, since the reduced size of the best matched residual can generally exceed the bits incurred encoding the motion vector or the quality achieved can be considered to justify the resources. However, not all motion pictures that can be encoded are videos (that is, captured with a camera). It should be recognized that when encoding (compressing) an image-in-image that was captured from a screen rather than from a camera, most motion vectors in the encoded bitstream generally point to whole pixels, while very few of them they tend to be found at fractional pixel positions. Of this
<img file="MX360925B_D0005.tif" />
mode, while encoders normally repr movement in bit streams of% pixel units, for screen sharing or for recording applications, bandwidth can in fact be saved without loss of quality by encoding the vector vectors. movement in units of only 1 pixel. Alternatively, even when motion vectors are represented in the bit stream on a fractional pixel scale, processing resources can be saved by restricting the motion vector search to pixel offsets.
Therefore, in accordance with an aspect described herein, an encoder is provided comprising an input for receiving mg moving data and a motion prediction module for use in encoding the mg moving data. The moving mg data comprises a frame sequence to be encoded and each frame is divided into a plurality of blocks in two dimensions, each block comprises a plurality of pixels in the two dimensions. For example, blocks can be divisions referred to as blocks or macroblocks in the H.26x standard as H.264 or H.265. The motion prediction module performs interframe coding by encoding each block (the target block) relative to the respective reference portion of another frame in the sequence (the reference "block") with the respective reference which is the displacement from the target block by a respective motion vector. Furthermore, in accordance with the present invention, moving image data
<img file="MX360925B_D0006.tif" />
of the plurality of frames comprise a screen display and the motion prediction module is configured to constrain each of the motion capture stream motion vectors by an integer number of pixels in at least one of the dimensions.
Therefore in some modes, the encoder may also comprise a controller that operates to switch the motion prediction module between two modes: a first mode and a second mode. In the first mode, the motion vector is not restricted to an integer number of pixels (in any dimension), but in the second mode, the motion vector is constrained to an integer number of pixels in at least one dimension (and in certain modalities, both). The controller may be configured to switch the motion prediction module to the second mode depending on the determination that the moving image data to be encoded comprises a screen capture stream.
For example, the moving image data may comprise the screen capture stream and a video stream (for example, these may be live streams of a call conducted over a packet-based network, such as the Internet, or may be stored currents for future reproduction. It may be that some frames of the motion picture frames are frames of the screen capture stream and other times, the frames of the motion picture data are video frames or they may be in different regions within each frame, comprising the screenshot and
<img file="MX360925B_D0007.tif" />
video streams respectively (eg di To adapt such cases, the controller may be configured to determine whether the current motion picture data to be encoded is the screen capture stream or video stream 5 and to adjust the Motion prediction module in the second mode for screen capture and the first mode for video. Alternatively, as another example, even when a screen capture stream and the video stream are included in different regions of some of the same frames, the controller may be configured to select the second mode when a frame does not contain any screen capture data and otherwise to select the first mode only when the box does not contain screen capture data or otherwise, it can be configured to switch to the second mode when the frame contains only screen capture and non-video data, and otherwise to select the first mode when the frame contains no video.
This summary is provided to present a selection of concepts in a simplified form that are also described in the detailed description. This summary is not intended to identify key characteristics or essential features of the claimed material, nor is it intended to be used to limit the scope of the claimed material.
Nor is the subject limited to implementations that address some or all of the disadvantages noted here.
Brief description of the drawings
To help understand the present invention and to show how the modalities can be put into effect, reference is made by way of example to the accompanying drawings, in which:
Figure 1 is a schematic representation of a video stream.
Figure 2 is a schematic block diagram of a communication system.
Figure 3 is a schematic representation of an encoded video stream.
Figure 4 is a schematic block diagram of an encoder.
Figure 5 is a schematic block diagram of a decoder.
Figure 6 is a schematic representation of an interframe coding scheme: and Figure 7 is a schematic representation of another interframe coding scheme.
Detailed description of the invention
Figure 1 provides a schematic illustration of an input video signal captured with a camera and divided into spatial divisions to be encoded by a video encoder to thereby generate an encoded bit stream. The signal comprises a moving video image divided in time into a plurality of frames (F), each frame representing the image in a time dlfer (... t-1, t, t + 1 ...). Within each frame, the frame is divided in space into a plurality of divisions, each representing a plurality of pixels. These divisions can be called as blocks. In certain schemes, the table is divided and subdivided into different levels of the block. For example, each frame can be divided into macroblocks (MB) and each macroblock can be divided into blocks (b), for example, each block represents an 8x8 pixel region within a frame and each macroblock represents a 2x2 block region ( 16x16 pixels). In certain schemes, each frame may also be divided into independently decodable slices (S), each comprising a plurality of macroblocks. Slices S, in general, can take any shape, for example each slice is one or more rows of macroblocks or in an irregularly or arbitrarily defined selection of macroblocks (eg corresponding to a region of interest, ROI, in the picture).
With respect to the term "pixel" below, the term is used to refer to samples and sample positions in the sample grid for image configuration (sometimes in the literature, the term "pixel" is used in another way to refer to the three color components corresponding to a single spatial position and sometimes used to refer to a single position or a single sample value I enter the single configuration). The resolution of the sample grid is often different between brightness and color sampling arrangements. In modalities, the following can be applied in a 4: 4: 4 representation, but potentially apply in 4: 2: 2 or 4: 2: 0, for example.
It should also be noted that any given standard may offer specific meanings in terms of block or macroblock, the term block is often used within the art to refer to dividing the frame at a level where operations are carried out encoding and decoding, such as interframe or interframe prediction and this meaning is what will be used, unless otherwise specified. For example, the blocks referred to here, in fact, may be the divisions called blocks or macroblocks in the H.26x standards and the different stages of encoding and decoding may operate at the level of such divisions as appropriate in the encoding mode, the application and / or the standard in question.
A block in the input signal as captured is usually represented in the spatial domain, where each color space channel is represented as a function of the spatial position within the block. For example, in the YUV color space, each of the luminance (Y) and chrominance (U, V) channels may be represented as a function of Cartesian "x" and "y" coordinates, Y (x, y ), U (x, y) and V (x, y) or in the RGB color space, each of the red (R), green (G) and blue (B) channels can be represented as a coordinate function R (x, y), G (x, y), B (x, y) Cartesian. In this representation, each block or portion is represented by a group of pixel values in different spatial coordinates, for example, "x" and "y" coordinates, so that each channel of the CI represented in terms of a respective magnitude of that channel in each of the discontinuous group of pixel locations.
However, prior to quantification, the block can be transformed into a representation of the transformation domain as part of the encoding process, typically a representation of the spatial frequency domain (sometimes referred to as the frequency domain). In the frequency domain, each color space channel in the block is represented as a function of the spatial frequency (domains of 1 / length) in each of the two dimensions. For example, this can be denoted by the number k<sub>x</sub> and k<sub>and</sub> of waves in the horizontal and vertical directions, respectively, so that the channels can be expressed as Y (k<sub>x</sub>, k<sub>and</sub>), U (k<sub>x</sub>, k<sub>and</sub>) and V (k<sub>x</sub>, k<sub>and</sub>) in the YUV or R (k space<sub>x</sub>, k<sub>and</sub>), G (k<sub>x</sub>, k<sub>and</sub>), B (k<sub>x</sub>, k<sub>and</sub>) in RGB space. Thus, instead of representing a color space channel in terms of magnitude in each discontinuous group of pixel positions, the transformation represents each color space channel in terms of a coefficient associated with each discontinuous group of color components. spatial frequency, which form the block, that is, the amplitude of each of the discontinuous group of spatial frequency terms corresponding to different frequencies of spatial variation across the block. The possibilities of such transformations include the Fourier transformation, the discontinuous cosine transformation (DCT), the Karhunen-Loeve transformation (KLT) or others.
The block diagram in Figure 2 provides an example of a
<img file="MX360925B_D0008.tif" />
IMPI communication system where the invention can be used. The communication system comprises a first transmission terminal 12 and a second reception terminal 22. For example, each terminal 12, 22 may comprise a mobile phone or a smartphone, a tablet, a laptop, a desktop computer, or other household appliance, such as a television set, an encoder box, a stereo system, etc. . Each of the first and second terminals 12, 22 is operatively coupled to the communication network 32 and the first transmitting terminal 12 is thus configured to transmit signals to be received by the second receiving terminal 22. Of course, transmission terminal 12 also has the ability to receive signals from receiving terminal 22 and vice versa, but for the purpose of description, transmission is described here from the perspective of first terminal 12 and reception is described here. from the perspective of the second terminal
22. Communication network 32 may comprise, for example, a packet-based network, such as the wide area Internet and / or the local area network and / or the mobile cellular network.
The first terminal 12 comprises a computer readable medium 14 such as a flash memory or other electronic memory, a magnetic storage device and / or an optical storage device. The first terminal 12 also comprises a processing apparatus 16 in the form of a processor or CPU having one or more execution units, a transceiver, such as a wired or wireless modem having a transmitter 18, a camera 15
IMPI
<img file="MX360925B_D0009.tif" />
video and a screen 17 (that is, a screen or single camera 15 and screen 17 may or may not be housed within the same enclosure as the rest of terminal 12 (and even transmitter 18 may be internal or external, for example, comprising a platform or a wireless router in the latter case). Each of the storage medium 14, the video camera 15, the display 17 and the transmitter 18 are operatively coupled with the processing apparatus 16 and the transmitter 18 is operatively coupled with the network 32 through a link wired or wireless. Similarly, the second terminal 22 comprises a computer readable storage medium 24, such as an electronic, magnetic and / or optical storage device and a processing apparatus 26 in the form of a CPU having one or more execution units. . The second terminal comprises a transceiver such as a wired or wireless modem having at least one receiver 28 and a display 25, which may or may not be housed within the same enclosure as the rest of terminal 22. Each of the storage medium 24, the display 25 and the receiver 28 of the second terminal is operatively coupled with the respective processing apparatus 26 and the receiver 28 is operatively coupled with the network 32 through a wired link or wireless.
The storage 14 in the first terminal 12 stores at least one encoder for encoding moving image data, the encoder is arranged to be executed in the respective processing apparatus 16. When run, the encoder receives a "raw" (uncoded) input video stream from the video camera 15, and operates to encode the video stream to a lower bit rate stream and outputs the encoded video for transmission through transmitter 18 and communication network 32 to receiver 28 of second terminal 22. Storage 24 in second terminal 22 stores at least one video decoder arranged to be executed on its own processing apparatus 26. When the decoder runs, it receives the encoded video stream from receiver 28 and decodes it for output to screen 25.
The encoder and decoder also operate to encode and decode other types of moving image data, including on-screen sharing streams. A screen sharing stream is Image data captured from screen 17 on the encoder side so that one or more other remote users can see what the user is seeing on the screen on the encoder side or for the user From that screen you can record what happens on the screen for playback by one or more users later. In the event of a call made between transmitting terminal 12 and receiving terminal 22, the motion content of screen 17 at transmitting terminal 12 will be encrypted and broadcast live (in real time) to be decoded and displayed on the screen 25 of the reception terminal 22. For example, the encoder-side user may want to share with another user what works on their desktop operating system or application.
It should be noted that when a current is captured for
<img file="MX360925B_D0010.tif" />
particular mechanism to do it. For example, the data can be read from the screen memory of the screen 17 or captured when receiving an instance of the same graphic data that is emitted from the operating system or from an application for display on the screen.
17.
Figure 3 provides a schematic representation of a bit stream 33 coded as it would be transmitted from the encoder operating at transmit terminal 12 to the decoder executing at receive terminal 22. Bitstream 33 comprises encoded Image data 34 for each frame or slice comprising the samples encoded for the blocks of that frame or slice along with any associated motion vector. In one application, the bit stream may be transmitted as part of a live (real-time) call, such as a VolP call between transmit and receive terminals 12, 22 (VolP calls may also include video and video sharing). screen). Bitstream 33 also comprises header information 35 associated with each frame or slice. In certain embodiments, the header 36 is arranged to include at least one additional element in the form of at least one flag 37 indicating the resolution of the motion vector, which will be described in more detail later.
Figure 4 is a block diagram illustrating an encoder such as can be implemented at transmission terminal 12. The encoder comprises a main encoding module 40, which comprises: a sewing transformation module 51, a quantizer 53, an inverted transformation module 61, an inverted quantizer 63, an intra-prediction module 41, an inter-prediction module 43, a switch 47, a step 49 of subtraction (-) and a step 65 of lossless decoding. The encoder also comprises a control module 50 coupled with the inter-prediction module 43. Each of these modules or stages can be implemented as a portion of the code stored in the storage medium 14 of the transmission terminal and is arranged to execute in its processing apparatus 16, although there is the possibility that some or all of them may be implemented in whole or in part on dedicated hardware circuitry.
The subtraction step 49 is arranged to receive an instance of an input signal comprising a plurality of blocks on a plurality of frames (F). The input current is received from camera 15 or captured from what will be displayed on screen 17. Intraprediction 41 or interprediction> 43 generates a predicted version of a current (target) block to be encoded based on another prediction, an already encoded block, or a correspondingly sized reference portion. The predicted version is supplied to an input of the subtraction stage 49, where it is subtracted from the input signal (i.e. the actual signal) in the spatial domain to produce a residual signal representing the difference between the predicted version of the block and the corresponding block in the actual input signal.
In intrapredicclone mode, module 4.
generates a predicted version of the actual (target) block to be encoded based on one prediction of another, a block already encoded in the same frame, typically the neighboring block. When performing intraframe encoding, the idea is only to encode and transmit the measurement in the way that a portion of the image data within the frame differs from another portion within the same frame. That chunk can be predicted on the decoder (due to certain absolute data to begin with) and so you only need to pass on the difference between the prediction and the actual data, rather than the actual data itself. The difference signal is typically less in magnitude, so it takes fewer bits to encode (due to the operation of lossless compression stage 65 - see below).
In interprediction mode, interprediction module 43 generates a predicted version of the actual (target) block to be encoded based on one prediction of another, a reference portion already encoded in a different frame than the actual block, the portion of reference has the same block size, but it is displaced relative to the target block in the spatial domain by a motion vector that is forecast by modulo 43 interprediction (interprediction is also referred to as motion prediction or motion estimate). Interprediction module 43 selects the optimal reference for a given target block, by searching, in the spatial domain, through a plurality of candidate reference portions displaced by a plurality of motion vectors
<img file="MX360925B_D0011.tif" />
Possible, respective in one or more squares, different selects the candidate that minimizes the residual with respect to the target block according to an appropriate metric. The inter-prediction module 43 is changed within the feedback path by the switch 47, instead of the intraframe prediction step 41 and thus the feedback loop is created between the blocks of one frame and another in order to encode the interframe in relation to those of another table. That is, the residual now represents the difference between the predicted block and the actual input block. This typically takes fewer bits to encode than intraframe encoding.
Samples of the residual signal (comprising the residual blocks after the input signal predictions are subtracted) are output from the subtraction step 49 through the transformation module 51 (DCT) (or other appropriate transformation) in where their residual values are converted into the frequency domain, then to the quantizer 53, where the transformed values become essentially essentially discontinuous quantization indices. The transformed, quantized indices of the residual are generated by the transformation and quantization modules 51, 53, as well as an indication of the prediction used in the prediction modules 41, 43 and any motion vector generated by the 43 n prediction module, all of them are emitted to be included in the encoded video stream 33 (see element 34 in figure 3), through another stage 65 of lossless encoding, such as a Golomb encoder or an entropy encoder, where the quantized motion vectors and transformed indices are also compressed with lossless encoding techniques known in the art.
An instance of the quantized, transformed signal is also fed back through inverted quantizer 63 and inverted transformation module 61 to generate a predicted version of the block (as seen in the decoder) for use by module 41 or 43 of prediction selected to forecast the block after being encoded, in the same way that the actual target block was predicted to be encoded based on an inverted quantized and inverted transform version of a previously encoded block. Switch 47 is arranged to pass the output of the reverse quantizer 63 to the input of either the intra-prediction module 41 or the inter-prediction module 43, as appropriate for the encoding used for the frame or block currently encoded.
Fig. 5 is a block diagram illustrating a decoder, such as can be implemented at the receiving terminal 22. The decoder comprises the inverted lossless encoding 95, an inverted quantization stage 83, an inverted DCT transformation stage 81, a switch 70 and an intraprediction stage 71 and a motion compensation stage 73. Each of these modules or steps can be implemented as a portion of a code stored in the storage medium 24 of the receiving terminal and is arranged to execute in its processing apparatus 26, although the possibility that some or all of them are fully or partially implemented in a dedicated hardware circuitry.
IΜ ΡI tNSTlTUTO MEXiCANÜ iffjjf ......
The inverted quantizer 81 is arranged pa encoded from the encoder, through the receiver 28 and the inverted lossless encoding step 95. The inverted quantizer 81 converts the quantization indices into a coded signal into dequantized samples of the residual signal (comprising the residual blocks) and passes the dequantized samples to the inverted DCT module 81, where they are transformed back from the frequency domain to the spatial domain. Switch 70 then passes residual spatial domain samples, dequantized to modulo 71 or 73 intraprediction or interprediction, as appropriate for the prediction mode used for the actual frame or block to be decoded and modulo 71, 73 intrapredicted or interpredicted. uses intraprediction or interprediction, respectively, to decode the blocks. The mode to be used is determined using an indication of the prediction and / or any motion vector received with the samples 34 encoded in the encoded bitstream 33. After this stage, the decoded blocks are broadcast to be played through screen 25 at the receiving terminal 22.
As mentioned, encoders-decoders in accordance with conventional standards carry out motion prediction at quarter pixel resolution, meaning that motion vectors are also expressed in terms of quarter pixel stages. An example of quarter-resolution resolution motion estimation is shown in Figure 6. In this example, the pixel “p in the upper left corner of the target block is
ΜΡΙ
<img file="MX360925B_D0012.tif" />
ΊΤϊϋΤΟ MEXICANO .ϋξχ, Λ lacinao forecast by interpolation between pixels a, pixels of the target block will also be forecast based on a similar interpolation between the respective groups of pixels in the reference frame, according to the offset between the target block in one box and the reference portion in another box (these blocks are shown with black dotted lines in Figure 6). However, performing motion estimation with this granularity has certain consequences, as described below.
With reference to encoder 65 and decoder 95 lossless, lossless encoding is a form of compression that works not by discarding information (such as quantization), but by using different lengths of a codeword to represent different values. depending on the probability of these values occurring or the frequency with which they occur, in the data to be encoded by the lossless encoding step 65. For example, the number of leading 0s in the encoder word before finding a 1 may indicate the length of the codeword, so 1 is the shortest codeword, then 010 and 011 are the shortest next, after 00100. .., and else. Thus, shorter codewords are much shorter than what is required when using a uniform codeword length, but longer codewords are longer than that. By assigning the most frequent or most likely values to the shortest codewords and only the least frequently or least likely values to the longest codewords, the resulting bit stream 33 may, on average, incur fewer bits per value encoded that when using a uniform codeword length and hence a, <i without discarding any information.
Much of the encoder 40 before lossless encoding step 65 is designed to attempt to make as many values as possible before going through lossless encoding step 65. Because they occur more frequently, smaller values will then incur a lower bit rate in the encoded bit stream 33 than higher values. This is why the residual is encoded opposite the absolute samples. This, too, is the rationale behind transformation 51, since many samples tend to transform into zero or small coefficients in the transformation domain.
A similar consideration can be applied in encoding motion vectors.
For example, in H.264 / MPEG-4, Part 10 and H.265 / HEVC, the
<td>vectors</td><td>of movement</td><td>are coded <</td><td>with Golomb encoding</td>
<td colspan="2">exponential. The next</td><td>table shows</td><td>the vector values of</td>
<td colspan="3">motion and encoded bits.</td><td></td>
<td>Value</td><td>Code word</td><td colspan="2">Number of bits incurred</td>
<td> 0</td><td> 1</td><td> 1</td><td></td>
<td> 1</td><td> 010</td><td> 3</td><td></td>
<td> 2</td><td> 011</td><td> 3</td><td></td>
<td> 3</td><td> 00100</td><td></td><td> 5</td>
<td> 4</td><td> 00111</td><td></td><td> 5</td>
<td> 5</td><td> 0001000</td><td> 7</td><td></td>
jVl i 1 wi o
From the previous table it can be observed if the values are used, more bits are used. This means that the higher the resolution of the motion vector, the more bits are incurred. For example, with a quarter pixel resolution, an offset of 1 pixel must be represented by a value of 4, which incurs 5 bits in the encoded bitstream.
When encoding video (captured with a camera) the cost of this resolution in the motion vector may be worth it, as the finer resolution may provide more opportunities in finding a lower cost residual reference. However, it is observed that for moving images captured from a screen, most spatial offsets tend to have full pixel offsets and fewer of them tend to be at fractional pixel locations, so most vectors Motion patterns tend to point to integer pixel values and very few tend to point to fractional pixel values.
On such a basis, it may be desirable to encode the motion vectors for the image data captured from a screen with a resolution of 1 pixel. By considering the fact that no bits need be spent on the fractional parts of motion vectors for such content, this means that the bit rate incurred in encoding such content can be reduced.
For example, although encoders normally interpret motion vectors into bit streams in units of% pixel offsets, in fact, the encoder often
<img file="MX360925B_D0013.tif" />
the
IMPW
INSTITUTO MEXICANO may have the ability to save the resolution bit rate and instead encode motion vectors for on-screen encoding applications in units of full pixel shifts. Although the precision of motion vectors will be reduced by a factor of four, such precision is generally not worth it for screen sharing or recording applications and this also reduces the number of bits required to encode the vectors. To forecast a current (target) block from the reference block, 1 pixel left of the target block, the motion vector will be (1.0) instead of (4.0). Using the Golomb encoding above, this means that the bits incurred to encode the motion vector change from (00111.1) to (010.1) so two bits are saved in this case.
Furthermore, in certain modalities, the reduced resolution motion vector can also reduce the complexity of the motion estimation carried out in the encoder by restricting the motion vector search to integer values, which reduces the processing resources incurred by the search. Alternatively, it would be possible to perform a normal search and round the resulting motion vectors to integer values.
Figure 7 shows an example of motion prediction restricted to full pixel resolution only, with the motion vector restricted to full pixel stage only. Contrary to Figure 6, pixel "p" is predicted from a single full pixel "a" without interpolation. Alternatively, they may have predicted from pixel "b", "c", "d" or another pixel, depending on the shift between the target block by a reference in another frame (shown again with black dotted lines), but due to the restriction you cannot predict the interpolation between pixels. It should be noted that: for any given block, the prediction of a quarter pixel as illustrated in Figure 6 as an example, can happen to generate a full pixel shift without interpolation, in case it gave the lowest residual . However, you haven't restricted yourself from doing this and it would be unlikely for an image with size to happen for all blocks.
When considering the values of the fractional motion vector it can also be very useful for the content captured by the camera, in certain modalities, the encoder 40 is provided with a controller 50 coupled with the motion prediction module 43 with the controller 50 configured to select resolution of the motion vector in a flexible way: When the source data is from a captured screen 17 and there is no fractional pixel motion, the motion vector is encoded and transmitted in units of only full pixels, but for camera content video, the motion vectors are also encoded and transmit with fractional pixel precision.
In order to accomplish this, controller 50 may be configured to measure performance heuristics, indicative of the fact that the type of content to be captured is screen content. In response, this then disables motion compensation from
I λ Λ ΟΙ ο
IJ V1 Γ ^ 1 11ΐΡ | β * ^<sup>; , η</sup>ν * ο
A JL JL JL A fig .............<sup>Α</sup> fractional pixel for alternative content encoding, controller 50 may receive an indication from an application or the operating system about the type of data it is supplying to the encoder for encoding and controller 50 can select from the mode on that basis. As another option, you can make the selection based on historical data. Selection can be done on a per frame basis or the mode can be individually selected for different regions within a frame, eg on a per slice basis.
In this way, before encoding a frame or slice, the encoder has the ability to decide the resolution of the motion vector based on factors, such as historical statistical data, the recognition of this type of application, analysis of multiple passes or some other technique. When the encoder decides to use full pixel motion estimation only, the fractional pixel search is skipped. When a scaled motion vector prediction has a fractional part, the prediction can be rounded to an integer value.
In other modes, the control can be optionally applied separately with the vertical or horizontal component of the vector. This can be useful for encoding screen video that is scaled horizontally or vertically.
In order to represent the motion vector on a reduced resolution scale in whole pixel units or stages and therefore, associated bit rate savings are achieved on encoders26
<img file="MX360925B_D0014.tif" />
IMPI
INSTITUTO MEXICANO conventional decoders to signal the neighbors will have to be updated with the H.265 standard (HEVC, High Efficiency Video Coding). To encode the captured screen content, the encoded data 34 format will have a reduced size motion vector field for each motion vector. For a screen capture stream encoded in an integer pixel mode, the relevant data 34 thus will comprise integer motion vectors in bit stream 33 and in modes only integer motion vectors in bit stream 33.
In modalities, this will be optional, with a flag 37 also included in header 36 to indicate whether the fractional pixel (for example,% pixel) or an integer pixel resolution is used in encoding the associated frame or slice (refer to another turn to figure 3). When horizontal and vertical resolutions can be selected separately, two flags 37 per frame or slice will be required.
Alternatively, in certain modalities it is not necessary to update the protocol of existing standards to implement integer pixel motion vectors. Instead, motion vectors may be restricted to integer offsets, but those integer motion vectors may be represented in bit stream 22 on the conventional fractional scale (eg,% pixel). Thus, in the case of% pixel resolution, a full pixel offset will still be represented in a conventional way by a value of 4 (eg codeword 00111), but due to the restriction applied in the encoder may not have the ability to be, i.e. represented by% of p>
(code word 00100). In this case, the savings in the bit rate of the motion vectors will not be realized, but the processing resources will also be saved by restricting the complexity of the motion vector search at integer offsets.
An exemplary modality based on an update to the H.265 standard is described below. The modification allows the motion vectors to be represented on an integer pixel scale, reduced in the encoded bitstream 33 and adds two flags 37 per slice in the header information 36 of the compressed stream in order to signal the resolution of motion vectors in their horizontal and vertical components.
Modification does not change the syntax or parsing process other than at the header level, but modifies the decoding process by interpreting motion vector differences as integers and rounding scaled MV predictors to integer values. The modification has been found to increase encoding efficiency by as much as 7% and on average about 2% for the tested screen content streams and can also reduce the complexity of the encoding and decoding processes.
A high level indicator (at the SPS, PPS and / or at the slice header level) is added to indicate the resolution for the interpretation of the motion vectors.
In the decoding process, when motion vectors are indicated to be at full pixel resolution
<img file="MX360925B_D0015.tif" />
MEXICAN STATUTE. DE and the fractional scale motion vector prediction, then in certain modalities, the prediction is rounded to an integer value. Movement vector differences are simply interpreted as integer offsets better than sample offsets of%. All other decoding processes remain the same. The analysis process (below the heading level) also remains unchanged. When the motion vectors are encoded with full sample precision and the input image data uses 4: 2: 2 or 4: 2: 0 sampling, the color motion vectors can be derived in the usual way , which will produce sample color movement offsets. Alternatively, color motion vectors can also be rounded to integer values.
Regarding the aforementioned escalation, this is something that can happen for example in HEVC (H.265). The idea is that when a motion vector is used to encode some other frame, the motion vector can be computed which will be equivalent in terms of the relative positioning displacement between: (i) the actual image and (ii) its reference image. This is based on the relative positioning of the displacement indicated by the motion vector in a co-located part of another image and based on the relative positioning displacement between (iii) that image and (iv) the image that was referenced as the image. reference. It should be noted that the time frame index of the encoded data may not always be constant and there may also be a difference between the order
............ p in which the Images are encoded in the stream <sup>1</sup> . . Which are captured or displayed, so that these temporal relationships can be computed and then used to scale the motion vector so that it represents basically the same speed of motion in the same direction. This is known as a prediction of the temporal motion vector.
Another possibility would be to disable temporal motion vector prediction each time integer motion is used only. There is already syntax in HEVC that allows the encoder to disable the use of such a feature. This will be a possible way to avoid requiring the decoder to have a special process that operates differently depending on whether the differences are encoded as integers or as fractional values. The gain obtained from the prediction of the temporal motion vector may be less (or zero) in these use cases, so disabling is not convenient.
Regarding the syntax change: a new two-bit indicator will be included in the PPS extension to indicate the modes of control of the resolution of the motion vector. This Indicator is referred to as movlmiento_vector_resoluc¡ón_ctl_¡dc. When the mode is 0, the motion vectors are encoded with% pixel precision, and all decoding processes remain unchanged. When the mode is 1, all motion vectors in the slices that refer to the PPS are encoded with full pixel precision. And when the mode is 2, the resolution of the motion vector is controlled on a slice-by-slice basis by a flag in the slice header.
<img file="MX360925B_D0016.tif" />
2, a call on the
IMPI
When motion_vector_resolution_ctl_idc is not inferred as 0.
When motion_vector_resolution_ctl_idc is the same, additional flag slice_movement_vector_resolution_flag_points the slice header. When the flag is zero, the motion vectors in this slice are encoded with a pixel pixel precision, and when the flag is 1, the motion vectors are encoded with full pixel precision. When the flag is not present, its value is inferred as equal to the value of motion_vector_resolution_ctl_¡dc.
The modified PPS syntax is illustrated as follows;
<td>image_parameter_adjustment_rbsp () {} []</td><td>Descriptor</td>
<td>pps_¡magen_parameter_adjustment_id</td><td>uc (v)</td>
<td>pps_sec_parameter_adjustment_id</td><td>ue (v)</td>
<td>depending_rebated_segments_permitted_flag</td><td>u (1)</td>
<td>flag_output_present_flag</td><td>u (1)</td>
<td>num_extra_rebanada_encabezado_bits</td><td>u (3)</td>
<td>s_gno_dato_oculto_perm ¡ti do_bandera</td><td>u (1)</td>
<td>cabac_inic_presente_bandera</td><td>u (1)</td>
<td> ...</td><td></td>
<td>Lists_modification_present flag</td><td>u (1)</td>
<td>register2_parallel_fuse_level_less2</td><td>ue (v)</td>
<td>slice_segment_heading_extension_present_band</td><td>u (1)</td>
<td>pps_extens¡ón1 _bandera</td><td>u (t)</td>
<td>when (pps_extens¡ón1_bandera) {</td><td></td>
<td>when (flag_allow_transformation_flag)</td><td></td>
<td>record2_max_transformation_stop_block_size_less 2</td><td>uc (v)</td>
<td>brightness_color_prediction_permitted_flag</td><td>u (1)</td>
<td>motion_vector_resolution_control_ide</td><td>u (2)</td>
<td>color_qp_adjusted_permitted_flag</td><td>u (1)</td>
<td>when (color_qp_adjusted_permitted_flag) {</td><td></td>
<td>d¡f_cu_color_qp_ajuste_p rotundity</td><td>ue (v)</td>
<td>color_qp_adjustment_table_size_less1</td><td>ue (v)</td>
<td>for (i = 0; i <= color_qp_adjustment_table_size_less1; i ++) {</td><td></td>
<td>cb_qp_ajuste [i]</td><td>be (v)</td>
<td>cr_qp_ajuete [i]</td><td>be (v)</td>
<td> }</td><td></td>
<td> }</td><td></td>
<td>pps_extensión2_bandera</td><td>u (1)</td>
<td> 1</td><td></td>
<td>when (pps_extensión2_bandera)</td><td></td>
<td>while (more_rbsp_data ())</td><td></td>
<td>flag_data_pps_extension</td><td>u (1)</td>
<td>rbsp_traseros_bits ()</td><td></td>
<td> }</td><td></td>
The modified slice header syntax is illustrated as follows:
<td>slice_heading_segment () {</td><td>Descriptor</td>
<td>first_segment_segment_in_image_flag</td><td>u (1)</td>
<td>when (nal_unldad_tipo> = BLA_W_LP && nal_unldad_tlpo <= RSVJRAP-VCL23)</td><td></td>
<td>no_salida_de_previas_imagenes_bandera</td><td>u (1)</td>
<td>slice_image_parameter_adjustment_id</td><td>ue (v)</td>
<td> ...</td><td></td>
<td>when (slice_type = = P || slice_type = = B) {</td><td></td>
<td>when (motion_vector_resolution_control_ldc == 2)</td><td></td>
<td>slice_movement_vector_resolution_flag</td><td>u (1)</td>
<td>num_ref_lndlce_activo_anular_bandera</td><td>u (1)</td>
<td>when (num_ref_lndice_actlvo_anular_bandera) {</td><td></td>
<td></td><td></td>
It will be appreciated that the above modalities have been described only by way of example.
For example, although the above has been described in terms of blocks, this does not necessarily limit divisions called blocks in any particular standard. For example, the blocks referred to here can be the divisions called blocks or macroblocks in the H.26x standards.
The scope of this invention limited to a particular encoder-decoder or standard and in general, the techniques described here can be implemented in any context of an existing standard or in an update of an existing standard, be it an H.26x type H264 standard. or
MPI
<img file="MX360925B_D0017.tif" />
H.265 or any other standard or can be implemented. 1 »« custom designed decoder. Furthermore, the scope of the invention is not specifically restricted to any particular representation of video samples, be it in terms of RGB, YUV or otherwise. Nor is the scope limited to any particular quantification or DCT transformation. For example, an alternative transformation can be used such as the Karhunen-Loeve (KLT) transformation or there may be no transformation. Furthermore, the invention is not limited to VolP communications or communications over any particular kind of network, it can be used in any network or medium with the ability to communicate data.
When motion vector scrolling is mentioned to be constrained or unconstrained to an integer number of pixels or the like, this may refer to motion estimation in any one or two color space channels or motion estimation in the three color channels.
Furthermore, the invention is not limited to selecting between integer pixel resolution and quarter pixel resolution. In general, the techniques described here can be applied to select between the integer pixel resolution and any fractional pixel resolution, for example, the resolution of<sup>1</sup>/ £ pixel or select between integer pixel resolution and a plurality of different fractional pixel modes, eg select between integer,% and% pixel modes.
Furthermore, the scope of the invention is not limited to an application where the encoded video and / or the screen is transmitted over a network, nor to the streams that are live. For example, in other applications, the current can be stored on a storage device such as an optical disk, a hard disk or other magnetic storage, or a flash memory stick or other electronic memory. It should be noted that the screen sharing stream therefore does not necessarily have to be shared live (although it is an option). Alternatively or additionally, it can be stored for sharing with one or more users later, or the captured image data is not shared, rather it is only recorded for the user who was using the screen at the time. In general, the screen capture can be any moving image data, which consists of captured encoder-side screen content, captured by any appropriate means (not necessarily when reading screen memory, although this is an option). , to be shared with one or more users (live or not) or simply recorded for the benefit of the user who captures or just to be archived (probably never seen again, as it happens).
It should be noted that the encoder-decoder is not necessarily limited to encoding only the screen capture data and the video. In certain modes, you may also have the ability to encode other types of moving image data, for example, an animation. Such other types of motion picture data can be encoded in fractional pixel mode or in whole pixel mode.
Furthermore, it should be noted that the encoding must always necessarily encode in relation to the previous frame, but more generally, some codes may allow encoding in relation to a frame other than a target encoder, before or in advance of the target frame (assuming an output memory).
Furthermore, as described above, it should be noted that the motion vectors themselves can be differentially encoded. In this case, when the motion vector is said to be pointed at an encoded bitstream it is constrained to an integer number of pixels or the like, this means that the differentially encoded shape of the motion vector is constrained as well (by example, the delta).
Furthermore, the decoder does not necessarily have to be implemented at the end-user terminal nor does it output the moving image data for Immediate consumption at the receiving terminal. In alternative implementations, the receiving terminal may be an intermediate terminal, such as a server running the decoder software, to output the motion picture data to another terminal in decoded or transcoded form, or to store the decoded data for consumption. future. Similarly, the encoder does not have to be implemented at the end user terminal, nor does it encode moving image data originating from the transmission terminal. In other embodiments, the transmission terminal, for example, may be an intermediate terminal, such as a server that runs the encoder software to receive the motion picture data in
IMPI unencoded or alternatively encoded form i encode or transcode this data for storage in the service or for sending to the receiving terminal.
In general, any of the functions described here can be implemented with the use of software, firmware, hardware (for example, fixed logic circuitry) or a combination of these implementations. The terms "module," "functionality," "component," and "logical" are generally used herein to represent software, firmware, hardware, or a combination thereof. In the case of a software implementation, the module, functionality, or logic represents the program code that performs specific tasks when they are run on a processor (for example, CPU or multiple CPUs). The program code can be stored on one or more computer readable memory devices. The characteristics of the techniques described below are platform independent, which means that the techniques can be implemented on a variety of commercial computing platforms that have a variety of processors.
For example, the terminals may include an entity (eg, software) that causes the hardware of the user's terminals to perform certain operations, for example, processor function blocks and so on. For example, the terminals may include a computer-readable medium that may be configured to maintain instructions that cause the user terminals, and more particularly, the operating system and associated hardware of the user terminals, to perform certain operations. Thus, the function of
<img file="MX360925B_D0018.tif" />
in the
IMPI instructions to configure the operating system and to carry out the operations and thus, transformation of the operating system and associated hardware to carry out the functions. Instructions can be provided by the computer readable medium to the terminals through a variety of different configurations.
Such a configuration of a computer readable medium is a signal carrying medium and is therefore configured to transmit the instructions (eg, the carrier wave) to the computing device, such as over the network. The computable readable medium may also be configured as a computer readable storage medium and therefore is not a signal-carrying medium. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disc, flash memory, hard disk memory, and other memory devices that may be magnetic. , optical and can use other techniques to store Instructions and other data.
Although the matter has been described in specific language for structural features and / or methodological actions, it should be understood that the material defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are described as exemplary ways to implement the claims.
Contents18
22 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22
76 members in 11 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 201461925090 | United States of America | P | |
| 201461925090 | United States of America | P | |
| 61925090 | United States of America | – | |
| 14530616 | United States of America | – | |
| 201414530616 | United States of America | A | |
| 201414530616 | United States of America | A | |
| 2014071331 | United States of America | W | |
| 2014071331 | United States of America | W | |
| 14530616 | – | – | – |
| 61925090 | – | – | – |
| PCTUS2014071331 | – | – | – |
| US201414530616 | – | – | – |
| US201461925090P | – | – | – |
| WO2014US71331 | – | – | – |
Members76
| Document | Office | Kind | |
|---|---|---|---|
| US2015195525A1 | United States of America | A1 | |
| US2015195557A1 | United States of America | A1 | |
| CA2935340A1 | Canada | A1 | |
| CA2935562A1 | Canada | A1 | |
| CA3118603A1 | Canada | A1 | |
| WO2015105661A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015105662A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2014376190A1 | Australia | A1 | |
| AU2014376189A1 | Australia | A1 | |
| CN105900419A | China | A | |
| CN105900420A | China | A | |
| KR20160106155A | Republic of Korea | A | |
| KR20160106703A | Republic of Korea | A | |
| MX2016009023A | Mexico | A | |
| MX2016009025A | Mexico | A | |
| EP3075153A1 | European Patent Office (EPO) | A1 | |
| EP3075154A1 | European Patent Office (EPO) | A1 | |
| JP2017508344A | Japan | A | |
| JP2017508348A | Japan | A | |
| BR112016015243A2 | Brazil | A2 | |
| BR112016015854A2 | Brazil | A2 | |
| US9749642B2 | United States of America | B2 | |
| US2017359587A1 | United States of America | A1 | |
| RU2016127410A | Russian Federation | A | |
| RU2016127312A | Russian Federation | A | |
| US9900603B2 | United States of America | B2 | |
| US9942560B2 | United States of America | B2 | |
| US2018131947A1 | United States of America | A1 | |
| AU2014376189B2 | Australia | B2 | |
| MX359698B | Mexico | B | |
| MX360925BThis record | Mexico | B | |
| AU2014376190B2 | Australia | B2 | |
| RU2679349C1 | Russian Federation | C1 | |
| RU2682859C1 | Russian Federation | C1 | |
| JP6498679B2 | Japan | B2 | |
| CN105900419B | China | B | |
| US10313680B2 | United States of America | B2 | |
| CN105900420B | China | B | |
| CN110099278A | China | A | |
| CN110149513A | China | A | |
| CN110177274A | China | A | |
| US2019281309A1 | United States of America | A1 | |
| JP6576351B2 | Japan | B2 | |
| EP3075154B1 | European Patent Office (EPO) | B1 | |
| MX373555B | Mexico | B | |
| US2020177887A1 | United States of America | A1 | |
| BR112016015854A8 | Brazil | A8 | |
| US10681356B2 | United States of America | B2 | |
| US10735747B2 | United States of America | B2 | |
| US2020329247A1 | United States of America | A1 | |
| EP3075153B1 | European Patent Office (EPO) | B1 | |
| KR20210073608A | Republic of Korea | A | |
| KR102270095B1 | Republic of Korea | B1 | |
| KR102271780B1 | Republic of Korea | B1 | |
| US11095904B2 | United States of America | B2 | |
| CA2935562C | Canada | C | |
| US2021337214A1 | United States of America | A1 | |
| KR102360403B1 | Republic of Korea | B1 | |
| KR20220019845A | Republic of Korea | A | |
| CN110099278B | China | B | |
| CN110149513B | China | B | |
| CN110177274B | China | B | |
| KR102465021B1 | Republic of Korea | B1 | |
| KR20220153111A | Republic of Korea | A | |
| CA2935340C | Canada | C | |
| BR112016015243B1 | Brazil | B1 | |
| BR112016015854B1 | Brazil | B1 | |
| BR122022001631B1 | Brazil | B1 | |
| US11638016B2 | United States of America | B2 | |
| US2023209070A1 | United States of America | A1 | |
| KR102570202B1 | Republic of Korea | B1 | |
| KR20230127361A | Republic of Korea | A | |
| US12108054B2 | United States of America | B2 | |
| KR102725421B1 | Republic of Korea | B1 | |
| KR20240161211A | Republic of Korea | A | |
| US2024414356A1 | United States of America | A1 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 360925
- Publication, DOCDB
- 360925
- Publication, EPODOC
- MX360925
- Application
- 2016009023
- Application, DOCDB
- 2016009023
- Application, EPODOC
- MX20160009023
Titles
- Spanish
- CODIFICACION DE VIDEO EN DATOS DE CONTENIDO DE PANTALLA.
Classification
- CPC, 8
- H04N19/109
- H04N19/523
- H04N19/42
- H04N19/136
- H04N19/174
- H04N19/503
- H04N19/70
- H04N19/90
- IPC, 6
- H04N19 51
- H04N19 42
- H04N19 109
- H04N19 136
- H04N19 174
- H04N19 523