Video coding device and video decoding device
9 claims: 9 independent, 0 dependent
- 1Videocodiervorrichtung, die codierten Daten eine hierarchische Struktur verleihen kann, mit:- einer Bereichsauswähl-Einrichtung (5) zum Auswählen eines spezifizierten Bereichs jedes Bilddatenrahmens;- einer Bereichs-Position/Form-Codiereinrichtung (6) zum Codieren der Position und der Form des ausgewählten Bereichs;- einer Untere-Ebene-Codiereinrichtung (4) zum Codieren eines Pixelwerts im ausgewählten Bereich in solcher Weise, dass er nur relativ niedriger Bildqualität entspricht;- einer Erste-Obere-Ebene-Codiereinrichtung (1) zum vorhersagenden Codieren eines Pixelwerts für ein gesamtes Bild jedes Bilddatenrahmens in solcher Weise, dass er von relativ niedriger Bildqualität ist, unter Verwendung des bereits decodierten Pixelwerts codierter Daten der unteren Ebene und eines bereits decodierten Pixelwerts codierter Daten der ersten oberen Ebene;- einer Zweite-Obere-Ebene-Codiereinrichtung (2) zum vorhersagenden Codieren eines Pixelwerts für den ausgewählten Bereich in solcher Weise, dass er von relativ hoher Bildqualität ist, unter Verwendung eines bereits decodierten Pixelwerts codierter Daten der unteren Ebene und eines bereits decodierten Pixelwerts decodierter Daten der zweiten oberen Ebene;und - einer Codierte-Daten-Integriereinrichtung (3) zum Integrieren codierter Daten, die durch die oben genannte Codiereinrichtung erhalten wurden, um den codierten Daten eine hierarchische Struktur zu verleihen.
- 2Videocodiervorrichtung, die codierten Daten eine hierarchische Struktur verleihen kann, mit:- einer Bereichsauswähl-Einrichtung (5) zum Auswählen eines spezifizierten Bereichs jedes Bilddatenrahmens;- einer Bereichs-Position/Form-Codiereinrichtung (6) zum Codieren der Position und der Form des ausgewählten Bereichs;- einer Untere-Ebene-Codiereinrichtung (4) zum Codieren eines Pixelwerts im ausgewählten Bereich in solcher Weise, dass er nur relativ niedriger Bildqualität entspricht, - einer Erste-Obere-Ebene-Codiereinrichtung (1) zum vorhersagenden Codieren eines Pixelwerts für einen anderen Bereich als den ausgewählten Bereich in solcher Weise, dass er von relativ niedriger Bildqualität ist, unter Verwendung des bereits decodierten Pixelwerts codierter Daten der unteren Ebene und eines bereits decodierten Pixelwerts codierter Daten der ersten oberen Ebene;- einer Zweite-Obere-Ebene-Codiereinrichtung (2) zum vorhersagenden Codieren eines Pixelwerts für den ausgewählten Bereich in solcher Weise, dass er von relativ hoher Bildqualität ist, unter Verwendung eines bereits decodierten Pixelwerts codierter Daten der unteren Ebene und eines bereits decodierten Pixelwerts decodierter Daten der zweiten oberen Ebene;und - einer Codierte-Daten-Integriereinrichtung (3) zum Integrieren codierter Daten, die durch die oben genannte Codiereinrichtung erhalten wurden, um den codierten Daten eine hierarchische Struktur zu verleihen.
- 3Videodecodiervorrichtung zum Decodieren eines Videobilds aus codierten Daten, die Folgendes beinhalten:Position/Form-Codes für einen ausgewählten spezifizierten Bereich jedes Bilddatenrahmens, einen Code einer unteren Ebene mit relativ niedriger Bildqualität für einen Pixelwert des ausgewählten Bereichs, einen Code einer ersten oberen Ebene mit relativ niedriger Bildqualität, der durch vorhersagendes Codieren eines Pixelwerts für ein gesamtes Bild jedes Bilddatenrahmens unter Verwendung eines bereits decodierten Pixelwerts der codierten Daten der unteren Ebene erhalten wurde, und einen Code einer zweiten oberen Ebene mit relativ hoher Bildqualität, der durch vorhersagendes Codieren eines Pixelwerts des ausgewählten Bereichs unter Verwendung eines bereits decodierten Pixelwerts der codierten Daten der unteren Ebene erhalten wurde, mit: - einer Codierte-Daten-Trenneinrichtung (7) zum gesonderten Entnehmen des Positionscodes und des Formcodes für den ausgewählten Bereich und des Codes der unteren Ebene aus den codierten Daten;- einer Bereichs-Position/Form-Decodiereinrichtung (8) zum Decodieren des Positionscodes und des Formcodes des ausgewählten Bereichs und - einer Untere-Ebene-Decodiereinrichtung (9) zum Decodieren des Codes der unteren Ebene und zum Erstellen eines decodierten Bilds für den ausgewählten Bereich, das von relativ niedriger Bildqualität ist.
- 4Videodecodiervorrichtung zum Decodieren eines Videobilds aus codierten Daten, die Folgendes beinhalten:Position/Form-Codes für einen ausgewählten spezifizierten Bereich jedes Bilddatenrahmens, einen Code einer unteren Ebene mit relativ niedriger Bildqualität für einen Pixelwert des ausgewählten Bereichs, einen Code einer ersten oberen Ebene mit relativ niedriger Bildqualität, der durch vorhersagendes Codieren eines Pixelwerts eines anderen Bereichs als des ausgewählten Bereichs jedes Bilddatenrahmens unter Verwendung eines bereits decodierten Pixelwerts der codierten Daten der unteren Ebene erhalten wurde, und einen Code einer zweiten oberen Ebene mit relativ hoher Bildqualität, der durch vorhersagendes Codieren eines Pixelwerts des ausgewählten Bereichs unter Verwendung eines bereits decodierten Pixelwerts der codierten Daten der unteren Ebene erhalten wurde, mit: - einer Codierte-Daten-Trenneinrichtung (7) zum gesonderten Entnehmen des Positionscodes und des Formcodes für den ausgewählten Bereich und des Codes der unteren Ebene aus den codierten Daten;- einer Bereichs-Position/Form-Decodiereinrichtung (8) zum Decodieren des Positionscodes und des Formcodes des ausgewählten Bereichs und - einer Untere-Ebene-Decodiereinrichtung (9) zum Decodieren des Codes der unteren Ebene und zum Erstellen eines decodierten Bilds für den ausgewählten Bereich, das von relativ niedriger Bildqualität ist.
- 5Videodecodiervorrichtung zum Decodieren eines Videobilds aus codierten Daten, die Folgendes beinhalten:Position/Form-Codes für einen ausgewählten spezifizierten Bereich jedes Bilddatenrahmens, und einen Code einer unteren Ebene mit relativ niedriger Bildqualität für einen Pixelwert des ausgewählten Bereichs, oder aus codierten Daten, die Folgendes enthalten: Positions- Form-Codes eines ausgewählten spezifizierten Bereichs jedes Bilddatenrahmens, einen Code einer unteren Ebene mit relativ niedriger Bildqualität eines Pixelwerts des ausgewählten Bereichs sowie einen Code einer oberen Ebene mit relativ niedriger Bildqualität, der durch vorhersagendes Codieren eines Pixelwerts des Gesamtbilds jedes Bilddatenrahmens unter Verwendung eines bereits decodierten Pixelwerts der codierten Daten der unteren Ebene erhalten wurde, mit: - einer Codierte-Daten-Trenneinrichtung (7) zum gesonderten Entnehmen des Positions- und des Formcodes für den Bereich, des Codes der unteren Ebene und des Codes der oberen Ebene aus den codierten Daten;- einer Bereichs-Position/Form-Decodiereinrichtung (8) zum Decodieren des Positionscodes und des Formcodes des ausgewählten Bereichs und - einer Untere-Ebene-Decodiereinrichtung (9) zum Decodieren des Codes der unteren Ebene und zum Erstellen eines decodierten Bilds für den ausgewählten Bereich, das von relativ niedriger Bildqualität ist;und - einer Obere-Ebene-Decodiereinrichtung (11) zum Decodieren des Codes der oberen Ebene unter Verwendung des decodierten Bilds der unteren Ebene und zum Erstellen eines gesamten Bilds von relativ niedriger Qualität;- wodurch sie ein decodiertes Bild des ausgewählten Bereichs, das durch die Untere-Ebene-Decodiereinrichtung erstellt wird, oder ein decodiertes Bild eines gesamten Bereichs reproduziert, das durch die Obere-Ebene-Decodiereinrichtung erstellt wird.
- 6Videodecodiervorrichtung zum Decodieren eines Videobilds aus decodierten Daten, die Folgendes enthalten:Positions- und Formcodes eines ausgewählten spezifizierten Bereichs jedes Bilddatenrahmens sowie einen Code einer unteren Ebene mit relativ niedriger Bildqualität eines Pixelwerts des ausgewählten Bereichs, oder aus codierten Daten, die Folgendes enthalten: Positions- und Formcodes eines ausgewählten spezifizierten Bereichs jedes Bilddatenrahmens, einen Code einer unteren Ebene mit relativ niedriger Bildqualität eines Pixelwerts des ausgewählten Bereichs, wobei der Code unter Verwendung eines bereits decodierten Pixelwerts der codierten Daten der unteren Ebene erhalten wird, und einen Code einer oberen Ebene mit relativ niedriger Bildqualität eines Pixelwerts eines anderen Bereichs als des ausgewählten Bereichs, wobei der Code durch vorhersagendes Codieren des Pixels unter Verwendung eines bereits decodierten Pixelwerts der codierten Daten der unteren Ebene erhalten wird, mit: - einer Codierte-Daten-Trenneinrichtung (10) zum gesonderten Entnehmen der Positions- und Formcodes des ausgewählten Bereichs, des Codes der unteren Ebene und des Codes der oberen Ebene aus den codierten Daten;- einer Bereichs-Position/Form-Decodiereinrichtung (9) zum Decodieren des Positionscodes und des Formcodes des ausgewählten Bereichs;- einer Untere-Ebene-Decodiereinrichtung (8) zum Decodieren des Codes der unteren Ebene und zum Erstellen eines decodierten Bilds des ausgewählten Bereichs so, dass es von relativ niedriger Bildqualität ist;- einer Obere-Ebene-Decodiereinrichtung (11) zum Decodieren des Codes der oberen Ebene unter Verwendung des decodierten Bilds der unteren Ebene und zum Erstellen eines decodierten Bilds eines anderen Bereichs als des ausgewählten Bereichs so, dass es von relativ niedriger Bildqualität ist;- wodurch sie ein decodiertes Bild des ausgewählten Bereichs, das durch die Untere-Ebene-Decodiereinrichtung decodiert wurde, oder ein decodiertes Bild des gesamten Bereichs, das ferner ein decodiertes erstelltes Bild eines anderen Bereichs enthält, das durch die Obere-Ebene-Decodiereinrichtung decodiert wurde, reproduziert.
- 7Videodecodiervorrichtung zum Decodieren eines Videobilds aus codierten Daten, die Folgendes enthalten:Positions- und Formcodes eines ausgewählten spezifizierten Bereichs jedes Vollbilds und einen Code einer unteren Ebene mit relativ niedriger Bildqualität eines Pixelwerts eines ausgewählten Bereichs, oder aus codierten Daten, die Folgendes enthalten: Positions- und Formcodes eines ausgewählten spezifizierten Bereichs jedes Bilddatenrahmens, einen Code einer unteren Ebene mit relativ niedriger Bildqualität eines Pixelwerts des ausgewählten Bereichs und einen Code einer unteren Ebene mit relativ hoher Bildqualität, der durch vorhersagendes Codieren eines Pixelwerts des ausgewählten Bereichs unter Verwendung des decodierten Bilds der unteren Ebene erhalten wurde, mit: - einer Codierte-Daten-Trenneinrichtung (10) zum gesonderten Entnehmen der Positions- und Formcodes des Bereichs, des Codes der unteren Ebene und des Codes der oberen Ebene aus den codierten Daten;- einer Bereichs-Position/Form-Decodiereinrichtung (9) zum Decodieren des Positionscodes und des Formcodes des ausgewählten Bereichs;- einer Untere-Ebene-Decodiereinrichtung (8) zum Decodieren des Codes der unteren Ebene und zum Erstellen eines decodierten Bilds für den ausgewählten Bereich mit relativ niedriger Bildqualität;und - einer Obere-Ebene-Decodiereinrichtung (13) zum Decodieren des Codes der oberen Ebene unter Verwendung des decodierten Bilds der unteren Ebene und zum Erstellen eines decodierten Bilds des ausgewählten Bereichs mit relativ hoher Bildqualität;- wodurch sie ein decodiertes Bild, das durch die Untere-Ebene-Decodiereinrichtung mit relativ niedriger Bildqualität erzeugt wurde, oder ein decodiertes Bild, das durch die Obere-Ebene-Decodiereinrichtung mit relativ hoher Bildqualität erzeugt wurde, reproduziert.
- 8Videodecodiervorrichtung zum Decodieren eines Videobilds aus codierten Daten, die Folgendes enthalten:Positions- und Formcodes eines ausgewählten spezifizierten Bereichs jedes Bilddatenrahmens, einen Code einer unteren Ebene mit relativ niedriger Bildqualität eines Pixelwerts des ausgewählten Bereichs, einen Code einer ersten oberen Ebene mit relativ niedriger Bildqualität eines Pixelwerts des gesamten Bilds jedes Bilddatenrahmens und einen Code einer zweiten oberen Ebene mit höherer Bildqualität eines Pixelwerts des ausgewählten Bereichs, mit: - einer Codierte-Daten-Trenneinrichtung (10) zum gesonderten Entnehmen der Positions- und Formcodes des Bereichs, des Codes der unteren Ebene, des Codes der ersten oberen Ebene und des Codes der zweiten oberen Ebene aus den codierten Daten;- einer Bereichs-Position/Form-Decodiereinrichtung (9) zum Decodieren des Positionscodes und des Formcodes des ausgewählten Bereichs;- einer Untere-Ebene-Decodiereinrichtung (8) zum Decodieren des Codes der unteren Ebene und zum Erstellen eines decodierten Bilds des ausgewählten Bereichs mit relativ niedriger Bildqualität;- einer Erste-Obere-Ebene-Decodiereinrichtung (11) zum Decodieren des Codes der ersten oberen Ebene unter Verwendung des decodierten Bilds der unteren Ebene und zum Erstellen eines decodierten Bilds für den gesamten Bereich mit relativ niedriger Bildqualität;- einer Zweite-Obere-Ebene-Decodiereinrichtung (13) zum Decodieren des Codes der zweiten oberen Ebene unter Verwendung des decodierten Bilds der unteren Ebene und zum Erstellen eines decodierten Bilds des ausgewählten Bereichs mit relativ hoher Bildqualität;- wodurch sie ein decodiertes Bild des ausgewählten Bereichs, das mit relativ niedriger Bildqualität von der Untere-Ebene-Decodiereinrichtung erstellt wurde, oder ein decodiertes Bild des gesamten Bereichs, das mit relativ niedriger Bildqualität durch die Erste-Obere-Ebene-Decodiereinrichtung erstellt wurde, oder ein decodiertes Bild des gesamten Bereichs, das ferner ein decodiertes Bild des ausgewählten Bereichs enthält, das durch die Zweite-Obere-Ebene-Decodiereinrichtung mit relativ hoher Bildqualität erzeugt wurde, reproduziert.
- 9Videodecodiervorrichtung zum Decodieren eines Videobilds aus codierten Daten, die Folgendes enthalten:Positions- und Formcodes eines ausgewählten spezifizierten Bereichs jedes Bilddatenrahmens, einen Code einer unteren Ebene mit relativ niedriger Bildqualität eines Pixelwerts des ausgewählten Bereichs, einen Code einer ersten oberen Ebene mit relativ niedriger Bildqualität eines Pixelwerts eines anderen Bereichs als des ausgewählten Bereichs jedes Bilddatenrahmens sowie einen Code einer zweiten oberen Ebene mit relativ hoher Bildqualität eines Pixelwerts des ausgewählten Bereichs, mit: - einer Codierte-Daten-Trenneinrichtung (10) zum gesonderten Entnehmen der Positions- und Formcodes des ausgewählten Bereichs, des Codes der unteren Ebene, des Codes der ersten oberen Ebene und des Codes der zweiten oberen Ebene aus den codierten Daten;- einer Bereichs-Position/Form-Decodiereinrichtung (9) zum Decodieren des Positionscodes und des Formcodes des ausgewählten Bereichs;- einer Untere-Ebene-Decodiereinrichtung (8) zum Decodieren des Codes der unteren Ebene und zum Erstellen eines decodierten Bilds des ausgewählten Bereichs mit relativ niedriger Bildqualität;- einer Erste-Obere-Ebene-Decodiereinrichtung (11) zum Decodieren des Codes der ersten oberen Ebene unter Verwendung des decodierten Bilds der unteren Ebene und zum Erstellen eines decodierten Bilds eines anderen Bereichs mit relativ niedriger Bildqualität;- einer Zweite-Obere-Ebene-Decodiereinrichtung (13) zum Decodieren des Codes der zweiten oberen Ebene unter Verwendung des decodierten Bilds der unteren Ebene und zum Erstellen eines decodierten Bilds des ausgewählten Bereichs mit relativ hoher Bildqualität;- wodurch sie ein decodiertes Bild des ausgewählten Bereichs, das mit relativ niedriger Bildqualität durch die Untere-Ebene-Decodiereinrichtung erzeugt wurde, oder ein decodiertes Bild des gesamten Bereichs, das ferner ein decodiertes Bild des anderen Bereichs, das durch die Erste-Obere-Ebene- Decodiereinrichtung mit relativ hoher Bildqualität erzeugt wurde, enthält, oder ein decodiertes Bild des gesamten Bereichs, das ferner ein decodiertes Bild des ausgewählten Bereichs enthält, das durch die Zweite-Obere-Ebene- Decodiereinrichtung mit relativ hoher Bildqualität erzeugt wurde, reproduziert.
Independent claims9
290 paragraphs in 4 sections, as filed
BACKGROUND OF THE INVENTION
The invention relates to the technical field of digital video processing, and more particularly relates to a video coding apparatus for coding video data with high efficiency, and a video decoding apparatus for decoding video coded data produced by the video coding apparatus with high efficiency.
A video coding method has been proposed, with which it is possible to code a specified area to have higher picture quality than other areas.
A video coding method described in ISO / IEC JTC1 / SC29 / WG11 MPEG95 / 030 is such that a specified range is selected (hereinafter referred to as a selected range) and made to be so by adjusting the quantization step sizes and the temporal resolution is coded to have higher image quality.
Another conventional method has an area selection section which is provided to select a specified area of a video image. In the case of a selection process, for. B. For a facial area of a video image on a display of a video phone, it is possible to select an area using a method as described in the reference "Realtime auto facetracking system" (The Institute of Image Electronics Engineers of Japan, 93-04-04, pp. 13-16, 1993).
An area position / shape coding section encodes the position and shape of a selected area. An optional form may be made using e.g. B. code codes are encoded. The coded position and shape are assembled into coded data and transmitted or accumulated by an encoded data integrating section.
An encoded parameter setting section sets a number of parameters usable for adjusting the image quality or the data amount in a video encoding process such that the area position / shape encoding section can encode a selected area to have higher image quality than has other areas.
A parameter coding section encodes a number of set parameters. The coded parameters are combined into coded data and transmitted or accumulated by an encoded data integrating section. The video coding section encodes input video data using a number of the parameters by a combination of conventional coding methods such as motion compensating prediction, orthogonal transform, quantization, and variable-length coding. The coded video data is composed into coded data by the coded data integrating section, and then the coded data is transmitted or accumulated.
Thus, the selected area is coded to have higher image quality than other areas.
As stated above, in the conventional technique, the quality of the image of a selected area is improved by assigning a larger amount of bits to the temporal resolution by setting parameters such as quantizer step sizes, spatial resolution. However, the conventional technique involves problems in that an image of a specified area can not be obtained by decoding a part of decoded data and / or an image of a decoded area having a relatively low quality can be obtained because a selected area and other areas in the same group are coded Data is included. Recently, much research has been done on the hierarchical structure of coded data, but no success has been achieved in creating a system that allows the selection of a specified range.
A video coding method has been studied which is designed to synthesize various types of video sequences.
A publication "Image coding using hierarchical representation and multiple templates", published in Technical Report of IEICE IE94-159, pp. 99-106, in 1995, describes an image synthesizing method which includes a video sequence representing a background video signal and a video synthesizer a foreground video signal (e.g. a figure image or a fish image cut out using the Chroma key technique) combines forming the partial video sequence to produce a new sequence.
In a conventional method, it is assumed that a first video sequence is a background video signal and a second video sequence is a sub-video signal. An alpha plane corresponds to weighted data as used when synthesizing a partial image having a background image in a sequence of a moving image (video). An example image of pixels weighted from 1 to 0 has been proposed. It is assumed that the data of the alpha level within a part have the value 1 and outside of a part have the value 0. The alpha data may have a value of 0 to 1 in a boundary portion between a part and the outside thereof to indicate a mixed state of pixel values in the boundary portion and the transparency of a transparent substance such as glass.
In the conventional method, a first video coding section encodes the first video sequence, and a second video coding section encodes the second video sequence according to an internationally standardized video coding system, e.g. MPEG or H.261. An alpha-level encoding section encodes an alpha plane. According to the above publication, this section uses the techniques of vector quantization and hair transformation. An encoded data integrating section (not shown) integrates encoded data received from the encoding sections, and performs accumulation or transmission of the integrated encoded data.
In a decoding apparatus according to the conventional method, a coded data decomposing section (not shown) performs a coded data decomposition into the coded data of the first video sequence, the coded data of the second video sequence and the alpha-plane coded data, which are then passed through a first video decoding section. a second video decoding section and an alpha-plane decoding section, respectively. Two decoded sequences are synthesized according to weighted averages by a first weighting section, a second weighting section and an adder. The first and second video sequences are combined according to the following equation:
f (x, y, t) = (1-α (x, y, t)) f1 (x, y, t) + α (x, y, t) f2 (x, y, t)
In this equation, (x, y) represents coordinate data of an intra-frame pixel position, t denotes a frame time, f1 (x, y, t) represents a pixel value of the first video sequence, f2 (x, y, t) represents a pixel value of the second video sequence, f (x, y, t) represents a pixel value of the synthesized video sequence and α (x, y, t) represents alpha-level data. D. that is, the first weighting section uses 1-α (x, y, t) as the weight, while the second weighting section uses α (x, y, t) as the weighting. As stated above, the conventional method generates a large number of encoded data because it has to encode alpha-level data.
To avoid this problem, an amount of information can be saved by digitizing alpha-level data, but this involves a visual defect in that at the boundary between an image part and a background as a result of the discontinuous change of pixel values around there a serrated line occurs.
A video coding method has been studied which is used to synthesize various types of video sequences.
An "image coding using hierarchical representation and multiple templates" described in Technical Report of IEICE IE94-159, pp. 99-106, 1995 describes an image synthesizing method in which a video sequence constituting a background video signal and a foreground video signal (e.g. a figure image or a fish image cut out using the Chroma key technique) can be combined to produce a new sequence.
A publication "Temporal Scalability based on image content" (ISO / IEC JTC1 / SC29 / WG11 MPEG95 / 211, (1995)) describes a technique for generating a new video sequence by synthesizing a high frame rate partial video sequence with a low frame rate video sequence. This system is for coding a lower frame rate lower-layer frame by a predictive coding method and for coding only a selected portion of a high-frame-rate upper-layer frame by predictive coding. The upper layer does not encode frames encoded in the lower layer, but uses a copy of the decoded lower layer image. The selected area may be considered part of an image to be considered, e.g. B. as a human form.
In a conventional method, on the coding side, an input video sequence is thinned by first and second thinning sections, and the thinned frame rate reduced video sequence is then transmitted to an upper layer encoding section and a lower layer encoding section, respectively. The coding portion of the upper layer has a higher frame rate than the coding portion of the lower layer. The lower layer coding section encodes the entire picture of each frame in the received video sequence using an internationally standardized video coding method such as MPEG, H.261, etc. The lower layer coding section also constructs decoded frames used for predictive coding and simultaneously a synthesizing section be entered.
In a code amount control section of a conventional coding section, a coding section encodes video frames using a method or a combination of methods such as motion-compensated prediction, orthogonal transform, quantization, variable-length coding, etc. A section determining the quantization width (step size) determines that in a coding section using quantization width (step size). A portion for determining the amount of coded data calculates the accumulated amount of generated coded data. In general, the quantization width is increased or decreased to prevent an increase or decrease in the amount of coded data.
The upper layer coding section encodes only a selected portion of each frame in a received video sequence based on area information using an internationally standardized video coding method such as MPEG, H.261, etc. However, coded frames in the lower layer coding section are not coded by the upper layer coding section coded. The area information is information indicating a selected area, e.g. B. indicates an image of a human figure in each video frame, which is a digitized image that takes the value 1 in the selected area and takes the value 0 outside it. The coding section of the upper layer also creates decoded selected areas of each frame, which are transmitted to the synthesizing section.
A region information coding section encodes region information using 8-directional quantization codes. An 8-directional quantization code is a numeric code that indicates the direction of a previous point and is generally used to represent digital graphics.
A synthesizing section outputs a decoded video frame of the lower layer which has been coded by the coding section of the lower layer and is to be synthesized. When there is a frame to be synthesized that has not been encoded in the lower layer encoding section, the synthesizing section outputs a decoded video frame generated using two encoded frames coded in the lower layer and before and after the missing frame of the lower one Layer and a decoded frame of the lower layer to be synthesized. The two frames of the lower layer are in front of and behind the frame of the upper layer. The synthesized video frame is input to the upper layer coding section to be used there for predictive coding. The image processing in the synthesizing section is the following:
First, an interpolation image is generated for two frames of the lower layer. The decoded image of the lower layer at time t is expressed as B (a, y, t), where x and y are coordinates defining the position of a pixel in space. When the two lower-layer decoded images are at times t1 and t2 and the upper-layer decoded image is at t3 (t1 <t3 <t2), the interpolation image I (a, y, t3) at time t3 becomes as follows Equation (1) calculates:
I (x, y, t3) = [(t2-t3) B (x, y, t1) + (t3-t1) B (x, y, t2)] / (t2-t1) (1)
Then, the upper-layer decoded image E is synthesized with the obtained interpolation image I using synthesizing weighting information W (x, y, t) generated from area information. A synthesized image 5 is defined according to the following equation:
S (x, y, t) = [1-W (x, y, t)] I (x, y, t) + E (x, y, t) W (x, y, t) (2)
The area information M (x, y, t) is a digitized image that takes the value 1 in a selected area and takes the value 0 outside of it. The weighting information W (x, y, t) can be obtained by processing the aforementioned digitized image several times with a low-pass filter. D. that is, the weighting information W (x, y, t) takes the value 1 in a selected range, takes the value of 0 outside the same, and takes a value of 0 to 1 at the boundary of the selected range.
The coded data generated in the lower layer coding section, the upper layer coding section and the area information coding section are integrated by an integrating section (not shown) and then transmitted or accumulated.
On the decoding side of the conventional system, a coded data decomposing section (not shown) separates the coded data into those of the lower layer, those of the upper layer and those of area information. These coded data are decoded by a lower-layer decoding section, upper-layer decoding section, and area information decoding section.
A synthesizing section on the decoding side has similar structure as the synthesizing section. It synthesizes an image using a decoded image of the lower layer and a decoded image of the upper layer according to the same method as described for the encoding side. The synthesized video frame is displayed on a display screen, and it is simultaneously input to the upper layer decoding section to be used for prediction.
The above-described decoding apparatus decodes frames of both the lower and upper layers, but a decoding apparatus which forms a lower-layer decoding section is also employed, omitting the upper-layer coding section and the synthesizing section. This simplified decoding apparatus can reproduce a part of coded data.
Problems to be solved by the invention are the following:
(1) As mentioned above, in the conventional art, an output image of two decoded lower layer images and one upper layer decoded image is obtained by preliminarily forming an interpolation image of two lower layer frames, thereby encountering a problem in that the output image is considerably impaired at a large disturbance occurring in it by a selected area, when the position of the selected area changes over time.
The above problem is described as follows:
Images A and C are two decoded frames of the lower layer, and an image B is a decoded upper layer frame. The images are displayed in chronological order A, B and C. As the selected area moves, an interpolation image determined from images A and B shows two selected areas that overlap each other. The image B is further synthesized using weighting information with the interpolation image. In the source image, three selected areas overlap each other. Two selected areas of the lower layer image appear like an afterimage around the image of the selected area in the upper layer, significantly affecting image quality. Since the frames of the lower layer are normal and only synthesized frames the above Having trouble, the video sequence may be displayed with a periodic, flicker-like disturbance that significantly degrades video quality.
(2) The conventional technique uses 8-directional guantizing codes for coding area information. In the case of encoding low-bit-range area information or a complicated-shape area, the amount of coded area information increases and takes a large proportion of the total amount of coded data, which may result in deterioration of the picture quality.
(3) In the conventional technique, weight information is obtained by making the area information pass through a low-pass filter several times. This increases the amount of processing.
(4) The conventional technique uses a predictive coding method. However, the predictive coding of lower layer frames can result in a large disturbance when a screen change occurs in a video sequence. Disturbance of any lower layer frame may propagate over related images of the upper layer, resulting in prolonged interference of the video signal.
(5) According to the conventional technique, each lower-layer frame is encoded using an internationally standardized video coding method (e.g., MPEG and H.261), whereby the picture of a selected area is little different in quality from other areas. In contrast, in each frame of the upper layer, only a selected area is coded to be of high quality, thereby temporally varying the quality of the image of the selected area. This is perceived as a flicker-like disorder that is a problem.
SUMMARY OF THE INVENTION
Accordingly, it is an object of the invention to provide encoding and decoding apparatus which can encode a selectively specified portion of a video image to be of relatively high picture quality throughout the system of encoded video data and which also provide for a hierarchical structure of the encoded data can do whatever it takes reproduce the specified area of the coded video image with a variation in image quality and / or any other area of relatively low image quality.
With the encoding and decoding devices thus constructed, a selected image area can be coded and decoded to have higher image quality than other areas, by distinguishing values of parameters such as spatial resolution, quantizer step sizes, and temporal resolution. The coding device can make coded data have respective hierarchical orders, and therefore the decoding device can easily decode a part of coded data.
Another object of the invention is to provide a coding apparatus and a decoding apparatus capable of producing a synthesized image from a reduced amount of coded data without impairing the quality of the synthesized image.
In the encoding and decoding apparatuses according to the invention, the decoding apparatus can generate weighting information for synthesizing a plurality of video sequences using a weighting means, eliminating the need to encode weighting information by the encoding apparatus.
The coded data is weighted, which can result in a total saving of the amount of data to be generated.
The weighting reversal carried out on the decoding side can produce weighted decoded data.
Another object of the invention is to provide a coding apparatus and a decoding apparatus free from the above-mentioned problems (described as problems to be solved in the prior art (1) to (5)) and the video frames with a reduced amount coded data without compromising image quality.
BRIEF DESCRIPTION OF THE DRAWINGS
Fig. 1 is a block diagram for explaining a prior art.
Fig. 2 is a view for explaining a concept of a coding method according to the invention.
Fig. 3 shows an example of a concept of a decoding method according to the invention.
Fig. 4 shows another example of a concept of a decoding method according to the invention.
Fig. 5 is a block diagram showing a coding apparatus representing one embodiment of the invention.
Fig. 6 shows an exemplary method of coding code data of a lower layer, a first upper layer and a second upper layer by a coding device according to the invention.
Fig. 7 shows another exemplary method for coding first code data of an upper layer by an encoding device according to the invention.
Fig. 8 shows another exemplary method of coding code data of a lower layer, a first upper layer and a second upper layer by a coding device according to the invention.
Fig. 9 shows another exemplary method of coding lower-layer code data, a first upper layer, and a second upper layer code data by a coding apparatus according to the invention.
Fig. 10 is a block diagram showing a decoding apparatus representing an embodiment of the invention.
Fig. 11 is a block diagram showing a decoding apparatus representing another embodiment of the invention.
Fig. 12 is a block diagram showing a decoding apparatus representing another embodiment of the invention.
Fig. 13 is a block diagram for explaining a conventional method.
Fig. 14 shows an example of an alpha plane according to a conventional method.
Fig. 15 is a block diagram for explaining an embodiment of the invention.
Fig. 16 shows an example of area information according to the invention.
Fig. 17 shows an example of a linear weighting function according to the invention.
Fig. 18 shows an example of creating an alpha plane according to the invention.
Fig. 19 is a block diagram for explaining another embodiment of the invention.
Fig. 20 is a block diagram for explaining an example of a video coding section in another embodiment of the invention.
Fig. 21 is a block diagram for explaining a video decoding section in another embodiment of the invention.
Fig. 22 is a block diagram for explaining another example of a video coding section in another embodiment of the invention.
Fig. 23 is a block diagram for explaining another example of a video decoding section in another embodiment of the invention.
Fig. 24 is a block diagram for explaining another example of a video coding section in another embodiment of the invention.
Fig. 25 is a block diagram for explaining another example of a video decoding section in another embodiment of the invention.
Fig. 26 is a block diagram for explaining an exemplary case that, in another embodiment of the invention, no area information is encoded.
Fig. 27 shows a concept of a conventional method.
Fig. 28 is a block diagram for explaining a conventional encoding and decoding system.
Fig. 29 is a block diagram for explaining a conventional method of controlling the number of codes.
Fig. 30 is a view for explaining an 8-directional quantization code.
Fig. 31 is a view for explaining problems in a conventional method.
Fig. 32 is a block diagram for explaining an embodiment of the invention.
Fig. 33 is a view for explaining effects of an embodiment of the invention.
Fig. 34 is a block diagram for explaining another embodiment of the invention.
Fig. 35 is a block diagram for explaining a coding page of another embodiment of the invention.
Fig. 36 is a block diagram for explaining a decoding side of another embodiment of the invention.
Fig. 37 shows an example of approximating area information using rectangles.
Fig. 38 is a block diagram for explaining another embodiment of the invention.
Fig. 39 shows an example method for generating weighting information according to the invention.
Fig. 40 is a block diagram for explaining another embodiment of the invention.
Fig. 41 is a block diagram for explaining another embodiment of the invention.
Fig. 42 is a view for explaining a target coefficient of codes usable for coding a selected area by a code amount control method according to the invention.
Fig. 43 is a view for explaining a target coefficient of codes usable for coding an area outside a selected area by a code amount control method according to the invention.
Fig. 44 is a view for explaining a target coefficient of codes usable for coding an area outside a selected area by a code amount control method according to the invention.
PREFERRED EMBODIMENT OF THE INVENTION
Fig. 1 is a block diagram showing a prior art for comparison with the invention. An area selection section 20 is to select a specified area of a video image. In the case of a selection process, e.g. B. For a facial area of a video image on a display of a video telephone, it is possible to select an area using a method described in the reference "Real-time Auto Face Tracking System" (The Institute of Image Electronics Engineers of Japan, Previewing Report of Society Meeting , Pp. 13-16, 1993).
In Fig. 1, an area position / shape coding section 21 encodes the position and shape of a selected area. It can be coded by an optional farm that z. For example, chain codes can be used. The coded position and the coded form are composed into code data and transmitted or accumulated by a code data integrating section 22.
A code parameter setting section 23 sets a number of parameters that are usable to set the picture quality or the data amount in the video coding so that the area position / shape coding section 21 can code a selected area to have higher picture quality than has the other areas.
The parameter coding section 24 encodes a number of set parameters. The code parameters are composed into code data and transmitted or accumulated by the code data integrating section 22. The video coding section 25 encodes input video data using a number of the parameters by a combination of conventional coding methods such as motion-compensated prediction, orthogonal transform, quantization and variable-length coding. The coded video data is composed into code data by the code data integrating section 22, and then the code data is transferred or accumulated. The concept of the invention will be described as follows.
Fig. 2 is a view for explaining the concept of the coding method according to the present invention. The hierarchical coding method of the invention uses a bottom layer (lower layer) and two top layers (higher layers). The lower layer encodes a selected area (shaded area) with relatively low image quality. A time to be noted is indicated by t, and a decoded image at time t is indicated by L (t). In the first upper layer, the entire image is encoded to have relatively low image quality. An encoded image of this layer is indicated by H1 (t). In this case, predictive coding is performed by using the decoded image of the lower layer L (t) and the decoded image of the first upper layer H1 (t-1). In the second upper layer, only the selected area is predictively coded to have higher image quality than that of the lower layer. The decoded image of this layer is labeled H2 (t). In this case, predictive coding is performed by using the decoded image of the lower layer L (t) and the decoded image of the second upper layer H2 (t-1).
Figures 3 and 4 are illustrative of the concept of the decoding method of the invention. Figure 3 shows decoding processes in three layers: decoding only the lower layer data, decoding only the first upper layer data, and decoding the data of all layers. In this case, by decoding the lower layer data, only one picture is reproduced for which the picture device has selected relatively low picture quality. By decoding the data of the first upper layer, an entire picture having relatively low picture quality is reproduced, and by decoding all the code data, the selected area of higher picture quality and all other areas of lower picture quality are reproduced. On the other hand, Fig. 4 shows the case where all the decoded signals are decoded after the data of the second upper layer has been decoded instead of the data of the first upper layer. In this case, the data of an intermediate layer (the second upper layer) is decoded to reproduce a selected image area only with higher image quality.
With the decoding apparatus of the present invention, only a selected picture area of lower picture quality is reproduced from lower layer corresponding code data, while the whole picture of lower picture quality or only a selected area of higher picture quality is reproduced from the upper layer corresponding code data. That is, any one of the two upper layers can be selected over a common lower layer.
An embodiment of the invention will be described as follows.
Fig. 5 is a block diagram showing an encoding apparatus embodying the invention.
In Fig. 5, an area selection section 5 and an area position / shape coding section 6 have similar functions to those in the prior art shown in Fig. 1.
In Fig. 5, a lower layer coding section 4 encodes only an area selected by the lower picture quality area selection section 5, creates lower layer code data, and generates a decoded picture from the code data. The decoded picture is used as a reference picture for predictive coding.
A first-layer coding section 1 encodes the entire picture which is to have lower picture quality, produces code data for the first layer, and generates a decoded picture from this code data. The decoded picture is used as a reference picture for predictive coding.
An encoding section 2 for the second layer encodes only the image of a selected area to be of higher image quality, it generates code data of the second layer, and generates a decoded image from this code data. The decoded picture is used as a reference picture for predictive coding.
A code data integrating section 3 integrates position / shape codes of a selected area, lower layer code data, first upper layer code data, and second upper layer code data.
There are several kinds of coding methods applicable to the lower layer coding section 4 1, the first layer coding section 1, and the second layer coding section 2, which are described as follows. Figs. 6 and 7 are illustrative of the technique of controlling the lower layer image quality and the upper layer image quality depending on quantization steps.
Fig. 6 (a) illustrates how to encode the image data of the lower layer. A hatched area represents a selected area. In the lower layer, a selected area of a first frame is intra-frame encoded, and selected areas of other, remaining frames are predictively encoded by a motion-compensating prediction method. As the reference image for the motion-compensating prediction, a selected area of a lower-layer frame which has already been coded and decoded is used. Although only forward prediction is shown in Figure 6 (a), use may be made in combination with backward prediction. Since the quantization step for the lower layer is made larger than that for the second upper layer, only a selected portion of an input image is encoded to have lower image quality (low signal / noise ratio). Accordingly, the lower layer image data is encoded using a smaller amount of code.
Fig. 6 (b) illustrates how to encode the image data of the first upper layer. An entire picture is encoded in this layer. For example, an entire image is encoded by predictive coding based on a decoded lower layer image and a first upper layer decoded image. In this case, the entire image of the first frame is coded by prediction from the decoded image of the lower layer (regions other than the ones selected are actually intra-frame encoded, since in practice the motion-compensating prediction method can not be used). Other frames may be encoded using the predictive coding in combination with the motion compensating prediction.
Also, a variation is applicable in which no selected area is coded by the predictive coding method shown in Fig. 6, but only other areas are coded. The encoding process is performed for areas other than the selected one.
Fig. 6 (c) illustrates how to encode the image data of the second upper layer. In a relatively small quantization step, only a selected image area is coded. In this case, objective data to be coded are difference data obtained between original image data and image data predicted from the lower layer image data. Although in FIG. 6 (c) shows only a prediction from the lower layer image data, application may be made in combination with a prediction from a second upper layer decoded frame.
Fig. 8 is a view for explaining a method of controlling the lower layer image quality and the upper layer image quality by using different temporal resolution values.
Fig. 8 (a) illustrates how to encode the image data of the lower layer. A hatched area represents a selected area. In the lower layer, a selected area of a first frame is intra-frame coded, and selected areas of other, remaining frames are predictively coded by motion-compensating prediction. As the reference image for the motion-compensating prediction, a selected area of a lower-layer frame which has already been coded and decoded is used. Although only forward prediction is shown in Figure 8 (a), use may be made in combination with backward prediction. The frame rate of the lower layer is lowered so that the temporal resolution is set to be lower than that for the second upper layer. It is also possible to encode frames with a smaller quantization interval, so that each frame can have a larger signal / noise ratio.
Fig. 8 (b) illustrates how to encode the image data of the first upper layer. An entire picture is encoded with a low time-picture resolution. In this case, it is possible to adopt a coding method similar to that shown in FIG. 6 (b) or FIG. 7.
Fig. 8 (c) illustrates how to encode the image data of the second upper layer. Only a selected range with higher temporal resolution is coded. In this case, a frame whose selected area has been coded in the lower layer is coded by prediction from the decoded picture of the lower layer, whereas all other frames are coded by motion compensating prediction from the already decoded upper layer frames. In the case of using a prediction from the decoded frame of the lower layer, it is not possible to encode any image data of the second upper layer using the decoded image of the lower layer as a decoded image of the second upper layer.
Fig. 9 is a view for explaining a method of image quality of the lower layer and image quality of the upper layer using different values of spatial resolution.
Fig. 9 (a) illustrates how to encode the image data of the lower layer. An original image is converted to a lower spatial resolution image by a low-pass filter or a thinning process. Only hatched, selected areas are coded. In the lower layer, a selected area of a first frame is intra-frame coded, and selected areas of other, remaining frames are predictively coded by motion-compensating prediction.
Fig. 9 (b) illustrates how to encode the image data of the first upper layer. An original image is converted to a lower spatial resolution image and the entire image is encoded at a higher temporal resolution. In this case, it is possible to adopt a coding method similar to that shown in FIG. 6 (b) or FIG. 7.
Fig. 9 (c) illustrates how to encode the image data of the second upper layer. Only a selected area with higher spatial resolution is coded. In this case, an encoded image of the lower layer is converted into an image having the same spatial resolution as that of an original image, and selected regions are obtained by prediction from the decoded image of the lower layer and by motion compensating prediction from the already decoded frame of the second upper layer coded.
The above-described image quality adjustment methods using the grayscale resolution (by the signal-to-noise ratio), the temporal resolution, and the spatial resolution may also be used in combination with each other.
For example, it is possible to adjust the image quality of the lower layer and the image quality of the upper layer using different spatial resolution and different temporal resolution in combination or a different quantization step and different temporal resolution in combination.
Thus, a selected area in an entire image is coded to have higher image quality than other areas. At the same time, the code data is assigned a respective one of three hierarchical layers (two upper layers and one lower layer).
Decoding devices that are preferred embodiments of the invention are described as follows.
Fig. 10 is illustrative of a first embodiment of a decoding apparatus according to the invention intended to decode only lower layer image data.
In Fig. 10, a code data separating section 7 is intended to divide code data into area position / shape code data and coded image data of a lower layer and selectively extracted desired code data.
An area position / shape decoding section 9 is to decode a position code and a shape code of a selected area.
A lower layer decoding section 8 is to decode lower layer code data for a selected area and to produce a lower quality decoded picture only for the selected area.
Accordingly, each image output from this decoding apparatus concerns image information of only a selected area displayed on a display screen as a window. The lower layer decoding section 8 may be provided with a spatial resolution converter to enlarge the selected area to the size of the full screen and display only this on the display screen.
With the illustrated embodiment, lower-quality decoded pictures can be obtained because only lower-layer data of a selected area is decoded, but it can be simple in hardware construction because the upper-layer decoding section is omitted, and it can easily decode the encoded picture thereby that it only processes a reduced amount of code data. The Fig. 11 Fig. 10 is illustrative of a second embodiment of a decoding apparatus according to the invention, in which a region position / shape decoding section 9 and a lower layer decoding section 8 have similar functions to those in the first embodiment.
In Fig. 11, the code data separating section 10 extracts from the code data separately area position / shape code data, code data for the lower layer of a region, and code data for the first upper layer.
A first upper layer decoding section 11 decodes code data of the first upper layer, and decodes an entire picture to lower quality using the area position / shape data, the decoded lower layer image, and the second upper layer decoded image. Thus, a decoded image is generated for the first upper layer.
Although the illustrated embodiment uses the code data of the first upper layer, it may also use the second upper layer instead of the first upper layer. In this case, the coded data separating section 10 separately extracts the area-position / shape coded data, the lower-layer coded data for one area, and the coded data of the second upper-layer separately from the coded data. The first upper layer code data separating section 11 is replaced by a second upper layer code data separating section that stores the second upper layer code data using the area position / shape data, the lower layer decoded image, and the decoded one Image of the second layer is decoded and only the selected higher quality image is decoded. A thus-generated decoded image of the second upper layer may be displayed as a window on a display screen, or it may be enlarged to the size of the full screen and then displayed thereon.
Fig. 12 is illustrative of a third embodiment of a coding apparatus according to the present invention, wherein a region position / shape decoding section 9 and a lower layer decoding section 8 are functionally similar to those shown in Fig. 2.
In Fig. 12, a code data separating section 10 separately takes out the code data of area position / shape data, lower layer code data, first upper layer code data, and second upper layer code data.
A first upper layer decoding section 11 decodes code data of the first upper layer, while a second upper layer decoding section 13 decodes second upper layer code data.
An upper layer synthesizing section 14 combines a decoded image of the second upper layer with a decoded image of the first upper layer to produce a synthesized image using information on the position and shape of the area. The synthesis for a selected area is performed using the decoded image of the second upper layer, while the synthesis is performed for other areas using the decoded image of the first upper layer. Therefore, an image output from the decoding apparatus relates to an overall image in which a selected area is specifically decoded to show higher quality in terms of parameters such as SRV (Signal to Noise Ratio), temporal and spatial resolution. A range selected by the encoder is accordingly decoded to be of higher quality than other ranges.
Fig. 13 is a block diagram showing a conventional apparatus for comparison with the invention. It is assumed that a first video sequence forms a background video signal and a second video sequence forms a sub-video signal. An alpha plane corresponds to weighting data as used when synthesizing a partial image having a background image in a moving image sequence (video). The Fig. 14 shows an exemplary image of pixels weighted with values from 1 to 0. It is assumed that the data of the alpha plane within a part have the value 1, while they have the value 0 outside one part. The alpha data may have a value of 0 to 1 in the boundary portion between a part and its exterior to indicate a mixed state of pixel values in the boundary portion and the transparency of a transparent substance such as glass.
Referring to Fig. 13, which illustrates the conventional method, a first video coding section 101 encodes the first video sequence, and a second video coding section 102 encodes the second video sequence according to an internationally standardized video coding system, e.g. MPEG or H.261. An alpha-plane encoding section 112 encodes an alpha plane. In the above publication, this section uses the techniques of vector quantization and hair transformation. A code data integrating section (not shown) integrates code data received from the coding sections and accumulates or sends the integrated code data.
In the decoding apparatus according to the conventional method, a coded data decomposing section (not shown) separates coded data into the coded data of the first video sequence, the coded data of the second video sequence and the alpha-plane coded data, which are then divided by a first video decoding section 105, a second video decoding section 106 and a second video decoding section 106, respectively Alpha plane decoding section 113 decoded. Two decoded sequences are synthesized according to weighted averages by a first weighting section 108, a second weighting section 109, and an adder 111. The first video sequence and the second video sequence are combined according to the following equation:
f (x, y, t) = (1-α (x, y, t)) f1 (x, y, t) + α (x, y, t) f2 (x, y, t)
In this equation, (x, y) represents coordinate data of an intraframe pixel position, t denotes a frame timing, f1 (x, y, t) represents a pixel value of the first video sequence, f2 (x, y, t) represents a pixel value of the second video sequence, f (x, y, t) represents a pixel value of the synthesized video sequence, and a (x, y, t) represents alpha-level data. D. That is, the first weighting section 108 uses the weighting 1- α (x, y, t), while the second weighting section 109 uses the weighting α (x, y, t).
As stated above, the conventional method generates a large number of code data because it has to encode data in the alpha plane.
To avoid this problem, the saving of information amount can be thought of by digitizing the data of the alpha plane, but it is accompanied by such a visual defect that at the boundary between a field and a background as a result of a discontinuous change of pixel values there tooth-shaped line occurs.
Fig. 15 is a block diagram showing a coding apparatus and a decoding apparatus embodying the invention. In the Fig. 15 For example, a first video coding section 101, a second video coding section 102, a first video decoding section 105, a second video decoding section 106, a first weighting section 108, a second weighting section 109, and an adder 111 have similar functions to those in the conventional apparatus, and therefore will not be explained further , In the Fig. 15 An area information coding section 103 encodes area information representing the shape of a field of a second video sequence, an area information decoding section 107 decodes the encoded area information, and an alpha level generating section 110 generates an alpha level using coded area information.
The functions of the coding device and the decoding device are as follows.
The coding device encodes the first and second video sequences by means of the first video coding section 101 and the second video coding section 102, respectively, and encodes area information by the area information coding section 103 according to a method to be described later. This coded data is integrated for further transmission or accumulation by the code data integrating section (not shown). On the other hand, the decoding apparatus separates the transmitted or accumulated coded data by the coded data separating section (not shown), and decodes the divided coded data by the first video decoding section 105, the second video decoding section 106, and the area information decoding section 107, respectively. The alpha plane generating section 110 creates from the decoded area information by an algorithm to be described later, an alpha plane. The first weighting section 108, the second weighting section 109 and the adder 111 may synthesize two decoded sequences using weighted averages corresponding to the created alpha plane.
FIG. 16 shows an example of area information corresponding to the area information of a sub video picture shown in FIG. 14. The area information is digitized using a threshold of 0.2. Thus, area information can be obtained by digitizing the alpha plane, or it can be determined by edge detection or other area division method. When an area is selected by a method as described in the reference "Real-time face image following-up method" (The Institute of Image Electronics Engineers of Japan, Previewing Society Meeting, 93-04-04, p. 16, 1993), the information to be used may be a rectangle. In this case, area information z. B. digitized so that it takes the value 1 within a body and outside the same 0.
A practical technique for coding area information, which will not be discussed in detail, may be run length coding and chain coding, since the area information corresponds to digitized data. If the region data represents a rectangle, only coordinate data for the origin, length, and width need to be encoded.
Various types of methods can be used to create an alpha plane depending on the shape representing the range information.
In the case of an area of rectangular shape, an alpha plane can be created by independently using the following linear weighting values in the horizontal and vertical directions of the rectangular area:
In the equation (1), M is "aN" and L is "NM" ("a" is a real value of 0 to 1). "N" represents the size of a rectangle area, and "a" represents the flatness of the weight to be applied to this area. Fig. 17 shows an example of a digital weighting function. An alpha-level corresponding to a rectangle is represented as follows:
α (x, y) = WNx, ax (x) WNy, ay (y) (2)
In the equation (2), the size of the rectangle is expressed by the number "Nx" of the pixels in the horizontal direction and the pixel "Ny" in the vertical direction, and the flatness of the weight is represented by "ax" in the horizontal direction and by "ay "expressed in the vertical direction.
Various combinations of linear weighting functions other than equation (1) for use may also be considered.
In the following, three different methods for creating an alpha-level for an area of any desired shape will be described by way of example.
The first method is to determine an enclosing rectangle of the area, and then apply the above-mentioned linear weighting functions to the enclosing rectangle in the horizontal and vertical directions.
The second method is for sequentially determining weighting values to be applied to an area from the circumference thereof, as shown in FIG. For example, pixels at the perimeter of the area are determined and each receive the weighting 0.2. Next, pixels at the periphery of a not-weighted part within the area are determined, and they are given the weighting 0.5, respectively. These operations are repeated until the perimeter pixels are weighted by one. The creation of an alpha layer is ended by applying a weight of 1.0 to the last unweighted area. The obtained alpha plane has a value of 1.0 in its central portion and a value of 0.2 at its peripheral portion. When weighting values are determined from the circumference of a region, it is possible to use a linear weighting function according to equation (1) or other linearly varying values. When sequentially changing a weighting value, the thickness of a perimeter pixel may correspond to a single or multiple pixels.
The third method is to provide the outside of a region with a weight of 0 and to weight the interior of the region 1 and then process such a digitized image via a low-pass filter to obtain a gradation in the region boundary portion. By changing the size and coefficient of a filter as well as the number of filter operations, different types of alpha planes can be created.
From the above, it can be seen that with the first embodiment, since the alpha plane is created by the decoder side, increased efficiency of data encoding can be achieved as compared with the conventional device, eliminating the need for encoding weighting information. In addition, the decoding apparatus constructs an alpha plane from the decoded area information, and synthesizes video sequences using the created alpha plane, thereby preventing the occurrence of such a visual defect that a serrated line appears at the boundary of one field in the background.
Other embodiments of the invention will be described as follows.
Fig. 19 is a block diagram showing an encoding apparatus and a decoding apparatus of the embodiment. In Fig. 19, a first weighting portion 108, a second weighting portion 109, and an adder 111 are similar to those of the conventional apparatus, and a further explanation thereof will be omitted. An area information coding section 103, an area information decoding section 107, and alpha-plane generating sections 120 and 121 have similar functions to those of the first embodiment, and therefore will not be explained further.
This embodiment is characterized in that the coding side is also provided with an alpha-plane generating section 120 for encoding a weighted value image for synthesizing a plurality of video sequences. Code data becomes smaller than the original data because the weight data is not larger than 1, and thus the code data amount can be reduced.
In Fig. 19, a first video coding section 122 and a second video coding section 123 encode pictures of video sequences by weighting based on respective alpha planes generated on the coding page. A first video decoding section 124 and a second video decoding section 125 decode the coded pictures of the video sequences by weight inversion based on the respective alpha planes created on the decoding side.
The first video coding section 122 or the second video coding section 123 may be constructed for transform coding, such as, e.g. B. shown in FIG. 20. A video sequence to be processed is the first or the second video sequence. A transform section 131 transforms an input image block by block using a DCT (discrete cosine transform), discrete Fourier transform and Weiblet transform method.
In Fig. 20, a first weighting section 132 weights a transform coefficient having an alpha level value. The value used for weighting may be a representation of an alpha plane within an image block to be processed. For example, the mean of the alpha plane within the block is used. Transformation coefficients of the first and second video sequences are represented by g1 (u, v) and g2 (u, v) and weighted according to the following equations:
gw1 (u, v) = (1 -) g1 (u, v) (3)
gw2 (u, v) = g2 (u, v)
In equation (3), gw1 (u, v) and gw2 (u, v) denote weighted transform coefficients, u and v denote the horizontal and vertical frequencies, and represents an alpha plane in a block.
In Fig. 20, a quantizing section 133 quantizes transform coefficients and a variable-length coding section 134 encodes the quantized transform coefficients with variable-run-length codes to generate code data.
A first video decoding section 124 or a second video decoding section 125 corresponding to the video coding section of FIG. 19 may be constructed as shown in FIG. 21. A variable-length-length decoding section 141 decodes code data, a reverse quantization section 142 performs inverse quantization of decoded data, and a weighting inversion section 143 performs a reverse operation on transformation coefficients to inverse the equation (2). D. that is, the transformation coefficients are weighted with weighting values inverse of those applied on the coding side according to the following equation:
1 (u, v) = w1 (u, v) / (1 -) (4)
2 (u, v) = w2 (u, v) /
In the equation (4), (roof) denotes decoded data; z. For example, gw1-roof is a weighted, decoded transform coefficient of the first video sequence.
Besides the above-mentioned weighting method, such a method is applicable in which not a direct component of a transform coefficient is weighted, but other transform coefficients are weighted according to the equation (2). In this case, the weighting is essentially done by combining a quantization step size used according to the international standard MPEG or H.261 using a representative value of the alpha plane within the block.
That is, a quantization step width changing section 38 is provided, as shown in Fig. 21, by which the quantization step size determined by a quantization step width determination section (not shown) is changed using the alpha plane data. In practice, a representative value (e.g. Average value) of the alpha plane within a block, and then the quantization step size is divided by a value (1 -) for the first video sequence or a value for the second video sequence to obtain a new quantization step size.
There are two weight reversal methods that correspond to the above-mentioned weighting method. The first method relates to the case that a quantization step width (without change by the quantization step width changing portion 138) is encoded by the encoding apparatus shown in FIG. In this case, the decoding apparatus of FIG. 23 provided with a quantization step width changing portion 148 that corresponds to that of the coding side of FIG. 22 corresponds to the quantization step width by a quantization step width decoding section (not shown), and then it changes the decoded quantization step width by the quantization step width changing section 148 according to the data of the alpha plane. The second method relates to the case where a quantization step size after being changed by the quantization step width changing section 138 by the quantization step width 138 shown in FIG. 22 illustrated encoding device is encoded. In this case, the decoder directly uses the decoded quantization step size and quantizes it in an inverse manner. This eliminates the use of a special weight reversing device (ie, the quantization step change section 108 of Fig. 23). However, it is believed that the second method has reduced flexibility in weighting compared to the first method.
The second embodiment described above uses transform coding. Therefore, a motion compensating coding section, as is characteristic of the MPEG or H.261 system, is omitted from Figs. However, this method can be applied to a coding system using motion compensating prediction. In this case, a transformation section 131 of FIG. 20 entered a prediction error for motion compensatory prediction.
Other weighting methods in the second embodiment are the following:
FIG. 24 shows an example of the first video coding section 122 or the second video coding section 123 of the coding apparatus shown in FIG. 19. That is, the coding section is provided with a weighting section 150 which performs a weighting operation before video coding by the standard MPEG or H.261 method according to the following equation:
fw1 (x, y) = (1 -) f1 (x, y) (5)
fw2 (x, y) = f2 (x, y)
In the equation (5), fw1 (x, y) is the first weighted video sequence, fw2 (x, y) is the second weighted video sequence, and a represents an alpha plane within a block.
The weighting can be performed according to the following equation:
fw1 (x, y) = (1-α (x, y)) f1 (x, y) (6)
fw2 (x, y) = α (x, y) f2 (x, y)
Fig. 25 shows a weighting reversing method of the decoding apparatus which corresponds to the above-mentioned weighting method. The weighting reversal section 161 weights the video sequence with a weighting inverse to that applied by the encoding device.
When the coding device has weighted the video sequence according to the equation (5), the weighting inversion section 61, the first weighting section 108 and the second weighting section 109 for synthesizing sequences as shown in Fig. 19 may be omitted from the decoding apparatus. That is, it is possible to use a coding device and a decoding device as shown in FIG. 26. A first video coding section 122 and a second video coding section 123 shown in Fig. 26 are constructed as shown in Fig. 24, and employ the weighting method of the equation (5). In this case, weighting information such as area information and alpha level data required to synthesize the video sequences is included in the video code data itself, and the weighting information need not be coded. Accordingly, sequences decoded by the decoding device can be added directly to each other to produce a synthesized sequence. Coding only data within a range is quite more effective than encoding an entire image when a video sequence 102 is a subpicture. In this case, it becomes necessary to code the area information by the coding apparatus and to decode the coded area information by the decoding apparatus.
The above description relates to an example of the weighting of each of a plurality of video sequences in the second embodiment of the invention. For example, the first video sequence is weighted with the value (1-) while the second video sequence is weighted with the value.
Although the embodiments in the case of synthesizing a background video sequence and a partial video sequence have been explained, the invention is not limited thereto but may be applied to composing a plurality of partial video sequences with a background. In this case, each area information corresponding to each field is coded.
The background image and the partial images can be coded independently or hierarchically, with the background image being regarded as the lower layer and the partial images being regarded as upper layers. In the latter case, each upper layer image can be effectively encoded by predicting its pixel value from that of the lower layer image.
A video coding method has been studied which is used to synthesize various types of video sequences.
The following description shows conventional devices for comparison with the invention.
The publication "Image coding using hierarchical representation and multiple templates" published in Technical Report of IEICE IE94-159, pp. 99-106, 1995 describes an image synthesizing method in which a video sequence constituting a background video signal and a partial video sequence constituting a foreground video signal (e.g. a figure image or a fish image cut out using the chrome key technique) can be combined to create a new sequence.
The publication "Temporal Scalability based on image content" (ISO / IEC JTC1 / SC29 / WG11 MPEG95 / 211, (1995)) describes a technique for creating a new video sequence by synthesizing a higher frame rate partial video sequence using a low frame rate video sequence. As shown in FIG. 27 This system is for coding a low-frame-rate lower-layer frame by a predictive coding method, and coding only a selected portion (hatched portion) of a high-frame-rate upper-layer frame by predictive coding. The upper layer does not encode frames encoded in the lower layer and uses a copy of the decoded lower layer image. The selected area may be considered part of the image to be considered, e.g. B. as a human form.
Fig. 28 is a block diagram showing a conventional method, on the coding side, in which an input video sequence is thinned by a first thinning section 201 and a second thinning section 202, and the thinned frame rate reduced video sequence is then thinned to an upper encoding section Layer or a coding section for the lower layer is transmitted. The upper layer coding section has a frame rate higher than that of the lower layer coding section.
The lower layer coding section 204 encodes the entire picture of each frame in the received video sequence using an internationally standardized video coding method such as MPEG, H.261, etc. The lower layer coding section 204 also creates decoded frames used for predictive coding and which are simultaneously input to a synthesizing section 205.
Fig. 29 is a block diagram of a code amount control section of a conventional coding section. In Fig. 29, a coding section 212 encodes video frames using a method or a combination of methods such as motion compensating prediction, orthogonal transform, quantization, variable-length coding, etc. A quantization width (step size) determining section 211 determines a quantization width (step size) to be used in a coding section 212. A code data amount determination section 213 calculates the accumulated amount of generated code data. In general, the quantization width is increased or decreased to avoid an increase or decrease in the code data amount.
In Fig. 28, the upper layer coding section 203 encodes only a selected portion of each frame in a received video sequence based on area information using an internationally standardized video coding method such as MPEG, H.261, etc. However, in the lower layer coding section 204 encoded frames are not encoded by the upper layer encoding section 203. The area information is such information that includes a selected area of e.g. 2, which is a digitized image having the value 1 in the selected region and the value 0 outside thereof. The upper layer coding section 203 also creates decoded selected areas of each frame, which are transmitted to the synthesizing section 205.
An area information coding section 206 codes area information using 8-directional quantization codes. An 8-directional quantization code is a digit code indicating the direction of a previous point, as shown in Fig. 30, and is generally used to represent digital graphics.
A synthesizing section 205 outputs a decoded video frame of the lower layer coded by the lower layer coding section and to be synthesized. When a frame to be synthesized has not been encoded in the coding portion of the lower layer, the synthesizing portion 205 outputs a decoded video frame which is generated by using two decoded frames coded in the lower layer and located respectively located behind the missing frame of the lower layer, as well as a decoded frame of the upper layer for synthesis. The two frames of the lower layer are in front of and behind the frame of the upper layer. The synthesized video frame is input to the upper layer coding section 203 to be used for predictive coding. The image processing in the synthesizing section 203 is the following.
First, an interpolation image is created for two frames of the lower layer. A decoded image of the lower layer at time t is represented as B (x, y, t), where x and y are coordinates defining the position of a pixel in space. When the two decoded images of the lower layer are present at times t1 and t2 and the upper layer decoded image is at t3 (t1 <t3 <t2), the interpolation image I (x, y, t3) at time t3 becomes as follows Equation (1) calculates:
I (x, y, t3) = [(t2-t3) B (x, y, t1) + (t3-t1) B (x, y, t2)) / (t2-t1) (1)
The decoded image E of the upper layer is then synthesized with the obtained interpolation image I by using the synthesizing weighting information W (x, y, t) made from the area information. A synthesis image 5 is defined according to the following equation:
S (x, y, t) = [1-W (x, y, t)] I (x, y, t) + E (x, y, t) W (x, y, t) (2)
The area information M (x, y, t) is a digitized image taking the value 1 in a selected area and the value 0 outside it. The weighting information W (x, y, t) can be obtained by processing the above digitized image several times with a low-pass filter. That is to say that the weighting information W (x, y, t) has the value 1 within a selected range, the value 0 outside it, and a value between 0 and 1 at the limit thereof.
The code data created by the lower layer coding section, the upper layer coding section and the area information coding section are integrated by an integrating section (not shown) and then transmitted or accumulated.
On the decoding side of the conventional system, a code data separating section (not shown) separates code data into those of the lower layer, those of the upper layer, and area information code data. These code data are decoded by a lower-layer decoding section 208, an upper-layer decoding section 207 and an area information decoding section 209, respectively.
A synthesizing section 210 on the decoding side has a similar structure as the synthesizing section 205. It synthesizes an image using a decoded image of the lower layer and a decoded image of the upper layer according to the same method as described for the encoding side. The synthesized video frame is displayed on a display screen, and at the same time, it is input to the upper layer decoding section 207 to be used for prediction.
The above-described decoding apparatuses decode frames of both the lower and upper layers, but a decoder section of a lower layer decoding section is also adopted, omitting the upper layer coding section 204 and the synthesizing section 210. This simplified decoding apparatus can reproduce a part of code data.
This embodiment of the invention is intended to solve a problem such as may occur in the synthesizing section 205 shown in FIG. This embodiment also relates to a video signal synthesizing apparatus capable of synthesizing an image from two decoded lower-layer frames and a decoded selected upper-layer portion or regions without having a replica-like image around the selected region or regions Disorder occurs. Fig. 32 is a block diagram showing an image synthesizing apparatus which is an embodiment of the invention.
In Fig. 32, a first area extracting section 221 for extracting a region relating to a first but not a second region is composed of first region information of a lower layer frame and second region information of a lower layer frame. In the Fig. 33 (a), the first area information is represented by a dotted line (0 within the dotted area and 1 outside thereof), and the second area information is represented by a dotted line (with similar numerical codes). Accordingly, an area to be extracted by the first area removal portion 221 is a hatched portion shown in FIG.
A second area extraction section 222 in Fig. 32 is intended to extract a region relating to the second but not the first region from the first region information of a lower layer frame and second region information of a lower layer frame. That is, the dotted area shown in Fig. 33 (a) is taken out.
In Fig. 32, a controller 223 controls a switch 224 in accordance with an output of the first and second area extraction sections. That is, the switch 221 is connected to a second decoding image side when the position of a pixel to be observed only concerns the first region, and it is connected to a first decoding image side when the pixel to be observed only concerns the second region. The switch is connected to the output of an interpolation image generation section 225 when the pixel to be observed does not concern either the first or the second region.
The interpolation image generation section 225 calculates an interpolation image between the first and second decoded images of the lower layer according to the above-defined equation (1). In the equation (1), the first decoded image is represented as B (x, y, t1), the second decoded image is represented as B (x, y, t2), and the interpolation image is represented as B (x, y, t3). played. "t1", "t2" and "t3" are times for the first decoded picture, the second decoded picture and the interpolation picture.
Referring to Fig. 33 (a), the interpolation image thus generated is characterized in that the hatched area is filled with a background image outside the selected area of the second decoded frame, a dotted area having a background image outside the selected area of the first one filled in the decoded frame and other portions are filled with the interpolation image between the first and the second decoded frame. The decoded image of the upper layer is then superimposed on the above-mentioned portion 226 as shown in Fig. 32 to produce a synthesized image shown in Fig. 33 (b) which has no after-image around the selected (hatched) area and is free from the disorder occurring in the prior art image. The weighted averaging section 226 combines the interpolation image with the upper layer decoded image using weighting measures. The weighted averaging method has been described above.
In the embodiment described above, it is also possible to use, instead of the average weighting section 225, pixel values of either the first decoded image B (x, y, t1) or the second decoded image B (x, y, t2) corresponding to the time t3 of FIG Image of the upper layer is closer in time. In this case, the interpolation image I can be reproduced using the frame number as follows:
I (x, y, t3) = B (x, y, t1) in the case t3-t1 <t1-t2 or
I (x, y, t3) = B (x, y, t2) in all other cases.
In the expressions, t1, t2, and t3 indicate times of the first decoded picture, the second decoded picture, and the upper layer decoded picture.
Another embodiment of the invention will be described as follows.
This embodiment relates to an image synthesizing apparatus based on the first embodiment and capable of producing a more detailed synthesized image in consideration of motion information of decoded lower layer images. Fig. 34 is a block diagram showing an apparatus for predicting a motion parameter and modifying area information of two corresponding frames.
In Fig. 34, a motion parameter estimation section 231 estimates information for movement from a first decoded lower layer image and a second decoded lower layer image by determining motion parameters, e.g. As the motion vector per block and the overall image movement (parallel shift, rotation, enlargement and reduction).
An area shape modifying section 232 modifies the first decoded image, the second decoded image, the first area information and the second area information according to respective predicted motion parameters based on the temporal positions of the synthesizable frames. For example, a motion vector (MVx, MVy) is determined from the first decoded picture to the second decoded picture as a motion parameter. MVx is a horizontal component, and MVy is a vertical component. A motion vector from the first decoded picture to the interpolation picture is determined according to the equation (t3-t1 / (t2-t1) (MVx, MVy) The first decoded picture is then shifted in accordance with the obtained vector If other motion parameters such as rotation, enlargement and reduction are used, the image is not only shifted, but also deformed. 34 it concerns deformation (modification) data sets "a", "b", "c" and "d" concerning the first decoded image, the second decoded image, the first region information and the second region information in Fig. 32, respectively. These data sets are input to the image synthesizing apparatus shown in Fig. 32 which produces a synthesized image. Although the embodiment described above predicts the motion parameters from two decoded pictures, it may also use one motion vector of each block of each picture, as generally included in coded data produced by predictive coding. For example, an average value of decoded motion vectors as a motion vector of an entire image from the first to the second decoded frame may be applied. It is also possible to determine a frequency distribution of decoded motion vectors and to use the vector with the greatest frequency as the motion parameter of an entire picture from the first to the second decoded frame. The above processing is independently performed in the horizontal and vertical directions.
Another embodiment of the invention is the following.
This embodiment relates to an area information coding apparatus capable of effectively coding area information. Figs. 35 and 36 are block diagrams of this embodiment, the coding side of which is shown in Fig. 35 and the decoding side thereof in Fig. 36.
In Fig. 35, an area information approaching section 241 performs approximation of area information using a plurality of geometric figures. Fig. 37 shows an example of the approximation of area information of a human shape (hatched portion) by two rectangles. A rectangle 1 represents the head of a person, and the other rectangle 2 represents the chest area of the person.
A proximity information coding section 242 encodes the approximate area information. An area approximated by rectangles as shown in Fig. 37 may be coded with a fixed-length code by coding the coordinates of the left upper point of each rectangle and the size of each rectangle by a fixed-length code. An area approximated by an ellipse may be coded by a fixed length code by coding the coordinates of the center, the length of the long axis, and the length of the short axis. The approximate area information and the code data are supplied to a selecting section 244.
Like the area information coding section 206 described with reference to Fig. 28, an area information coding section 243 in Fig. 35 encodes area information using an 8-directional quantization code without approximation. The area information and the code data are supplied to a selecting section 244.
The selecting section 244 selects one of the two output signals 242 and 243. When the output signal 243 is selected, the code data of the approximate area information having single bit (e.g., 1) selection information is supplied to a code data integrating section (not shown), and approximate area information is supplied to a synthesizing section (not shown). When the output signal 344 is selected, the code data of the unequal area information having one bit (e.g., 1) of selection information is supplied to a code data integrating section (not shown), and the non-approximate area information is supplied to a synthesizing section, in accordance with FIG the invention takes place.
The selection section may be, for. For example, it may operate to select the one output that produces the smaller amount of code data, or select the output 244 if the code data of the non-approximate information does not exceed a threshold, but selects the output 242 if that amount exceeds the threshold. This makes it possible to reduce the code data amount, thereby preventing area information from being destroyed.
The operation on the decoding side in this embodiment is the following.
In Fig. 36, a selection section 251 selects which kind of area information to use, approximated or not, based on the one-bit selection information in the received code data.
In Fig. 36, a proximity information decoding section 252 decodes the approximate area information, whereas an area information decoding section 253 decodes the non-approximate area information. A switch 254 is controlled by a signal from the selecting section 251 to select approximate or non-approximate area information as an output for a synthesizing section.
Thus, either approximate or non-approximate range information is adaptively selected, encoded and decoded. When area information is complicated and can generate a large code data amount, the approximate area information is selected to encode the area information with a small amount of information.
In the above case, the non-approximate range information is coded using 8-directional quantization codes, but it can be coded more effectively using a combination of 8-directional quantization with predictive coding. A code for 8-directional quantization takes eight values from 0 to 7, as shown in Fig. 30, which are distinguished by predictive coding of -7 to 7. However, the difference may be limited to the range of -3 to 4 by adding 8 if the difference is -4 or less and subtracting 8 if the difference is greater than 4. In decoding, an original value of 8-directional coding can be obtained by first adding the difference to the previous value and then subtracting or adding 8 if the result is a negative value or greater than 7. An example is given below:
Value from 8-directional quantization 1, 6, 2, 1, 3 ...
Difference 5, -4, -1, -2 ...
converted value -3, 4, -1, 2 ...
decoded value 1, 6, 2, 1, 3 ...
For example, the difference between the quantization value 6 and the previous value is 5, of which 8 is subtracted to obtain the result -3. When decoding, -3 is added to the previous value 1, and the value -2 is obtained which is negative, therefore it is increased by adding 8 to finally obtain the decoded value 6. Such predictive coding is done using the cyclic feature of 8-directional coding.
Although in this embodiment approximate area information of each picture is coded independently, it is possible to increase the coding efficiency by using the result of predictive coding, since video frames generally have high inter-frame correlation. That is, only the difference of approximate area information of two consecutive frames is encoded when the approximate area information is continuously encoded between two frames. If z. B. When an area is approximated by a rectangle, a rectangle of a previous frame is represented by its upper left position (19, 20) and its size (100, 150), and the rectangle of the current frame is represented by its upper left position (13, 18 ) and its size (100, 152) are reproduced, and the upper left position difference (3, 2) and the size difference (0.2) for the current frame are coded. When the change of the area shape is small, code data amount for the area information can be considerably saved by using entropy coding, e.g. Huffman coding, since differences in a small change in area shape are close to zero. When a rectangle does not change frequently, it is effective to code single-bit information as information for changing a rectangle regarding a current frame. D. that is, single-bit information (e.g., 0) is encoded for a current frame whose rectangle does not change, whereas single-bit information (e.g., 1) and difference information are encoded for frames whose rectangle varies.
Hereinafter, another embodiment of the invention is set forth.
This embodiment relates to weighting information generating apparatus for generating weighting information having many values of area information. Fig. 38 is a block diagram of this embodiment.
In Fig. 38, a horizontal weighting generating section 261 scans area information horizontally, and detects the value of 1 therein, and then calculates a corresponding weighting function. In practice, first, the abscissa x0 of the left-side point and the horizontal length N of the region are determined, and then a horizontal weighting function is calculated, as shown in Fig. 39 (a). The weighting function can be created by combining straight lines or combining a line with a trigonometric function. An example of the latter case will be described below. If N> W (W is the width of a trigonometric function), the following weighting functions can be used:
- sin [(x + 1/2) π / (2w)] xsin [(x + 1/2) π / (2W)] when 0 ≤ x <W;
- 1 if W ≤ x <N - W;
- sin [(x-N + 2W + 1/2) π / (2W)] xsin [(x-N + 2W + 1/2) π / (2W)]
if 0 ≤ x <W;
- sin2 [x + 1/2) π / N] xsin [(x + 1/2) π / N] when N ≤ 2W.
In the above case, the point x0 at the left end of the range is set to 0.
In Fig. 38, a vertical weighting generation section 502 vertically scans the area information, and detects the value 1 in it, and then calculates a corresponding vertical weighting function. In practice, the ordinate y0 of the upper-end point and the vertical-length M of the region are determined, and then a vertical weighting function is calculated as shown in Fig. 39 (b).
A multiplier 263 multiplies an output 261 by an output 262 for each pixel position to produce weighting information.
By the above method, weighting information with a reduced number of operations adapted to the shape of the area information can be obtained.
Now, another embodiment of the invention is set forth.
This embodiment relates to a method of adaptively switching the coding mode from interframe prediction to intraframe prediction, and vice versa, in predictive coding of lower and upper layer frames. Fig. 40 is a block diagram of this embodiment.
In Fig. 40, a mean value calculating section 271 determines the average of pixel values in an area corresponding to an input original image and input area information. The mean value is input to a differentiator 273 and a memory 272.
The differentiator 273 determines the difference between a previous average stored in the memory 272 and a current average output from the average calculating section 271.
A discrimination section 274 compares the absolute value of the difference calculated by the differentiator 273 with a predetermined threshold and outputs mode selection information. If the absolute value of the difference is greater than the threshold value, the discrimination section 273 judges that a scene change has occurred in a selected area, and generates a mode selection signal to always execute intraframe prediction encoding.
The mode selection thus executed by judging a scene change in a selected area is effective for obtaining high-quality coded pictures even if e.g. B. a person from behind a cover emerges or anything is turned. The stated embodiment may be applied to a system for coding a selected area apart from other areas when coding lower-layer frames. In this case, area information is input to the lower layer coding section. This embodiment can also be applied to coding only a selected portion of the frame of the upper layer.
Another embodiment of the invention is the following.
This embodiment relates to a method of controlling the amount of data in the case of coding a separate area separately from other areas of each lower layer frame. Fig. 41 is a block diagram of this embodiment.
In Fig. 41, a coding section 283 separates and codes a selected area from other areas. An area discriminating section 281 receives area information and detects whether the encodable area is inside or outside the selected area. A code data amount estimating section 285 estimates the code data amount in each area based on the above-mentioned discrimination result. A distribution ratio calculating section 282 calculates distribution ratios of a target amount of codes per frame assigned to areas. The method of determining distribution ratios will be described later. A quantization width calculating section determines a quantization step size corresponding to the target amount of code data. The method of determining the quantization step size is the same as the conventional method.
The method of determining a code distribution ratio by the target code assignment calculating section is the following.
A target code amount Bi for a frame is calculated according to the following equation:
Bi = (number of usable bits - number of bits used to encode previous frames) / number of remaining frames
This target number Bi for bits is distributed at a specified ratio to pixels within a selected area and pixels outside thereof. The ratio is determined using an appropriate fixed ratio RO and a complexity ratio Rp for the previous frame. The complexity ratio Rp for the previous frame is calculated by the following equation:
Rp = (gen_bitF * avg_qF) / (gen_bitF * avg_qF + gen_bits * avg_gE)
where gen_bitF = number of bits to encode pixels in a selected area of a previous frame; gen_bitB = number of bits to encode pixels outside the selected area of a previous frame; avg_qF = average quantization step size in the selected area of a previous frame; and avg_qB = average quantization step size outside the selected area of a previous frame. In order to encode a selected region of high image quality, it is desirable to set the quantization step size so that the average quantization step size is kept slightly smaller in the selected region than outside it, and simultaneously followed by image change in a moving image sequence. In general, a fixed-ratio distribution RO is used to maintain a substantially constant relationship for the quantization step size between pixels in and outside of the selected area, while a distribution having the complexity ratio Rp for a previous frame is used to change an image in one Sequence of moving pictures to follow. Accordingly, in the invention, it is intended to use a combination of the advantages of both methods by making a target bit rate distribution ratio an average of the fixed ratio RO and the complexity ratio Rp for the previous frame. That is, the distribution ratio Ra is determined as follows: Ra = (RO + Rp) / 2
In Fig. 42, there are two exemplary curves plotted with dashed lines representing the fixed ratio RO and the complexity ratio Rp for the previous frame in a selected area for an entire video sequence. In this example, a solid curve in FIG. 42 the attainable ratio Ra for distributing a target code data amount that does not deviate too much from the fixed-ratio curve and reflects, to some extent, the change of an image in a video sequence. With a fixed ratio (1-RO) and a complexity ratio (1-Rp) for the previous frame for the outside of the selected range, an average ratio that is the target bit-amount distribution ratio (1-Ra) for pixels outside the selected range corresponds to the course of the bold dashed line in FIG. The total value of the two target bit-amount distribution ratios for pixels within and outside a selected range is 1.
Thus, the quantization step size can be adjusted adaptively. However, the bit rate of an entire video sequence may sometimes exceed a predetermined value because the number of bits used exceeds the target value Bi in some frames. In this case, the following procedure can be used.
As described above, the target bit-amount distribution ratio Ra for coding pixels in a selected area is an average of the fixed ratio RO and the complexity ratio Rp for the previous frame, whereas the target bit-amount distribution ratio Rm for coding pixels outside the selected range is the minimum value Rm of the fixed ratio (1-RO) and the complexity ratio (1-Rp) for the previous frame for coding pixels outside the selected range. In this case, the target bit-amount distribution ratio (1-Ra) for coding pixels outside the selected range may vary as exemplified by the solid line in FIG. 44 is shown. Since Ra + Rm ≤ 1, the bit count may be reduced for a frame or frames in which excessive bits occur. In other words, the bit rate of an entire video sequence can be kept within the predetermined limit by reducing the target bit amount of a background area of a frame or frames.
With the video coding and video decoding apparatuses according to the invention, it is possible to code a selected area of an image to be of higher quality than the other areas.
It is possible to decode only a selected area of lower image quality if only lower-layer code data is decoded.
When decoding code data of an upper layer, it is possible to select whether the first or the second upper layer is decoded. An entire picture is decoded with lower picture quality when the first layer is selected, whereas only a selected area with high picture quality is decoded when the second upper layer is selected.
When decoding all code data, an image may be decoded in such a way that a selected area of the image has higher image quality than all other areas thereof.
Although in the above-described preferred embodiments of the invention, it is assumed that the decoding apparatus receives all the code data, it may also be arranged so that, in a video communication system, a decoding terminal requests the coding page to transmit a limited amount of data, e.g. B. Code data for position and shape of a region, lower layer code data, and first layer code data for communication over a low bandwidth transmission line. D. that is, the present invention realizes data communication in which only lower layer data is transmitted over a very small bandwidth transmission line, or any one of two types of upper layer data is selectively transmitted over a slightly wider bandwidth line or all Types of data can be transmitted over a line of even greater bandwidth.
With the video coding apparatus according to the present invention, it is possible to lower the code data amount because weighted averaged information is generated from digitized information by inputting a plurality of partial video sequences on a background video sequence and using weighted averages. Since the weighted average data produced from the digitized information takes a value of 0 to 1, the boundary between the sub-images and the background images can be uniformly synthesized without any detectable disturbance.
In weighting all uncoded data using weighting values to be used for synthesizing video sequences, the code data amount may be lowered, or the quality of the decoded image may be improved over known devices for the same code data size.
The video coding device according to the invention is intended to carry out the following:
(1) synthesizing an unencoded lower-layer frame from a previous and subsequent lower-layer frame by weighted averaging of two lower-layer frames existing before and after the synthesizable frame, for an overlapping portion of a first portion having a second portion or an area not belonging to the first subarea and the second subarea, using a frame of the lower layer, which is behind the synthesizable frame in time, for a part of only the first portion, and using a frame of the lower layer which is earlier in time than the synthesizable frame, for a part of only the second portion, thereby forming a high-quality synthesized image even then without interference when an object is moving;
(2) synthesizing the lower-layer frame (1) using a lower-layer frame located near the synthesizable frame in time for an overlapping portion of a first portion having a second portion or a portion not belonging to the first portion and the second portion or using only a first frame of the lower layer or only a second frame of the lower layer, thereby obtaining a synthesized image of high quality without double display of the synthesized background image even when the background image is moving;
(3) synthesizing the lower-layer frame (1) by modifying (deforming) the first lower-layer frame, the second lower-layer frame, the first partial region, and the second partial region by motion-compensating motion parameters based on the temporal position of the synthesizable frame the lower layer to thereby obtain a synthesized image of high quality wherein the movement of a background image is followed in the frame of the lower layer;
(4) synthesizing the lower-layer frame (3) using motion vector information obtained by motion-compensating predictive coding to thereby obtain a motion parameter with reduced processing overhead than in the case of re-prediction of a motion parameter;
(5) adaptively selecting either approximation of the area information by a plurality of geometric figures or coding without approximation to thereby effectively encode and decode area information;
(6) converting area information (5) into 8-directional quantization data, determining the difference between the 8-directionally quantized data, and encoding and decoding the difference data by variable-length coding to thereby more efficiently perform reversible coding and decoding of area information;
(7) further efficiently coding and decoding approximate area information (5) by determining the inter-frame difference of information from geometric figures, encoding and decoding by a variable-length coding method, and adding information indicating no change of area information without coding other area information when the difference data is all 0;
(8) horizontally scanning area information to detect the length of each line in it and determine a horizontal weighting function; vertically scanning the area information to detect the length of a line therein and determining a vertical weighting function; Generating multi-valued weighting information to thereby efficiently generate weighting information by a weighting information producing device when a top-layer sub-image is synthesized by a weighted averaging method having a bottom-layer frame;
(9) Coding and decoding video frames using area information indicating the shape of an object or the shape of a part, determining an average of pixels in an area from the input image and, accordingly, area information, calculating the difference between average values of a previous frame and the current frame, comparing the difference with a specified value and selecting the intraframe coding, if the difference exceeds the specified value to thereby make it possible to correctly switch the coding mode from predictive (interframe) coding to the intraframe coding when a scene change occurs, and to ensure the coding and decoding of high quality images;
(10) separating a video sequence into background image areas and a plurality of foreground fields, and separately coding each separate background area and each field area by determining whether code data and encodable blocks exist within or outside a partial area, which is done by separately calculating the code data amount in the sub-image area and the code data amount in the background image area and determining target bit-amount distribution ratios for the sub-image area and the background image area to thereby provide correct distribution of the target bit number to obtain high quality coded images.
Contents4
42 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42
35 members in 4 offices
Priority claims12
| Document | Office | Kind | Date |
|---|---|---|---|
| 17864295 | Japan | A | |
| 17864295 | Japan | – | |
| 17864395 | Japan | A | |
| 17864395 | Japan | – | |
| 27550195 | Japan | A | |
| 27550195 | Japan | – | |
| 17864295 | – | – | – |
| 17864395 | – | – | – |
| 27550195 | – | – | – |
| JP19950178642 | – | – | – |
| JP19950178643 | – | – | – |
| JP19950275501 | – | – | – |
Members35
| Document | Office | Kind | |
|---|---|---|---|
| EP0753970A2 | European Patent Office (EPO) | A2 | |
| JPH0937240A | Japan | A | |
| JPH0937260A | Japan | A | |
| JPH09121346A | Japan | A | |
| EP0753970A3 | European Patent Office (EPO) | A3 | |
| US5963257A | United States of America | A | |
| US5986708A | United States of America | A | |
| EP0961496A2 | European Patent Office (EPO) | A2 | |
| EP0961497A2 | European Patent Office (EPO) | A2 | |
| EP0961498A2 | European Patent Office (EPO) | A2 | |
| US6023299A | United States of America | A | |
| US6023301A | United States of America | A | |
| US6084914A | United States of America | A | |
| US6088061A | United States of America | A | |
| JP3098939B2 | Japan | B2 | |
| JP3101195B2 | Japan | B2 | |
| EP0961496A3 | European Patent Office (EPO) | A3 | |
| EP0961497A3 | European Patent Office (EPO) | A3 | |
| EP0961498A3 | European Patent Office (EPO) | A3 | |
| JP2001036913A | Japan | A | |
| JP2001036914A | Japan | A | |
| EP0753970B1 | European Patent Office (EPO) | B1 | |
| DE69615948D1 | Germany | D1 | |
| DE69615948T2This record | Germany | T2 | |
| EP0961497B1 | European Patent Office (EPO) | B1 | |
| DE69626142D1 | Germany | D1 | |
| EP0961496B1 | European Patent Office (EPO) | B1 | |
| DE69628467D1 | Germany | D1 | |
| DE69626142T2 | Germany | T2 | |
| JP3494617B2 | Japan | B2 | |
| DE69628467T2 | Germany | T2 | |
| JP3526258B2 | Japan | B2 | |
| EP0961498B1 | European Patent Office (EPO) | B1 | |
| DE69634423D1 | Germany | D1 | |
| DE69634423T2 | Germany | T2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| No opposition during term of oppositionOpposition8364 | 8364 |
Numbers
- Publication
- 69615948
- Publication, DOCDB
- 69615948
- Publication, EPODOC
- DE69615948T
- Application
- 69615948
- Application, DOCDB
- 69615948
- Application, EPODOC
- DE1996615948T
Titles2
- German
- Hierarchischer Bildkodierer und -dekodierer
- English
- Hierarchical image coder and decoder
Classification
- CPC, 8
- H04N19/577
- H04N19/107
- H04N19/17
- H04N19/20
- H04N19/30
- H04N19/37
- H04N19/53
- H04N19/61
- IPC, 4
- G06T9 00
- H04N7 26
- H04N7 46
- H04N7 50
