Encoding device and encoding method
14 claims: 2 independent, 12 dependent
- 1CLAIMS REIVINDICAÇÕES 1. Coding apparatus, comprising:1. Aparelho de codificação, que compreende: a base layer encoding section that encodes an input signal to acquire base layer encoded data;uma seção de codificação de camada de base que codifica um sinal de entrada para adquirir os dados codificados de camada de base;a base layer decoding section that decodes the base layer encoded data to acquire a base layer decoded signal;and an enhancement layer encoding section that encodes a residual signal that represents a difference between the input signal and the base layer decoded signal, to acquire the enhancement layer encoded data, wherein the layer encoding section improvement program comprises: uma seção de decodificação de camada de base que decodifica os dados codificados de camada de base para adquirir um sinal decodificado de camada de base;e uma seção de codificação de camada de melhoramento que codifica um sinal residual que representa uma diferença entre o sinal de entrada e o sinal decodificado de camada de base, para adquirir os dados codificados de camada de melhoramento, em que a seção de codificação de camada de melhoramento compreende: a division section that divides the residual signal into a plurality of sub-bands;uma seção de divisão que divide o sinal residual em uma pluralidade de sub-bandas;a first shape vector encoding section that encodes the plurality of sub-bands to acquire the first shape-encoded information, and which calculates the target gains from the plurality of sub-bands;uma primeira seção de codificação de vetor de forma que codifica a pluralidade de sub-bandas para adquirir as primeiras informações codificadas de forma, e que calcula os ganhos-alvo da pluralidade de subbandas;a gain vector formation section that forms a gain vector using the plurality of target gains;and a gain vector encoding section that encodes the gain vector to acquire the first encoded gain information. uma seção de formação de vetor de ganho que forma um vetor de ganho utilizando a pluralidade de ganhos-alvo;e uma seção de codificação de vetor de ganho que codifica o vetor de ganho para adquirir as primeiras informações codificadas de ganho.
- 14Coding method comprising:14. Método de codificação que compreende: dividing the transform coefficients acquired by transforming an input signal into a frequency domain, into a plurality of sub-bands;dividir os coeficientes de transformada adquiridos pela transformação de um sinal de entrada em um domínio de frequência, em uma pluralidade de sub-bandas;codificar os coeficientes de transformada da pluralidade de subbandas para adquirir as primeiras informações codificadas de forma e calcular os ganhos-alvo dos coeficientes de transformada da pluralidade de subbandas;encode the transform coefficients of the plurality of sub-bands to acquire the first shape-coded information and calculate the target gains of the transform coefficients of the plurality of sub-bands;form a gain vector using the plurality of target earnings;and encode the gain vector to acquire the first encoded gain information. formar um vetor de ganho utilizando a pluralidade de ganhosalvo;e codificar o vetor de ganho para adquirir as primeiras informações codificadas de ganho. 1/33 1/33 COEFFICIENTS OF ω COEFICIENTES DE ω Η co Η co
Independent claims2
262 paragraphs in 7 sections, as filed
(54) Title: CODING DEVICE AND (57) Summary:
CODING METHOD (30) Unionist Priority: 02/03/2007 jp 2007-053502,
05/18/2007 JP 2007-133545, 07/13/2007 JP 2007-185077, 02/26/2008
JP 2008-045259, 07/13/2007 JP 2007-185077, 05/18/2007 JP 2007133545, 02/26/2008 JP 2008-045259 (73) Owner (s): Panasonic Corporation (72) Inventor (s): Masahiro Oshikiri, Tomofumi Yamanashi, Toshiyuki Morii (74) Attorney (s): Dannemann, Siemsen, Bigler & Ipanema Moreira (86) International Application: pct JP2008000408 of 29/02/2008 (87) International Publication: wo 2008 / i20440de 09 / 10/2008
<img file="BRPI0808428A2_D0001.tif" />
Descriptive Report of the Invention Patent for CODING DEVICE AND CODING METHOD.
TECHNICAL FIELD
The present invention relates to an encoding apparatus and an encoding method used in a communication system that encodes and transmits input signals such as speech signals. TECHNICAL FUNDAMENTALS
It is demanded in a mobile communication system that the voice signals are compressed to low bit rates to transmit in order to use radio wave resources efficiently, and so on. On the other hand, it is also demanded that a quality improvement in telephone call voice and a high fidelity call service can be performed, and, to meet these demands, it is preferable not only to provide quality voice signals but also to encode other quality signals than voice signals, such as wider band quality audio signals.
The technique of integrating a plurality of layered coding techniques is promising for these two contradictory demands. This technique combines in layers the base layer to encode the input signals in a form suitable for speech signals at low bit rates and an enhancement layer to encode the differential signals between the input signals and the decoded signals of the layer base in a form suitable for signals other than the voice. The technique of performing a layered encoding in this way has the characteristics of providing scalability in bit streams acquired from a coding device, that is, acquiring the decoded signals from part of bit stream information, and therefore is generally referred to as scalable encoding (layered encoding).
The scalable encoding scheme can flexibly support communication between networks at variable bit rates thanks to its characteristics, and, consequently, is suitable for a future network environment where several networks will be integrated by IP (Internet Protocol).
For example, Non-Patent Document 1 describes a technique for performing scalable coding using the technique that is standardized by MPEG-4 (Moving Picture Experts Group phase-4). This technique uses CELP (Linear Excited Prediction in Code) encoding suitable for voice signals, in the base layer, and uses transform encoding such as AAC (Advanced Audio Encoder) and TwinVQ (Weighted Interleaving Vector Quantization of Transformed Domain) with respect to residual signals that subtract the base layer decoded signal from the original signal, in the enhancement layer.
Furthermore, to flexibly support a network environment in which the transmission speed fluctuates dramatically due to the transfer between different types of networks and the occurrence of congestion, scalable encoding of small bit rate scales needs to be performed and, consequently, needs be configured by providing multiple layers of lower bit rates.
Patent Document 1 and Patent Document 2 describe a transform encoding technique for transforming a signal which is the target to be encoded in the frequency domain and encoding the resulting frequency domain signal. In such a transform encoding, first, an energy component of a frequency domain signal, that is, the gain (ie, the scale factor) is calculated and quantized on a per-band basis, and a thin component of the frequency domain signal above, that is, a shape vector, is calculated and quantized.
Non-Patent Document 1: All about MPEG-4, written and edited by Sukeichi MIKI, the first edition, Kogyo Chosakai Publishing, Inc., September 30, 1998, pages 126 to 127.
Patent Document 1: Japanese Translation of PCT Application Open to Public Inspection Number 2006-513457.
Patent Document 2: Japanese Patent Application Open to Public Inspection Number HEI7-261800 DESCRIPTION OF THE INVENTION
PROBLEMS TO BE SOLVED BY THE INVENTION
However, when two successive parameters are quantized in order, the parameter that is quantized afterwards is influenced by the quantization distortion of the parameter that is quantized before, and therefore is inclined to show an increased quantization distortion. Therefore, there is a trend that generates, in the transform coding described in Patent Document 1 and Patent Document 2, to quantize a gain and a shape vector in order, the shape vectors show an increased quantization distortion and are incapable to represent the precise spectral form. This problem produces a significant quality deterioration with respect to strong tonal signals such as vowels, that is, signals that have spectral characteristics in which multiple peak forms are observed. This problem becomes more distinct when a lower bit rate is implemented.
It is therefore an object of the present invention to provide a coding apparatus and a coding method to precisely encode spectral forms of strong tonal signals such as vowels, that is, spectral forms of signals that have spectral characteristics in which multiple forms of peak are observed, and to improve the quality of decoded signals such as the sound quality of decoded signals.
MEANS TO SOLVE THE PROBLEM
The coding apparatus according to the present invention employs a configuration which includes: a base layer coding section that encodes an input signal to acquire the coded base layer data; a base layer decoding section that decodes the base layer encoded data to acquire a base layer decoded signal; and an enhancement layer encoding section that encodes a residual signal that represents a difference between the input signal and the base layer decoded signal, to acquire the enhancement layer encoded data, and in which the encoding section of enhancement layer has: a division section that divides the residual signal into a plurality of sub-bands; a first shape vector encoding section that encodes the plurality of sub-bands to acquire the first shape-encoded information, and which calculates the target gains from the plurality of sub-bands; a gain vector formation section that forms a gain vector using the plurality of target gains; and a gain vector encoding section that encodes the gain vector to acquire the first encoded gain information.
The encoding method according to the present invention includes: dividing the transform coefficients acquired by transforming an input signal into a frequency domain, into a plurality of sub-bands; encode the transform coefficients of the plurality of sub-bands to acquire the first shape-coded information and calculate the target gains of the transform coefficients of the plurality of sub-bands; form a gain vector using the plurality of target earnings; and encode the gain vector to acquire the first encoded gain information.
ADVANTAGE EFFECTS OF THE INVENTION
The present invention can more accurately encode spectral forms of strong tonal signals such as vowels, that is, spectral forms of signals that have spectral characteristics in which multiple peak forms are observed, and improve the quality of decoded signals such as the sound quality of decoded signals.
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1 is a block diagram showing the main configuration of a voice coding apparatus according to Mode 1 of the present invention;
Figure 2 is a block diagram showing the configuration within a coding section of the second layer according to Mode 1 of the present invention;
Figure 3 is a flowchart showing the steps for processing the encoding of the second layer in the encoding section of the second layer according to Mode 1 of the present invention;
Figure 4 is a block diagram showing the configuration within a shape vector coding section according to Mode 1 of the present invention;
Figure 5 is a block diagram showing the configuration within a gain vector coding section according to Mode 1 of the present invention;
Figure 6 illustrates in detail the operation of the target gain arrangement section according to Mode 1 of the present invention;
Figure 7 is a block diagram showing the configuration within a gain vector encoding section according to Mode 1 of the present invention;
Figure 8 is a block diagram showing the main configuration of a voice decoding apparatus according to Mode 1 of the present invention;
Figure 9 is a block diagram showing the configuration within a decoding section of the second layer according to Mode 1 of the present invention;
Figure 10 illustrates a shape vector codebook according to Mode 2 of the present invention;
Figure 11 illustrates the multiple shape vector candidates included in the shape vector code book according to Mode 2 of the present invention;
Figure 12 is a block diagram showing the configuration within the coding section of the second layer according to Mode 3 of the present invention;
Figure 13 illustrates a range selection processing in a range selection section in accordance with Mode 3 of the present invention;
Figure 14 is a block diagram showing the configuration within the decoding section of the second layer according to Modality 3 of the present invention;
Figure 15 shows a variation of the range selection section according to Modality 3 of the present invention;
Figure 16 shows a variation of a range selection method in the range selection section in accordance with Modality 3 of the present invention;
Figure 17 is a block diagram showing a variation of the configuration of the range selection section according to Modality 3 of the present invention;
Figure 18 illustrates how the lane information is formed in the lane information formation section in accordance with Mode 3 of the present invention;
Figure 19 illustrates the operation of a variation of an error transform coefficient generation section of the first layer according to Modality 3 of the present invention;
Figure 20 shows a variation of the range selection method in the range selection section according to Modality 3 of the present invention;
Figure 21 shows a variation of the range selection method in the range selection section according to Modality 3 of the present invention;
Figure 22 is a block diagram showing the configuration within the coding section of the second layer according to Mode 4 of the present invention;
Figure 23 is a block diagram showing the main configuration of the voice coding apparatus according to Mode 5 of the present invention;
Figure 24 is a block diagram showing the main configuration within the coding section of the first layer according to Mode 5 of the present invention;
Figure 25 is a block diagram showing the main configuration within the decoding section of the first layer according to Mode 5 of the present invention;
Figure 26 is a block diagram showing the main configuration of the voice decoding apparatus according to the Modality of the present invention;
Figure 27 is a block diagram showing the main configuration of the voice coding apparatus according to Mode 6 of the present invention;
Figure 28 is a block diagram showing the main configuration of the voice decoding apparatus according to the Modality of the present invention;
Figure 29 is a block diagram showing the main configuration of the voice coding apparatus according to Mode 7 of the present invention;
Figure 30 illustrates the processing of selecting the range which is the target to be encoded in the encoding processing in the speech coding apparatus according to Mode 7 of the present invention;
Figure 31 is a block diagram showing the main configuration of the voice decoding apparatus according to the Modality of the present invention;
Figure 32 illustrates a case where the target to be encoded is selected from band candidates arranged at equal intervals, in the coding processing in the speech coding apparatus in accordance with Modality 7 of the present invention; and
Figure 33 illustrates a case where the target to be encoded is selected from candidates of the band arranged at equal intervals, in the coding processing in the speech coding apparatus according to Modality 7 of the present invention.
BEST MODE FOR CARRYING OUT THE INVENTION
Hereinafter, the modalities of the present invention will be explained in detail with reference to the accompanying drawings. A speech coding apparatus / speech decoding apparatus will be used below as an example of a coding apparatus / decoding apparatus according to the present invention for explanation.
(Mode 1)
Figure 1 is a block diagram showing the main configuration of a voice coding apparatus 100 according to Mode 1 of the present invention. An example will be explained where the voice coding apparatus and the voice decoding apparatus according to the present modality employ a scalable two-layer configuration. In addition, the first layer constitutes the base layer and the second layer constitutes the enhancement layer.
In figure 1, the voice coding apparatus 100 has a frequency domain transformation section 101, a first layer coding section 102, a first layer decoding section 103, a subtractor 104, a second coding section layer 105 and a multiplexing section 106.
The frequency domain transformation section 101 transforms a time domain input signal into a frequency domain signal, and outputs the resulting input transform coefficients to the first layer coding section 102 and subtractor 104.
The encoding section of the first layer 102 performs an encoding processing with respect to the input transform coefficients received from the frequency domain transformation section 101, and outputs the resulting encoded data from the first layer to the decoding section of the first layer 103 and the multiplexing section 106.
The decoding section of the first layer 103 performs decoding processing using the encoded data from the first layer received from the encoding section of the first layer 102, and outputs the resulting decoded transform coefficients of the first layer to subtractor 104.
Subtractor 104 subtracts the decoded transform coefficients of the first layer received from the decoding section of the first layer 103 from the input transform coefficients received from the frequency domain transformation section 101, and outputs the error transform coefficients of the first layer resulting to the encoding section of the second layer 105. The encoding section of the second layer 105 performs encoding processing with respect to the error transform coefficients of the first layer received from subtractor 104, and outputs the resulting encoded data from the second layer to the multiplexing section 106. In addition, the Second layer 105 encoding will be described in detail later.
The multiplexing section 106 multiplexes the encoded data of the first layer received from the encoding section of the first layer 102 and the encoded data of the second layer received from the encoding section of the second layer 105, and outputs the resulting bit stream to a transmission channel. .
Figure 2 is a block diagram showing the configuration within the coding section of the second layer 105.
In Figure 2, the second layer 105 coding section has a subband forming section 151, a shape vector coding section 152, a gain vector forming section 153, a vector coding section of gain 154 and a multiplexing section 155.
Subband forming section 151 divides the first layer error transform coefficients received from subtractor 104, into M subbands, and outputs the resulting M subband transform coefficients to the vector coding section of form 152. Here, when the error transform coefficients of the first layer are represented as ei (k), om— subband transform coefficients e (m, k) (where 0 <m <M-1) are represented by equation 1 below.
e (m, k) = e<sub>}</sub>(k + F (mf) (0 <k <F (m +1) - F ^ rri))
Equation 1
In equation 1, F (m) represents the frequency at the limit of each subband, and the ratio of 0 <F (0) <F (1) <... <F (M) <FH is true. Here, FH represents the highest frequency of the error transform coefficients of the first layer, and assumes an integer of 0 <m <M-1.
The shape vector coding section 152 performs shape vector quantization with respect to the M subband transformation coefficients sequentially received from the subband forming section 151, to generate the shape coded information from the M subbands and calculates the target gains of the M subband transform coefficients. The shape vector coding section 152 outputs the shape coded information generated for the multiplexing section 155, and issues the active gains for the gain vector forming section 153. Also, the shape vector coding section 152 will be described in detail later.
The gain vector formation section 153 forms a gain vector with the M target gains received from the shape vector encoding section 152, and outputs this gain vector to the gain vector encoding section 154. Also, the gain vector formation section 153 will be described in detail later.
The gain vector encoding section 154 performs vector quantization using the gain vector received from the gain vector forming section 153 as a target value, and outputs the resulting encoded gain information to the multiplexing section 155. In addition, the gain vector encoding section 154 will be described in detail later.
The multiplexing section 155 multiplexes the shape coded information received from the shape vector coding section 152 and the gain coded information received from the gain vector coding section 154, and outputs the resulting bit stream as the coded data of the second layer for multiplexing section 106.
Figure 3 shows a flowchart showing the steps of the second layer coding processing in the second layer coding section 105.
First, in step (hereinafter, abbreviated as ST) 1010, subband forming section 151 divides the error transform coefficients of the first layer into M subbands to form M subband transform coefficients.
Then, on ST 1020, the encoding section of the second bed11 of 105 initializes a subband counter m that counts the subbands to 0.
Next, in ST 1030, the shape vector coding section
152 performs shape vector encoding with respect to the m— subband transform coefficients to generate the m— subband encoded information and generate the target gain of the m— subband transform coefficients.
Then in ST 1040, the encoding section of the second layer 105 increments the subband counter m by one.
Then, in ST 1050, the encoding section of the second layer 105 decides whether m <M is true or not.
In ST 1050, when deciding that m <M is true (ST 1050; YES), the encoding section of the second layer 105 returns the processing step to ST 1030.
In contrast to this, in ST 1050, when deciding that m <M is not true (ST 1050: NO), the gain vector formation section
153 forms a gain vector using M target gains in ST 1060.
Next, in ST 1070, the gain vector coding section
154 performs vector quantization using the gain vector formed in the gain vector formation section 153 as a target value to generate the encoded gain information.
Next, in ST 1080, the multiplexing section 155 multiplexes the shape coded information generated in the shape vector encoding section 152 and the gain coded information generated in the gain vector encoding section 154.
Figure 4 is a block diagram showing the configuration within the shape vector coding section 152.
In figure 4, the shape vector coding section 152 has a shape vector codebook 521, a cross correlation calculation section 522, an autocorrelation calculation section 523, a search section 524 and a target gain calculation 525.
The shape vector codebook 521 stores a plurality of shape vector candidates that represent the shape of the error transform coefficients of the first layer, and outputs the vector candidates sequentially to the cross-correlation calculation section 522 and the autocorrelation calculation section 523 based on a control signal received from research section 524. Yet, generally, there are cases where a shape vector code book adopts the way to actually secure a storage space and store shape vector candidates, and there are cases where a shape vector code book forms candidates shape vector according to predetermined processing steps. In the latter cases, it is not really necessary to secure storage space. Although any of the shape vector code books can be used in the present mode, the present mode will be explained below assuming that the shape vector code book 521 that stores the shape vector candidates shown in figure 4 is provided . From here on, hi<sup>and</sup> shape vector candidate in the plurality of shape vector candidates stored in the shape vector code book 521, is represented as c (i, k). Here, k represents ok<sup>s</sup> element of a plurality of elements that form a shape vector candidate.
The cross-correlation calculation section 522 calculates the cross-correlation ccor (i) between m— subband transform coefficients received from subband forming section 151 and i<sup>s</sup> shape vector candidate received from shape vector code book 521, according to equation 2 below, and issues the ccor (i) cross-correlation for research section 524 and target gain calculation section 525.
F (m + 1) -F (m) —1 ccor (i) = F, e (m, k) -c (i, k) Equation 2 jt = O
The autocorrelation calculation section 523 calculates the autocorrelation according to (i) the vector candidate of form c (i, k) received from the form vector code book 521, according to the following equation 3, and issues the autocorrelation according to (i) for research section 524 and target gain calculation section 525.
F (m + 1) -F (zn) -1 according to (f) = c (i, k)<sup>2</sup> Equation 3 i = 0
Research section 524 calculates a contribution a represented by equation 4 below, using the cross correlation ccor (i) received from the cross correlation calculation section 522 and the autocorrelation according to (i) received from the autocorrelation calculation section 523, and issues a control signal for the vector codebook of form 521 until the maximum value of contribution A is found. Research section 524 issues index i<sub>op</sub>t of the vector candidate form when contribution A maximizes, as an optimal index, for the target gain calculation section 525, and issues the index i<sub>op</sub>t as the form-encoded information for multiplexing section 155.
[4]
A ~ ccor (í)<sup>2</sup> acor (í)
Equation 4
The target gain calculation section 525 calculates the target gain according to the following equation 5 using the cross-correlation ccor (i) received from the cross-correlation calculation section 522, the autocorrelation according to (i) received from the calculation section of autocorrelation 523 and the optimal index i<sub>op</sub>t received from research section 524, and issues this target gain to the gain vector formation section 153.
[5] ccor (i) gain = --— Equation 5 acor (i<sub>opt</sub>)
Figure 5 is a block diagram showing the configuration within the gain vector formation section 153.
In figure 5, the gain vector formation section 153 has the disposition position determination section 531 and the target gain disposition section 532.
The disposition positioning section 531 has a counter that takes 0 as an initial value, increments the value in the counter by one each time a target gain is received from the vector decoding section of form 152 and, when the counter value reaches the total number of subbands M, sets the counter value to zero again. Here, M is also the vector length of a gain vector formed in the gain vector formation section 153, and the processing in the counter provided in the disposition position determination section 531 is equal to dividing the value in the counter by the length vector icon of gain and find your rest. That is, the value in the counter assumes an integer between 0 and M-1. Each time the value on the counter is updated, the disposal position determination section 531 issues the updated value on the counter as the disposal information for the target gain disposal section 532.
The target gain disposition section 532 has M temporary stores that assume 0 as an initial value and a key that disposes of the target gain received from the 152 vector decoding section, in each candidate temporary storage, and this key has the target gain received from the 152 vector decoding section, in a temporary store that is designated as a value number shown by the disposition information received from the disposition determination section 531.
Figure 6 illustrates the operation of the target gain arrangement section 532 in detail.
In figure 6, when the layout information entered in the key shows 0, the target gain is displayed at 0<sup>2</sup> storage and, when the disposition information shows M-1, the target gain is arranged in (M-1)<sup>2</sup> storage. When the target gains are arranged in all temporary stores, the target gain arrangement section 532 issues a gain vector formed with the target gains arranged in M stores, to the gain vector encoding section 154.
Figure 7 is a block diagram showing the configuration within a gain vector encoding section 154.
In figure 7, the gain vector coding section 154 has a gain vector codebook 541, an error calculation section
542 and a research section 543.
The gain vector code book 541 stores a plurality of gain vector candidates representing a gain vector, and issues the gain vector candidates sequentially to the error calculation section 542, based on the received control signal of research section 543. Yet, generally, there are cases where a gain vector code book adopts a way to actually secure a storage space and store the gain vector candidates, and there are cases where a gain vector code book forms the candidate candidates. gain vector according to predetermined processing steps. In the latter cases, it is not really necessary to secure storage space. Although any of the gain vector code books can be used in the present mode, the present mode will be explained below assuming that the gain vector code book 541 that stores the gain vector candidates shown in figure 7 is provided . From now on, oj<sup>s</sup> gain vector candidate of the plurality of gain vector candidates stored in the gain vector code book 541, is represented as g (j, m). Here, m represents m<sup>s</sup> element of M elements that form a gain vector candidate.
The error calculation section 542 calculates the error EQ) according to the following equation 6 using the gain vector received from the gain vector formation section 153 and the gain vector candidate received from the gain vector code book 541, and issues error E (j) for research section 543. [6] <sup>AND</sup>(j) = Σ - g (J> m))<sup>2</sup> Equation 6 »i = 0
In equation 6, m represents the number of sub-bands, and gv (m) represents a gain vector received from the gain vector formation section 153.
The search section 543 emits a control signal to the gain vector code book 541 until a minimum value of the error E (j) received from the error calculation section 542 is found, search by the index jopt of when the error E (j) is minimized, and emits the index j<sub>op</sub>t as the encoded gain information for the multiplexing section 155.
Fig. 8 is a block diagram showing the main configuration of the voice decoding apparatus 200 in accordance with the present embodiment.
In figure 8, the voice decoding apparatus 200 has a demultiplexing section 201, a first layer decoding section 202, a second layer coding section 203, an adder 204, a switching section 205, a transformation section time domain 206 and a post-filter 207.
The demultiplexing section 201 demultiplexes the bit stream transmitted from the voice coding apparatus 100 through a transmission channel, in the encoded data of the first layer and in the encoded data of the second layer, and outputs the encoded data of the first layer and the data encoded from the second layer to the decoding section of the first layer 202 and the encoding section of the second layer 203, respectively. However, there are cases, depending on the state of the transmission channel (for example, the occurrence of congestion), where part of the encrypted data such as the encrypted data of the second layer or the encrypted data that includes the encrypted data of the first layer and the encoded data from the second layer is lost. The demultiplexing section 201 then decides whether only the encoded data from the first layer is included in the received encoded data or both the encoded data from the first layer and the encoded data from the second layer are included, and outputs 1 as the layer information in the first case, and issues 2 as the layer information in the latter case. Also, when deciding that all encrypted data including the encrypted data from the first layer and the encrypted data from the second layer are lost, the demultiplexing section 201 performs predetermined compensation processing to generate the encrypted data from the first layer and the encrypted data from the second layer, outputs the encoded data from the first layer and the encoded data from the second layer to the decoding section of the first layer 202 and the decoding section of the second layer 203, respectively, and outputs 2 as the layer information, to the switching section 205 .
The decoding section of the first layer 202 performs decoding processing using the encoded data from the first layer received from the demultiplexing section 201, and outputs the decoded transform coefficients from the first layer to the adder 204 and the switching section 205.
The second layer decoding section 203 performs decoding processing using the second layer encoded data received from the demultiplexing section 201, and outputs the decoded transform coefficients from the first layer to the adder 204.
Adder 204 adds the decoded transform coefficients of the first layer received from the decoding section of the first layer 202 and the error transform coefficients of the first layer received from the decoding section of the second layer 203, and outputs the decoded transform coefficients of the second resulting layers for switching section 205.
The switching section 205 outputs the decoded transform coefficients of the first layer as the decoded transform coefficients for the time domain transformation section 206 when the layer information received from the demultiplexing section 201 shows 1, and output the decoded transform coefficients of the second layer as the decoded transform coefficients for the time domain transformation section 206 when the layer information shows 2.
The time domain transformation section 206 transforms the decoded transform coefficients received from the switching section 205 into a time domain signal, and outputs the resulting decoded signal to the post filter 207.
Post-filter 207 performs post-filter processing such as format emphasis, step emphasis and spectral slope adjustment, with respect to the decoded signal received from the 206 time domain transformation section, and outputs the result as a decoded voice. .
Figure 9 is a block diagram showing the configuration within the decoding section of the second layer 203.
In figure 9, the decoding section of the second layer 203 has a demultiplexing section 231, a shape vector code book 232, a gain vector code book 233, and an error transform coefficient generation section of the first layer 234.
Demultiplexing section 231 further demuitiplexes the encoded data from the second layer received from demultiplexing section 201 into encoded information and gain encoded information, and outputs encoded information and gain encoded information to the vector codebook of form 232 and the gain vector code book 233, respectively.
Shape vector code book 232 has shape vector candidates identical to a plurality of shape vector candidates provided in shape vector code book 521 in Figure 4, and outputs the shape vector candidate shown by the information coded form received from the demultiplexing section 231, for the section of generation of error transform coefficient of the first layer 234.
The gain vector code book 233 has gain vector candidates identical to a plurality of gain vector candidates provided in the gain vector code book 541 in figure 7, and issues the gain vector candidate shown by the information gain codes received from the demultiplexing section 231, for the first layer 234 error transform coefficient generation section.
The first layer 234 error transform coefficient generation section multiplies the shape vector candidate received from the shape vector code book 232 by the gain vector candidate received from the gain vector code book 233 to generate the error transform coefficients of the first layer, and issues the error transform coefficients of the first layer to the adder 204. To be more specific, the om<sup>2</sup> element of the M elements that make up the gain vector candidate received from the gain vector code book 233, that is, the target gain of the m— subband transform coefficients, is multiplied by the m<sup>2</sup> vector candidate sequentially received from shape vector codebook 232. Here, as described above, M represents the total number of sub-bands.
In this way, the present modality employs a coding configuration of the spectral form of a target signal (that is, the error transform coefficients of the first layer with the present modality) on a subband basis (shape vector encoding) ), then calculating a target gain (ie, an ideal gain) that minimizes the distortion between the target signal and an encoded vector and encoding the target gain (target gain encoding). Hereby, compared with the scheme as a conventional technique of encoding the energy component of a target signal on a sub-band basis (gain encoding or scale factor) the normalization of the target signal using the encoded energy component and then encoding the spectral shape (shape vector encoding), the present invention that encodes the target gain to minimize distortion with respect to a target signal, can essentially minimize encoding distortion. In addition, the target gain is a parameter that can be calculated after the shape vector is encoded as shown in equation 5, and therefore, despite the coding scheme as a conventional technique to perform vector encoding in a manner that is temporally subsequent to the encoding gain information is unable to use the target gain as the target for the encoding gain information, the present modality makes it possible to use the target gain with the target to encode the gain information and can also minimize the coding distortion.
In addition, the present modality employs a configuration of forming and encoding a gain vector using the target gains of a plurality of adjacent sub-bands. The energy information between adjacent sub-bands of a target signal is similar, and the similarity of target gains between adjacent sub-bands is similarly high. Therefore, a non-uniform density distribution of gain vectors is produced in the vector space. By making the gain vector candidates included in the gain code book to be adapted for this non-uniform density distribution, it is possible to reduce the coding distortion of the target gain.
In this way, according to the present modality, it is possible to reduce the coding distortion of the target signal and, consequently, to improve the sound quality of the decoded voice. Furthermore, the present modality can encode precisely the spectral forms for the signal spectra with a strong tone such as the voice vowels and the music signals.
Also, with a conventional technique, the spectral amplitude is controlled by using two parameters, the subband gain and the shape vector. This can be interpreted as the spectral amplitude is represented separately by two parameters, the subband gain and the shape vector. In contrast to this, with the present modality, the spectral amplitude is controlled only by a target gain parameter. Furthermore, this target gain is an ideal gain that minimizes the coding distortion in relation to the vector in coded form. Consequently, it is possible to perform the encoding efficiently compared to a conventional technique and to produce a high quality sound even when the bit rate is low.
Yet, although a case has been explained with the present modality as an example where the frequency domain is divided into a plurality of sub-bands by the sub-band formation section 151 and coding is performed on a per-sub basis. band, the present invention is not limited to this. By performing vector encoding temporally before gain vector encoding, a plurality of sub bands can be collectively encoded, so that, similar to the present modality, it is possible to provide an advantage of more precisely encoding the spectral forms of strong signals. key such as vowels. For example, a configuration may be possible where the shape vector encoding is performed first, then the shape vector is divided into sub-bands and the target gains are calculated on a per-band basis to form a gain vector. and the gain vector is encoded.
Yet, although a case has been explained with the present embodiment as an example where the encoding section of the second layer 105 has a multiplexing section 155 (see figure 2), the present invention is not limited to this, and the shape vector encoding section 152 and the gain vector encoding section 154 can output the shape encoded information and the encoded gain information directly to the multiplexing section 106 of the voice encoding apparatus 100 (see figure 1). In contrast, the decoding section of the second layer 203 may not include the demultiplexing section 231 (see figure 9), and the demultiplexing section 201 of the voice decoder 200 (see figure 8) may demultiplex and output the shape coded information and the encoded gain information using a bit stream, directly into the shape vector codebook 232 and the gain vector codebook 233, respectively.
Yet, although a case has been explained with the present modality as an example where cross-correlation calculation section 522 calculates the cross-correlation ccor (i) according to equation 2, the present invention is not limited to this, and the cross-correlation calculation section 522 can calculate the cross-correlation ccor (i) according to the following equation 7 to increase the contribution of a perceptually important spectrum. [7] ccor (i) = y, c (iJc)
Equation 7
In equation 7, w (k) represents a weight relative to the characteristics of human perception and increases when a frequency has a higher importance in perceptual characteristics.
Also, similarly, the autocorrelation calculation section 523 can calculate the autocorrelation according to (i) according to the following equation 8 to increase the contribution of a perceptually important spectrum by applying a large weight to the perceptually important spectrum [8] according to (i) = 22 <sup>i = 0</sup> Equation 8
Also, similarly, the error calculation section 542 can calculate error E (j) according to the following equation 9 to increase the contribution of a perceptually important spectrum by applying a large weight to the perceptually important spectrum [9]
Af-1 <sup>AND</sup>(f) = Σ <sup>M</sup>'(<sup>w</sup>0 '(Mw) - g (j, ni))<sup>2</sup><sup>m</sup>= ° Equation 9
Like the weights in equation 7, in equation 8 and in equation 9, for example, weights can be found and used using the characteristics of human perceptual intensity or a perceptual masking limit calculated based on an input signal or a decoded signal from a lower layer (that is, a decoded signal from the first layer).
Yet, although a case has been explained with the present modality as an example where the shape vector coding section 152 has an autocorrelation calculation section 523, the present invention is not limited to this, and, when the coefficients of autocorrelation according to (i) calculated according to equation 3 or the autocorrelation coefficients according to (i) calculated according to equation 8 become constant, autocorrelation according to (i) can be calculated in advance and used without providing the autocorrelation calculation section 523.
(Mode 2)
The voice coding apparatus and the voice decoding apparatus according to Mode 2 of the present invention employ the same configuration and perform the same operation as the voice coding apparatus 100 and the voice decoding apparatus 200 described in the Modality 1, and Mode 2 differs from Mode 1 only in the shape vector code book.
To explain the vector codebook of form according to the present modality, figure 10 illustrates the spectrum of the Japanese vowel as an example of a vowel.
In figure 10, the horizontal geometric axis is the frequency and the vertical geometric axis is the logarithmic energy of the spectrum. As shown in figure 10, in the spectrum of a vowel, multiple peak forms are observed, showing a strong tone, yet, Fx is the frequency at which one of the multiple peak forms is placed.
Figure 11 illustrates a plurality of shape vector candidates included in the shape vector code book according to the present embodiment.
In figure 11, among the shape vector candidates, (a) illustrates a sample (that is, a pulse) that has an amplitude value of +1 or -1 and (b) illustrates a sample that has an amplitude value of 0 A plurality of vector candidates as shown in Figure 11 includes a plurality of pulses placed at arbitrary frequencies. Consequently, by searching for shape vector candidates shown in Figure 11, it is possible to more precisely encode a strong tone spectrum shown in Figure 10. To be more specific, a shape vector candidate is searched and determined with respect to a strong hue shown in figure 10 so that the amplitude value that corresponds to the frequency at which a peak shape is placed. For example, the amplitude value at the Fx position shown in figure 10 assumes +1 or -1 (that is, the sample (a) shown in figure 11) and the frequency amplitude value other than the peak form takes 0 ”(Ie, the sample (b) shown in figure 11).
With a conventional technique of performing gain coding temporarily before shape vector encoding, a subband gain is quantized, a spectrum is normalized using the subband gain and then the thin component (ie, the vector spectrum form is encoded). When the subband gain quantization distortion becomes significant by making the bit rate lower, the normalization effect becomes small and the dynamic range of the normalized spectrum cannot be reduced much. Hereby, the quantization step in the next vector section needs to be done grossly and therefore the quantization distortion increases. Due to the influence of this quantization distortion, the peak form of a spectrum attenuates (that is, loss of the true peak form), and the spectrum which does not form a peak form is amplified and appears as the peak form (ie ie, appearance of a false peak shape). In this way, the frequency position of the peak shape changes, causing a deterioration in sound quality in a vowel portion of a voice signal with a strong peak and a music signal.
In contrast to this, the present modality employs a configuration of first determining a vector, then calculating a target gain and quantizing this target gain. When some vector elements include a shape vector represented by a +1 or -1 pulse as in the present modality, determining the shape vector first means first determining the frequency position in which this pulse is under. The frequency position at which a pulse rises can be determined without the influence of gain quantization, and, consequently, the phenomenon where the true peak shape is lost or a false peak shape appears does not occur, so it is possible to prevent the problem described above with the conventional technique.
Thus, the present modality employs a configuration of determining the shape vector first to perform the shape vector encoding using the shape vector codebook formed with the shape vector that includes a pulse, so that it is possible to specify the frequency of the spectrum that has a strong peak and increase a pulse at this frequency. By this means, it is possible to encode signals that have strong tone spectra such as voice signal vowels and high quality music signals.
(Mode 3)
Modality 3 of the present invention differs from Modality 1 in selecting a band (i.e., region) with a strong tone in the voice signal spectrum and encoding only the selected band.
The speech coding apparatus according to Mode 3 of the present invention employs the same configuration as the speech coding apparatus 100 according to Mode 1 (see figure 1), and differs from speech encoding apparatus 100 only in the inclusion of the second layer 305 coding section instead of the second layer 105 coding section. Therefore, the total configuration of the voice coding device according to the present modality is not shown, and its detailed explanation will be omitted.
Figure 12 is a block diagram showing the configuration within the coding section of the second layer 305 according to the present embodiment. Furthermore, the coding section of the second layer 305 employs the same basic configuration as the coding section of the second layer 105 described in Modality 1 (see figure 1), and the same components will be assigned the same reference numbers and their explanation will be omitted.
The coding section of the second layer 305 differs from the coding section of the second layer 105 in accordance with Modality 1 in additionally including a section selection strip 351. Also the vector coding section of form 352 of the coding section of the second layer 305 differs from the shape vector coding section 152 from the second layer 105 coding section in part of the processing, and different reference numbers will be assigned to show this difference.
The band selection section 351 forms a plurality of bands using an arbitrary number of adjacent sub bands of M subband transform coefficients received from the subband forming section 151, and calculates the hue in each band. The range selection section 351 selects the range of the strongest hue, and outputs the range information showing the selected range, for the multiplexing section 155 and the shape vector encoding section 352. In addition, the track selection processing in the track selection section 351 will be explained in detail later.
The shape vector coding section 352 differs from the shape vector coding section 152 according to Mode 1 only in the selection of the subband transform coefficients included in a range of subband transform coefficients received from subband formation section 151, based on track information received from track selection section 351, and in the execution of the vector quantization in relation to the selected subband transform coefficients, and its detailed explanation will be omitted here.
Figure 13 illustrates the range selection processing in the range selection section 351.
In figure 13, the horizontal geometric axis is frequency and the vertical geometric axis is logarithmic energy. Still, figure 13 illustrates a case where the total number of sub-bands M is 8, a track 0 is formed using the 0-sub-band for the third sub-band, track 1 is formed using the second sub-band to the fifth subband and track 2 is formed using the fourth subband to the seventh subband. As an indicator for evaluating hue in a predetermined range, range selection section 351 calculates the spectral monotony measure (SFM) represented using the ratio of the geometric mean and the arithmetic mean of a plurality of subband transform coefficients included in a predetermined range. SFM assumes a value between 0 and 1 and the closest value to 0 shows a strong hue. Consequently, SFM is calculated for each range and the range that has the SFM closest to 0 is selected.
The voice decoding apparatus according to the present modality employs the same configuration as the voice decoding apparatus 200 according to Modality 1 (see figure 8), and differs from the voice decoding apparatus 200 only by including a section of voice decoding. decoding of the second layer 403 instead of the decoding section of the second layer 203. Therefore, the total configuration of the voice decoding apparatus will not be illustrated, and its detailed explanation will be omitted.
Figure 14 is a block diagram showing the configuration within the decoding section of the second layer 403 in accordance with the present embodiment. Also, the second layer decoding section
403 it uses the same basic configuration as the decoding section of the second layer 203 described in Mode 1, and the same components will be assigned the same reference numbers and their explanation will be omitted.
The demultiplexing section 431 and the error transform coefficient generation section of the first layer 434 of the decoding section of the second layer 403 differs from the demultiplexing section 231 and the error transform coefficient generation section of the first layer 234 of decoding section of the second layer 203 in part of the processing, and different reference numbers will be assigned to show this difference.
Demultiplexing section 431 differs from demultiplexing section 231 described in Modality 1 in demultiplexing and emitting range information in addition to the encoded shape information and the encoded gain information, for the first layer error coefficient generation section 434, and its detailed explanation will be omitted.
The error transform generation section of the first layer 434 multiplies the shape vector candidate received from the shape vector code book 232 by the gain vector candidate received from the gain vector code book 233 to generate the error transform coefficients of the first layer, it has these error transform coefficients of the first layer in the subband included within the range shown by the range information and outputs the result to the adder 204.
In this way, according to the present modality, the voice coding device selects the range of the strongest key and encodes the vector temporally before the gain of each subband in the selected range. Hereby, the spectral forms of signals with strong tonality such as voice vowels or music signals are encoded more precisely and the encoding is performed only within the selected range, so that it is possible to reduce the encoding bit rate .
Yet, although a case has been explained with the present modality as an example where an SFM is calculated as an indicator to assess the hue within each predetermined range, the present invention is not limited to this. For example, taking advantage of the high association between the average energy within the predetermined range and the intensity of the shade, the average energy of transform coefficients included in the predetermined range can be calculated as the shade assessment indicator. By this means, it is possible to reduce the computational complexity compared to the case where an SFM is calculated.
To be more specific, range selection section 351 calculates energy E<sub>R</sub>(j) the error transform coefficients of the first layer ei (k) included in range j, according to the following equation 10. [10]
FRfi (j) <sup>AND</sup>R <J) = Σ <sup>Ç</sup>1<sup>(/</sup><)<sup>2</sup> k = FRL <j) Equation 10
In this equation, j represents the identifier to specify the range, FRL (j) represents the lowest frequency in the range j and FRH (j) represents the highest frequency in the range j. The range selection section 351 calculates the energies E<sub>R</sub>(j) of the ranges in this way, then specify the range where the energy of the error transform coefficients of the first layer is the highest, and encode the error transform coefficients of the first layer included in this range.
In addition, the energy of the error transform coefficients of the first layer can be calculated according to the following equation 11 by carrying out a weighting taking into account the characteristics of human perception. [11]
FRHíj) <sup>and</sup>r (j) = Σ (<sup>k) 2</sup> k = FRíoy Equation 11
In such a case, the weight w (k) is increased higher for a frequency of higher importance in perceptual characteristics so that the range that includes this frequency is likely to be selected, and the weight w (k) is decreased for the frequency of minor importance so the range that includes this frequency is not likely to be selected. Hereby, a perceptually important band is likely to be preferentially selected, so that it is possible to improve the sound quality of decoded voice. Like this weight w (k), weights can be found and used using the characteristics of human perceptual intensity or a perceptual masking limit calculated based, for example, on an input signal or a lower layer decoded signal ( that is, a decoded signal from the first layer).
In addition, the band selection section 351 can be configured to select a band of bands arranged at lower frequencies than a predetermined frequency (i.e., the reference frequency).
Figure 15 illustrates a method for selecting in the band selection section 351 a range of bands arranged at frequencies lower than a predetermined frequency (i.e., the reference frequency).
Figure 15 shows the case as an example where eight selection band candidates are arranged in bands lower than the predetermined reference frequency Fy. These eight tracks are each formed with a band of a predetermined length starting from F1, F2, ..., and F8 as the base point, and track selection section 351 selects a track from these eight candidates based on the selection method described above. Hereby, the bands positioned at frequencies lower than the predetermined frequency Fy are selected. In this way, the advantages of performing coding emphasizing the low frequency band (or the medium - low frequency band) are as follows.
In the harmonic structure which is a characteristic of a voice signal (or is referred to as a harmony structure), that is, in the structure in which the spectrum shows peaks at given frequency ranges, the peaks appear high in a low frequency band compared to a high frequency band. Similar peaks are seen in the quantization error (that is, the error spectrum or error transform coefficients) produced in the coding processing, and the peaks appear high in a low frequency band compared to a high frequency band. Therefore, when the energy of an error spectrum in a low frequency band is lower than in a high frequency band, the peaks of an error spectrum are acute and therefore the error spectrum is likely to exceed a threshold. of perceptual masking (a threshold at which people can perceive a sound), causing a perceptual deterioration in sound quality. That is, even when the energy of the error spectrum is low, the perceptual sensitivity in a low frequency band is higher than in a high frequency band. Consequently, the band selection section 351 employs a configuration of selecting a band of candidates willing to lower frequencies than a predetermined frequency, so that it is possible to specify the band which is the target to be encoded, of a band of low frequency at which the peaks of the error spectrum are acute and improve the sound quality of decoded voice.
Also, as a method to select the range which is the target to be encoded, the range of the current frame can be selected in association with the range selected in the past frame. For example, there are methods of (1) determining the range of the current frame from bands positioned in the vicinity of the range selected in the previous frame, (2) re-arranging the range candidates to the current frame in the vicinity of the range selected in the previous frame to determine the current frame of the redistributed band candidates, and (3) transmitting the track information once every several frames and using the track shown by the track information transmitted in the past within the frame in which the track information is not transmitted (discontinuous transmission of track information).
In addition, track selection section 351 can divide an entire band into a plurality of partial bands in advance as shown in figure 16 to select a track from each partial band and concatenate the selected tracks from each partial band to make this track concatenated. target to be encoded. Figure 16 illustrates a case where the number of partial bands is two, and the partial band 1 is configured to cover a low frequency band and the partial band 2 is configured to cover a high frequency band. In addition, partial band 1 and partial band 2 are each formed with a plurality of bands. The track selection section 351 selects a track from each of the partial band 1 and the partial band 2. For example, as shown in figure 16, track 2 is selected on partial band 1 and track 4 is selected on partial band 2. Henceforth, information showing the selected band of partial band 1 is referred to as track information first band, and information showing the selected track for partial band 2 is referred to as second band partial band information. Next, track selection section 351 concatenates the selected track from partial band 1 and the selected track from partial band 2 to form a concatenated track. This concatenated track becomes the track selected in the track selection section 351, and the outside vector encoding section 352 performs shape vector encoding with respect to this concatenated track.
Figure 17 is a block diagram showing the configuration of the band selection section 351 that supports the case where the number of partial bands is N. In figure 17, the subband transform coefficients received from the sub-formation section band 151 are provided for the partial band selection section 1 511-1 up to the partial band selection section N 511-N. Each section selection of partial band n 511-n (where n = 1 to N) selects a band from each partial band n, and emits the information that shows the selected band, that is, the band information of the umpteenth partial band, for the band information formation section 512. The band information formation section 512 acquires the band concatenated by concatenating the bands shown by each band information n<sup>The</sup> partial band (where n = 1 to N) received from the partial band selection section 1 511-1 to the partial band selection section N 511-N. Then, the strip information formation section 512 outputs the information that shows the concatenated strip as the strip information, for the shape vector coding section 352 and the multiplexing section 155.
Figure 18 illustrates how the lane information is formed in the lane information formation section 512. As shown in figure 18, the lane information formation section 512 forms the lane information by arranging the first partial lane lane information. (ie bit A1) up to the partial band N-band information (ie bit AN) in order. Here, the An bit length of each n track information<sup>The</sup> partial band is determined based on the number of candidate bands included in each partial band and can take on a different value.
Figure 19 illustrates the operation of the error transform coefficient generation section of the first layer 434 (see figure 14) that supports the range selection section 351 shown in figure 17. Here, a case will be explained as an example where the number of partial bands is two. The error transform coefficient generation section of the first layer 434 multiplexes the shape vector candidate received from the shape vector code book 232 as the gain vector candidate received from the gain vector code book 233 then, the section of generation of error transform coefficient of the first layer 434 sets the vector candidate up after the gain multiplication, in each band shown by each band information of partial band 1 and partial band 2. The signal found in this way is output as the error transform coefficients of the first layer.
The selection method shown in figure 16 determines a range for each partial band and can have at least one decoded spectrum within each partial band. Consequently, by determining in advance a plurality of bands for which the sound quality needs to be improved, it is possible to improve the decoded voice quality compared to the band selection method of selecting only one band from the total band. For example, the range selection method shown in figure 16 is effective when, for example, quality improvement in both the low frequency and high frequency bands needs to be performed at the same time.
Also, as a variation of the range selection method shown in figure 16, a fixed range can be selected all the time in a specific partial band as shown in figure 20. With the example shown in figure 20, range 4 is selected all the time within the partial band 2 and forms part of the concatenated band. Similar to the effect of the track selection method shown in figure 16, the track selection method shown in figure 20 can determine in advance a band for which the sound quality needs to be improved and, for example, the band track information partial band 2 is not required, so it is possible to reduce the number of bits to represent the band information.
Yet, although figure 20 shows a case as an example where a fixed band is selected all the time in a high frequency band (partial band 2), the present invention is not limited to this, and the fixed band can be selected as all the time in a low frequency band (that is, partial band 1) and also a fixed band can be selected all the time in the partial band of the medium frequency band that is not shown in figure 20.
Also, as variations of the range selection methods shown in figure 16 and figure 20, the bandwidths of candidate ranges included in each partial band can be different. Figure 21 illustrates a case where the bandwidth of the candidate band included in the partial band 2 is shorter than the candidate bands included in the partial band 1.
(Mode 4)
Modality 4 of the present invention decides the degree of hue on a per frame basis, and determines the order of shape vector encoding and gain encoding depending on the outcome of the decision.
The voice coding device according to Mode 4 of the present invention employs the same configuration as the voice coding device 100 according to Mode 1 (see figure 1), and differs from the voice coding device 100 only in the inclusion of the second layer 505 coding section instead of the second layer 105 coding section. Therefore, the total configuration of the voice coding apparatus according to the present invention is not shown, and its detailed explanation will be omitted.
Figure 22 is a block diagram showing the configuration within the coding section of the second layer 505. In addition, the coding section of the second layer 505 employs the same basic configuration as the coding section of the second layer 105 shown in figure 1 , and the same components will be assigned the same reference numbers and their explanation will be omitted.
The coding section of the second layer 505 differs from the coding section of the second layer 105 in accordance with Modality 1 in that it additionally includes a tone decision section 551, a switching section 552, a gain coding section 553, a section standardization 554, a shape vector coding section 555 and a switching section 556. Still, in figure 22, the shape vector encoding section 152, the gain vector forming section 153, and the gain vector encoding section 154 constitute the encoding sequence (a) and the encoding section of gain 553, normalization section 554 and shape vector coding section 555 constitute the coding sequence (b).
The shade decision section 551 calculates an SFM as an indicator to assess the shade of the first layer error transform coefficients received from subtractor 104, and issues high as the tone decision information for switching section 552 and switching section 556 when the calculated SFM is less than the predetermined threshold and issues low as the tone decision information for switching section 552 and switching section 556 when the calculated SFM is equal to or greater than the predetermined limit.
Although, although the present modality is explained using SFM as an indicator to evaluate the shade, the present invention is not limited to this, and the decision can be made using another indicator such as the variance of the error coefficients of the first layer . Furthermore, the decision can be made using another signal such as an input signal to decide the key. For example, a height analysis results from an input signal or a result of encoding the input signal in a lower layer (that is, the encoding section of the first layer with the present modality) can be used.
The switching section 552 sequentially outputs M subband transform coefficients received from the subband forming section 151, to the shape vector coding section 152 when the tone decision information received from the tone decision section 551 shows high, and sequentially issues M subband transform coefficients received from subband forming section 151, for gain vector coding section 553 and normalization section 554 when tone decision information received from tone decision section 551 shows low.
The gain vector coding section 553 calculates the average energy of M subband transform coefficients received from switching section 552, quantizes the calculated average energy and outputs the quantized index as the coded gain information, for the switching 556. In addition, the gain vector encoding section 553 performs a gain decoding processing using the encoded gain information, and outputs the resulting decoded gain to the normalization section 554.
The normalization section 554 normalizes the M subband transform coefficients received from the switching section 552 using the decoded gain received from the gain coding section 553, and outputs the resulting normalized form vector to the vector coding section of form 555.
The shape vector coding section 555 performs coding processing against the normalized shape vector received from normalization section 554, and outputs the resulting coded information to the switching section 556.
The switching section 556 emits the encoded shape information and the encoded gain information received from the shape vector encoding section 152 and the gain vector encoding section 154, respectively, when the tone decision information received from the tone decision section 551 show high, and output the coded form information and the coded gain information received from the gain coding section 553 and the form vector coding section 555, respectively, when the tone decision information received from the tone decision section 551 shows low.
As above, the voice coding apparatus according to the present modality performs vector coding temporally before gain coding using sequence (a) in the case where the tonality of the error transform coefficients of the first layer is high, and executes the gain encoding temporally before the shape vector encoding using the sequence (b) in the case where the tonality of the error transform coefficients of the first layer is low.
In this way, the present modality adaptively changes the order of gain coding and vector coding of form according to the tone of the error transform coefficients of the first layer and, consequently, can suppress both the gain coding distortion and the shape vector encoding distortion according to an input signal which is the target to be encoded, so that it is possible to further improve the decoded voice sound quality.
(Mode 5)
Figure 23 is a block diagram showing the main configuration of the voice coding apparatus 600 according to Mode 5 of the present invention.
In figure 23, the voice coding apparatus 600 has a first layer coding section 601, a first layer decoding section 602, a delay section 603, a subtractor 604, a frequency domain transformation section 605, a second layer coding section 606 and a multiplexing section 106. Among these components, the multiplexing section 106 is the same as the multiplexing section 106 shown in figure 1, and therefore its detailed explanation will be omitted. Furthermore, the coding section of the second layer 606 differs from the coding section of the second layer 305 shown in figure 12 in part of the processing, and different reference numbers will be assigned to show this difference.
The first layer 601 coding section encodes an input signal, and outputs the first layer coded data generated to the first layer 602 decoding section and the multiplexing section 106. The first layer 601 coding section will be described later in detail.
The first layer decoding section 602 performs decoding processing using the first layer encoded data received from the first layer coding section 601, and outputs the first layer decoded signal generated to subtractor 604. The first decoding section layer 602 will be described in detail later.
The delay section 603 applies a predetermined delay to the input signal and outputs the input signal to subtractor 604. The delay duration is equal to the delay duration produced in the processing in the coding section of the first layer 601 and in the decoding section of the first layer 602.
Subtractor 604 calculates the difference between the delayed input signal received from delay section 603 and the first layer decoded signal received from the first layer decoding section 602, and outputs the resulting error signal to the domain transformation section of frequency 605.
The frequency domain transformation section 605 transforms the error signal received from subtractor 604 into a frequency domain signal, and outputs the resulting error transform coefficients to the second layer coding section 606.
Figure 24 is a block diagram showing the main configuration within the coding section of the first layer 601.
In figure 24, the coding section of the first layer 601 has a resolution reduction section 611 and a core coding section 612.
Resolution reduction section 611 reduces the resolution of the time domain input signal to convert the sample rate of the time domain signal to a desired sample rate, and outputs the reduced resolution time domain signal to the core coding section 612.
Core encoding section 612 performs encoding processing with respect to the input signal converted to the desired sample rate, and outputs the encoded data from the first layer generated to the decoding section of the first layer 602 and multiplexing section 106.
Figure 25 is a block diagram showing the main configuration within the decoding section of the first layer 602.
In figure 25, the decoding section of the first layer 602 has a core decoding section 621, a step-up section 622 and a high frequency band component addition section 623, and replaces an approximate signal for a band high frequency. This is based on a technique of making an improvement in sound quality of fully decoded voice representing a high frequency band of low perceptual importance with an approximate signal and instead increasing the number of bits to be allocated in a low frequency band ( or a medium - low frequency band) perceptually important to improve the fidelity of this band with respect to the original signal.
The core decoding section 621 performs decoding processing using the encoded data from the first layer received from the encoding section of the first layer 601, and outputs the resulting core decoded signal to the step-up section 622. Also, the section core decoding 621 emits the coefficients of
Decoded LPC found in decoding processing, for the 623 high frequency band component addition section.
Resolution-increasing section 622 increases the resolution of the decoded signal received from core decoding section 621 to convert the sample rate of the decoded signal to the same sample rate as the input signal, and outputs the decoded core signal from increased resolution for the 623 high frequency band component addition section.
Using an approximate signal, the high frequency band component addition section 623 compensates for a high frequency band component which was missing due to downsampling processing in the downsampling section 611. As a method for generating an approximate signal, a method for forming a synthesis filter with the decoded LPC coefficients found in decoding processing in core decoding section 621 and sequentially filtering out a noise signal from which energy is adjusted, for example through the synthesis filter and the bandpass filter, is known. The high frequency band component acquired in this method contributes to improving the perceptual feel of a band but has a completely different waveform from the high frequency band component of the original signal, and therefore the energy in the high frequency band the error signal acquired on the subtractor increases.
When the encoding processing of the first layer includes such characteristics, the energy in a high frequency band of the error signal increases, so that a low frequency band that essentially has a high perceptual sensitivity is not likely to be selected. Consequently, the coding section of the second layer 606 according to the present modality selects a range of candidates arranged at lower frequencies than a predetermined frequency (that is, a reference frequency), so that it is possible to prevent the above problem. described caused by an increase in energy of the error signal in a high frequency band. That is, the coding section of the second layer 606 performs the selection processing shown in figure 15.
Figure 26 is a block diagram showing the main configuration of the voice decoding apparatus 700 according to Mode 5 of the present invention. Meanwhile, the voice decoder 700 has the same basic configuration as the voice decoder 200 shown in Figure 8, and the same components will be assigned the same reference numbers and their explanation will be omitted.
The decoding section of the first layer 702 of the voice decoding apparatus 700 differs from the decoding section of the first layer 202 of the speech decoding apparatus 200 in part of the processing, and therefore different reference numbers will be assigned. In addition, the configuration and operation of the decoding section of the first layer 702 are the same as in the decoding section of the first layer 602 of the voice coding apparatus 600, and therefore its detailed explanation will be omitted.
The time domain transformation section 706 of the voice decoding apparatus 700 differs from the time domain transformation section 206 of the voice decoding apparatus 200 only in the disposition positions, but performs the same processing, and therefore different reference numbers will be assigned and your detailed explanation will be omitted.
In this way, the present modality replaces an approximate signal such as noise for a high frequency band in the encoding processing in the first layer, instead of increasing the number of bits to be allocated in a low frequency band (or a frequency band). medium - low frequency) perceptually important to improve fidelity with respect to the original signal of this band, still preventing a problem due to an increase in the energy of the error signal in a high frequency band that uses the lower range than a predetermined frequency as the target to be encoded in the second layer encoding processing and executing the vector encoding temporally41 before gain encoding, so that it is possible to more accurately encode spectral forms of strong tonal signals such as vowels, further reducing the gain vector encoding distortion without increasing the bit rate and, consequently, further improving the decoded voice sound quality.
Yet, although a case has been explained as an example where subtractor 604 finds the difference between time domain signals, the present invention is not limited to this and subtractor 604 can find the difference between domain transform coefficients frequency. In such a case, the input transform coefficients are found by arranging the frequency domain transformation section 605 between the delay section 603 and the subtractor 604, and the decoded transform coefficients of the first layer are found by providing another transformation section of frequency domain between the decoding section of the first layer 602 and the subtractor 604. Then, subtractor 604 finds the difference between the input transform coefficients and the decoded transform coefficients of the first layer, and supplies these error transform coefficients directly to the coding section of the second layer. This configuration allows adaptive subtraction processing to find the difference in a given band and not to find the difference in other bands, so that it is possible to further improve the sound quality of decoded voice.
Yet, although a configuration has been explained with the present modality as an example where information relating to a high frequency band is not transmitted to the voice decoding device, the present invention is not limited to this, and a configuration can be possible where a signal from a high frequency band is encoded at a low bit rate compared to a low frequency band and is transmitted to a voice decoding device.
(Mode 6)
Fig. 27 is a block diagram showing the main configuration of the voice coding apparatus 800 according to Modalida42 of 6 of the present invention. In addition, the voice coding device 800 employs the same basic configuration as the voice coding device 600 shown in figure 23, and the same components will be assigned the same reference numbers and their explanation will be omitted.
The voice coding apparatus 800 differs from the speech coding apparatus 600 in that it also includes a weighting filter 801.
The 801 weighting filter performs a perceptual weighting by filtering an error signal, and outputs the error signal after weighting, to the frequency domain transformation section 605. The 801 weighting filter attenuates (makes white) the spectrum of a input signal or change it to spectral characteristics for the attenuated spectrum. For example, the weighting filter transfer function w (z) is represented by the following equation 12 using the decoded LPC coefficients acquired in the decoding section of the first layer 602. [12]
NP r (.-) = ι-J «O) · / · - '<sup>Í = 1</sup> Equation 12
In equation 12, (i) are the LPC coefficients, NP is the order of the LPC coefficients, and γ is a parameter to control the degree of attenuation (turn white) of the spectrum and assumes values in the range of 0 <y < 1. When γ is greater, the degree of attenuation becomes greater, and 0.92, for example, is used for γ.
Figure 28 is a block diagram showing the main configuration of the voice decoding apparatus 900 according to Mode 6 of the present invention. In addition, the voice decoder 900 has the same basic configuration as the voice decoder 700 shown in figure 26, and the same components will be assigned the same reference numbers and their explanation will be omitted.
The voice decoder 900 differs from the voice decoder 700 in that it also includes a synthesis filter 901.
The synthesis filter 901 is formed with a filter that has spectral characteristics opposite to the weighting filter 801 of the speech coding apparatus 800, and performs filtering processing with respect to a signal received from the time domain transformation section 706 and issues the result. The transfer function B (z) of the synthesis filter 901 is plotted using equation 13 below. [13]
<img file="BRPI0808428A2_D0002.tif" />
<img file="BRPI0808428A2_D0003.tif" />
Equation 13
In equation 13, (i) are the LPC coefficients, NP is the order of the LPC coefficients, and γ is a parameter to control the degree of attenuation (turn white) of the spectrum and assumes values in the range of 0 <γ < 1. When γ is greater, the degree of attenuation becomes greater, and 0.92, for example, is used for γ.
As described above, the weighting filter 801 of the speech coding apparatus 800 is formed with a filter having a spectral characteristic opposite to the spectral envelope of an input signal, and the synthesis filter 901 of the speech decoding apparatus 900 is formed with a filter that has characteristics opposite to the weighting filter. Consequently, the synthesis filter has characteristics similar to the spectral envelope of the input signal. Generally, greater energy appears in a low frequency band than in a high frequency band and in the spectral envelope of a voice signal, so that even when the low frequency band and the high frequency band have a distortion of equal coding of a signal before this signal passes through the synthesis filter, the coding distortion becomes greater in the low frequency band after this signal passes through the synthesis filter. Although, ideally, the weighting filter 801 of the voice coding device 800 and the synthesis filter 901 of the voice decoding device 900 are introduced so that the coding distortion is not heard thanks to the perceptual masking effect, when the encoding distortion cannot be reduced due to the low bit rate, the perceptual masking effect does not work much and the encoding distortion is likely to be noticed. In such a case, the synthesis filter 901 of the voice decoding apparatus 900 increases the energy in a low frequency band that includes coding distortion and, therefore, a deterioration in quality and is likely to appear distinctly. With the present modality, as described in Modality 5, the coding section of the second layer 606 selects a range, which is the target to be coded, of candidates arranged at lower frequencies than a predetermined frequency (that is, the frequency reference), so that it is possible to alleviate the problem described above of emphasizing coding distortion in a low frequency band and improving the sound quality of decoded voice.
Thus, the present modality provides a weighting filter in the voice coding device, performs a quality improvement by providing the synthesis filter in the voice decoding apparatus and using a perceptual masking effect and uses a lower range than a predetermined frequency as the target to be encoded in the second layer encoding processing to alleviate a problem of increasing energy in a low frequency band including encoding distortion and performing vector encoding temporally before encoding gain, so that it is possible to more precisely encode the spectral forms of strong tonal signals such as vowels, reduce the gain vector encoding distortion and, consequently, further improve the sound quality of encoded voice.
(Mode 7)
The selection of the range which is the target to be encoded in each enhancement layer will be explained with Modality 7 of the present invention in the case where the voice coding device and the voice decoding device are configured to include three or more layers formed with a base layer and a plurality of enhancement layers.
Fig. 29 is a block diagram showing the main configuration of the voice coding apparatus 1000 according to Mode 7 of the present invention.
The voice coding apparatus 1000 has a frequency domain transformation section 101, a first layer coding section 102, a first layer decoding section 602, a subtractor 604, a second layer coding section 606, a second layer decoding section 1001, an adder 1002, a subtractor 1003, a third layer coding section 1004, a third layer decoding section 1005, an adder 1006, a subtractor 1007, a fourth layer coding section 1008 and a multiplexing section 1009, and is formed with four layers. Among these components, the configurations and operations of the frequency domain transformation section 101 and the coding section of the first layer 102 are shown in figure 1, the configurations and operations of the decoding section of the first layer 602, of subtractor 604 and the coding section of the second layer 606 are as shown in figure 23, and the configurations and operations of the blocks that have the numbers 1001 to 1009 are similar to the configurations and operations of blocks 101, 102, 602, 604 and 606 can be estimated and, therefore, the detailed explanation will be omitted here.
Fig. 30 illustrates the processing of selecting the range which is the target to be encoded in the coding processing on the voice coding apparatus 1000. Fig. 30A through Fig. 30C illustrates the processing of selecting the ranges in the coding of the second layer in second layer coding section 606, third layer coding in third layer coding section 1004 and fourth layer coding in fourth layer coding section 1008.
As shown in figure 30A, the selection band candidates are arranged in bands lower than the reference frequency of the second layer Fy (L2) in the encoding of the second layer, selection band candidates are arranged in bands lower than the third layer reference frequency Fy (L3) in the third layer coding and selection range candidates are arranged in bands lower than the reference frequency of fourth layer Fy (L4) in fourth layer encoding. Furthermore, the relationship of Fy (L2) <Fy (L3) <Fy (L4) is true between the reference frequencies of the improvement layers. The number of candidate range candidates in each enhancement layer is the same, and a case where the number of candidate range is four will be described as an example. That is, at a lower layer of a lower bit rate (for example, the second layer), the range which is the target to be encoded is selected from low frequency bands of perceptually higher sensitivities, and in a higher layer or a higher bit rate (for example, the fourth layer), the range which is the target to be encoded is selected from wider bands that include up to a high frequency band. By using such a configuration, a lower layer emphasizes a low frequency band and a higher layer covers a wider band, so that it is possible to produce the sound quality of voice signals.
Fig. 31 is a block diagram showing the main configuration of the voice decoding apparatus 1110 according to Modality 7 of the present modality.
In figure 31, the voice decoder 1110 has a demultiplexing section 1101, a first layer decoding section 1102, a second layer decoding section 1103, a summing section 1104, a third layer decoding section 1105 , a sum section 1106, a fourth layer decoding section 1107, a sum section 1108, a switching section 1109, a time domain transformation section 1110, and a post-filter 1111, and is formed with four layers. Meanwhile, the configurations and operations of these blocks are similar to the configurations and operations of blocks in the voice decoding apparatus 200 shown in figure 8 and can be estimated, and therefore their detailed explanation will be omitted.
In this way, according to the present modality, the scalable voice coding apparatus selects the range to be encoded, from low frequency bands of higher perceptual sensitivities in a lower layer of a lower bit rate. lowers and selects the range that is the target to be encoded, from wider bands that include even a high frequency band in a higher layer of a higher bit rate, to emphasize the low frequency band in the lowest layer and cover the widest bands in the highest layer and perform a vector encoding of shape temporally before the gain encoding, so that it is possible to more precisely encode the spectral forms of strong tonality such as vowels, further reduce the gain vector encoding distortion without increasing the bit rate and additionally improve the sound quality of decoded voice.
Still, although a case has been explained with the present modality as an example where the target to be coded is selected from the range selection candidates shown in figure 30 in the coding processing in each improvement layer, the present invention is not limited to this, and the target to be encoded can be selected from the range candidates arranged at equal intervals as shown in figure 32 and figure 33.
Fig. 32A, Fig. 32B and Fig. 33 illustrate a track selection processing in second layer coding, third layer coding and fourth layer coding. As shown in figure 32 and figure 33 the number of candidates for the selection range varies between the improvement layers, and a case will be illustrated here where the numbers of candidates for the selection range are four, six and eight. In such a configuration, the range to which the target is to be encoded is determined from low frequency bands, in a lower layer, and the number of selection band candidates is less compared to a higher layer, so that it is possible to reduce computational complexity and bit rate.
Also, as a method to select the range which is the target to be encoded by each enhancement layer, the range of the current layer can be selected in association with the range selected in the lowest layer. For example, there are methods of (1) determining the current layer range of the bands located in the vicinity of the selected range48 in the lowest layer, (2) redistribute the strip candidates to the current layer in the vicinity of the selected strip in the lowest layer to determine the current layer strip of the redisposed strip candidates and (3) transmit the strip information once every several frames and use the track shown by the track information transmitted in the past, within the frame in which the track information is not transmitted (discontinuous transmission of track information).
The modalities of the present invention have been explained.
Yet, although a scalable two-layer configuration has been explained as an example of the configuration of the voice coding apparatus and the voice decoding apparatus, the present invention is not limited to this, and the scalable configuration of three or more layers it may be possible. Furthermore, the present invention is also applicable to a voice coding apparatus that does not employ a scalable configuration.
Furthermore, the modalities described above can use the CELP method as the first layer encoding method.
The frequency domain transformation section in the above modalities is implemented by FFT, DFT (Discrete Fourier Transform), DCT (Discrete Cosine Transform), MDTC (Modified Discrete Cosine Transform), a sub-filter band and so on.
Although the modalities described above assume the speech signal as decoded signals, the present invention is not limited to this and, for example, decoded signals may be possible as audio signals.
Also, although cases have been described with the above modality as examples where the present invention is configured by hardware, the present invention can also be realized by software.
Each function block used in the description of each of the aforementioned modalities can typically be implemented with an LSI consisting of an integrated circuit. These can be individual chips or partially or completely contained in a single chip. LSI is adopted here but this can also be referred to as IC, system LSI, super LSI, or ultra-LSI depending on different integration extensions.
In addition, the circuit integration method is not limited to LSIs, and implementation using a dedicated circuit or general purpose processors is also possible. After manufacturing the LSI, the use of a programmable FPGA (Field Programmable Port Network) or a reconfigurable processor where the connections and circuit cell settings within an LSI can be reconfigured is also possible.
In addition, if an integrated circuit technology were to replace LSIs as a result of the advancement of semiconductor technology or another derived technology, it is naturally also possible to perform function block integration using this technology. The application of biotechnology is also possible.
Japanese Patent Application Number 2007053502, filed on March 2, 2007, Japanese Patent Application Number 2007-133545, filed on May 18, 2007, Japanese Patent Application Number 2007-185077, filed on March 13, 2007 July 2007, and Japanese Patent Application Number 2008-045259, filed on February 26, 2008, including specifications, drawings and abstracts are hereby incorporated by reference in their entirety. INDUSTRIAL APPLICABILITY
The voice coding apparatus and the voice coding method according to the present invention are applicable to a wireless communication terminal device, a base station device and so on in a mobile communication system.
Contents7
33 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33
37 members in 11 offices
Priority claims19
| Document | Office | Kind | Date |
|---|---|---|---|
| 2007053502 | Japan | – | |
| 2007053502 | Japan | A | |
| 2007133545 | Japan | – | |
| 2007133545 | Japan | A | |
| 2007185077 | Japan | – | |
| 2007185077 | Japan | A | |
| 2008045259 | Japan | – | |
| 2008045259 | Japan | A | |
| 2008000408 | Japan | W | |
| 2007053502 | – | – | – |
| 2007133545 | – | – | – |
| 2007185077 | – | – | – |
| 2008045259 | – | – | – |
| 2008000408 | – | – | – |
| JP20070053502 | – | – | – |
| JP20070133545 | – | – | – |
| JP20070185077 | – | – | – |
| JP20080045259 | – | – | – |
| WO2008JP00408 | – | – | – |
Members37
| Document | Office | Kind | |
|---|---|---|---|
| AU2008233888A1 | Australia | A1 | |
| WO2008120440A1 | World Intellectual Property Organization (WIPO) | A1 | |
| JP2009042734A | Japan | A | |
| JP2009042740A | Japan | A | |
| KR20090117890A | Republic of Korea | A | |
| EP2128857A1 | European Patent Office (EPO) | A1 | |
| CN101622662A | China | A | |
| US2010017204A1 | United States of America | A1 | |
| RU2009132934A | Russian Federation | A | |
| JP2011175278A | Japan | A | |
| JP4871894B2 | Japan | B2 | |
| SG178727A1 | Singapore | A1 | |
| SG178728A1 | Singapore | A1 | |
| CN102411933A | China | A | |
| MY147075A | Malaysia | A | |
| RU2471252C2 | Russian Federation | C2 | |
| AU2008233888B2 | Australia | B2 | |
| JP5236040B2 | Japan | B2 | |
| EP2128857A4 | European Patent Office (EPO) | A4 | |
| US8554549B2 | United States of America | B2 | |
| US2013325457A1 | United States of America | A1 | |
| US2013332154A1 | United States of America | A1 | |
| JP5403949B2 | Japan | B2 | |
| RU2012135696A | Russian Federation | A | |
| RU2012135697A | Russian Federation | A | |
| CN101622662B | China | B | |
| CN102411933B | China | B | |
| CN103903626A | China | A | |
| BRPI0808428A2This record | Brazil | A2 | |
| KR101414354B1 | Republic of Korea | B1 | |
| US8918314B2 | United States of America | B2 | |
| US8918315B2 | United States of America | B2 | |
| RU2579662C2 | Russian Federation | C2 | |
| RU2579663C2 | Russian Federation | C2 | |
| BRPI0808428A8 | Brazil | A8 | |
| CN103903626B | China | B | |
| EP2128857B1 | European Patent Office (EPO) | B1 |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Dismissal acc. art. 36, par 1 of ipl - no reply within 90 days to fullfil the necessary requirementsB11B | B11B | |
| Preliminary requirement: requests with searches performed by other patent offices: procedure suspended [chapter 6.21 patent gazette]B06U | B06U | |
| Objections, documents and/or translations needed after an examination request according [chapter 6.6 patent gazette]B06F | B06F | |
| Others concerning applications: alteration of classificationB15K | B15K | |
| Requested transfer of rights approvedB25A | B25A |
Numbers
- Publication
- PI0808428
- Publication, DOCDB
- PI0808428
- Publication, EPODOC
- BRPI0808428
- Application
- 8428
- Application, DOCDB
- PI0808428
- Application, EPODOC
- BR2008PI08428
Titles2
- Portuguese
- DISPOSTIVO DE CODIFICAÇÃO E MÉTODO DE CODIFICAÇÃO
- English
- CODING DEVICE AND CODING METHOD
Classification
- CPC, 9
- G10L19/24
- G10L19/00
- G10L19/02
- G10L19/0208
- G10L19/038
- G10L19/083
- G10L25/18
- G10L19/005
- G10L19/06
- IPC, 4
- G10L19 02
- G10L11 00
- G10L19 083
- G10L19 16
