Method of speech coding
6 claims: 4 independent, 2 dependent
- 1Patenttivaatimukset 1. Menetelmä puhesignaalin laadun parantamiseksi lineaarisen ennustavan koodauksen yhteydessä, jossa dekoodaus koostuu koodausparametrien eli LPC-suodatinmallin (LPC, Linear Predictive Coding) kertoimien ja herätesignaalin demultipleksauksesta ja dekvantisoinnista sekä puhesignaalin syntesoimisesta synteesisuodattimessa, jonka sisääntuloon viedään vastaanotettu herätesignaali ja jonka kerroinarvoiksi on asetettu vastaanotetut LPC-parametrit, tunnettu siitä, että - dekoodatut puheen lyhytaikaista spektrikäyttäytymistä kuvaavat suodatinkertoimet käsitellään epälineaarisessa muokkauslohkossa (205), joka suorittaa niille epälineaarisen käsittelyn mediaanioperaation avulla, - suodatinkerrointen epälineaarista muokkausta (205) ohjataan siten, että muokkaus (205) aktivoidaan vain kun suodatinkertoimia kuvaavissa parametreissä on merkittävästi siirtovirheitä .
- 2Patenttivaatimuksen 1 mukainen menetelmä, tunnettu siitä, että epälineaarisen muokkauslohkon (301) sisääntuloon (300) tuodaan LPC-parametriesitys ja perättäisen N:n parametriarvon kesken suoritetaan lajitteluoperaatio, joka antaa ulostulonaan (302) mediaanin kyseisistä N:stä arvosta, ja että epälineaarinen muokkaus suoritetaan erikseen kullekin dekoodatulle LPC-kertoimelle.
- 3Patenttivaatimuksen 1 tai 2 mukainen menetelmä, tunnettu siitä, että epälineaarisessa muokkauslohkossa (401) käytetään rekursiivista mediaanioperaatiota, jolloin lajittelijan (403) sisäänmenoista muokkauslohkon (401) sisääntulosta (400) päin katsoen (k+2):nteen sisäänmenoon viedään edellinen lajittelijan (403) ulostuloarvo (402).
- 4Jonkin edellä olevan patenttivaatimuksen mukainen menetelmä, tunnettu siitä, että muokkauslohkossa (501) kutakin LPC-parametrijoukkoa käsitellään samanaikaisesti vektorina (503) ja jolloin ulostulovektori muodostetaan LPC-parametri10 vektorien X„, X nl '···' Χη.κ avulla siten, että lasketaan kunkin vektorin X, etäisyys muihin K:hon vektoriin ja etsitään minimietäisyyden muihin antava vektori, joka valitaan dekooderin synteesisuodatuksessa käytettäväksi.
- 5Jonkin edellä olevan patenttivaatimuksen mukainen menetelmä, tunnettu siitä, että vain riippuvuutta lähimpiin puhesignaalin näytearvoihin kuvaavat LPC-parametrit käsitellään epälineaarisessa muokkauslohkossa (205) ja muut välitetään synteesisuodattimelle (203) ilman käsittelyä muokkaus lohkossa (205) .
- 6Digitaalinen dekooderi, jossa on demultiplekseri (201) lineaarisen ennustavan koodauksen koodausparametrien ja herätesignaalin demultipleksaamiseksi ja dekvantisoijat (204, 202) näiden dekvantisoimiseksi sekä synteesisuodatin (203) puhesignaalin syntetisoimiseksi, jolloin dekooderin sisääntuloon viedään vastaanotettu herätesignaali, ja suodattimen kerroinarvoiksi on asetettu vastaanotetut LPC-parametrit, ja jolloin dekooderissa vastaanotettu bittivirta (200) on sovitettu johdettavaksi demultiplekserille (201), ja demultiplekseriltä (201) saatava LPC-parametriesitys on sovitettu dekvantisoitavaksi dekvantisoijassa (204), tunnettu epälineaarisesta muokkauslohkosta (205), jossa puheen lyhytaikaista spektrikäyttäytymistä kuvaavat suodatinkertoimet käsitellään mediaanioperaation avulla, jolloin LPC-parametrit on sovitettu johdettavaksi dekvantisoijasta (204) edelleen muokkauslohkoon (205), josta saadut käsitellyt parametriarvot viedään synteesisuodattimelle (203) kertoimiksi ja ennustusvirhesignaali, joka dekvantisoidaan dekvantisoijassa (202), on sovitettu johdettavaksi herätteeksi synteesisuodattimelle (203), jonka ulostulosta (206) saadaan dekoodattu puhesignaali, ja jolloin muokkauslohko (205) aktivoidaan vain kun suodatinkertoimia kuvaavissa parametreissä on merkittävästi siirtovirheitä.
Independent claims6
32 paragraphs, as filed
Speech signal quality improvement method for a coding system using linear prediction - A method for improving the quality of a coding system using a linear prediction system
The invention relates to a method for improving the quality of speech coding methods using linear prediction.
Linear Predictive Coding (LPC) is a widely used and known method of speech coding.
The prior art will be described below with reference to the accompanying Figure 1, which shows the implementation of a solution according to the prior art.
Figure 1 shows a block diagram of a prior art speech prediction encoder based on linear prediction. In the encoder, the incoming signal s (n) 100 is processed block by block. The block length N is generally selected to be about 10-30 ms in length. The sampling frequency of the speech signal 100 is usually 8 kHz, in which case a degree of 8 ... 12 is sufficient for the linear prediction model. In the LPC analyzer 103, LPC parameters, i.e. filter coefficients, are calculated for each block of the speech signal 100. These can be the coefficients ai of the direct filter model; i = 1, 2, ..., P, where P is the degree of the LPC model used. The filters of the LPC model are often implemented as a lattice-structured filter, for which the direct-form filter coefficients are converted into so-called reflection coefficients rci, i = 1, 2, ..., P. The calculated filter coefficients are quantized and applied to block 106, which performs multiplexing and error correction coding.
The speech signal 100 to be encoded is applied to the analysis filter 101 so that each block of the speech signal 100 is filtered in the analysis filter 101 using the filter coefficient values computed from that block in the LPCanalyzer 103. The analysis filter 101 uses synteesisuodatukselle. The output of quantization block 104 is applied to dequantization block 105 and further to analysis filter 101 for use as filter coefficients. The output of the analysis filter 101 is the so-called a prediction error for that 100 blocks of the speech signal. This prediction error signal is quantized by the quantizer 102 and is also applied to the multiplexer 106 for further transmission to the communication channel 107.
Depending on how the prediction error of the LPC model is transmitted to the decoder, several different coding methods can be derived for the speech signal. When quantizing each sample of prediction error at a time, it is referred to as Residual Excited Predictive Coding (REPC, see, e.g., U.S. Patent No. 4,220,819). The most efficient methods based on linear prediction use the so-called an analysis-synthesis technique in which a suitable quantized representation of a prediction error is sought by performing synthesis of a speech signal in the encoder with different excitation possibilities, i.e. quantized error signals, and selecting the stimulus producing the best synthesis result for transmission to the decoder.
When a representation containing only a small number of non-zero sample values is retrieved for a prediction error by an analysis-synthesis search, we speak of Multi Pulse Coding (MPC, see, e.g., U.S. Pat. No. 4,472,832). In Code Excited Linear Prediction (CELP), see e.g., U.S. Patent No. 4,817,157) uses a vector representation of each prediction error block, wherein the stimulus optimized by the analysis-synthesis technique may contain a large number of non-zero sample values while limiting the number of different stimulus combinations to the small number required for low transfer rate.
The quality of the speech signal transmitted by the coding methods based on linear prediction is clearly degraded if transmission errors occur on the transmission channel. Especially on the noisy channels of mobile radio traffic, the best possible ability of the coding method to cope with transmission errors is essential in order to achieve the best possible speech signal quality. Transmission errors can be protected to some extent by the use of special error correction coding. In this case, in addition to the parameters representing the speech signal, additional bits used for error correction are transmitted to the receiver. However, the transmission of such additional error correction information reduces the number of bits available for the actual speech coding and thus increases the speech signal distortion caused by the speech coding itself. On the other hand, not all transmitted coding parameters can be effectively protected by error correction coding. Thus, it would be desirable to provide a reduction in the effect of transmission errors by means of the coding parameters themselves, which could be performed without the transmission of additional information decreasing the channel capacity. Such a reduction in the effects of transmission errors could work either as such or in combination with separate error correction coding.
It is an object of the present invention to provide a method for improving the quality of a speech signal in the context of linear predictive coding, by means of which the above-mentioned deficiencies and problems could be solved. To achieve this, the invention is characterized in that the Decoded filter coefficients describing the short-term spectral behavior of speech are processed in a nonlinear modification block which performs nonlinear processing on them by a median operation, and that the nonlinear modification of the filter coefficients is controlled.
Media operations per se are described, for example, in J. Astola, P. Heinonen, Y. Neuvo, Vector Media Filters, Proc. IEEE, Vol. 78, no. 4, April 1990, pages 678689, and P. Haavisto, M. Gabbouj, Y. Neuvo, Media Based Idempotent Filters, Journal of Circuits and Systems and Computers, Vol. 2, 1991, pages 125-148.
The method according to the invention can be applied in all Encoders using LPC modeling, in which the Prediction coefficients of the model are transmitted to the receiver in a transmission channel producin g transmi ssion errors.
The invention will now be described in detail with reference to the accompanying drawings, in which: Figure 1 shows a block diagram of a prior art linear prediction speech signal encoder, Figure 2 shows a block diagram of a decoder according to the invention, Figure 3 shows a block diagram of a nonlinear processing block of a speech coder according to the invention, Figure 4 shows a block diagram of a non-linear modification block activity.
Figure 1 is described above. The solution according to the invention is described below with reference to Figures 2-5, which show the implementation of the solution according to the invention.
Figure 2 shows a block diagram of a decoder according to the invention. The decoder functions similarly to the use of non-linear modification, except for the decoder based on linear prediction according to the prior art. In the decoding part of the encoder based on the prior art linear prediction, the inverse operations are performed for the encoding of Fig. 1. The various encoding parameters are demultiplexed from the bit stream applied to the decoder and dequantized. The speech signal is synthesized in the decoder using an inverse synthesis filter for the encoder analysis filter model. The dequantized prediction error signal is used as an excitation for a synthesis filter whose coefficients are obtained by dequantizing the transmitted prediction coefficients. A synthesized speech signal is obtained from the output of the synthesis filter.
The bit stream 200 received in the decoder is applied to a demultiplexer 201. The LPC parameter representation from the demultiplexer 201 is dequantized in a dequantizer 204. The LPC parameters are further applied to an editing block 205, from which the processed parameter values are applied to the synthesis filter 203 as coefficients. In addition to the LPC parameters, a prediction error signal is obtained from the demultiplexer 201, which is dequantized in the dequantizer 202 and applied to the synthesis filter 203. The decoded speech signal s' (n) is obtained from the output 206 of the synthesis filter 203.
By using the modification block 205 according to the invention, the effect of transmission errors generated in connection with the spectral parameters on the quality of the speech signal to be synthesized in the decoder can be reduced. By means of nonlinear modification, parameters containing transmission errors can thus be used in synthesis filtering to produce a good quality speech signal.
The operation of the editing block 205 is controlled by the information from the error correction decoding on the number of channel transmission errors. The modification block 205 is activated only if the number of transmission errors in the spectral parameters becomes significantly large. The modification operation is not performed, i.e., the dequantized LPC parameters are passed directly to the synthesis filter 203 for use if the transmission link is error-free or its errors in the LPC parameters do not substantially degrade the speech signal quality.
The operation of the modification block 205 is based on the identification of values containing transmission errors and their replacement with usable values by a median operation. The modification is performed by means of the LPC parameter values of several consecutive speech frames, and this procedure is explained in more detail in the exemplary embodiments to be presented later.
Using the method for LPC parameters, the so-called the number of frames classified as poor can be reduced and thus the replacement of bad frames by a separate compensation procedure is seldom necessary.
The method does not require the transmission of additional error correction information and thus does not cause a strain on the transmission capacity. The method can therefore be easily incorporated for use in speech prediction-based codecs by implementing it in the decoding section of the LPC parameters as shown in Figure 2.
Figure 3 shows a block diagram of a non-linear editing block of a speech encoder according to the invention. The processing is based on the median operation. An LPC parameter representation obtained from the dequantizer is applied to the input 300 of the editing block 301. A sorting operation is performed between the N consecutive parameter values of each LPC parameter. The sorting block 303 gives as its output 302 the median value of these N input values of the sorter 303, i.e. when N = 2k + 1, then the output 302 gives the (k + 1) largest value of the sorter inputs li, I<sub>2</sub>,. . . , l2t + i. The non-linear processing according to the figure is performed in parallel separately for each LPC coefficient transmitted in the transmission channel. It should be noted that the unit delay symbols 304 refer to the calculation frequency of the LPC parameters and not to the sampling frequency of the speech signal.
Figure 4 shows an alternative implementation of a non-linear editing block of a speech encoder according to the invention. The processing is based on a recursive median operation. In this case, the output 402 of the sorter 403 is passed on to the sorting block 403 for processing. The LPC parameter value to be processed is input to the input 400 of the modification block 401. In recursive processing from the inputs of the sorter 403 to the left, i.e. from the input 400 of the editing block 401, to the (k + 2) th input, the previous output value 402 of the sorter 403 is applied and not the previous value of the (k + 1) th input of the sorter 403.
The recursive processing makes the operation of the editing block 401 more efficient, whereby a short sorting operation can be used and the delay caused by the editing can be considered reasonable. In this case, too, the processing is performed separately for each LPC parameter. Even a three-input sorting operation provides a good editing result in the decoder. Recursive processing also keeps the computational load due to modification low.
The computational load caused by the method can be further reduced by processing only the most important LPC parameter vector values in the modification block 401, i.e. by processing only the LPC parameters describing the dependence on the nearest speech signal sample values and passing other LPC parameters without modification to the synthesis filters. For example, when using 8-stage modeling, almost as good a result is obtained by processing the three or four lowest LPC parameters in the modification block 401 as by processing all eight parameters.
Figure 5 shows a block diagram of a nonlinear modification block of a vector type according to the invention. The modification method implements vector processing of LPC parameters. Because the prediction coefficients are a set of parameters computed simultaneously for each block of the input signal, they are inherently vector-type. In each frame n, a prediction vector X can naturally be formed<sub>of</sub>, which, for example, when using the reflection coefficient representation, contains the reflection coefficient values (rcjn), rc<sub>2</sub>(n), ..., rc<sub>p</sub>(of)).
Each set of LPC parameters is processed into a vector, which is applied to the input 500 of the vector editing block 501. In terms of speech quality, the dequantized reflection coefficient vector X „
503 better speech quality than direct use in a channel with transmission errors is obtained by applying to the synthesis filter the output 502 of the output block 501 of the vector Y<sub>of</sub> processed reflectance values contained in.
In vector modification, the output vector is formed by X<sub>of</sub>,
X „-i,. . , 2L-K using a reflection coefficient vector by performing a vector media operation. The vector media operation is performed by calculating the distance X <of each vector to the other K vectors and finding the vector giving the minimum distance to the others. The distance of the vectors is calculated as the sum of the distances of the components of the vectors. The distance measurements can be weighted so that the lower components of the reflection coefficient vector take on a higher importance. The vector media operation can also be performed recursively by including the previous output vector of the editing block 501 at the input of the sorter.
The method according to the invention can be utilized in all methods using linear prediction, i.e. in linear predictive coding methods. By using the nonlinear modification method according to the invention, the probability of interrupting the speech signal is reduced.
By means of the modification method according to the invention, the prediction coefficients according to the LPC model can be used to synthesize the speech signal even if they contain significant transmission errors. The method allows the bit stream otherwise classified as unusable in the transmission connection to be utilized at the receiver to synthesize the speech signal.
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
11 members in 7 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 921250 | Finland | A | |
| 921250 | – | – | – |
| FI19920001250 | – | – | – |
Members11
| Document | Office | Kind | |
|---|---|---|---|
| EP0562777A1 | European Patent Office (EPO) | A1 | |
| AU3537693A | Australia | A | |
| FI90477B | Finland | B | |
| JPH0612099A | Japan | A | |
| FI90477CThis record | Finland | C | |
| US5432884A | United States of America | A | |
| AU666172B2 | Australia | B2 | |
| EP0562777B1 | European Patent Office (EPO) | B1 | |
| DE69329568D1 | Germany | D1 | |
| DK0562777T3 | Denmark | T3 | |
| DE69329568T2 | Germany | T2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Publication of examined applicationBB | BB | |
| Name/ company changed in applicationHC | HC |
Numbers
- Publication, DOCDB
- 90477
- Publication, EPODOC
- FI90477C
- Application
- 921250
- Application, DOCDB
- 921250
- Application, EPODOC
- FI19920001250
Titles3
- Finnish
- Puhesignaalin laadun parannusmenetelmä lineaarista ennustusta käyttävään koodausjärjestelmään
- Swedish
- En metod för förbättring av kvaliteten vid ett kodningssystem som använder lineär prognostisering
- English
- The quality of the speech signal enhancement method uses a linear predictive Coding
Classification
- CPC, 1
- G10L19/06
- IPC, 1
- G10L19 06
