Multi-pulse excited linear predictive speech coder.
Abstract
An LPC-synthesizer (1) produces a synthetic speech signal (s(n)) of which the difference (2) from the reference speech signal (s(n)) is perceptually weighted (4). In response to the weighted error signal (e(n), ε (n)) error minimizing means (5) control the multi-pulse excitation signal generator (6). The error minimizing procedure is accelerated by affecting the minimizing operation only in the region of the maximum of an auxiliary function (Mk(n) which is a measure of the energy of the weighted error signal.

Term
Term ended
Projected expiry passed 17 August 2004, 22.1 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
1 claim: 1 independent, 0 dependent
- 1A multi-pulse excited linear predictive speech coder canpri- sing a multi-pulse excitation signal generator, means for perceptually weighting the difference between a signal synthesized by means of a synthesizing operation from the multi-pulse excitation signal and the multi-pulse excitation signal itself, respectively, and the reference speech signal and a residual signal derived from the reference speech signal by means of an analysing operation which is the inverse of the said synthesizing operation, respectively, for generating a weighted error signal, and means for controlling the multi-pulse excitation generator in response to the weighted error signal in order to reduce the error signal, characterized in that in order to determine the position of the k th- pulsse in a given interval in the multi-pulse excitation signal an auxiliary function (M k (n)) is determined, which is a measure of the energy of the weighted error signal determined on the basis of a multi-pulse excitation signal of which (k-1) pulses have been determined, that means are present for determining the value n' k of n for which the auxiliary function (Mk(n)) is the maximum, that means are present for determining a reduced interval shorter than the predetermined interval, in the region of n' k , and means for determining the position of the k th pulse of the multi-pulse excitation signal in the reduced interval.
35 paragraphs, as filed
0001The invention relates to a multi-pulse excited linear predictive speech coder, comprising a multi-pulse excitation signal generator, means for perceptually weighting the difference between a signal synthesized by means of a synthesizing operation from the multi-pulse excitation signal and the multi-pulse excitation signal itself, respectively, and the reference speech signal and a residual signal derived from the reference speech signal by means of an analysing operation which is the inverse of the said synthesizing operation, respectively, for generating a weighted error signal and means for controlling the multi-pulse excitation generator in response to the weighted error signal, in order to reduce the error signal.
0002Such a speech coder is disclosed in the Proceedings of the ICASSP - 82, Paris, April 1982, pages 614-617.
0003Figure 1 shows the block diagram of such a multi-pulse excited speech coder (vocoder), which functions in accordance with the analysis-by-synthesis principle. In response to a multi-pulse signal r(n) a linear-predictive speech synthesizer 1 (LPC - SNT) produces synthetic speech samples s(n) which, in a difference producer 2, are compared with the reference speech samples s(n) which are applied to an input terminal 3. The difference s(n) - s(n) is perceptually weighted in block 4 (PRC-WGH) and the result is a weighted error signal e(n).
0004In response to the error signal e(n), block 5 (R-MN) effects a control of the multi-pulse excitation signal generator 6, which pro- duces the multi-pulse signal r(n), such that the synthetic speech signal s(n) reproduces the reference speech signal s(n) to the best possible extent. The procedure followed in block 5 is called the error-minimizing procedure.
0005Perceptually weighting the difference signal s(n) - s(n) in block 4 is effected by means of a transfer function denoted by W(z) in the Z-transform notation. This transfer function can be formed in such manner, that comparatively large errors are allowed in the formant areas as compared to the intermediate areas.
0006Let Ap(z) in the Z-transform notation represent the transfer function of the inverse LPC-filter. In terms of the inverse filter coefficients ap <sub>k</sub> the inverse filter transfer function is given<maths id="math0001" num=""><img file="EP0137532A2_D0001.tif" /></maths>
0007A suitable choice for W(z) is given by:<maths id="math0002" num=""><img file="EP0137532A2_D0002.tif" /></maths>where 0≤γ≤1 and q≤p.
0008The synthesizer 1 may be considered to be a filter having a transfer function S(z) which is given by S(z) = 1/Ap(z). The expressions shown in Figure 2a then hold for the combination of synthesizer 1 and the perceptual error weighting arrangement 4. They change into those of Figure 2b for the case in which the numerator function Ap(z) is split-off from transfer function W(z) of block 4 and is shifted to the input side of difference producer 2 emerging as block 8 on the one hand and disappearing in the combination with the synthe- sizer function S(z) = 1/A<sub>p</sub>(z) of block 1 on the other hand. In block 7 is left the transfer function G(z) = 1/Aq,γ (z).
0009In Figure 2b the filtering operation on the reference speech signal s(n) by the inverse LPC-filter Ap( z) produces the residual signal r(n). This signal is compared with the multi-pulse model r̂(n) thereof in the difference producer 2 and the difference is weighted in block 7 in accordance with the filter function 1/A<sub>q</sub>,γ. (z). The result is the error signal E (n) which has a strong correlation with the error signal e(n).
0010The reproduced speech will increase in quality by the insertion of a pitch predictor filter 9 into the lead to difference producer 2 carrying the signal r(n) and having the transfer function 1/P(z) wherein P(z) = 1-η3z<sup>-M</sup>.
0011In the above transfer function 1/P(z) the factor has an absolute value smaller than 1 and M represents the distance between the pitch pulses in number of samples. These values may be calculated for seg-ments of suitable length, say N from the speech correlation<maths id="math0003" num=""><img file="EP0137532A2_D0003.tif" /></maths> M is the value of k≠0 for which r(k) reaches a maximum value end η is proportional to r(M). The range of values of M at a sample frequency of 8 <sub>K</sub>Hz is typically from 16 to 160.
0012The effect of the inclusion of the inverse pitch predictor as represented by block 9 in Fig. 2b is shown in Fig. 6 wherein the signal-to-noise ratio of the reproduced speech is represented in dB versus time per segment of 10 msec. for a sequence of such segments. The drawn line is without the pitch predictor and the dashed line with the pitch predictor.
0013The Figures 1 and 2a represent the prior art as shown in the above-mentioned article or, as for the case represented in Figure 2b, extensions thereof.
0014In addition, the Figures 2a and 2b represent alternative methods of calculating a significant error signal e(n) or e (n), the latter having the advantage of a simple structure.
0015The complexity of the speech coder shown in Figure 1 is determined to an important extent by the procedure represented by block 5, i.e. the error minimizing procedure, in accordance with which the position and the amplitude of the pulses in the multi-pulse excitation signal r(n) are determined.
0016According to the prior art, in a given interval having a given number of possible pulse positions that position is determined, pulse for pulse, which minimizes a mean square error (m.s.e.) function or square distance function E<sub>k</sub>(b,t), where k is the number, b the amplitude and ℓ the position of the pulse under consideration. The number of function calculations will then be approximately equal to the product of the number of pulses to be determined and the number of pulse positions possible in the given interval.
0017The invention has for its object to provide a speech coder of the type specified in the preamble with a reduced complexity.
0018According to the invention, the speech coder is characterized in that in order to determine the position of the k<sup>th</sup> pulse in a given interval in the multi-pulse excitation signal an auxiliary function (M<sub>k</sub>(n)) is determined, which is a measure of the energy of the weighted error signal determined on the basis of a multi-pulse excitation signal of which (k-1) pulses have been determined, that means are present for determining the value n'<sub>k</sub> of n for which the auxiliary function (M<sub>k</sub>(n)) is the maximum, that means are present for determining a reduced interval shorter than the predetermined given interval, in the region of n'<sub>k</sub>, and means for determining the position of the k<sup>th</sup> pulse of the multi-pulse excitation signal in the reduced interval.
0019The auxiliary function M<sub>k</sub>(n) can be chosen such that it can be calculated in a simple way. The number of distance functions to be calculated by means of the method according to the invention is equal to the product of the number of pulses of the excitation signal to be determined in the given interval and the number of possible pulse positions in the reduced interval. As the reduced interval can be of a much shorter length than the predetermined given interval, the number of necessary calculations is significantly reduced and thus the complexity of the speech coder is reduced.
0020The invention will now be described in greater detail by way of example with reference to the accompanying Figures and an embodiment. <ul id="ul0001" list-style="none"><li>Figure 1 shows a block diagram of a prior art speech coder (vocoder).</li><li>Figure 2a and 2b show alternative methods for the determination of a weighted error signal:</li><li>Figure 3 shows a time scale (n) along which a multi-pulse excitation signal<maths id="math0004" num=""><img file="EP0137532A2_D0004.tif" /></maths>is plotted.</li><li>Figures 4a and 4b illustrate the relations between the different intervals.</li><li>Figures 5a and 5b illustrate a typical error signal and a typical distance function, respectively.</li><li>Figure 6 illustrates the signal-to-noise ratio of the reproduced speech with and without the use of a pitch predictor.</li></ul>
0021In the speech coder according to the invention which will be described hereafter the weighted error signal ( E (n)) will be calculated in accordance with the method as shown in Figure 2b at first without block 9. Herein:<maths id="math0005" num=""><img file="EP0137532A2_D0005.tif" /></maths>and<maths id="math0006" num=""><img file="EP0137532A2_D0006.tif" /></maths>
0022In block 5 (Figure 1) a distance function d(r,r):<maths id="math0007" num=""><img file="EP0137532A2_D0007.tif" /></maths><maths id="math0008" num=""><img file="EP0137532A2_D0008.tif" /></maths>is calculated between the residual signal r(n) - Fourier transform R(e<sup>jθ</sup>) - and the multi-pulse excitation signal r(n) - Fourier transform R̂(re<sup>jθ</sup>) -.
0023The error minimizing procedure of block 5 controls excita- tation signal generator 6 in such manner, that the synthetic speech signal s(n) (Figure 1) is obtained from a multi-pulse excitation signal r̂(n) A for which the distance function d(r,r) is at a minimum.
0024The error signal ε (n) (Figure 2b) is given by: ε (n) = (r(n) - r̂(n)) ∗ g(n) (7) where g(n) is the impulse response of the filter 7 with the transfer function G(z) and ∗ respresents the convolution operation.
0025As is illustrated in Figure 3, the multi-pulse excitation signal is divided into segments of the length L1. This length is less than or equal to the length L of the interval over which the distance A function d(r,r) (6) is calculated (L1 ≤ L). The number of possible pulse positions within a segment of the length L1 is, for example, 80, whereas within each segment the positions and amplitudes of, for example, 8 pulses must be determined which minimize the distance function.
0026According to the invention, the search for a suitable pulse position is always limited to a reduced interval or search interval of the length L<sup>e</sup><sub>I</sub> which is less than the length L1(L<sup>e</sup><sub>1</sub> < L1), preferably much less, comprising, for example, 5 to 10 possible pulse positions. The positons of the search intervals of the length L<sup>e</sup><sub>1</sub> within an interval of the length L1 are generally different for different pulses of the multi-pulse excitation signal. The above-mentioned ratios are illustrated in Figures 4a and 4b. As is illustrated in Figure 4b the positions of the search interval of the length L<sup>e</sup><sub>I</sub> will be in the region of A the minimum of the square of the distance function d(r,r).
0027The invention is based on the recognition that there is a high degree of correlation between the local minimum of the distance function d(r,r) and the local concentration of energy in the error signal which is optimized by the preceding pulse determinations. The distance function for the k<sup>th</sup> pulse determination is indicated by d<sub>k</sub>(r,r̂). Instead of an energy calculation, use is made of an average magnitude auxiliary function M<sub>k</sub>(n) which is given by:<maths id="math0009" num=""><img file="EP0137532A2_D0009.tif" /></maths>where m is the length of the integration interval, k is the number of the pulse of the multi-pulse excitation signal r(n) and E <sub>k</sub>(n) is the weighted error signal in accordance with the method shown in Figure 2b when k pulses of the multi-pulse excitation signal have been determined.
0028Figures 5a and 5b, respectively show by way of illustration a typical error signal E <sub>k-1</sub>(n) and a typical distance function d<sub>k</sub>(r,r̂) in a mutual relationship.
0029The procedure for the determination of a pulse in the multi-pulse exitation signal is as follows. When M<sub>k</sub>-<sub>1</sub>(n) reaches its maximum at n=n'<sub>k</sub>, then the distance function d<sub>k</sub>(r,r) is calculated for each available pulse position in the search interval, of the length L<sup>e</sup><sub>1</sub>, which is situated in the region of n'<sub>k</sub>. The suitable value for L<sup>e</sup><sub>1</sub> will depend on the length of m the integration interval and on the specific nature of the impulse response of the synthesis filter. In this example fixed-length search intervals are used. In the search interval the pulse position is then determined corresponding to the minimum of the distance function (Figure 4b).
0030This procedure is repeated until the desired number of pulse positions in the given interval of length L1 has been determined, whereafter a sub-sequent interval is proceeded to.
0031The following details can be given by way of illustration: <ul id="ul0002" list-style="none"><li>- sample frequency: 8KHz;</li><li>- L<sup>e</sup><sub>l</sub>: 5 to 10 possible pulse positions;</li><li>- L1: 80 possible pulse positions;</li><li>- number of pulse positions to be determined within interval L1: 8 to 10;</li><li>- integration interval, m<sup>=</sup>4.</li></ul>
0032The position of the search interval of length L<sup>e</sup> relative to the maximum of the auxiliary function M<sub>k</sub>(n) will adequately be such that it precedes this maximum with, optionally, a suitable shift (offset) relative to this maximum.
0033The auxiliary function M<sub>k</sub>(n) can be realised by an integrator to which the magnitude of the error signal E <sub>k</sub>(n) is applied and which integrates it over m pulse positions.
0034As has been indicated with respect to figure 2b, the quality of the synthesized speech will considerably improve when a pitch predictor 9 is inserted in the lead for the multi-pulse excitation signal r̂(n).
0035For the purpose of this specification the term multi-pulse excitation signal is considered generic for the multi-pulse excitation signal r(n) as indicated in the figures and the signal appearing at the output of the pitch predictor 9 in figure 2b when such predictor is in fact included and the multi-pulse excitation signal r(n) is applied thereto.
12 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US5899968A | Cited by | United States of America | Search report |
| US5974377A | Cited by | United States of America | Search report |
| US4864621A | Cited by | United States of America | Search report |
| WO9306590A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| DE19920501A1 | Cited by | Germany | Search report |
| EP0721180A1 | Cited by | European Patent Office (EPO) | Search report |
| WO8802165A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO9621219A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| FR2729244A1 | Cited by | France | Search report |
| US4944013A | Cited by | United States of America | Search report |
| US5963898A | Cited by | United States of America | Search report |
| WO9621219A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| GB2110906A | Cites | United Kingdom | Search report |
11 members in 7 offices; this record represents the family
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 8302985 | Netherlands (Kingdom of the) | A | |
| 8302985 | Netherlands (Kingdom of the) | – | |
| NL19830002985 | – | – | – |
| 8302985 | – | – | – |
Members11
| Document | Office | Kind | |
|---|---|---|---|
| AU3237884A | Australia | A | |
| NL8302985A | Netherlands (Kingdom of the) | A | |
| EP0137532A2This record | European Patent Office (EPO) | A2 | |
| JPS6070500A | Japan | A | |
| EP0137532A3 | European Patent Office (EPO) | A3 | |
| CA1213059A | Canada | A | |
| US4736428A | United States of America | A | |
| AU574708B2 | Australia | B2 | |
| EP0137532B1 | European Patent Office (EPO) | B1 | |
| DE3475664D1 | Germany | D1 | |
| JPH0562760B2 | Japan | B2 |
33 legal events, as 2 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Notification of lapseLapsedST | ST | FR | |
| Se: european patent has lapsedLapsedEUG | EUG | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Gb: european patent ceased through non-payment of renewal feeCeasedGBPC | GBPC | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Change of name or company nameCD | CD | FR | |
| It: changes in ownership of a european patentITPR | ITPR | EP | |
| Se: european patent in force in swedenEAL | EAL | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Be: lapsedLapsedBERE | BERE | EP | |
| It: last paid annual feeITTA | ITTA | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Fr: translation filedET | ET | EP | |
| It: translation for a ep patent filedITF | ITF | EP | |
| It: translation for a ep patent filedITF | ITF | EP | |
| Corresponds to:REF | REF | EP | |
| Designated contracting statesAK | AK | EP | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Designated contracting statesAK | AK | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 0137532
- Publication, DOCDB
- 0137532
- Publication, EPODOC
- EP0137532
- Application
- 84201194
- Application, DOCDB
- 84201194
- Application, EPODOC
- EP19840201194
Titles6
- German
- Linearer Prädiktionssprachcodierer mit Mehrimpulsanregung.
- English
- Multi-pulse excited linear predictive speech coder.
- French
- Codeur à prédiction linéaire pour signal vocal avec excitation par impulsions multiples.
- German
- Linearer Prädiktionssprachcodierer mit Mehrimpulsanregung
- English
- Multi-pulse excited linear predictive speech coder
- French
- Codeur à prédiction linéaire pour signal vocal avec excitation par impulsions multiples
Classification
- CPC, 1
- G10L19/10
- IPC, 1
- G10L19 10
Designated states6
- Contracting states, 6
- Belgium
- Germany
- France
- United Kingdom
- Italy
- Sweden