Method and apparatus for encoding/decoding of background sounds
10 claims: 4 independent, 6 dependent
- 1PATENTKRAV 1. Förfarande för kodning och/eller avkodning av bakgrundsljud i en digital rambaserad talkodare och/eller -avkodare innehållande en signalkälla ansluten till ett filter, varvid filtret definieras av en uppsättning filterparametrar för varje ram, i och för reproducering av den signal som skall kodas och/eller avkodas, vilket förfarande innefattar stegen:(a) detektering av huruvida signalen som förs till kodaren/avkodaren representerar primärt tal eller bakgrundsljud;och (b) om signalen som förs till kodaren/avkodaren representerar primärt bakgrundsljud, begränsning av den temporala variationen mellan på varandra följande ramar och/eller domänen av åtminstone några filterparametrar i uppsättningen .
- 2Förfarande enligt krav 1 i vilket den temporala variationen av filterparametrarna begränsas genom lågpassfiltrering av filterparametrarna över flera ramar.
- 3Förfarande enligt krav 2 i vilket den temporala variationen av filterparametrarna begränsas genom medelvärdesbildning av filterparametrarna över flera ramar.
- 4Förfarande enligt krav 1, 2 eller 3 i vilket domänen av filterparametrarna modifieras för förflyttning av filtrets poler närmare origo i det komplexa planet.
- 5Förfarande enligt något av föregående krav i vilket signalen som erhålles från källan och filtret med modifierade parametrar modifieras ytterligare av ett postfilter för undertryckning av förutbestämda frekvensområden däri.
- 6Anordning för kodning och/eller avkodning av bakgrundsljud i en digital rambaserad talkodare och/eller avkodare, innehållan 470 577 de en signalkälla ansluten till ett filter, varvid filtret definieras av en uppsättning filterparametrar för varje ram, i och för reproducering av den signal som skall kodas och/eller avkodas, vilken anordning innefattar:(a) organ (16, 34) för detektering av huruvida signalen som förs till kodaren/avkodaren representerar primärt tal eller bakgrundsljud;och (b) organ (18, 36) för begränsning av den temporala variationen mellan på varandra följande ramar och/eller domänen av åtminstone några filterparametrar i uppsättningen om signalen som förs till kodaren/avkodaren representerar primärt bakgrundsljud.
- 7Anordning enligt krav 6 i vilken den temporala variationen av filterparametrarna begränsas av ett lågpassfliter (54) som filtrerar filterparametrarna över flera ramar.
- 8Anordning enligt krav 7 i vilken den temporala variationen av filterparametrarna begränsas av ett lågpassfilter som medelvärdesbildar filterparametrarna över flera ramar.
- 9Anordning enligt krav 6, 7 eller 8 i vilken domänen av filterparametrarna modifieras i organ (52) som förflyttar filtrets poler närmare origo i det komplexa planet.
- 10Anordning enligt något av föregående krav 6-9 i vilken signalen som erhålls från källan och filtret med modifierade parametrar modifieras ytterligare genom ett postfilter (44) för undertryckning av förutbestämda frekvensområden däri. 470 577 (ap) IC/AhI (av) |(X)j?| (a?) |(Z)j/| 470 577 2/3 470 577 3/3
Independent claims10
65 paragraphs in 5 sections, as filed
(54) NAME Procedure and apparatus for coding and / or decoding of background noise (56) PUBLICATIONS Cited: - - - (57) SUMMARY:
A method and apparatus for encoding and / or decoding background noise in a digital frame-based speech encoder and / or decoder include a signal source connected to a filter, the filter being defined (12) by a set of filter parameters for each frame, and reproducing it. signal to be encoded and / or decoded. The following steps are performed: detecting (16) whether the signal passed to the encoder / decoder represents primary speech or background noise, and, if the signal represents primary background noise, limiting (18) the temporal variation between consecutive frames and / or the domain of at least some filter parameters in the set.
<img file="SE470577B_D0001.tif" />
The numbers in brackets indicate international identification code, INID code. Letters in clamps indicate international document code.
470 577
TECHNICAL FIELD
The present invention relates to a method and apparatus for encoding / decoding background noise in a digital frame-based speech encoder and / or decoder including a signal source connected to a filter, the filter being defined by a set of filter parameters for each frame, for reproducing the signal to be is coded and / or decoded.
BACKGROUND OF THE INVENTION
Many modern speech encoders belong to a large class of speech encoders called LPC (Linear Predictive Coders). Examples of encoder belonging to this class are: 4.8 Kbit / s CELP from the US Department of Defense, the RPE-LTP encoder according to the European digital cellular mobile phone system GSM, the VSELP encoder according to the corresponding American system ADC, and the VSELP encoder according to the Japanese digital cellular system JDC.
These encoders all use a source-filter concept in the signal generation process. The filter is used to form a model of the short-term spectrum of the signal to be reproduced, while the source is assumed to handle all other signal variations.
A common feature of these source-filter models is that the signal to be reproduced is represented by parameters defining the source output and filter parameters defining the filter. The term linear predictive refers to the method commonly used for estimating the filter parameters. Thus, the signal to be reproduced is partially represented by a set of filter parameters.
The method of using a combination of source filters as a signal model has been found to work relatively well for speech signals. When the user of a mobile phone is silent and the input signal is made up
470 However, 577 of background sounds, the coders known at the moment have difficulties in handling this situation, since they are optimized for speech signals. A listener on the other hand can easily be annoyed by the fact that well-known background sounds cannot be recognized because they have been malfunctioned by the encoder.
SUMMARY OF THE INVENTION
An object of the present invention is a method and a device for encoding / decoding background sounds in such a way that the background sounds are accurately encoded and reproduced.
The above objects are achieved by a method comprising the steps of:
(a) detecting whether the signal transmitted to the encoder / decoder represents primary speech or background noise; and (b) when the signal transmitted to the encoder / decoder represents primarily background noise, limiting the time variation between successive frames and / or the domain of at least some filter parameters in the set.
The device comprises:
(a) means for detecting whether the signal transmitted to the encoder / decoder represents primary speech or background noise; and (b) means for limiting the time variation between successive frames and / or the domain of at least some parameters in the set when the signal transmitted to the encoder / decoder represents primarily background noise.
470 577
BRIEF DESCRIPTION OF THE DRAWINGS
The invention and further objects and advantages thereof are best understood by reference to the following description and the accompanying drawings, in which:
Figures 1 (a) - (f) are frequency spectrum diagrams for 6 consecutive frames of the transmission function of a filter representing background noise, which filter has been estimated by a prior art encoder;
Fig. 2 is a block diagram of a speech encoder for carrying out the method according to the present invention;
Fig. 3 is a block diagram of a speech encoder for carrying out the method according to the present invention;
Figures 4 (a) - (c) are frequency spectrum diagrams similar to the diagrams of Figure 1, but for an encoder performing the method of the present invention;
Fig. 5 is a block diagram of the parameter modifier of Fig. 2; and
Fig. 6 is a flow chart illustrating the method of the present invention.
DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
In a linear predictive encoder, the synthetic number is generated from a source represented by its z-transform G (z), followed by a filter represented by its z-transform H (z), resulting in the synthetic number S (z) = G (z) H (z). Often, the filter is selected as a filter H (z) = 1 / A (z), containing only poles, where
470 577
Μ
Α (ζ) = 1 + Σ a<sub>IB</sub>z ~<sup>IA</sup> where Μ is the order of the filter.
The filter represents the short-term correlation of the speech signal. The filter parameters a<sub>n</sub> is assumed to be constant under each speech frame. Typically, the filter parameters are updated every 20 ms. If the sampling frequency is 8 kHz, each such frame corresponds to 160 samples. These samples, possibly in combination with samples from the end of the previous and the beginning of the next frame, are used for estimating the filter parameters in each frame in accordance with standard procedures. Examples of such procedures are the Levinson-Durbin algorithm, the Burg algorithm, the Cholesky decomposition (Rabiner, Schafer: Digital Processing of Speech Signals, Chapter 8, Prentice-Hall, 1978), the ScKur algorithm (Strobach: New Forms of Levinson and Schur Algorithms, IEEE SP Magazine, Jan 1991, pages 12-36), The Le Roux-Gueguen Algorithm (Le Roux, Gueguen: A Fixed Point Computation of Partial Correlation Coefficients, IEEE Transactions of Acoustics, Speech and Signal Processing, Vol. ASSP-26, No. 3, p. 257-259, 1977). It will be appreciated that a frame may consist of either more or fewer samples than mentioned above, depending on the application. In an extreme case, a frame can even consist of only one sample.
As mentioned above, the encoder is designed and optimized for handling voice signals. This has resulted in poor coding of sounds other than speech, e.g. background noise, music, etc. In the absence of a voice signal, these encoders therefore have poor performance.
Fig. 1 shows the magnitude of the filter transfer function (in dB) as a function of frequency (z = <sub>e</sub>i2nf / Fs cfg<sub>r</sub> gp§ consecutive frames in case background noise has been encoded by conventional coding methods. Although the background sound should have a uniform character over time (the background sound has a uniform
470 577 texture), the filter parameters a<sub>m</sub> to vary substantially from frame to frame when estimating under still images of only 21.25 ms (including samples from the end of the previous and the beginning of the next frame), as illustrated by the 6 frames (a) - (f) of Fig. 1. For listener at the other end, this coded sound will have a swirling or swirling character. Although the sound as a whole has a fairly uniform texture or uniform statistical properties, these short snapshots during filter estimation can provide very different filter parameters from frame to frame.
Fig. 2 shows an encoder according to the invention intended to solve the above problems.
On an input line 10, an input signal is applied to a filter estimator 12, which estimates the filter parameters in accordance with standard procedures as mentioned above. The filter estimator 12 outputs the filter parameters for each frame. These filter parameters are fed to an excitation analyzer 14, which also receives the input signal on line 10. Excitation analyzer 14 determines the best source and excitation parameters in accordance with the standard procedure. Examples of such procedures are VSELP (Gerson, Jasiuk: Vector Sum Excited Linear Prediction (VSELP), in Atal et al., Advances in Speech Coding, Kluwer Academic Publishers, 1991, pp. 69-79), TBPE (Salami, Binary Pulse Excitation: A Novel Approach to Low Complexity CELP Coding, pages 145-156 in previous reference), stochastic codebook (Campbell et al: The DoD4.8 KBPS Standard (Proposed Federal Standard 1016), p. 121-134 in previous reference), ACELP (Adoul, Lamblin: A Comparison of Some Algebraic Structures for CELP Coding of Speech, Proc. International Conference on Acoustics, Speech and Signal Processing 1987, pp. 1953-1956). These excitation parameters, the filter parameters, and the input of the line 10 are fed to a speech detector 16. The detector 16 determines whether the input signal contains primary speech or background noise. One possible detector is, for example, the voice activity detector according to the GSM system (Voice Activity Detection, GSM Recommendation 06.32,
470 577
ESI / PT 12). A suitable detector is disclosed in EP, A, 335 521 (BRITISH TELECOM PLC). The speech detector 16 generates an output signal indicating whether or not the encoded input contains substantially speech. The output along with the filter parameters is applied to a parameter modifier 18.
The parameter modifier 18, which will be further described with reference to Fig. 5, modifies the determined filter parameters in the event that a speech signal is present in the input of the encoder. If a speech signal is present, the filter parameters pass through the parameter modifier 18 without change. The possibly changed filter parameters and excitation parameters are fed to a channel encoder 20 which generates on the line 22 the bit stream transmitted over the channel.
The parameter modifier in parameter modifier 18 can be performed in several ways.
One possible modification is a bandwidth expansion of the filter. This means that the filter poles are moved towards the origin in the complex plane. Assuming that the original filter H (z) = 1 / A (z) is given by the above expression, when the poles are moved by a factor r, 0 <r <1, the bandwidth-expanded version will be defined by A (z / r) or:
A (-f) = 1 + £ <sup>r</sup> jn-1
Another possible modification is low-pass filtering of the filter parameters in the temporal domain. That is, rapid variations of the filter parameters from frame to frame are attenuated by low-pass filtering of at least some of the parameters. A special case of this method is the averaging of the filter parameters over several frames, e.g. 4-5 frames.
Parameter modifier 18 can use a combination of these two methods, e.g. perform a bandwidth expansion followed by a low pass
470 577 filtration. It is also possible to start with a low pass filtering and then add the bandwidth expansion.
In the embodiment of Fig. 2, the speech detector 16 is located after the filter estimator 12 and the excitation analyzer 14. This means that in this embodiment, the filter parameters are first estimated, after which they are modified in the absence of a speech signal. Another possibility would be to directly detect the presence / absence of a speech signal, e.g. by using two microphones, one for speech and one for background sound. In such an embodiment, it would be possible to directly modify the filter estimator to obtain appropriate filter parameters even in the absence of a speech signal.
In the above explanation of the invention, it has been assumed that the parameter modification is performed in the encoder in the transmitter. However, it will be appreciated that a similar procedure can also be performed in the receiver's decoder. This is illustrated by the embodiment of Fig. 3.
In Fig. 3, a bit stream is received from the channel on the input line 30. This bit stream is decoded by the channel decoder 32. The channel decoder 32 outputs filter parameters and excitation parameters. In this case, it is assumed that these parameters have not been modified in the transmitter's encoder. The filter and excitation parameters are fed to a speech detector 34, which analyzes these parameters to determine whether or not the signal to be reproduced by these parameters contains a speech signal. The output of the speech detector 34 is applied to a parameter modifier 36 which also receives the filter parameters. If the speech detector 34 has determined that there is no speech signal in the received signal, the parameter modifier 36 performs a modification similar to the modification performed by the parameter modifier 18 in Fig. 2. If a speech signal is present, no modification occurs. The possibly modified filter parameters and excitation parameters are applied to a speech decoder 38 which produces a synthetic output on line 40. The speech decoder 38 uses these excitation parameters to generate the above source signals and the possibly modified
470 577 the filter parameters to define the filter in the source-filter model.
As mentioned above, the parameter modifier 36 modifies the filter parameters in a similar manner to the parameter modifier 18 of FIG.
2nd The possible modifications are thus a bandwidth expansion, low pass filtering or a combination of these.
In a preferred embodiment, the decoder of Fig. 2 also contains a post filter calculator 42 and a post filter 44. A post filter in a speech decoder is used to accentuate or suppress certain parts of the spectrum of the generated speech signal. If the received signal is dominated by background noise, an improved signal can be obtained by tilting the spectrum of the output of the line 40 to reduce the amplitude of the higher frequencies. In the embodiment of FIG. 3 Therefore, the output of the speech detector 34 and the output parameters of the parameter modifier 36 are applied to the post filter 42. In the absence of a speech signal in the received signal, the post filter calculator 42 calculates a suitable slope of the spectrum of the output on the line 40 and sets the post filter 44 accordingly. The final output is obtained on line 46.
From the above description, it can be seen that the filter parameter modification can be performed either in the transmitter's encoder or in the receiver's decoder. This feature can be used to implement the parameter modification in the encoder and decoder in a base station. In this way, it would be possible to take advantage of the improved coding performance for background noise obtained by the present invention without modifying the encoder (s) in the mobile stations. When a background-containing signal is received by the base station over the land system, the parameters at the base station are modified so that already modified parameters will be received by the mobile station, in which no further action is taken. On the other hand, if the mobile station transmits a signal containing primary background noise to the base station can
470 The 577 filter parameters that characterize this signal are modified in the base station decoder and then output to the land system.
Another possibility would be to divide the filter parameter modification between the encoder at the transmitter end and the decoder at the receiver end. For example, the poles of the filter in the encoder could be partially moved closer to the origin in the complex plane and moved closer to the origin in the decoder. In this embodiment, a partial improvement in performance would be obtained in mobiles without parameter modification equipment and the entire improvement would be obtained in mobiles with this equipment.
To illustrate the improvements obtained by the present invention, Figure 4 shows the spectrum of the filter's transfer function in three consecutive frames containing primary background noise. Figures 4 (a) - (c) have been produced with the same input as Figures 1 (a) - (c). However, in Fig. 4, the filter parameters have been modified in accordance with the present invention. It is obvious that spectrum varies very little from frame to frame in Figure 4.
Fig. 5 shows a schematic diagram of a preferred embodiment of the parameter modifier 18, 36 used in the present invention. A switch 50 directs the unmodified filter parameters either directly to the output or to the blocks 52, 54 for parameter modification, depending on the control signal from the speech detector 16, 34. If the speech detector 16, 34 has detected the primary number, the switch 50 directs the parameters directly to the parameter modifier 18, 36 output. . If the speech detector 16, 34 detected primary background noise, switch 50 directs the filter parameters to an allocation block 52.
The allocation block 52 performs a bandwidth expansion on the filter parameters by multiplying each filter coefficient a<sub>B</sub>(k) with a factor r „, 0 <r <1 and k denotes the current frame, and assigning these new values to each a<sub>B</sub>(K). Before
470 577 is in the range of 0.85-0.96. A suitable value is 0.89.
The new values a<sub>B</sub>(k) from block 52 is led to assignment block 54, where the coefficients a<sub>B</sub>(k) low pass filtered in accordance with formula ga<sub>B</sub>(K) + (Ig) A<sub>B</sub>(k), where 0 <g <1 and a<sub>B</sub>(kl) denotes the filter coefficients in the previous frame. Preferably, the range is 0.92-0.995. A suitable value is 0.995. These modified parameters are then passed to the output of the parameter modifier 18, 36.
In the described embodiment, the bandwidth expansion and low-pass filtering were performed in two separate blocks. However, it is also possible to combine these two steps into a single step, according to formula a<sub>B</sub>(k) <- ga<sub>B</sub>(kl) + (lg) a<sub>B</sub>(K) r ". Furthermore, the low-pass filtering contains only the current and a previous frame. However, it is also possible to include older frames, e.g. 2-4 previous frames.
Fig. 6 shows a flow chart illustrating a preferred embodiment of the method according to the present invention. The procedure starts in step 60. In step 61, the filter parameters are estimated according to one of the above mentioned methods. These filter parameters are then used to estimate the excitation parameters in step 62. This is done according to any of the above methods. In step 63, the filter parameters and excitation parameters, and possibly the input signal itself, are used to determine whether the input signal is a speech signal or not. If the input signal is a speech signal, the procedure proceeds to the final step 66 without modifying the filter parameters. If the input signal is not a speech signal, the procedure proceeds to step 64, in which the bandwidth of the filter is expanded by moving the filter poles closer to the origin of the complex plane. Next, the low-pass filter parameters are filtered in step 65, e.g. by forming the mean of the current filter parameters from step 64 and filter parameters from previous signal frames. Finally, the procedure proceeds to the final step 66.
470 577
In the above description, the filter coefficients a<sub>m</sub> to illustrate the method of the present invention. However, it will be appreciated that the same basic ideas can be applied to other parameters that describe or are related to the filter, e.g. filter reflection coefficients, log area ratios (log area ratios), polynomials, autocorrelation functions (Rabiner, Schafer: Digital Processing of Speech Signals, Prentice-Hall, 1978), arcsin of the reflection coefficients (Gray, Markel: Quantization and Bit Allocation in speech Processing, IEEE Transactions on Acoustics, Speech and Signal Processing, vol ASSP-24, No. 6, 1976), 1-line Spectrum Pairs (Soong, Juang: Line Spectrum Pair (LSP) and Speech Data Compression, Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing 1984, p. 1.10.11.10.4).
Another modification of the described embodiment would be an embodiment with no mail flutter in the receiver. Instead, the gradient of spectrum would be obtained already in the modification of the filter parameters, either in the transmitter or in the receiver. This can be done, for example, by varying the so-called reflection coefficient 1.
Those skilled in the art will recognize that various modifications and changes can be made to the present invention without departing from its basic idea and framework, as defined by the appended claims.
470 577
Contents5
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
33 members in 22 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 9300290 | Sweden | A | |
| SE19930000290 | – | – | – |
Members33
| Document | Office | Kind | |
|---|---|---|---|
| SE9300290D0 | Sweden | D0 | |
| CA2133071A1 | Canada | A1 | |
| SE9300290L | Sweden | L | |
| WO9417515A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU5981394A | Australia | A | |
| SE470577BThis record | Sweden | B | |
| NO943584D0 | Norway | D0 | |
| NO943584L | Norway | L | |
| FI944494A | Finland | A | |
| FI944494L | Finland | L | |
| EP0634041A1 | European Patent Office (EPO) | A1 | |
| KR950701113A | Republic of Korea | A | |
| CN1101214A | China | A | |
| JPH07505732A | Japan | A | |
| TW262618B | Taiwan Province of China | B | |
| AU666612B2 | Australia | B2 | |
| NZ261180A | New Zealand | A | |
| US5632004A | United States of America | A | |
| SG46992A1 | Singapore | A1 | |
| PH31235A | Philippines | A | |
| EP0634041B1 | European Patent Office (EPO) | B1 | |
| AT168809T | Austria | T | |
| ATE168809T1 | Austria | T1 | |
| DE69411817D1 | Germany | D1 | |
| DK0634041T3 | Denmark | T3 | |
| ES2121189T3 | Spain | T3 | |
| DE69411817T2 | Germany | T2 | |
| BR9403927A | Brazil | A | |
| CN1044293C | China | C | |
| KR100216018B1 | Republic of Korea | B1 | |
| HK1015183A1 | Hong Kong, China | A1 | |
| NO306688B1 | Norway | B1 | |
| MY111784A | Malaysia | A |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent has lapsedLapsedNUG | NUG | |
| Patent in forceNAL | NAL |
Numbers
- Publication, DOCDB
- 470577
- Publication, EPODOC
- SE470577
- Application
- 9300290
- Application, DOCDB
- 9300290
- Application, EPODOC
- SE19930000290
Titles2
- English
- Method and apparatus for encoding and / or decoding background noise
- Swedish
- Förfarande och anordning för kodning och/eller avkodning av bakgrundsljud
Classification
- CPC, 2
- G10L25/78
- G10L19/028
- IPC, 1
- G10L25 78
