Method and apparatus in coding digital information
Abstract
A speech encoder (100) receives speech signals (S) which are encoded and transmitted on a communication channel (120). Silence in the speech is utilized by a data encoder (101) to transmit data on the speech frequency band via the channel (120). A signal classifier (103) switches between the encoders (100, 101). The speech encoder has synthesis filter (115) with state variables in a delay line, predictor adaptor (116), gain predictor (113, 114) and excitation codebook (112). The data encoder (101) has delay line with state variables stored and updated in a buffer (192). On switching (103, 102, 193) from data to speech, the buffer state variables are fed into the synthesis filter delay line via an input (144) for smooth transition in the speech encoding. Coefficient values in the synthesis filter (115) and an excitation signal (ET(1...5)) are generated. Thereby a buffer in the gain predictor (113, 114) is preset and its predictor coefficients and gain are generated. The incoming speech signal (S) newly detected is encoded (CW) by the values generated in the speech encoder (100), which is successively adapted. The receiver side has corresponding speech and data decoders.

Term
No projected expiry on record.
- Priority and filed
- Granted
- Today
24 claims: 7 independent, 17 dependent
- 1PATENTKRAV 1. Förfarande i ett transmissionssystem för att överföra signaler över en kommunikationskanal (120), vilket system omfattar :- en första bakåtkopplad adaptiv kodare (100) innefattande ett syntetiseringsfilter (115) med dels element (140) för tillståndsvärden (SB(1...105)), dels koefficientelement (141) för prediktorkoefficienter (A 2 ...A 51 );- en andra bakåtkopplad adaptiv kodare (101) med element för tillståndsvärden (VSB(1...105)) ;och - en omställningskrets (103) för omställning mellan de nämnda första och andra kodarna (100, 101) vid val av den av kodarna som skall utnyttjas vid överföringen;varvid förfarandet omfattar: - överförande av signaler via den andra kodaren (101) och lagring av dess tillståndsvärden (VSB(1...105)) i en buffert (192) ;- omställning för överföring via den första kodaren (100) med hjälp av omställningskretsen (103);- förinställande av åtminstone en andel av den första kodarens (100) tillståndsvärden (SB(1...105)) med de nämnda lagrade tillståndsvärdena (VSB(1...105));- framtagande av åtminstone en andel av prediktorkoefficienterna (Α 2 ·..Α 51 ) i den första kodaren (100);och - genererande av en utsignal (SD) från syntetiseringsfiltret i beroende av de framtagna prediktorkoefficienterna (Α 2 ·..Α 51 ).
- 2Förfarande enligt patentkrav 1, varvid den andra kodaren (101) har koefficientelement för prediktorkoefficienter (B2...B51), svarande mot koefficientelementen (141) hos den första kodaren (100), och förfarandet ytterligare omfattar:- lagring av prediktorkoefficienterna (B2....B51) hos den andra kodaren (101) i den nämnda bufferten (192) ;och - framtagandet av prediktorkoeff icienterna (Α 2 ·..Α 51 ) första kodaren (100) utföres genom överföring av de lagrade prediktorkoefficienterna (B2...B51) syntetiseringsfiltrets (115) koefficientelement (141). i den nämnda till 504 010
- 3Förfarande enligt patentkrav 1, varvid framtagandet av prediktorkoefficienterna (A 2 ...A gl ) i den första kodaren (100) utföres med hjälp av dess förinställda tillståndsvärden (SB(1...105)).
- 4Förfarande enligt patentkrav 3, varvid endast en andel (A 2 . .A 11 ) av prediktorkoefficienterna (Α 2 ·..Α 51 ) framtages.
- 5Förfarande enligt patentkrav 1, 2, 3 eller 4, vilket ytterligare omfattar följande förfarandesteg:- genererande av vektorer (ZINR(1...5)), innefattade i ett svar på en nollvärdesinsignal (0) till syntetiseringsfiltret (115), med hjälp av tillstånden (SB(1...105)) och prediktorkoeff icienterna (Α 2 ·..Α 51 ) hos syntetiseringsfiltret (115);- genererande av vektorer (ZSTR(1...5)) för nolltillståndssvar genom att subtrahera vektorerna (ZINR(1...5)) för svaret på nollvärdesinsignalen (”0”) från de motsvarande tillståndsvärdena, uppdelade som tillståndsvektorer (SB (1...5)), för syntetiseringsfiltret (115);och genererande av en exciteringssignal (ET(1...5)) för syntetiseringsfiltret (115) med hjälp av vektorerna (ZSTR(1. ..5)) för nolltillståndssvaret.
- 6Förfarande enligt patentkrav 5, varvid den första kodaren (100) har en förstärkningsprediktor (134) med dels element (150) för tillståndsvärden (SBLG), dels koefficientelement (151) för prediktorkoefficienter (GP 2 ....GP 11 ), vilket förfarande ytterligare omfattar:- generering och förinställning av förstärkningsprediktorns (134) tillståndsvärden (SBLG) med utnyttjande av den genererade exciteringssignalen (ET(1...5)) ;- generering av förstärkningsprediktorns (134) koefficienter (GP 2 ...GP 1;l ) med hjälp av förstärkningsprediktorns tillståndsvärden (SBLG);och - genererande av en predikterad förstärkningsfaktor (GAIN') för syntetiseringsfiltrets (115) första exciteringssignal (ET(1...5)) efter ett initieringsskede för den första kodaren (100). 504 010 A?
- 7Förfarande i ett transmissionssystem för att mottaga signaler över en kommunikationskanal (120), vilket system omfattar :- en första bakåtkopplad adaptiv avkodare (200) innefattande ett syntetiseringsfilter (215) med dels element (140) för tillståndsvärden (SB(1...105)), dels koefficientelement (141) för prediktorkoefficienter (Α 2 ·..A 51 );- en andra bakåtkopplad adaptiv avkodare (290) med element för tillståndsvärden (VSB(1...105));och - en omställningskrets (103) för omställning mellan de nämnda första och andra avkodarna (200, 290) vid val av den av avkodarna som skall utnyttjas vid mottagandet;varvid förfarandet omfattar: - mottagande av signaler via den andra avkodaren (101) och lagring av dess tillståndsvärden (VSB(1...105)) i en buffert (292) ;- omställning för mottagande via den första avkodaren (200) med hjälp av omställningskretsen (103);- förinställande av åtminstone en andel av den första avkodarens (200) tillståndsvärden (SB(1...105)) med de nämnda lagrade tillståndsvärdena (VSB(1...105));- framtagande av åtminstone en andel av prediktorkoefficienterna (Α 2 ·..Α 51 ) i den första avkodaren (200);och - genererande av en utsignal (SD) från syntetiseringsfiltret i beroende av de framtagna prediktorkoefficienterna (Α 2 ·..Α 51 ).
- 8Förfarande enligt patentkrav 7, varvid den andra avkodaren (290) har koefficientelement för prediktorkoeff icienter (B2...B51), svarande mot koefficientelementen (141) hos den första avkodaren (200), och förfarandet ytterligare omfattar:- lagring av prediktorkoefficienterna (B2....B51) hos den andra avkodaren (290) i den nämnda bufferten (292);och - framtagandet av prediktorkoeff icienterna (Α 2 ·..Α 51 ) i den första avkodaren (200) utföres genom överföring av de nämnda lagrade prediktorkoefficienterna (B2...B51) till syntetiseringsfiltrets (215) koefficientelement (141). 504 010
- 9Förfarande enligt patentkrav 7, varvid framtagandet av prediktorkoefficienterna (Α 2 ·..Α 51 ) i den första avkodaren (200) utföres med hjälp av dess förinställda tillståndsvärden (SB(1...105)).
- 10Förfarande enligt patentkrav 9, varvid endast en andel (Α 2 ·..Α 11 ) av prediktorkoefficienterna (A 2 .A 51 ) framtages.
- 11Förfarande enligt patentkrav 7, 8, 9, eller 10, vilket ytterligare omfattar följande förfarandesteg:- genererande av vektorer (ZINR(1...5)), innefattade i ett svar på en nollvärdesinsignal (O) till syntetiseringsfiltret (215), med hjälp av tillstånden (SB(1...105)) och prediktorkoeff icienterna (Α 2 ·..Α 51 ) hos syntetiseringsfiltret (215);- genererande av vektorer (ZSTR(1...5)) för nolltillståndssvar genom att subtrahera vektorerna (ZINR(1...5)) för svaret på nollvärdesinsignalen (0) från de motsvarande tillståndsvärdena, uppdelade som tillståndsvektorer (SB (1...5)), för syntetiseringsfiltret (215);och genererande av en exciteringssignal (ET(1...5)) för syntetiseringsfiltret (215) med hjälp av vektorerna (ZSTR(1.. .5)) för nolltillståndssvaret.
- 12Förfarande enligt patentkrav 11, varvid den första avkodaren (200) har en förstärkningsprediktor (134) med dels element(150) för tillståndsvärden (SBLG), dels koefficientelement (151) för prediktorkoefficienter (GP 2 ·.. .GP 1;L ), vilket förfarande ytterligare omfattar:- generering och förinställning av förstärkningsprediktorns (134) tillståndsvärden (SBLG) med utnyttjande av den genererade exciteringssignalen (ET(1...5));- generering av förstärkningsprediktorns (134) koefficienter (GP 2 ...GP 11 ) med hjälp av förstärkningsprediktorns tillståndsvärden (SBLG);och - genererande av en predikterad förstärkningsfaktor (GAIN 1 ) för syntetiseringsfiltrets (215) första exciteringssignal (ET(1...5)) efter ett initieringsskede för den första avkodaren (200). i 504 010 ίΐ
- 13Anordning i ett transmissionssystem för att överföra signaler över en kommunikationskanal (120), vilken anordning omfattar :- en första bakåtkopplad adaptiv kodare (100) innefattande ett syntetiseringsfilter (115) med dels element (140) för tillståndsvärden (SB(1...105)), dels koefficientelement (141) för prediktorkoefficienter (A 2 ·..A 51 );- en andra bakåtkopplad adaptiv kodare (101) med element för tillståndsvärden (VSB(1...105));- en omställningskrets (103) med omkopplare (98,102) för inkoppling av en av de nämnda första och andra kodarna (100, 101) till kommunikationskanalen (120);- en buffert (192) för lagring av den andra kodarens (101) tillståndsvärden (VSB(1...105)) vid överföring av signaler via denna andra kodare;- anordning (193, 144) för inmatande av åtminstone en andel av de nämnda lagrade tillståndsvärdena (VSB(1. .. 105)) i elementen (140) för tillståndsvärden (SB(1...105)) hos den första kodaren (100) vid inkoppling för överföring via denna första kodare (100);- anordning (116;49, 50, 51 146), ansluten till ingångar (139) hos koefficientelementen (141), för framtagande av åtminstone en andel av prediktorkoefficienterna (Α 2 ·. ,A 51 ) i den första kodaren (100);och - anordning (142, 143) ansluten till koefficientelementen (141) för att generera en utsignal (SD) från syntetiseringsfiltret (115).
- 14Anordning enligt patentkrav 13, varvid den andra kodaren (101) har koefficientelement för prediktorkoefficienter (B2...B51) svarande mot koefficientelementen (141) hos den första kodaren (100);- den nämnda bufferten (192) är anordnad att lagra prediktorkoefficienterna (B2....B51) hos den andra kodaren (101);och - anordningen för framtagandet av prediktorkoefficienterna (Α 2 ·..Α 51 ) i den första kodaren (100) omfattar medel (193, 139) för att överföra de nämnda lagrade prediktorkoefficienterna (B2...B51) till syntetiseringsfiltrets (115) koefficientelement (141). 504 010
- 15Anordning enligt patentkrav 13, varvid anordningen för framtagandet av prediktorkoefficienterna (Α 2 ·..Α 51 ) omfattar medel (116;48, 49, 50, 51, 146) för att, med hjälp av de nämnda lagrade tillståndsvärdena (VSB(1...105)) i elementen (140) för tillståndsvärden (SB(1...105)) hos den första kodaren (100), generera prediktorkoefficienterna (A 2 ...A 51 ).
- 16Anordning enligt patentkrav 15, varvid de nämnda medlen (116;48, 49, 50, 51, 146) för att generera prediktorkoefficienterna är anordnade att generera endast en andel (Α 2 ·..Α 11 ) av prediktorkoef f icienterna (Α 2 ·..Α 51 ).
- 17Anordning enligt något av patentkraven 13-16 omfattande:medel (147) för att generera vektorer (ZINR(1...5)), innefattade i ett svar på en nollvärdesinsignal (0) till syntetiseringsfiltret (115), med hjälp av tillstånden (SB(1...105)) och prediktorkoefficienterna (Α 2 ·..Α 51 ) hos syntetiseringsfiltret (115);- medel (148) för att generera vektorer (ZSTR(1...5)) för nolltillståndssvar omfattande en subtraherare (148) vilken subtraherar vektorerna (ZINR(1...5)) för svaret på nollvärdesinsignalen (0”) från de motsvarande tillståndsvärdena, uppdelade som tillståndsvektorer (SB(1.. .5)) , för syntetiseringsfiltret (115);och - medel (149) för att generera en exciteringssignal (ET(1...5)) för syntetiseringsfiltret (115) med hjälp av vektorerna (ZSTR(1...5)) för nolltillståndssvaret.
- 18Anordning enligt patentkrav 17, varvid den första kodaren (100) har en förstärkningsprediktor (134) med dels element (150) för tillståndsvärden (SBLG), dels koefficientelement (151) för prediktorkoefficienter (GP 2 ... •GP 11 ) , vilken anordning omfattar:medel (152, 155) för generering och förinställning av förstärkningsprediktorns (134) tillståndsvärden (SBLG) med utnyttjande av den genererade exciteringssignalen (ET(1...5));- medel, anslutna dels till elementen (150) för tillståndsvärden, dels till koefficientelementen (151), för att generera förstärkningsprediktorns (134) koefficienter (GP 2 — GP n) med 3/ 504 010 hjälp av förstärkningsprediktorns tillståndsvärden (SBLG);och - medel (153, 156) för att generera en predikterad förstärkningsfaktor (GAIN') för syntetiseringsfiltrets (115) första exciteringssignal (ET(1...5)) efter ett initieringsskede för den första kodaren (100).
- 19Anordning i ett transmissionssystem för att mottaga signaler över en kommunikationskanal (120), vilken anordning omfattar:- en första bakåtkopplad adaptiv avkodare (200) innefattande ett syntetiseringsfilter (215) med dels element (140) för tillståndsvärden (SB(1...105)) , dels koefficientelement (141) för prediktorkoefficienter (A 2 ·..A 51 );- en andra bakåtkopplad adaptiv avkodare (290) med element för tillståndsvärden (VSB(1...105));- en omställningskrets (103) med omkopplare (203,198) för inkoppling av en av de nämnda första och andra avkodarna (200, 290) till kommunikationskanalen (120);- en buffert (292) för lagring av den andra avkodarens (290) tillståndsvärden (VSB(1...105)) vid mottagande av signaler via denna andra avkodare;- anordning (293, 145) för inmatande av åtminstone en andel av de nämnda lagrade tillståndsvärdena (VSB(1... 105)) i elementen (140) för tillståndsvärden (SB(1... 105)) hos den första avkodaren (200) vid inkoppling för mottagande via denna första avkodare (200);- anordning (116;49, 50, 51 146), ansluten till ingångar (139) hos koefficientelementen (141) för framtagande av åtminstone en andel av prediktorkoefficienterna (Α 2 ·..Α 51 ) i den första avkodaren (200);och - anordning (142, 143) ansluten till koefficientelementen (141) för att generera en utsignal (SD) från syntetiseringsfiltret (215).
- 20Anordning enligt patentkrav 19, varvid den andra avkodaren (290) har koefficientelement för prediktorkoefficienter (B2...B51) svarande mot koefficientelementen (141) hos den första avkodaren (200);- den nämnda bufferten (292) är anordnad att lagra prediktorkoef ficienterna (B2....B51) hos den andra avkodaren (290);och 504 010 - anordningen för framtagandet av prediktorkoef ficienterna (Α 2 ·..Α 51 ) i den första avkodaren (200) omfattar medel (193, 139) för att överföra de nämnda lagrade prediktorkoefficienterna (B2...B51) till syntetiseringsfiltrets (215) koefficientelement (141).
- 21Anordning enligt patentkrav 19, varvid anordningen för framtagandet av prediktorkoefficienterna (Α 2 ·..Α 51 ) omfattar medel (116;48, 49, 50, 51, 146) för att, med hjälp av de nämnda lagrade tillståndsvärdena (VSB(1...105)) i elementen (140) för tillståndsvärden (SB(1...105)) hos den första avkodaren (200), generera prediktorkoefficienterna (Α 2 ·..Α 51 ).
- 22Anordning enligt patentkrav 21, varvid de nämnda medlen (116;48, 49, 50, 51, 146) för att generera prediktorkoefficienterna är anordnade att generera endast en andel (Α 2 ·..Α 11 ) av prediktorkoef ficienterna (A 2 ...a 51 ).
- 23Anordning enligt något av patentkraven 19-22 omfattande:medel (147) för att generera vektorer (ZINR(1...5)), innefattade i ett svar på en nollvärdesinsignal (0) till syntetiseringsfiltret (215), med hjälp av tillstånden (SB(1...105)) och prediktorkoefficienterna (A 2 ...A 51 ) hos syntetiseringsfiltret (215);- medel (148) för att generera vektorer (ZSTR(1...5)) för nolltillståndssvar omfattande en subtraherare (148) vilken subtraherar vektorerna (ZINR(1...5)) för svaret på nollvärdesinsignalen (’O) från de motsvarande tillståndsvärdena, uppdelade som tillståndsvektorer (SB(1. ..5)), för syntetiseringsfiltret (215);och - medel (149) för att generera en exciteringssignal (ET(1...5)) för syntetiseringsfiltret (215) med hjälp av vektorerna (ZSTR(1...5) ) för nolltillståndssvaret.
- 24Anordning enligt patentkrav 23, varvid den första avkodaren (200) har en förstärkningsprediktor (134) med dels element (150) för tillståndsvärden (SBLG), dels koefficientelement (151) för prediktorkoefficienter (GP 2 ....GP 11 ), vilken anordning omfattar:504 010 - medel (152, 155) för generering och förinställning av förstärkningsprediktorns (134) tillståndsvärden (SBLG) med utnyttjande av den genererade exciteringssignalen (ET(1...5));- medel, anslutna dels till elementen (150) för tillståndsvärden, 5 dels till koefficientelementen (151), för att generera förstärkningsprediktorns (134) koefficienter (GP 2 ...GP 1X ) med hjälp av förstärkningsprediktorns tillståndsvärden (SBLG);och - medel (153, 156) för att generera en predikterad förstärkningsfaktor (GAIN·) för syntetiseringsfiltrets (215) första exci- 10 teringssignal (ET(1...5)) efter ett initieringsskede för den första avkodaren (200). 1/8 504 010 Insignal l· Sändare Mottagare
Independent claims24
141 paragraphs in 9 sections, as filed
(54)
PATENT HOLDER Telefonaktiebolaget LM Ericsson, 126 25 Stockholm SE
INVENTOR'S OFFICE NAME (56)
CALLED PUBLICATIONS:
Rudi Hofmann, Forchheim DE
Lövgren T
Method and apparatus for predictive coding of speech data signals and
US A 5,339,384 (395: 2.2), US A 5,313,554 (395: 2.28)
CCITT Rec. G. 728, Nov. 1992, Coding of Speech at 16 kbit / s using low delay code excited linear prediction pp. 17, 18, 24-29
IEEE Global Telecommunications Conference, vol.3, dec. 1991, A Backward Adaptive 8 kbit / s Speech Coder using Conditional Pitch Prediction ”, A. Kataoka, pp. 1889-1893 (57) SUMMARY:
A speech encoder (100) or eight speech signals (S) is encoded (CW) to be transmitted on an audio communication edge (120). The silence 1 speech signal is used to transmit an encoder (101) to the data in the speech frequency band of the channel (120). A signal rating device (102) provides the cellan encoder (100,101). The speech encoder has synthesizing chips (115) by state warning labels in a delay line, a predictor adaptation circuit (116), a drying predictor (113,114) and an excitation codebook (112). The data encoder (101) has a delay line that state variables are stored and updated in a buffer (192). In oases (103, 102, 193, 102) from data coding to speech coding, the states of the buffer (192) to the delay line of the synthesizing filter are entered via an input (144) to obtain a smooth transition in the coding of spoken information. Coefficient values in the synthesis filter (115) are calculated and a value of a synthesis signal (£ 1 (1, 5)) is generated. By means of this, a buffer is pre-set in the gain predictor (113, 114), whose prediction coefficients and gain are generated. , which is subsequently adapted. Get the music receiver page there are speech decoders and data decoders »corresponding to the two encoders (100, 101) in the figure.
<img file="SE504010C2_D0001.tif" />
The numbers in brackets indicate International Identification Code, INID code. Letters in clips indicate International Document Code.
504 010
TECHNICAL FIELD
The invention relates to speech coding techniques and general processing of spoken information. More specifically, it relates to speech coding methods based on algorithms for analysis by synthesis in combination with backward adaptation technology.
BACKGROUND OF THE ART
A system based on analysis by synthesis and backward adaptation was used in, for example, a speech encoder / decoder called Low-Delay Code Excited Linear Prediction (LD-CELP), recently standardized by the International Telecommunication Union (ITU) in the publication CODING OF SPEECH AT 16 kbit / s USING LOWDELAY CODE EXCITED LINEAR PREDICTION, copyright ITU 1992, Recommendation G.728. This speech signal compression algorithm has gradually become well known by voice coding experts around the world.
Digital networks are used to transmit digitally encoded signals. In the past, it was mainly voice signals that were transmitted. Nowadays, it is the data traffic caused by a widespread use of electronic mail networks that is growing more and more throughout the world. From an economic point of view, the number of connected users must be maximized without the network being blocked by traffic. One consequence of this is that speech compression algorithms have been developed and especially optimized by using noise-masking effects. Unfortunately, these coding algorithms are not very well suited for transmitting data signals within the speech frequency band. One idea here is to also use algorithms for signal classification and to use algorithms for compression of data signals within the voice frequency band (Voiceband Data Signal Compression, VDSC) when data signals are detected. Currently, a transmission system with 16 kb / s is being standardized which utilizes this idea, a so-called
504 010
Digital Circuit Multiplication Equipment (DCME) system. The said LD-CELP encoder / decoder will be used to transmit the spoken information while a new coding algorithm for transmitting data in the voice frequency band is under development within ITU.
In practical applications, the signal classification algorithms can make misinterpretations, resulting in a more or less tight switching between different coding methods. If the next coding method were to always start from its reset state, this would probably not be critical when transmitting data signals on the voice frequency band. However, when spoken information is to be transmitted more generally, this should result in rather disruptive effects.
To overcome this problem in 16 kb / s DCMC systems, it was proposed to provide the LD-CELP architecture also for compressing data signals in the speech frequency band. Only the bit rate would be increased, for example, by providing larger waveform codebooks to ensure a sufficiently accurate quantification. With such a method, continuous shaping of the signal in the schedule would be guaranteed upon switching from one coding mode to the other.
The disadvantage of this solution is twofold: On the one hand, the amount of calculation would be significantly increased during transmission at higher bit rates. This makes realizations less attractive, since a conventional LD-CELP encoder / decoder requires almost all of the computational capacity of the digital signal processors currently offered in the market. On the other hand, it is very likely that the coding of data signals in the speech frequency band can be made much more efficient with specially optimized architectures, resulting in bit rates below 40 kb / s or higher performance. To date, the required bit rate seems to be 40 kb / s for VDSC algorithms. It is trivial to mention that this switching problem also occurs if pre-existing signal compression algorithms are used in combination with LD-CELP type encoder / decoders. Known systems, for example, use the algorithms in
504 010 in accordance with ITU recommendations G.711 (64 kb / s) or G.726 (32 kb / s or 40 kb / s) when transmitting data signals in the speech frequency band.
In this context, one may mention a coding algorithm, called ADPCM, whose structure has similarities to the LD-CELP algorithm in that it contains forwarded error correction. Hereby reference is made to a document Digital Communications by Simon Haykin, John Wiley & Sons, 1988.
U.S. Patent 5,233,660 discloses a small delay digital speech coder and decoder based on coded linear prediction (LD-CELP). The coding includes partly backward adaptive adjustment of both codebook gain and parameters in short-term synthesis filter, and partly forward-adaptive adaptation of parameters in long-term synthesis filter. An efficient, short-term delayed derivation and quantization of pitch parameters allow a total delay which is a fraction of the previous coding delay at uniform speech quality.
U.S. Patent No. 5,339,384 also discloses a CELP encoder for voice and audio transmission. The encoder is adapted for short-term delayed coding by performing spectral analysis of a portion of a previous simulated-decoded frame to determine a synthesis filter of a much higher order than is commonly used for decoding synthesis and then transmits only the index of the vector gives the weakest internal error signal. Modified perceptual weight parameters and a new use of post-filtration improve the tandem operation of a number of encodings and decodings while maintaining high quality of rendering.
The patent US 5,228,076 is also of interest, since it involves the use of the above-mentioned coding algorithm ADPCM.
504 010
DISCLOSURE OF THE INVENTION
During the transmission of spoken information, a significant part of the transmission time is pure silence. During these silent intervals, it is possible to use the transmission link for data transfer. Data and spoken information are coded with different codes and a problem is to switch between different encoders and to avoid discontinuities in the number after switching. This is especially the case when backward adaptive coding algorithms are used. Even when transmitting other types of information than speech, time intervals can be used which can be used to transmit alternative information on the same channel.
Discontinuities in the output signal can be eliminated if the states of the coding algorithm to be activated are preset with the same values as if this coding algorithm had already been active before. The problem with this is that the generation of corresponding initial values for the state variables is not trivial when the encoder / decoder is based on reverse feedback adaptive algorithms, as is the case for LD-CELP-type coding algorithms. The predictor coefficients depend on the previously quantized output signal, for example coefficients in a synthesized filter in the LD-CELP-type coding algorithm. Additional states and predictor coefficients depend on the preceding quantized excitation signal, for example, as coefficients in a gain predictor depend on an excitation signal for a synthesizer filter in said LDCELP encoder / decoder. More specifically, the problem is that this previous excitation signal is not available when the encoder / decoder is to be switched on. Although the state variables can be recovered, tremendous instantaneous signal processing power would be required at the tip point when the encoder / decoder is to be initiated. This signal processing would overload all currently available digital signal processors (DSPs) in the market.
The present invention discloses the technique required to recover the state variables and shows the ways in which the
504 The required signal processing or data power is reduced to allow for practical implementations. The problem is solved by using output samples from an encoder / decoder that is switched off to preset the states in the encoding algorithm for a parallel encoder / decoder which is turned on.
More in detail, the problem is solved by generating coefficient values from the preset state variables and restoring a signal sequence (vector) from these coefficient values and the signal sequence. This signal sequence (vector) is utilized to directly generate the decoded output signal, for example spoken information, in the decoder and also in the encoder and is normally generated successively during transmission. By restoring the signal sequence (vector), the encoder / decoder is quickly run.
In a simplified embodiment, the coefficient values are not generated in the encoder / decoder but are transmitted directly from the parallel encoder / decoder which is switched off. The transmitted coefficients were used to restore the signal sequence (vector).
It is an object of the present invention to provide suitable devices and methods that allow reverse feedback adaptive coding algorithms, such as the LD-CELP type speech encoder / decoder, to maintain a continuous design of the reconstructed output. Modifications are also presented such that the load map signal processing around the initialization phase can be kept at a suitably low level.
The advantage of the invention is that only a moderate signal processing power is required when switching to an encoder / decoder and switching can be performed without difficult discontinuities in the output signal. When spoken information and data are transmitted on the same communication channel, no inconvenient effects are observed in the spoken information when switching to the speech coder.
504 010
DESCRIPTION
Fig. 1 illustrates in a high-level block diagram a transmission system comprising two different encoders / decoders used for different purposes.
Fig. 2 illustrates in a high-level block diagram a general speech coding algorithm based on backward adaptation technique.
Fig. 3a shows a block diagram of an LD-CELP encoder.
Fig. 3b shows a block diagram of an LD-CELP decoder.
Fig. 4 illustrates in more detail the contents of the local decoder shown in Fig. 2.
Fig. 5 illustrates, in a low-level block diagram, the backward adaptation of the synthesizing filter and the corresponding predictor coefficients.
Fig. 6 illustrates in a low level block diagram the backward adaptation of the gain predictor and the corresponding predictor coefficients.
Figures 7a and b illustrate the procedure for performing the operations of the synthesizing filter in the LD-CELP type speech encoder.
Fig. 8 shows in a flow diagram the method of "heating" the states in an LD-CELP-type speech encoder.
Fig. 9 shows a block diagram for generating an excitation vector.
PREFERRED EMBODIMENT
In order to describe the preferred embodiment of the invention, it is expedient to explain certain details of backward-coupled adaptive speech coding algorithms as used in, for example, the LD-CELP algorithm.
Fig. 1 illustrates, in the form of a block diagram, a transmission system with different coding algorithms for speech signals and data signals on the speech frequency band. On the transmitter side there is an encoder 100 for LD-CELP encoding of spoken information and a data encoder 101 according to the VDSC algorithm. An incoming line 99 is connected to the encoder through a switch 98 and the output of the encoder is connected to a communication channel 120 through a switch 102. A signal rating device 103 is connected to incoming line 99 and controls switches 98 and 102. On the receiver side, there is a decoder for speech decoding and a data decoder 290. The decoders are connected to the communication channel through a switch 203 and their outputs are connected to an outgoing line 219. through a switch 198. Signal rating device 103 is connected to switches 203 and 198 through a separate signaling channel 191 and controls these switches in parallel with the switches on the transmitter side. A buffer 192 is connected to an auxiliary output of the data encoder 101 and is connected to an input 144 of the speech encoder 100 via a switch 193. This switch is activated by the signal classifier 103. On the receiver side there is a corresponding buffer 292 and a switch 293. As an example of one embodiment, speech encoder 100 is of the LD-CELP type and is used when encoding spoken information, while another encoding algorithm according to VDSC is used in data encoder 101 when data signals on the speech frequency band are available. The information about the currently used compression algorithm is usually sent from the transmitter to the receiver over the separate signaling channel 191. The invention is related to the situation when the coding algorithm VDSC has been active and the signal classification device has just detected the presence of spoken information. This results in the LDCELP-type speech encoders 100 and 200 being activated.
Fig. 2 illustrates at a very high level the basic principle of a reverse feedback adaptive speech coding algorithm such as
504 010 is used, for example, in the encoder LD-CELP. On the transmitter side there is a search unit 130 for a codebook as well as a local decoder 95. The local decoder 95 is connected to an input of the code register, which also has an input for an incoming signal (Input signal). An output from the search unit to the code register is connected to the input of the local decoder. The transmitter transmits a code vector CW to the receiver. On the receiver side, a local decoder 96 is connected to an after filter 217, which in turn is connected to the output 219. Both on the transmitter and receiver side, the quantized output signal is reconstructed in a block 'Local Decoder' 95 and 96 respectively. of the preceding reconstructed signal to be able to find optimized parameters for a current segment of spoken information to be encoded, as will be described in detail below.
Fig. 3a shows a simplified block diagram of the LD-CELP encoder 100 and also the VDSC encoder 101. The switches 102 and 98 for selecting encoder 100 or 101 and the signal rating circuit 103 controlling the switches 98 and 102 are also shown as well as the buffer 192 and switch 193. The incoming signal S is connected to the signal rating circuit 103 and to the LD-CELP encoder 100. The LD-CELP encoder comprises a PCM converter 110 which is connected to a vector buffer 111. Encoder 100 also includes a first excitation codebook 112 which is connected to a first gain setting unit 113 with a first gain switching adapter 114. The output of the first gain setting unit 113 is connected to a first synthesizing filter 115 having input 144 and is connected to a reverse gain switching circuit 116th The output of the synthesis filter 115 is connected to a differential circuit 117 to which also the vector buffer 111 is connected. The differential circuit 117 is in turn connected to a perceptual weight filter 118, the output of which is connected to a circuit 119 which calculates errors according to the square mean method. The latter is connected to the exit code register and to the communication channel 120 which connects LD-CELP encoder 100 to LD-CELP decoder 200 on the receiver side of the transmission, as shown in Fig. 3b.
504 010
Fig. 3b shows VDSC encoder 290 with switches 198 and 203 and also buffer 292 with switch 293. The LD-CELP decoder comprises a second excitation codebook 212 connected to communication channel 120 and to a second gain setting circuit 213 with a second reverse switching adapter 214. The second gain setting circuit 213 is connected to a second reverse predictor adapter circuit 216. An adaptive after filter 217 is connected with its input to the synthesizing filter 215 and with its output to a PCM converter 218 with an A-layer or μ-layer output.
The LD-CELP encoder works as follows. The signal S, which is converted according to PCM A layer or μ layer, is converted to uniform PCM signal in converter 110. The input signal is then divided into blocks of five consecutive input samples, called input signal vectors, and stored in weight or buffer 111. For each input vector allows the encoder to pass through each of 128 proposed codebook vectors stored in code register 112 through the first gain setting unit 113. In this unit, each of the vectors is multiplied by eight different gain factors and the resulting vectors 1024 are passed through the first synthesizing filter 115. An error value generated in the differential circuit 117 between each of the input weights and the 1024 the proposed vectors, are frequency weighted in the weighting filter 118 and squared and averaged in the circuit 119. The encoder identifies the best code vector, i.e. the vector that minimizes the squared and averaged error of one of the input vectors and a 10-bit codebook index CW of the best code vector is transmitted to decoder 200 over channel 120. The best code vector may also pass through the first gain setting unit 113 and the first synthesizing filter 115 for to establish the correct filter memory to be prepared for the encoding of the next ins ignal vector. The identification of the best code vector and the update of the filter memory is repeated for all the ins vector. The coefficients of the synthesizer filter and gain in the first gain setting unit are periodically updated by the adapter circuits 116 and 114, respectively.
504 010 in a backward adaptive manner, based on the previously quantized signal and gain-tuned excitation.
The decoding in the decoder 200 is also performed based on a stepwise procedure. After receiving each of the 10-bit codebook index CW on channel 120, the decoder performs a table lookup to separate the corresponding code vector from excitation code register 212. The separated code vector is then passed through the second gain setting circuit 213 and the second synthesizer 21 to create a valid decoded signal vector. The coefficients of the second synthesizing filter 215 and the gain in the second gain setting circuit 213 are then updated in the same way as in the encoder 100. The decoded signal vector is then passed through the after filter 217 to improve the reception quality. The coefficients of the after filter are periodically updated by utilizing the available information in the decoder 200. The five samples of the after-filter signal vector are then passed to the PCM converter 218 and converted to five A-layer or μ-layer PCM output samples. Of course, both encoder 100 and decoder 200 utilize one and the same of the two mentioned PCM laws.
Fig. 4 illustrates in more detail the generation of the quantized output or the reconstructed signal in the local decoder 95 and 96. In Fig. 3a, the local decoder includes the synthesis filter 115 and the gain tuning unit 113 with its gain adapter 114. More in detail, the excitation code register 112 includes waveform codebook 130 and a gain codebook 131 and circuits 113 & 114 include multipliers 132 and 133 and a gain predictor 134. The latter generates a gain factor GAIN ', the so-called excitation vector, and the gain code register generates a gain factor GF2. In the multiplier 133, a total gain factor GF3 is generated. In other words, the gain factor consists of the predicted portion GAIN * and the innovation portion GF2 which is selected from eight possible values stored in the gain code register 131. In the local decoder, the transmitted codebook index CW from Fig. 3 is divided into a waveform codebook11.
504 010 index SCI (7 bits) and a gain codebook index GCI (3 bits). The selected excitation vector from the waveform codebook 130 is multiplied by the gain factor GF3 to the excitation signal ET (1 ... 5) and fed through the synthesis filter 115. The energy of this excitation signal ET (1 ... 5) is used to predict the gain of the next excitation vector. GAIN ·. Therefore, the gain factor GF2 retrieved from the gain code register is utilized only to correct any incorrectly predicted gain factor GAIN *.
Fig. 5 illustrates in detail the basic principles of reverse feedback adaptive linear prediction as used in, for example, the LD-CELP encoder / decoder. A delay line has delay elements 140, each of which has a delay of a sampling period T. The outputs of the delay elements are connected to each coefficient element 141 with predictor coefficients A<sub>2</sub> to A<sub>51</sub>, whose outputs are connected to a summed or 142. This summator is in turn connected to a differential element 143 which has an input to the sequence of the excitation signal ET (1 ... 5) and which is connected to the first delay element 140 in the delay line. Each of the delay elements is connected to an LPC analyzing unit, which is the reverse feedback predictor adapter circuit 116 of Fig. 3. The delay elements are also connected to an input 144. The adapter circuit 116 is connected to the respective coefficient elements 141. The connection between the differential element 143 and the delay line has an output for a quantized output signal, which is the decoded speech signal SD. The previously reconstructed speech signal samples of the signal SD are stored in the delay line element 140, in which T denotes a delay of a sampling period. The most recent samples in this delay line are weighted by the predictor coefficients (Α<sub>1</sub>· ... Α<sub>51</sub>, A<sub>1</sub>= l) and together with the sequence of excitation signals ET (1 ... 5) produces the quantized output signal or, in other words, the decoded speech signal SD. The last-generated samples SD are then shifted into the delay line. The corresponding predictor coefficients A<sub>2</sub> to A<sub>51</sub> derived from the past history of the decoded spoken information by being well known
504 010
LPC technology is applied in the reverse feedback predictor adaptation circuit 116. As indicated in Fig. 5, the elements 141 are connected through inputs 139 to the outputs of the predictor adaptation circuit 116. In Recommendation G.728, the entire delay line, consisting of 105 samples, is called 'Speech Buffer' (buffer). for spoken information) and is noted as a group 'SB (1 ... 105)' in the pseudocode. The most recent part of this buffer is called the Synthesis Filter, 'Synthesis Filter', and is noted as 'STATELPC (1 ... 50)' in the pseudocode.
Fig. 6, which corresponds to the reverse amplifier adapter 114 and partially amplifier setting unit 113 of Fig. 3, illustrates in detail the situation of amplifier predictor 134. An energy value generating unit 152 is connected to a delay line with delay elements 150, each of which has a delay sampling periods noted with 5T in the elements. Some of the delay elements 150 are connected to coefficient elements 151 with predictor coefficients GP<sub>2</sub> to GP 2. The coefficient elements are connected to a summator 153, which has an output for the signal GAIN '. All the delay elements 150 are connected to a prediction adapter 154, the outputs of which are connected to the coefficient elements 151. The energy value of the excitation signal ET (1 ... 5) is shifted into the delay line. Here, too, the most recently added values of the energy are weighted with the predictor coefficients (GP ^ .- GP ^, GP ^ 1) and the sum generated in summator 153 gives the gain factor GAIN 'predicted for the next incoming signal vector to be encoded. Here, too, the corresponding prediction coefficients are derived from the past history of the energy of the excitation signal ET (1 ... 5) by applying well-known LPC technology in the prediction adapter 154. In parentheses, in the LD-CELP encoder / decoder, the state variables in the gain predictor (134) are represented in logarithmic domain as indicated by units 155 and 156. This may be different in other backward adaptive algorithms.
Finally, it is of value with some knowledge of the procedure performed to find the optimal excitation signal.
504 010
ET (1 ... 5). Referring to Figs. 7a and 7b, which show portions of the synthesizing filter (115) of Fig. 5. Figs. 7a and 7b show the synthesizing filter as operated in various states described in ITU Recommendation G.728, page 39 and also set forth in FIG. its FIGURE 2 / G.728 through separate blocks 22 and 9 of the synthesis filter. For example, in the LD-CELP encoder / decoder, five consecutive samples are collected which form the vector to be encoded. If a vector is complete, five samples are computed and subtracted by the synthesis filter ring from this incoming speech signal vector to give the target vector. The ringing, or response to a Zero Input Response (ZINR) input signal (1 ... 5), is made by feeding the synthesis filter with zero input samples as shown in Fig. 7b. This signal can also be considered as the predicted samples for the current speech signal vector. In the encoder, all the 1024 possible excitation signals from the waveform codebook 130 combined with the gain code register 131 are fed through the synthesizer filter, starting from a zero value state for each new vector to give a zero state response (Zero State Response) ZSTR (1 ... 5) according to Fig. 7a. The resulting five samples for each excitation signal are compared with the target vector. Finally, the one that produces the smallest error is selected. Once the optimal excitation vector is found, the synthesis filter is updated. This means that the zero state response associated with the selected excitation signal is added to the zero value input which results in five new samples of the decoded speech information or five new state values in the synthesis filter. This update is performed in the local decoder on both the transmitter and receiver side.
It should be carefully noted that the detailed description above of Figures 4, 5, 6 and 7 is made for the transmitter side but should equally be applied to the receiver side, as shown in the description to Figures 1, 2, 3a and 3b.
Having described, as described above, both an overview of the invention and the most important details of the speech coding algorithm of LD-CELP, the detailed
504 The description of a preferred embodiment of the invention will be described. When a backward adaptive speech encoder / decoder such as the LD-CELP speech encoder / decoder is to be activated, no states for this encoder / decoder are available, i.e., no values are available in the delay elements 140 in the delay line of FIG. 5 or in the elements 150 of FIG. 6. Only the quantized signal generated by the previously working coding algorithm can be utilized. Therefore, in order to effect smooth transitions, a recovery of the LDCELP states is performed by using the history of the past output as a basis. In the above example, this history of the past output is taken from the encoder / decoder VDSC, and the history is stored in the buffers 192 and 292 of Fig. 1. It should be noted that an encoder / decoder for compressing a data signal on the frequency band, as shown in the example
VDSC encoder / decoders 101 and 290, have delay lines with delay elements similar to the elements 140 of the encoder / decoder in Fig. 5. It is the delay line states of the VDSC encoder / decoder which
The LD-CELPi is stored in buffers 192 and 292 and updated when processing in
The VDSC encoder / decoder continues. The values in the buffers are fed parallel to the elements 140 via their respective input 144.
From Fig. 5- · it can be seen that the states in the synthesis filter contain the history of the gone reconstructed output signal. This applies to the above described LD-CELP encoder / decoder and also applies to VDSC encoder / decoder. When the signal classifier 103 of Fig. 1 indicates spoken information on line 99 and switches from the VDSC coders / decoders 101 and 290 to the LDCELP coders / decoders 100 and 200, the updating of buffers 192 and 292 is interrupted. Switches 193 and 293 are activated for a brief moment by circuit 103 and the state values of the buffers are loaded into the delay elements 140 of the delay line of the synthesizer filter via the inputs 144. the encoders / decoders 100 and 200 are preset with these buffer values. The remaining task is to find it
504 010 excitation signal ET (1 ... 5) that would have generated these states, if the LD-CELP encoder / decoder had worked previously. Once this excitation signal ET (1 ... 5) is found, it would be easy to preset the states of the gain predictor described in connection with Fig. 6.
The details of the algorithms are explained below by specifying pseudocode as used in ITU Recommendation G.728 Coding of Speech at 16 kbit / s Using Low-Delay Code Excited Linear Prediction. Signals and coefficients are listed in accordance with TABLE 2 / G.728 of the Recommendation.
The description of how the states in the gain predictor are generated begins with the procedure for updating the synthesis filter, as it is performed in LD-CELP when operating in its normal mode. Five samples of the excitation signal ET (1 ... 5) are input to the synthesis filter as follows: First, five samples of the response to the zero value input ZINR (1 ... 5) of Fig. 7b are calculated. This is the output of the synthesis filter when fed with a zero value (ringing) signal. Second, the five samples in the zero-state response ZSTR (1 ... 5) are calculated according to Fig. 7a. Note that only five of the states differ from zero. Therefore, only these first five states are shown in Fig. 7a. ZSTR (1 ... 5) is the output vector of the zero-state synthesis filter fed by the excitation signal ET (1 ... 5). Thereafter, the five new values are generated by the STATELPC synthesis filter (1: 5) by adding the previously generated components:
STATELPC (i) = ZINR (i) + ZSTR (i); i = l, ..., 5
With this procedure in mind, we can now derive the method for obtaining the excitation signal ET (1 ... 5). When switching from the second encoder / decoder, for example the encoder / decoder VDSC in Fig. 1 to the LD-CELP encoder / decoder, only the samples in the STATELPC group (1, ..., 50) are known by placing reconstructed the signal in the correct positions of the group STATELPC (1, ..., 50) or group SB (1, ..., 105), whereby
504 01 0
STATELPC (1, ..., 50) can be considered as part of the group SB (1, ..., 105) Fig. 5. The excitation signal ET (1 ... 5) is hidden in the values of the zero state response stored in ZSTR ( 1 .. .5) which must be separated first. To this end, the response to the zero value input ZINR (1 ... 5) must be generated by feeding the synthesis filter with five zero value samples ». Then, the zero state response can be separated by generating:
ZSTR (i) = STATELPC (i) -ZINR (I); i = l, ..., 5
ZSTR (i) is the zero-output synthesis filter output when fed with the excitation vector ET (1 ... 5). This vector can now be derived by applying the inverse filter operation to this zero-state response. The excitation vector ET (1 ... 5) is perfectly reconstructed because the samples in the zero-state response do not contain all the components of a continuous rolling convolution process with fifty predictor coefficients. This final step of recovering the excitation vector ET (1 ... 5) from the zero-state response ZSTR (1 ... 5) can be more clearly recognized when the corresponding operations are explained using a piece of pseudocode. Table 1, left column, shows the pseudo code for calculating the zero state response as performed in accordance with Recommendation G.728. In the right column, the corresponding inverse operations are shown to return the excitation vector such as the inverse filter operation.
Table 1: inverse operation of 'calculation of the zero-state response' Calculation of the zero-state response-. Inverse filter operation
1) ZSTR (l) = ET (1) 1) ET (1) = ZSTR (1)
2) ZSTR (2) = ET (2) -A<sub>2</sub>* ZSTR (l) 2) ET (2) = ZSTR (2) + A<sub>2</sub>· ZSTR (l)
3) ZSTR (3) = ET (3) -A<sub>3</sub>· ZSTR (l) - - »3) ET (3) = ZSTR (3) 4-Aj · ZSTR (l) +
A<sub>2</sub>'ZSTR (2) A<sub>2</sub>-ZSTR (2)
Once you have the excitation signal ET (1 ... 5), the corresponding values for the gain predictor can be generated as recommended for example in Block 20 in G.728 l-vector delay, RMS calculator and logarithm calculator ”. Thus, all signals are
504 010 available which is required to obtain a smooth transition from any other encoder / decoder to the LD-CELP type speech encoder / decoder. This generation of amplifier states will be briefly repeated below. The excitation vector ET (1 ... 5) is fed to the energy value generating unit 152 of FIG. 6, the delay elements 150 are filled with the states of the gain predictor, the coefficients GP in the coefficient elements 151 are generated and the excitation vector of the gain GAIN * is generated. Just at the beginning of the speech transmission, the codebook index CW is generated and switched back to the exit code register 112, a new value for the excitation vector ET (1 ... 5) is generated as described in connection with FIG. 4, the states of the synthesizer filter are updated as well as the predictor coefficients A1 to A51 of the synthesizer filter in coefficient elements 141 and a new value SD of the decoded spoken information is generated. A new value of the GAIN 'gain vector is generated for the next codebook index. In this way, the states in LD-CELP are updated successfully for the voice transmission.
An overview of the method according to the invention will now be described in connection with the flow diagram in Fig. 8. The flow diagram illustrates the method of switching between two different speech encoder so that a smooth transition in the decoded output signal is obtained. The method starts in block 300 with the signal classifier 103 sensing whether spoken information is transmitted. In an alternative NO, the VDSC encoder / decoder continues to encode data to be transmitted according to a block 301. In an alternative YES, the buffer for spoken information, the delay elements 140, is preset in the LD-CELP encoder / decoder with state values VSB (1 ... 10) from the VDSC encoder / decoder stored in the buffer 192, according to a block 302. The predictor coefficients Α<sub>1</sub>· ... Α<sub>51</sub> in the synthesis filter is generated according to a block 303. The excitation signal ET (1 ... 5) is recreated, block 304, and in a block 305, the buffer is preset in the gain predictor, i.e., the delay elements 150 in Fig. 6. The coefficients GP<sub>1</sub> to GP ^ in the gain predictor is generated in a block 306 and the excitation vector GAIN * for the gain is generated in a block 307. The LD-CELP encoders
504 010
100 and 200 operate on block 308 and spoken information is transmitted between the transmitter and the receiver. Block 309 shows that the signal classifier 103 continuously senses if data in the speech frequency band is transmitted. In an alternative NO (no to the data in the speech frequency band!), The LD-CELP coders / decoders continue to work on. In an alternative YES, the VDSC encoders / decoders are connected to the communication channel 120 and begin to encode the detected data information to be transmitted.
It should be noted that the coding algorithm for the VDSC encoder / decoder may also be a reverse feedback adaptive coding algorithm. In such a case, the VDSC encoder / decoder can be started by presetting the state values in the VDSC encoder / decoder with the state values SB (1, ..., 105) of the LDCELP encoder / decoder. This is indicated by a block 310 in FIG.
Eighth In this way, the invention can be utilized for both the voice information and data encoders / decoders on a transmission line. Other encoders / decoders with backward adaptive coding algorithms can also utilize the invention.
The generation of the excitation signal ET (1 ... 5) will now be described in connection with Fig. 9 before the very detailed description in pseudocode is performed below. The state values from the VDSC encoder / decoder are stored in parallel in the elements 140 of the speech signal buffer, the group SB (1 ... 10). A temporary copy of a portion of the speech signal buffer is stored in a memory 145 and a signal TEMP is output following a signal processing described in more detail below in pseudocode. The complete content of the speech signal buffer SB (1 ... 10) is transmitted to a hybrid window unit 49 via a connection 48. By hybrid windowing in the unit 49, processing with Levinson's recursion method in a unit 50 and the bandwidth expansion in a block 51 are generated and stored the predictor coefficients A<sub>2 </sub>to A-, in a memory 146. Values A<sub>O</sub>.... A<sub>C</sub>, are transmitted to the respective coefficient elements 141 via the inputs 139. Response values of the zero value input ZINR (1 ... 5) are generated in a unit 147 by means of the signal TEMP and the A coefficients from the memory 146. Values of the zero state response ZSTR (1 ... 5 ) is generated in a difference forming unit 148 and in a unit 149 the values are generated
504 010 on the excitation signal ET (1 ... 5). These values are transmitted to the energy value generating unit 152. Values for the decoded speech signal SD can now be generated, at the beginning of the process, using the A values from memory 146 stored in coefficient elements 141 and with the states of the VDSC encoder / decoder 101. which states are stored in the delay elements 140.
In a simplified embodiment of the invention, coefficient values A are generated<sub>2</sub> to A<sub>51</sub> not in units 49, 50, 51 and 146. Instead, the corresponding coefficients B are transferred<sub>2</sub> to B<sub>51</sub> in Figs. 3a and 3b, in VDSC, the encoder / decoder encodes the LD-CELP encoder / decoder and is input to coefficient elements 141 via the inputs 139.
When transmitted according to DCME algorithms, it is known that incorrect decisions in the signal classification algorithm could result in switching from one coding algorithm to the other at intervals of 2.5 ms. If the second coding algorithm were as costly as the LD-CELP algorithm, there would be no chance of distributing the data power available within 5 ms between the two coding algorithms, since both the state preset operations and the normal working mode calculations must be performed. Therefore, when the LD-CELP encoder / decoder is switched on, the available data power within 2.5 ms must be shared between the initialization phase and the next working phase. Both together should not require more data power than was used in the normal working mode. The following describes methods for reducing the complexity during the startup process and also during the first adaptation cycle.
During the initialization stage, the computational load required to copy previous samples into the state variables of the synthesis filter is negligible. Updating the states in the gain predictor can be slightly more costly. However, considerably more computational capacity is required for calculating the predictor coefficients A<sub>1</sub> to A<sub>51</sub> in the synthesis filter. Hybrid windows and Levinson recursion would require a huge amount
504 01 0 point action of processor power.
One way to reduce the complexity of this is to change the synthesizer in the predictor order to values around ten during the initialization phase, so that only coefficients up to Α<sub>1χ</sub> generated. Periods of a slightly degraded number can hardly be detected as long as the signal is slightly affected for only a few milliseconds. This is the case here, since the speech signal buffer SB (1 ... 105) can immediately be filled with worn samples. A first complete set of fifty predictor coefficients is available after a maximum of 30 samples or 3.75 ms. A reduced filter order has the advantage that the complexity is low in the calculation of the zero state response during the initialization process. For each new sample of the zero-state response, fifty multiplication addition operations must be performed as can be seen in Fig. 7b. This computational cost is reduced by a factor of 5 if a reduced filter order of size 10 is applied.
Another method would be to use the coefficients, corresponding to coefficients A, in the LD-CELP encoder / decoder, previously generated by the second coding algorithm VDSC. This saves a significant amount of computational power required for window computation, AFC coefficients and Levinson recursion.
In addition, the computational power required for updating the coefficients during the first adaptation cycle after the LD-CELP encoder / decoder is started can be stolen and transferred to the initialization portion. The predictor coefficients calculated in advance are retained during the first or two first adaptation cycles. The resulting deterioration in speech quality is negligible, however, the increase in computational power is significant.
Further reduction of complexity can be obtained in the portion of the LD-CELP encoder / decoder containing the gain predictor.
504 010
The states of the gain predictor in the elements 150 of the LDCELP encoder / decoder consist of ten filter pins. Therefore, at least ten consecutive vectors of the excitation signal ET (1 ... 5) should be derived from the states of the synthesis filter. In addition, predictor coefficients should be GP<sub>2</sub>... GP<sub>11 </sub>is designed to predict the gain of the first vector in the first adaptation cycle following the initialization stage. Fortunately, the conditions in the gain predictor are less sensitive to minor disturbances. This allows a preset with only roughly estimated values. Therefore, the following modifications can be made to reduce the complexity during the initialization stage:
Calculate the gain GAIN 'for only the latest excitation signal ET (1 ... 5) and assume that this would be the mean of the past and also of the predicted value of the first vector of the first adaptation cycle. It is also true that a new set of predictor reinforcements has already been calculated during the calculation of the first vector of the first adaptation cycle. Therefore, one should set for setting GP<sub>2</sub> . ..GP ^ = 0 be sufficient.
A slightly more costly method would be to calculate a few of the most recent log reinforcements and take the mean of the results for the current and past reinforcements.
Now, the preferred embodiment of one of many possible combinations is explained in detail using pseudocode which is also applied in Recommendation G.728. This step is displayed when switching from any other coding algorithm to the LD-CELPal algorithm.
Let's assume that the alternative coding algorithm has previously generated quantized output samples VS and the history of this signal is stored in a group labeled VSB (1 ... 105), with VSB (105) containing the oldest and VSB (l) the latest sample. . All other labels mentioned below are the same ones used in Recommendation G.728. Thus, when the LD-CELP encoder / decoder
504 010 is in turn, the following operations are performed in advance:
1st Copy samples from the VSB group (1 ... 105) to the SB group (1 ... 105); SB (1 ... 50) is identical to the state variables for the synthesis filter stored in STATELPC (1 ... 50) with the latest sample being stored in STATELPC (l).
2nd Calculate 51 predictor coefficients A<sub>1</sub>- .. A<sub>51</sub> wherein A11 by running the hybrid window module (block 49), Levinson's unit one (block 50) and the bandwidth expanding unit (block 51). These coefficients were used during the initialization stage to calculate the response to the zero value input and during the first adaptation cycle.
3rd The gain predictor states are preset by computing only the log gain vector and by copying this SBLG () or GSTATEO.
a) Calculate five samples of the response to the zero value input:
For k = 1,2, .., 50
TEMP (k) = SB (k + 5)
FOR k = l, 2, ..., 5 {ZINR (k) = 0
For i = 2.3, ..., 50 {ZINR (k) = ZINR (k) -TEMP (k + i-2) · A ^ TEMP (i) = TEMP (i-1)}
ZINR (k) = ZINR (k) -TEMP (k + 49) · A<sub>51 </sub>TEMP (1) = ZINR (k)} for the latest excitation value in other locations by making a temporary copy STATELPC () can be executed so that it is part of group SB ().
So instead of STATLEPC (), only the group SB () was used in the following.
b) Calculate five samples for the zero-state response:
For k = l, 2, ..., 5
ZSTR (k) = SB (k) -ZINR (k) inversion vector
c) Calculate five samples for filter function:
504 010
ET (1) = ZSTR (1)
For k = 2.3, ..., 5 <ET (k) = ZSTR (k)
For i = 2, .., k
ET (k) = ET (k) + ZSTR (k-i + 2) · A |}
d) Blocks 76, 39.40 (calculation of log gain)
ETRMS = ET (1) · ET (1)
For k = 2.3, .., 5
ETRMS = ETRMS + ET (k) · ET (k)
ETRMS = ETRMS DIMINV IF (ETRMS <1) ETRMS = 1 ETRMS = 10 · log<sub>10</sub>(ETRMS)
e) Fill in the states reinforcement:
of the gain predictor with logFor i = 1,2, .., 33 SBLG (i) = ETRMS-GOFF
GAINLG = SBLG (33) + GOFF gmn- = io '<sup>gainlg</sup>/<sup>2O</sup>>
(f) On the encoder side only: Perform convolution with the waveform code vector and energy table calculation (blocks 12, 14, 15):
For the calculation of the impulse response, the weight filter is not required at this time. Therefore, the contributions of AWZ () and AWP () of block 12 can be removed.
504 010
This proposed procedure in combination with the operations performed during the first adaptation cycle is no more costly than the computational load would be, but for the setting. This is especially true if the Levinson recursion (block 50) is spread over several vectors as is usually done in practical implementations.
The ITU Recommendation G.728, as referred to above, is attached to the description.
504 010
K
Contents9
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
25 members in 13 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 9500452 | Sweden | A | |
| SE19950000452 | – | – | – |
Members25
| Document | Office | Kind | |
|---|---|---|---|
| SE9500452D0 | Sweden | D0 | |
| SE9500452L | Sweden | L | |
| CA2211347A1 | Canada | A1 | |
| WO9624926A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU4682396A | Australia | A | |
| WO9624926A3 | World Intellectual Property Organization (WIPO) | A3 | |
| SE504010C2This record | Sweden | C2 | |
| FI973270A | Finland | A | |
| FI973270A0 | Finland | A0 | |
| MX9705890A | Mexico | A | |
| BR9607033A | Brazil | A | |
| CN1179848A | China | A | |
| KR19980702044A | Republic of Korea | A | |
| JPH10513277A | Japan | A | |
| US6012024A | United States of America | A | |
| EP0976126A2 | European Patent Office (EPO) | A2 | |
| AU720430B2 | Australia | B2 | |
| CN1110791C | China | C | |
| KR100383051B1 | Republic of Korea | B1 | |
| EP0976126B1 | European Patent Office (EPO) | B1 | |
| DE69633944D1 | Germany | D1 | |
| DE69633944T2 | Germany | T2 | |
| CA2211347C | Canada | C | |
| FI117949B | Finland | B | |
| JP4111538B2 | Japan | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Patent has lapsedLapsedNUG | NUG |
Numbers
- Publication, DOCDB
- 504010
- Publication, EPODOC
- SE504010
- Application
- 9500452
- Application, DOCDB
- 9500452
- Application, EPODOC
- SE19950000452
Titles2
- Swedish
- Förfarande och anordning för prediktiv kodning av tal- och datasignaler
- English
- Method and apparatus for predictive coding of speech and data signals
Classification
- CPC, 3
- G10L19/18
- G10L13/00
- G10L2019/0003
- IPC, 3
- G10L19 18
- H03M7 30
- H04B3 06