Frequency domain speech coding.
Abstract
Adaptive bit allocation to the channels of a sub-band coder (or to the coefficients of a transform coder) is simplified by using a fixed set of numbers of bits, only the selection of those channels to which the available bits are assigned being varied.

Term
Term ended
Projected expiry passed 23 August 2005, 21.1 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
7 claims: 4 independent, 3 dependent
- 1A coder for speech signals comprising separation means for receiving speech signals and generating series of values each series representing respective portions of the frequency spectrum of the input signal encoding means for digitally encoding each series, and bit allocation means for varying the number of bits used for encoding the respective series in dependence on the relative energy contents thereof, characterised in that the number of series to which any given number of bits is allocated is constant, only the selection of the series to which respective numbers of bits are allocated being varied.
- 5A method of coding a speech signal in which the signal is divided into separate channels representing respective portions of the frequency spectrum of the input signal, and the channels are each encoded using a variable number of bits dependent upon the relative energy contents of the channels, characterised in that the number of channels to each of which any given number of bits is allocated is constant, only the selection of the channels to which respective numbers of bits are allocated being varied.
- 6A coder for speech signals substantially as herein described with reference to the accompanying drawings.
- 7A method of coding speech signals substantially as herein described with reference to the accompanying drawings.
Independent claims4
26 paragraphs, as filed
This invention concerns frequency domain speech coding, such as sub-band coding in which the frequency spectrum of an input signal is separated into two or more sub-bands which are then coded individually, or transform coding in which a block of input samples are converted to a set of transform coefficients.
Sub-band coding has been shown to be an effective method of reducing the bit-rate required for the transmission of signals - see, for example Crochiere, Webber and Flanagan "Digital Coding of Speech in Sub-bands", Bell System Technical Journal, Vol.55 pp.1069-1085 (October 1976) and Crochiere, "On the Design of Sub-band Coders for Low Bit-Rate Speech Communications, ibid Vol.56, pp747-779 (May-June 1977).
The technique involves splitting the broadband signal into two or more frequency bands and encoding each band separately. Each band can then be down-sampled, and coding efficiency improved by using different coding strategies, which can be optimised to the statistics of the signal. This is especially advantageous for speech transmission because the coders can exploit certain perceptual effects connected with hearing. In particular, provided that appropriate quantisers are used, the technique will result in the quantisation noise at the output of the codec having a similar power spectral distribution to that of the uncoded signal; it is well-known that the human ear is relatively tolerant to noise in the parts of the spectrum occupied by high level wanted signals. Additionally, the higher frequency components can be represented with reduced accuracy because the ear is less sensitive to their absolute content. Although benefits can be obtained using a fixed coding scheme, an adaptive scheme which takes account of the changing properties of the signal with time, is preferred. After transmission, the individual sub-bands are up-sampled and passed through interpolation filters prior to recombination.
In transform coding, a block of (say) 128 input samples is subjected to a suitable transformation such as the discrete cosine transform to produce a set of 128 coefficients; again efficiency can be improved by adaptive coding.
When adaptive bit allocation is used, the coding characteristics of the system are matched to the short-term spectrum of the input signal. One such proposal (J.M.Tribolet and R.E.Crochiere, "Frequency Domain Coding of Speech" IEEE Trans. on ASSP, Vol. ASSP-27, No.5, Oct 1979) utilises a fully adaptive assignment of bits to each sub-band signal. The algorithm proposed is: <chemistry id="chem0001" num="0001"><img file="EP0176243A2_D0001.tif" /></chemistry> where R is the total number of bits available divided by the number of sub-bands N in the system. σ<sup>2</sup><sub>i</sub> is the energy of the ith sub-band signal within the time interval under consideration, and R<sub>i</sub> is the number of bits allocated to the ith sub-band. γ may be varied for noise shaping.
The problem with the bit allocation strategy of equation (1) is that the value of R<sub>i</sub> is usually fractional and often negative. The process of i) rounding R<sub>i</sub> to an integer value, ii) restricting the maximum value of R<sub>i</sub> to (typically) 5 and iii) setting negative values of R<sub>i</sub> to 0, may result in the total number of bits allocated exceeding or falling below the number of bits available. In this case readjustments have to be made by reapplying equation (1) or by an arbitrary process of adding or subtracting bits to or from the bands. Moreover, the amount of computation involved is substantial.
According to the present invention there is provided a coder for speech signals comprising separation means for receiving speech signals s and generating series of values each series representing respective portions of the frequency spectrum of the input signal, encoding means for digitally encoding each series, and bit allocation means for varying the number of bits used for encoding the respective series in dependence on the relative energy contents thereof, characterised in that the number of series to which any given number of bits is allocated is constant, only the selector of the series to which respective numbers of bits are allocated being varied.
In another aspect the invention provides a method of coding a speech signal in which the signal is divided into separate channels s representing respective portions of the frequency spectrum of the irpout signal, and the channels s are each encoded using a variable number of bits dependent upon the relative energy contents s of the channels, characterised in that the number of channels s to each of which ary given number of bits is allocated is constant, only the selection of the channels to which respective number of bits are allocated being varies
The series or channel referred to may be the channels s of a sub-band coder or the transform coefficients of a transform coder.
Where a scaling factor is applied to the signals, preferably the bit allocation is performed as a function of the scaling factor, so that decoding can be carried out by reference to the scale factors, avoiding the necessity for transmission of additional side information.
An embodiment of the invention will now be described, by way of example, with reference to the accompanying drawings, in which Figure 1 is a block diagram of a sub-band coder according to the invention; and Figure 2 illustrates the bit allocation process of the apparatus of Figure 1.
Figure 1 shows a 14 band 32kbit/s sub-band coder system. The input signal having a nominal bandwidth of 7kHz is sampled at 14kHz - illustrated schematically by a switch 1 - and the full band spectrum is divided into fourteen uniform bands by a four-stage tree-structured filter bark 2 employing quadrature mirror filters. 32-tap finite impulse response filters are suggested though lower order filters could be employed at the higher stages of the filter bank. The filter outputs are, as is conventional, down samples (by means not shown) to 1kHz.
Laplacian forward adaptive quantisers are employed for the quantisation of the sub-band signals. Essentially there are two stages here; firstly (normalisation stages Nl...N14) the signal is normalised by dividing by a scaling factor which is defined every 16ms from estimates of the energy of the relevant sub-band. Basically this is the rms value of the signal over that period. 16 samples (for sub-band 1, x<sub>1j</sub>- j=l,...16) are buffered in a register 3, and the scaling factor or step size Δl calculated at 4 using the relation Δ<sub>k</sub>=<maths id="math0001" num=""><img file="EP0176243A2_D0002.tif" /></maths>The scaling factors are quantised to 5-bit accuracy in a quantiser 5 and the quantitised value a transmitted as side information to the receiver. Thus the side information accounts for almost 4.5kbit/s and thus approximately 27.5kbit/s is available for transmission of the samples themselves. These processes are carried out for each of the fourteen sub-bands. The normalised samples (S<sub>kj</sub>=x<sub>kj</sub>/Δ<sub>k</sub>) for each sub-band are then fed to a quantiser 6 which encodes them using the desired number of bits B<sub>K</sub> prior to transmission via multiplexer 7. Bit allocation is indicated in Figure 1 schematically as unit 8.
In n the prior proposal of Tribolet and Crochiere, equation (1) was used to define the bit allocation. In the present proposal, equation (1) is used to determine the bit allocation patterns for all the 16msec frames in an input training sequence which is free of any silent intervals.γ is set to -0.3 and the maximum number of bits M allowed in the allocation is set to 5. Let N<sub>i</sub> represent the total number of times in the training sequence that i bits are allocated, where i = 1....,5.
Next, we define f<sub>i</sub> as:<maths id="math0002" num=""><img file="EP0176243A2_D0003.tif" /></maths> where N<sub>T</sub> is the total number of bits available for allocation throughout the training sequence. f<sub>i</sub> therefore represents the portion of N<sub>T</sub> used in allocating i bits for the coding of sub-band signals. If there are N<sub>t</sub> bits available for allocation the expected number of bands n<sub>i</sub> that receive i bits can be calculated according to:<maths id="math0003" num=""><img file="EP0176243A2_D0004.tif" /></maths> for i = 1, ..., M.
A time-invariant bit allocation pattern is thus obtained using the n<sub>i</sub> estimates, i.e. (n<sub>5</sub><sup>*</sup> 5 bits, n<sub>4</sub><sup>*</sup>4 bits,...,n<sub>1</sub><sup>*</sup> 1 bit; 0 bit for the remaining bands), assuming M is equal to five. This means that, within a 16 msec frame, n<sub>5</sub> sub-bands receive 5 bits, n<sub>4</sub> sub-bands receive 4 bits and so on. Manual adjustment is normally required to ensure that the total number of bits in the invariant allocation pattern gives the desired total transmission bit rate. For the 14-band coder, the 27.5kbit/sec capacity and lkHz sampling rate permit 27 bits, and the bit pattern obtained was given by: (1<sup>*</sup>5, 1<sup>*</sup>4, 3<sup>*</sup>3, 2<sup>*</sup>2, 5<sup>*</sup>1, 2<sup>*</sup>0)
Though the pattern is fixed, the allocation is based on the scale factors of the sub-bands signals. For each frame of 16 msec the band with the largest scaling factor is allocated 5 bits, the 2nd largest 4 bits and so on. The processing requirements of this algorithm are considerably reduced when compared with those of the fully adaptive scheme, since once the invariant allocation pattern has been derived, it is fixed for a given coder. Also because the allocation of these bit groups to the particular sub-bands is determined by reference to the scaling factors, the transmission of further side information to the receiver is not necessary.
Considering now transform coding, in this example an adaptive transform coder using the discrete cosine transform employs a blocksize of 128 samples. An estimation of the 16 primary coefficients of the basis spectrum (R Zelinski and P Noll, "Adaptive transform coding of speech signals", IEEE Trans. on ASSP, Vol. ASSP-25, No. 4, pp 299-309, Aug. 1977) is carried out every 8 msec although the average of two sets of these coefficients, from adjacent frames, is used to define the step-sizes of the transform coefficient quantisers and the bit allocation pattern. 3 bits Gaussian quantisers are used to quantise the 16 primary values of the average basis spectrum. Normalisation of the input samples is also carried out using a normalisation parameter which is evaluated every 256 samples. The normalisation parameter is quantised using a 5 bit Gaussian quantiser.
The problem of efficiently coding the resulting 128 coefficients is similar to that of coding the sub-band samples in the previous examples. Here equations 1 to 3 are applied to a training sequence to obtain a bit allocation pattern (y=-0.2) of: <ul id="ul0001" list-style="none"><li>(1<sup>*</sup>7, 4<sup>*</sup>6, 5<sup>*</sup>5, 9<sup>*</sup>4, 20<sup>*</sup>3, 25<sup>*</sup>2, 28<sup>*</sup>1, 36<sup>*</sup>0),</li></ul> that is, out of the 128 transform coefficients, 1 coefficient is quantised with 7 bits, 4 coefficients with 6 bits etc.
The advent of digital signal processing (DSP) devices has facilitated the real-time implementation of a number of otherwise difficult to implement speech coding algorithms. A sub-band coder for example, can be conveniently implemented using possibly a DSP chip. The implementation complexity of a coder depends to an extent on the number of multiplications/divisions, additions/substractions and on the size of memory required for storing the intermediate variables of the coding algorithm. Table 1 illustrates the computational requirements, including delays, of the coders considered. SBC and ATC indicate sub-band and transform coding respectively, whilst ABA indicates adaptive bit allocation according to equation (1) and SBA the simplified bit allocation as described above. <tables id="tabl0001" num="0001"><img file="EP0176243A2_D0005.tif" /></tables>
A fast algorithm for the cosine transform was assumed in deriving the above estimates. Note that the adaptive transform coders also requires additional 1<sub>092</sub> and inverse 1<sub>092</sub> look-up Tables. For the sub-band coder, the higher stages of the quadrature mirror filter analysis bank can be implemented using lower order FIR filters to reduce the memory size and coder delay. Excluded in the estimation is the memory required for the programme instructions of the coding algorithm. Currently, due to their stringent real-time and memory requirements, large blocksize transform coders can be more conveniently implemented using array processors.
The performance of the coders described has been assessed by computer simulation in terms of <ul id="ul0002" list-style="none"><li>(1) average segmental signal-to-noise ratio</li><li>(2) long-term average spectral density plot of the output noise and</li><li>(3) informal subjective listening tests.</li></ul>
The input data used in our computer simulation experiments consisted of two sentences of male speech and two sentences of female speech. Table 2 shows the average segmental SNR performance (in dBs) of the coders. <tables id="tabl0002" num="0002"><img file="EP0176243A2_D0006.tif" /></tables>
The two sub-band coding schemes offer the best SNR measurements of 19.88 and 19.24 dB. Informal subjective listening tests indicate that the SBC/ABA system produces an excellent quality recovered speech. This is due to the fact that the output noise level is low enough and is masked by the speech energy in each band. Also, the use of the simplified bit allocation algorithm did not affect the subjective quality of the 14-band coder though there is a drop of 0.6 dB in SNR.
The next scheme, in order of merit, is adaptive transform coding employing the full algorithm. The distortion due to inter-block discontinuities can be substantially reduced by smoothing. It should be noted that subjectively the difference between sub-band and transform coding is not as significant as suggested by their large difference in SNR values. The transform coder employing the simplified bit allocation algorithm was found to have an SNR reduction of 1 dB compared to the one with the fully adaptive algorithm. The block-end distortion becomes more pronounced and the recovered speech is also degraded by a "whispery" noise. This means that as the noise level, at this bit rate, is just at the threshold of audibility, the use of the full adaptation algorithm becomes necessary. However, if more bits are allowed for the transform coder, the SBA algorithm might prove to be a valuable method in reducing the coder complexity.
In general, some degradation in the quality of the ATC speech at 32 kbits/sec is caused by interblock discontinuities. Though the underlying speech can be very good, the effect of discontinuities is perceptually unacceptable. One suggested solution to this problem is to apply 10 percent overlap between adjacent blocks. Another method is to employ either median filtering or a moving average filtering process to a few samples at both ends of each block. The 10 percent overlap scheme is found to be the least effective because fewer bits are available for the quantisation of the transform coefficients which in turn increases the amount of block-end distortion. The method of median filtering is found to give some subjective improvement while the best performance is obtained from the moving averaging method. In its use, 10 samples x<sub>l</sub>, x<sub>2</sub>,..., x<sub>10</sub> (the last five samples of the previous block and the first five samples of the present block) were replaced by Y<sub>1</sub>, Y<sub>2</sub>, .....,y<sub>10</sub>, where y<sub>i</sub> = 1/3(x<sub>i-1</sub><sup>+</sup> x<sub>i</sub><sup>+</sup> x<sub>i-1</sub>), and i = 1,...., 10..
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US5984514A | Cited by | United States of America | Search report |
| EP3913628A1 | Cited by | European Patent Office (EPO) | Search report |
| US5222189A | Cited by | United States of America | Search report |
| US5563960A | Cited by | United States of America | Search report |
| EP3128514A4 | Cited by | European Patent Office (EPO) | Search report |
| WO9302508A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| EP0511692A3 | Cited by | European Patent Office (EPO) | Search report |
| US5631978A | Cited by | United States of America | Search report |
| EP0511692A2 | Cited by | European Patent Office (EPO) | Search report |
| US10468035B2 | Cited by | United States of America | Applicant |
| WO9410758A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| EP0522219A1 | Cited by | European Patent Office (EPO) | Search report |
| EP0508581A2 | Cited by | European Patent Office (EPO) | Search report |
| US5142656A | Cited by | United States of America | Search report |
| EP0560413A3 | Cited by | European Patent Office (EPO) | Search report |
| US11676614B2 | Cited by | United States of America | Applicant |
| WO9009064A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| AU677856B2 | Cited by | Australia | Search report |
| EP0522219A1 | Cited by | European Patent Office (EPO) | Search report |
| EP0208712A1 | Cited by | European Patent Office (EPO) | Examiner |
| EP1396938A4 | Cited by | European Patent Office (EPO) | Search report |
| US10909993B2 | Cited by | United States of America | Applicant |
| EP0513860A2 | Cited by | European Patent Office (EPO) | Search report |
| US5740317A | Cited by | United States of America | Search report |
| US5706309A | Cited by | United States of America | Search report |
| US4914701A | Cited by | United States of America | Search report |
| US5412741A | Cited by | United States of America | Search report |
| US5109417A | Cited by | United States of America | Search report |
| US5752225A | Cited by | United States of America | Search report |
| US11688406B2 | Cited by | United States of America | Applicant |
| US5235671A | Cited by | United States of America | Search report |
| EP0513860A3 | Cited by | European Patent Office (EPO) | Search report |
| EP0497413A1 | Cited by | European Patent Office (EPO) | Search report |
| WO9009064A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US4790016A | Cited by | United States of America | Search report |
| US4899384A | Cited by | United States of America | Search report |
| WO9502928A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| EP0560413A2 | Cited by | European Patent Office (EPO) | Search report |
| EP0508581A3 | Cited by | European Patent Office (EPO) | Search report |
| US5838377A | Cited by | United States of America | Search report |
| US4926482A | Cited by | United States of America | Search report |
| EP1396938A1 | Cited by | European Patent Office (EPO) | Search report |
| EP0145332A2 | Cites | European Patent Office (EPO) | Search report |
| DE3102822A1 | Cites | Germany | Search report |
10 members in 7 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 8421498 | United Kingdom | A | |
| 8421498 | United Kingdom | – | |
| 8421498 | – | – | – |
| GB19840021498 | – | – | – |
Members10
| Document | Office | Kind | |
|---|---|---|---|
| GB8421498D0 | United Kingdom | D0 | |
| EP0176243A2This record | European Patent Office (EPO) | A2 | |
| JPS61112433A | Japan | A | |
| EP0176243A3 | European Patent Office (EPO) | A3 | |
| CA1248234A | Canada | A | |
| EP0176243B1 | European Patent Office (EPO) | B1 | |
| AT50896T | Austria | T | |
| DE3576444D1 | Germany | D1 | |
| US4949383A | United States of America | A | |
| JPH0525408B2 | Japan | B2 |
42 legal events, as 3 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Nl: ceased due to reaching the maximum lifetime of a patentCeasedNLV7 | NLV7 | EP | |
| Se: european patent has lapsedLapsedEUG | EUG | EP | |
| Patent expired after termination of 20 yearsExpiredPE20 | PE20 | GB | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Be: lapsedLapsedBERE | BERE | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| European patent in force as of 2002-01-01IF02 | IF02 | GB | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Se: european patent in force in swedenEAL | EAL | EP | |
| Lu: last paid annual feeEPTA | EPTA | EP | |
| It: last paid annual feeITTA | ITTA | EP | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Patent ceasedCeasedPL | PL | CH | |
| Fr: translation filedET | ET | EP | |
| Corresponds to:REF | REF | EP | |
| It: translation for a ep patent filedITF | ITF | EP | |
| It: translation for a ep patent filedITF | ITF | EP | |
| Designated contracting statesAK | AK | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Corresponds to:REF | REF | EP | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP3 | RAP3 | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for examination filed17P | 17P | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Designated contracting statesAK | AK | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 0176243
- Publication, DOCDB
- 0176243
- Publication, EPODOC
- EP0176243
- Application
- 85306015
- Application, DOCDB
- 85306015
- Application, EPODOC
- EP19850306015
Titles6
- German
- Sprachkodierung im Frequenzbereich.
- English
- Frequency domain speech coding.
- French
- Codage de la parole dans le domaine des fréquences.
- German
- Sprachkodierung im Frequenzbereich
- English
- Frequency domain speech coding
- French
- Codage de la parole dans le domaine des fréquences
Classification
- CPC, 1
- H04B1/667
- IPC, 4
- H03M3 04
- H03M1 12
- H04B1 66
- H04B14 04
Designated states1
- Contracting states, 1
- Sweden