AU5205101A

Robust parameters for noisy speech recognition

Abstract

The invention concerns a method for automatically processing noisy speech comprising the following steps: sensing and digitising speech (1); retrieving several frames (15) corresponding to said signal (10); breaking down each frame (15) using an analysing system (20, 40) into at least two different frequency bands to obtain two vectors of representative parameters (45); converting (50) the vectors (45) into second vectors relatively insensitive to noise (55), each converter system (50) being associated with a frequency band. The knowledge acquisition of said converter systems (50) being produced on a noise-contaminated speech corpus (102).

Term

Term ended

Projected expiry passed 25 April 2021, 5.4 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

12 claims: 2 independent, 10 dependent

  1. 1
    CLAIMS 1. A method of automatic processing of noiseaffected speech comprising at least the following steps:5 - capture and digitising of the speech in the form of at least one digitised signal (1), - extraction of several time-based sequences or frames (15), corresponding to said signal, by means of an extraction system (10), 10 - decomposition of each frame (15) by means of an analysis system (20, 40) into at least two different frequency bands so as to obtain at least two first vectors of representative parameters (45) for each frame (15), one for each frequency band, and 15 - conversion, by means of converter systems (50) , of the first vectors of representative parameters (45) into second vectors of parameters relatively insensitive to noise (55), each converter system (50) being associated with one frequency band and converting the first vector 20 of representative parameters (45) associated with said same frequency band, and the learning of said converter systems (50) being achieved on the basis of a learning corpus which corresponds to a corpus of speech .contaminated by noise (102) . 25
  2. 2
    The method according to Claim 1, characterised in that it further comprises a step of concatenation of the second vectors of representative parameters which are relatively insensitive to noise (55), associated with the different frequency bands of the same frame (15) so as to 30 have no more than one single third vector of concatenated parameters (56) for each frame (15) which is then used as input in an automatic speech-recognition system (60).
  3. 3
    The method according to Claim 1 or 2, characterised in that the conversion, by means of converter 35 systems (50) , is achieved by linear transformation or by non-linear transformation.
  4. 4
    The method according to one of Claims 1 to 3, characterised in that the converter systems (50) are artificial neuronal networks.
  5. 5
    The method according to Claim 4, characterised 5 in that the said artificial neuronal networks are of the multi-layer perceptron type and each comprises at least one hidden layer.
  6. 6
    The method according to Claim 5, characterised in that the learning by the said artificial neuronal 10 networks of the multi-layer perceptron type relies on targets corresponding to the basic lexical units for each frame of the learning corpus, the output vectors of the last hidden layer or layers of the said artificial neuronal networks being used as vectors of representative parameters 15 which are relatively insensitive to the noise.
  7. 7
    An automatic speech-processing system comprising at least:- an acquisition system for obtaining at least one digitised speech signal (1), 20 - an extraction system (10) , for extracting several timebased sequences or frames (15) corresponding to said signal (1), . - means (20, 40) for decomposing each frame (15) into at least two different frequency bands so as to obtain at 25 least two first vectors of representative parameters (45), one vector for each frequency band, and - several converter systems (50) , each converter system (50) being associated with one frequency band and making it possible to convert the first vector of representative 30 parameters (45) associated with this same frequency band into a second vector of parameters which are relatively insensitive to noise (55), and the learning by the said converter systems (50) being achieved on the basis of a corpus of speech corrupted by 35 noise (102) .
  8. 8
    The automatic speech-processing system according to Claim 7, characterised in that the converter systems (50) are artificial neuronal networks, preferably of the multi-layer perceptron type. 5
  9. 9
    The automatic speech-processing system according to Claim 7 or Claim 8, characterised in that it further comprises means allowing the concatenation of the second vectors of representative parameters which are relatively insensitive to noise (55), associated with
  10. 10
    10 different frequency bands of the same frame (15) so as to have no more than one single third vector of concatenated parameters (56) for each frame (15), said third vector then being used as input into an automatic speech-recognition system (60). 15 10. Use of the method according to one of Claims 1 to 6 and/or of the system according to one of Claims 7 to 9 for speech recognition.
  11. 11
    Use of the method according to one of Claims 1, 3 to 6 and/or of the system according to one of Claims 7 20 or 8 for speech coding.
  12. 12
    Use of the method according to one of Claims 1, 3 to 6 and/or of the system according to one of Claims 7 or 8 for removing noise from speech.