EP0484455A1

A method and apparatus for language and speaker recognition.

Abstract

An audio source (100) is amplified (102), filtered (104) and digitized (108) so that a Fourier transform can be performed by a digital signal processor (112). The interesting frequency elements are then shaped into a histogram for a duration of the order of 5 minutes, the words being sampled every 16 ms. The histogram and the identification of the audio source are performed by a computer (114) controlled by an appropriately programmed algorithm.

Term

Term ended

Projected expiry passed 20 July 2010, 16.2 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

54 claims: 9 independent, 45 dependent

  1. 1
    Claims of equivalent WO 9102347 A1 26 WHAT IS CLAIMED IS:1. A method for recognizing an aspect of speech, comprising the steps of: creating energy distribution diagrams, indicative of spectral content of speech, for each of a plurality of known aspects of speech;receiving a segment of unknown speech whose aspect is to be recognized;creating an energy distribution diagram indicative of spectral content of speech for said unknown speech;determining differences between said energy distribution diagram for said unknown speech and each of said energy distribution diagrams for said known aspects;and recognizing an aspect by determining which energy distribution diagram of a known aspect is closest to said energy distribution diagram for said unknown speech, the closest one indicating a recognition if the difference is less than a predetermined amount.
  2. 9
    A method of creating a database from which an aspect of a sound can be identified, comprising the steps of:first determining spectral distributions for each of a plurality of samples of input sounds in which said aspect is known;determining most commonly occurring ones of said spectral distributions;creating a composite basis set, including said most commonlyOccurring ones of said spectral distributions for all of said samples;second determining numbers of occurrences of spectral distributions included in said composite basis set of spectral distributions, for a plurality of samples for which said aspect is known;and creating a relation of said numbers of occurrences of said spectral distributions included in said composite basis set for each of said plurality of samples.
  3. 20
    A method of determining an aspect of a particular sound from a plurality of aspects, comprising the steps of:determining a number of most common spectral distributions occurring in each of said plurality of aspects;receiving an unknown sample in which said aspect is to be determined;determining spectral distributions of said sample;determining which of said most common spectral distributions is closest to each of said spectral distributions of said sample, and creating a histogram of frequency of occurrence of said most common spectral distributions for said unknown sample;and . comparing said histogram with prestored histograms for each of said plurality of aspects.
  4. 28
    A method of creating a database from which a particular aspect of a sound from a plurality of aspects can be identified, comprising the steps of:determining a number of most common spectral distributions occurring in each of said plurality of aspects;creating a composite basis set including each said most common spectral distributions in each of said plurality of aspects;analyzing each of a plurality of known aspects, to determine numbers of occurrences of each element of said composite basis set;and creating histograms, for each said aspect, indicative of said occurrences of said elements in said composite basis set.
  5. 30
    A method of determining an aspect of a particular sound from a plurality of aspects, comprising the steps of:comparing each incoming spectral distribution for each of a plurality of samples of input sounds in which said aspect is known with stored spectral distributions to determine if said incoming spectral distribution is similar to a previously obtained spectral distribution by taking a dot product between incoming spectral distributions and previously stored spectral distributions, and recognizing them to be similar if the result of the dot product is less than a predetermined amount;storing said incoming spectral distribution if it is not similar to any of said previously obtained spectral distributions;incrementing a number of occurrences of a particular spectral distribution if said incoming spectral distribution is similar to said particular spectral distribution;taking a weighted average between the incoming spectral distribution and the particular spectral distribution if said incoming spectral distribution is similar to said particular spectral distribution;forming a basis set of said spectral distributions and said number of occurances;determining most commonly occurring ones of said spectral distributions in said basis set;creating a composite basis set, including said most commonly occurring ones of said spectral distributions for all of said samples;second determining numbers of occurrences of spectral distributions included in said composite basis set of spectral distributions, for a plurality of samples for which said aspect is known;creating histograms of said numbers of occurrences of said spectral distributions included in said composite basis set for each of said plurality ofj. samples;receiving an unknown sample in which said aspect is to'be determined;determining spectral distributions of said sample;determining which of said most common spectral distributions in said composite basis set is closest to said spectral distributions of said sample, and creating a histogram of frequency of occurrence of said most common spectral distribution for said unknown sample;and comparing said histogram with prestored histograms for each of said plurality of aspects by determining euclidean distance, and recognizing one of said plurality of aspects which has the minimum euclidean distance. ,
  6. 33
    An apparatus for recognizing an aspect of speech, comprising:means for receiving a plurality of speech samples;and processing means, for: a) creating energy distribution diagrams, indicative of spectral content, for each of a plurality of known aspects of speech;b) receiving a segment of unknown speech whose aspect is to be recognized from said receiving means;c) creating an energy distribution diagram for said unknown speech;and d) determining differences between said energy distribution diagram for said unknown speech and each of said energy distribution diagrams for said known aspects;and e) recognizing an aspect by determining which energy distribution diagram to a known aspect is closest, the closest one indicating a recognition if the difference is less than a predetermined amount.
  7. 41
    An apparatus for creating a database from which an aspect of a sound can be identified, comprising:means for receiving input sounds;means for fast fourier transforming said input sounds, for first determining spectral distributions for each of a plurality of samples of input sounds in which said aspect is known and second determining numbers of occurrences of spectral distributions for a plurality of samples for which said aspect is known;and processing means for: a) determining most commonly occurring ones of said spectral distributions determined in said firεft determining;b) creating a composite basis set, including said most commonly occurring ones of said spectral distributions for all of said samples;and c) creating a histogram between numbers of occurrences of said spectral distributions included in said composite basis set for each of said plurality of samples.
  8. 47
    An apparatus for determining an aspect of a particular sound from a plurality of aspects, comprising:memory means for storing a number of most common spectral distributions occurring in each of said plurality of aspects and storing prestored histograms for each of said plurality of aspects;means for receiving an unknown sample in which said aspect is to be determined;- means for determining spectral distributions of said sample;processing means for a) determining which of said most common spectral distributions is closest to each of said spectral distributions of said sample, and creating a histogram of frequency of occurrence of said most common spectral distribution for said unknown sample;and b) comparing said histogram with prestored histograms for each of said plurality of aspects.
  9. 52
    An apparatus for determining an aspect of a particular sound from a plurality of aspects, comprising:means for receiving a plurality of incoming signals;means for A/D converting and FFTing said incoming signals to produce spectral distributions thereof;and . processing means, for: a) comparing each incoming spectral distribution for each of a plurality of samples of input sounds in which said aspect is known with stored spectral distributions to determine if said incoming spectral distribution is similar to a previously obtained spectral distribution by taking a dot product between incoming spectral distributions and previously stored spectral distributions, and recognizing them to be similar if the result of the dot product is less than a predetermined amount;b) storing said incoming spectral distribution if it is not similar to any of said previously obtained spectral distributions;> c) incrementing a number of occurrences of a particular spectral distribution if said incoming spectral distribution is similar to said particular spectral distribution;d) taking a weighted average between the incoming spectral distribution and the particular spectral distribution if said incoming spectral distribution is similar to said particular spectral distribution;e) forming a basis set of said spectral distributions and said number of occurrences;f) determining most commonly occurring ones of said spectral distributions in said basis set;g) creating a composite basis set, including said most commonly occurring ones of said spectral distributions for all of said samples;h) second determining numbers of occurrences of spectral distributions included in said composite basis set of spectral distributions, for a plurality of samples for which said aspect is known;i) creating histograms of said numbers of occurrences of said spectral distributions included in said composite basis set for each of said plurality of samples;j) receiving an unknown sample in which said aspect is to be determined;k) determining spectral distributions of said unknown sample;1) determining which of said most common spectral distributions in said composite basis set is closest to said spectral distributions of said sample, and creating a histogram of frequency of occurrence of said most common spectral distribution for said unknown sample;and m) comparing said histogram with prestored histograms for each of said plurality of aspects by determining euclidean distance, and recognizing one of said plurality of aspects which has the minimum euclidean distance.