Speaker recognition method
Abstract
The invention relates to a speaker recognition method, comprising comparing a model calculated on the basis of samples derived from a speech signal with a stored model of at least one known speaker. In the method according to the invention, the averages of the cross-sectional areas or other cross-sectional dimensions of portions (C1 - C8) of a lossless tube model of the speaker's vocal tract, calculated by means of the samples derived from the speech signal, are compared with the corresponding averages of the portions of the stored vocal tract model of at least one known speaker.

Term
No projected expiry on record.
- Priority and filed
- Granted
- Today
7 claims: 1 independent, 6 dependent
- 1Patentkrav Patenttivaatimukset The claims 1. A method for identifying a speaker, comprising comparing a model calculated from samples of a speech signal to a stored model of at least one known speaker, characterized by comparing the averages of the cross-sectional areas or other cross-dimensions of at least one known speaker corresponding averages. 1. Förfarande för identifiering av en talare, vilket förfarande omfattar en jämförelse av en modell som uträknats pä basis av prov tagna pä en talsignal med en lagrad modell för ätminstone en känd talare, känneteckn a t av att medeltalen för tvärsnittsareor eller övriga tvärsnittsmätt hos de delar av ett förlustfritt rör som utgör en modell för talarens ljudkanal, vilka uträknats pä basis av talsignalproven, jämförs med motsvarande medeltal för delar av en lagrad 1judkanalsmodell för ätminstone en känd talare. 1. Menetelmä puhujan tunnistamiseksi, joka menetelmä käsittää puhesignaalista otettujen näytteiden perusteella lasketun mallin vertaamisen ainakin yhden tunnetun puhujan tallennettuun malliin, tunnettu siitä, että verrataan puhesignaalin näytteistä laskettuja, puhujan ääniväylää mallintavan häviöttömän putken osien poikkipinta-alojen tai muiden poikkimittojen keskiarvoja ainakin yhden tunnetun puhujan tallennetun ääniväylämallin osien vastaaviin keskiarvoihin.
29 paragraphs, as filed
Method for speaker identification
The invention relates to a method for identifying a speaker, which method comprises comparing a model calculated on the basis of samples taken from a speech signal with a stored model of at least one known speaker.
One known way to identify and authenticate a user in various systems, such as computer systems or telephone systems, is to identify the user based on a speech signal. All known speaker recognition methods seek to find speech features that can be used to automatically identify and differentiate speakers. In this case, on the basis of a sample taken from the speech of each speaker, certain models with speech-specific parameters are formed, which are stored in the memory of the speaker recognition system. When an anonymous speaker is then to be identified, a sample of his speech signal is formed into a model with the same parameters, which is compared with reference models in the system's memory. If the model generated from the recognizable speech signal corresponds with sufficient accuracy according to a predetermined criterion to a model of a known speaker stored in the memory, the anonymous speaker is identified as the person from whose speech signal the reference reference is formed. In general, all known speaker recognition systems follow the general principle outlined above, but the parameters and solutions used in speaker voice modeling are very different. Examples of speaker identification methods and systems are disclosed in U.S. Patent Nos. 4,720,863 and 4,837,830, GB Patent Application 2,169,120, and EP Patent Application 0,369,485.
It is an object of the invention to provide a new type of speaker recognition method which seeks to identify a speaker on the basis of an arbitrary speech signal with a better and> 25 simpler algorithm.
This is achieved by a method of the type described in the introduction, which compares the averages of the ice paths of the spark areas characterized according to the invention by the averages of the speech model lossless tube sections or other cross-dimensions with the averages of the recorded sound path model of at least one known speaker.
The basic idea of the invention is to identify the speaker on the basis of the voice path characteristic of the speaker. In this context, the voice path refers to the voice of the human vocal cords, the voice channel formed by the throat. The exact shape of the speaker and the sound path of the head, pharynx, mouth, and lips that allows a person to form a bus change over time. However, the method according to the invention does not require the exact shape of the audio25 bus. In the invention, the speaker's voice path is modeled in a so-called with a lossless tube model whose shape is characteristic of the speaker. In addition, although the profile of the speaker's audio path and, at the same time, the lossless tube design is constantly changing during speaking, the extremes and mean of the audio path and lossless tube design are instead constant for the speaker. Therefore, in the method according to the invention, the speaker can be identified with moderate accuracy on the basis of the average shape or shapes of the lossless tube modeled by the speaker. In one embodiment of the invention, in addition to the average cross-sectional areas of the cylindrical parts of the lossless tube model, the extremes, i.e. the maximum and minimum values of the cross-sectional areas of the cylindrical parts, are used in the identification.
In another embodiment of the invention, the accuracy of personal identification is further improved by averaging the lossless tube model for individual sounds. During a given sound, the shape of the voice path remains almost unchanged and better describes the speaker's voice path. When multiple sounds are used for recognition, very accurate recognition is obtained.
The cross-sectional areas of the cylindrical parts of the lossless tube model used in the invention can be easily calculated from the so-called so-called speech coding algorithms formed in conventional speech coding algorithms. reflection coefficients. Of course, another cross-sectional dimension, such as radius or diameter, can be defined as the reference parameter for the surface area. On the other hand, the cross-section of the tube may have some other shape instead of a circular shape.
The invention will now be described in more detail by means of an exemplary embodiment with reference to the accompanying drawing, in which Figures 1 and 2 illustrate the modeling of a speaker's voice path by a lossless tube , and Fig. 5 shows a block diagram, which illustrates speaker recognition at the sound level.
Reference is now made to Figure 1, which is a perspective view of a model of a lossless tube consisting of successive cylinder sections C1-C8, which forms a rough model for the human voice path. The model of the lossless tube in Figure 1 can be seen in a side view in Figure 2. The human voice pathway generally refers to the voice pathway formed by the human vocal cords, throat, pharynx and lips, through which a person generates speech sounds. In Figures 1 and 2, the cylinder portion C1 depicts the shape of the voice path portion immediately after the glottis, the blade portion C8 depicts the shape of the voice path at the lips, and the intermediate cylinder portions C2-C7 depict the shape of the discrete voice path portions between the vocal cords and lips. The shape of the audio bus is characterized by the fact that it varies continuously during speech as different sounds are generated. Similarly, the diameters and areas of the discrete cylinders C1-C8 depicting different parts of the audio path also vary during speech. However, the inventor has found that the average voice path shape calculated from a relatively large number of instantaneous voice path shapes is a constant specific to each speaker that can be used to identify the speaker. Similarly, the long-term averages of the C1-C8 cross-sections of the cylindrical portions calculated from the instantaneous values of the cross-sectional areas of the cylinders C1-C8 of the lossless tube model modeling the sound path are relatively closely constant. Furthermore, the extremes of the cross-dimensions of the cylinders are also determined by the extremes of the actual sound path and are thus relatively accurate constants characteristic of the speaker.
The method according to the invention utilizes the so-called intermediate results formed in linear predictive coding (LPC), which are well known in the art. reflection coefficients, i.e. the so-called PARCOR coefficients r<sub>k</sub>, which have a certain connection to the shape and structure of the audio bus. Reflection coefficients r<sub>k</sub> and the cylinder portions C of the lossless tube model representing the sound path<sub>k</sub> areas A<sub>k</sub> the connection between is in accordance with Equation (1)
A (k + 1) - A (k) _ <sub>r (k)</sub> « ----------- (1)
A (k + 1) + A (k) where k = 1,2,3, ....
The LPC analysis producing the reflection coefficients used in the invention is utilized in many known speech coding methods. A preferred application of the method according to the invention is the identification of subscribers> 1> 25 in radiotelephone systems, in particular in the pan-European digital radiotelephone system GSM. GSM Recommendation 06.10 defines very precisely the RPE-LTP (Regular Pulse Excitation-Long Term Prediction) method used in the system. The use of the method according to the invention in connection with this speech coding method is advantageous because the reflection coefficients required in the invention are obtained as an intermediate result in the above-mentioned RPE-LPC coding method. In a preferred embodiment of the invention, all steps of the method up to the calculation of the reflection coefficients follow said speech coding algorithm according to the GSM 06.10 recommendation, and for details of these steps reference is made to said recommendation. In the following, these method steps will be described only in general terms essentially for an understanding of the invention with reference to the flow chart of Figure 3.
In Fig. 3, in block 10, the input signal IN is sampled at a sampling frequency of 8 kHz and a sequence of 8-bit samples s is formed.<sub>o</sub>. In block 11, a direct component (dc component) is removed from the samples to remove any interfering side noise that may be generated during coding. Then, in block 12, the sample signal is pre-emphasized by weighting the high signal frequencies with a first order FIR filter. In block 13, the samples are segmented into 160 sample frames, with a frame duration of about 20 ms.
In block 14, the spectrum of the speech signal is modeled by performing an LPC analysis on each frame with the degree p = 8 by the autocorrelation method. In this case, p + 1 values of the autocorrelation function ACF are calculated from the frame
160
ACF (k) = Σ s (i) s (ik) (2) i = l where k = 0,1, ..., 8.
Instead of the autocorrelation function can be used
<img file="FI91925B_D0001.tif" />
other suitable function, such as, for example, a covariance function. From the obtained values of the autocorrelation function, the eight so-called short-term analysis filters used in the speech coder are calculated by Schur recursion or other suitable recursion method. reflection coefficient r<sub>k </sub>set of values. Schur recursion produces new reflection coefficients every 20 ms. In a preferred embodiment of the invention, the coefficients are 16-bit and there are 8 of them. By extending the Schur recursion for a longer period, the number of reflection coefficients can be increased if desired.
In block 16, the reflection coefficients r calculated from each frame are calculated<sub>k</sub> the sound path of the speaker with cylindrical sections modeling the lossless tube of each cylinder section C<sub>k</sub> area A<sub>k</sub>. Since the Schur recursion produces new reflection coefficients every 20 ms, the areas for each cylinder part C<sub>k</sub> 50 pcs / s are obtained. After calculating the cylinder areas of the lossless tube for n pieces of frames, in step 17, the C of the cylinder parts of the N lossless tube model thus obtained is calculated.<sub>k</sub> area averages A<sub>k ave</sub> and is determined for each cylinder section C<sub>k</sub> maximum cross-sectional area A<sub>k nax</sub>, which has occurred during these N frames. Then, in step 18, the C parts of the cylinder parts of the lossless tube modeling the speaker sound path thus obtained are compared.<sub>k</sub> average areas A<sub>k ave</sub> and maximum areas A<sub>k max</sub> a lossless tube model of at least one known speaker stored in the average and maximum areas of the cylinder parts. If the calculated average lossless tube shape based on the compared parameters corresponds to one of the stored models, the decision block 19 proceeds to block 21, where the speaker is confirmed as an identified person indicated by that model. If the calculated parameters do not correspond to the corresponding parameters of any of the stored models, the decision block 19 proceeds to block 20, where the speaker is assigned as unknown.
For example, in a radiotelephone system, block 21 may allow, for example, the establishment of a connection or the use of a service, and block 20 accordingly prevents these operations.
Calculating and storing new models for identification can be performed in a substantially similar procedure to the flowchart in Figure 3, except that after calculating the average and maximum areas in block 18, these areas are stored as a speech-specific file in system memory along with other necessary personal information such as name, phone number, etc. , with.
In another embodiment of the invention, the analysis used in the identification is refined to the sound level so that the averages of the cross-sectional areas of the lossless tube model modeling the audio path are calculated from the cylinder areas of instantaneous lossless tube models generated from the speech signal to be analyzed. The duration of one sound is quite long, so several, even dozens of time-lossless tube models can be calculated from one sound in a speech signal. This is illustrated in Figure 4, which shows four time-sequential instantaneous lossless tube models S1-S4. It can be clearly seen from Figure 4 that the radii (and cross-sectional areas) of the individual cylinders of the lossless tube change over time. For example, the instantaneous models S1, S2, and S3 could, roughly classified, be formed during the same sound, in which case they could be averaged. On the other hand, the model S4 is clearly different and related to a different sound and is therefore not taken into account in the calculation of the average.
Audio level recognition will now be described with reference to the block diagram of Figure 5. Although recognition can already be done on the basis of one sound, at least two different sounds, e.g. vowels and / or consonants, are preferably used in the identification, which correspond to the corresponding lossless tube patterns recorded by a known speaker. combination table 58. The simple combination table may contain, for example, the average areas of the cylinders of the lossless tube models calculated for the three tones a, e and i, i.e. three different average lossless tube models. This combination table is stored in said speaker-specific file. When the instantaneous lossless tube model calculated from the speech by the reflection coefficients (block 51) is identified (quantized) roughly corresponding to one of these predetermined models (block 52), it is stored in memory (block 53) for later averaging. When a sufficient number of instantaneous lossless tube models have been obtained for each sound, the averages of the cross-sectional areas A of the cylindrical parts of the lossless tube model are calculated separately for each sound.<sub>i; |</sub> (block 55), which are then compared to the cylinder cross-sectional areas A of the corresponding models stored in the combination table 58.<sub>Id</sub>. Each model in the combination table 58 has its own comparison function 56 and 57, e.g., a cross-correlation function to estimate the agreement or correlation between the model calculated from speech and that stored model in the combination table 58. An unknown speaker is defined as identified if, in the case of all or a sufficient number of voices, the calculated model and the stored model correlate with sufficient accuracy.
The instantaneous lossless tube pattern 59 formed from the speech signal can be identified in block 52 as corresponding to a particular sound if the cross section of each cylinder portion of the instantaneous lossless tube pattern 59 is within predetermined stored limits for the corresponding sound of a known speaker. These tone-specific and cylinder-specific limit values are stored in the so-called quantize91925 to title table 54. In Fig. 5, reference numerals 60 and 61 illustrate how said tone- and cylinder-specific limit values form a mask or pattern for each tone, for which the instantaneous voice path pattern 59 to be identified is allowed in the allowable areas 60A and 61A 5 (unshaded areas). In Figure 5, the instantaneous voice path model 59 fits the sound mask 60 but clearly does not fit the sound mask 61. Block 52 thus acts as a kind of audio filter that sorts the audio bus models into the correct voice groups a, e, i, etc.
The method can in practice be implemented, for example, programmatically in a conventional signal processor.
The figures and the related description are intended to illustrate the present invention only. In particular, the method according to the invention may vary within the scope of the appended claims.
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
17 members in 9 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 912088 | Finland | A | |
| FI19910002088 | – | – | – |
Members17
| Document | Office | Kind | |
|---|---|---|---|
| FI912088A | Finland | A | |
| WO9220064A1 | World Intellectual Property Organization (WIPO) | A1 | |
| NO924782D0 | Norway | D0 | |
| AU1653092A | Australia | A | |
| NO924782L | Norway | L | |
| EP0537316A1 | European Patent Office (EPO) | A1 | |
| JPH05508242A | Japan | A | |
| FI91925BThis record | Finland | B | |
| FI91925C | Finland | C | |
| AU653811B2 | Australia | B2 | |
| US5522013A | United States of America | A | |
| EP0537316B1 | European Patent Office (EPO) | B1 | |
| AT140552T | Austria | T | |
| DE69212261D1 | Germany | D1 | |
| DE69212261T2 | Germany | T2 | |
| NO306965B1 | Norway | B1 | |
| JP3184525B2 | Japan | B2 |
Numbers
- Publication, DOCDB
- 91925
- Publication, EPODOC
- FI91925B
- Application
- 912088
- Application, DOCDB
- 912088
- Application, EPODOC
- FI19910002088
Titles3
- Finnish
- Menetelmä puhujan tunnistamiseksi
- Swedish
- Förfarande för identifiering av en talare
- English
- A method for identifying the speaker
Classification
- CPC, 1
- G10L17/02
- IPC, 3
- G10L15 10
- G10L15 02
- G10L17 02