Voice quality converting method
Abstract
(57) A summary and the purpose The vocal quality conversion method which controls vocal quality is offered maintaining audio quality. Composition Step 41 which conducts the spectrum analysis of the input voice signal, and Step 42 which vector-quantizes the LPC parameter obtained at Step 41 based on the input speaker code book 14 created beforehand, The conversion rule corresponding to the code vector obtained at Step 42, the which shows the feature of the data under voice 13 for input speaker study -- the 1* -- the which shows the feature of formant F1*F4 and the speaker-voices data 23 for conversion of four -- it choosing from the spectrum conversion rule 33 which matched the 1* 4th formant F'1*F'4, and, It consists of Step 43 which changes the FFT parameter (spectrum) of the input voice signal acquired at Step 41 using this conversion rule.
Term
No projected expiry on record.
- Priority and filed
- Published
- Today
2 claims: 1 independent, 1 dependent
- 1[Claims] 1. In a voice quality conversion method for converting a voice input by an input speaker into a voice having a voice quality of a speaker to be converted different from that of the input speaker. A spectrum analysis process for spectral analysis of the waveform of the input voice and A vector quantization process in which the analysis results obtained in the spectrum analysis process are vector-quantized based on a codebook of an input speaker created in advance, and a vector quantization process. The conversion rule corresponding to the code vector obtained in the vector quantization process is selected from the spectral conversion rules in which the characteristics of the input voice and the characteristics of the voice of the speaker to be converted are associated with each other using a statistical method. Then, using this conversion rule, it consists of a spectrum conversion process that converts the spectrum of the waveform of the input voice obtained in the spectrum analysis process. A voice quality conversion method characterized in that voice corresponding to the spectrum converted in the spectrum conversion process is output. 【特許請求の範囲】 【請求項1】 入力話者による入力音声を、前記入力話者と異なる変換対象話者の声質を有する音声に変換する声質変換方法において、 前記入力音声の波形をスペクトル分析するスペクトル分析過程と、 前記スペクトル分析過程で得られた分析結果を、予め作成しておいた入力話者のコードブックに基づいてベクトル量子化するベクトル量子化過程と、 前記ベクトル量子化過程で得られたコードベクトルに対応する変換規則を、前記入力音声の特徴と前記変換対象話者の音声の特徴とを統計的な手法を用いて対応付けたスペクトル変換規則から選択し、この変換規則を用いて、前記スペクトル分析過程で得られた前記入力音声の波形のスペクトルを変換するスペクトル変換過程とからなり、 前記スペクトル変換過程で変換されたスペクトルに応じた音声が出力されることを特徴とする声質変換方法。
79 paragraphs, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Industrial application field]
The present invention relates to a voice quality conversion method for converting a voice of an input speaker into a voice having a voice quality of a desired speaker.
【0002】
[Conventional technology]
Conventionally, various parameters representing speech spectrum entrainment characteristics have been calculated based on a linear predictive analysis / synthesis method (hereinafter referred to as LPC (Linear Predictive Coding) analysis / synthesis method) as a voice quality conversion method for speech. A method of changing the voice quality of voice by changing parameters, a voice waveform between the source speaker (hereinafter referred to as input speaker) and the conversion destination speaker (hereinafter referred to as conversion target speaker), or There is known a method of obtaining a correspondence between spectra in advance and converting the voice generated by the input speaker into the voice of the speaker to be converted according to the correspondence.
【0003】
Here, the outline of the voice quality conversion method based on the LPC analysis / synthesis method will be described. In the method based on the conventional LPC analysis / synthesis method, a linear predictive coefficient (hereinafter referred to as LPC parameter) representing the characteristics of the vocal tract from the vocal cords to the lips, a pulse representing the sound source (vibration of the vocal cords), a Rosemberg wave, etc. Parameters are collected for the input speaker and the speaker to be converted, and the correspondence between the parameters is grasped experimentally or empirically from appropriate sample data to determine the conversion rule of voice quality.
【0004】
Then, when converting the input voice of the input speaker, each of the above parameters is calculated from the input voice signal, each parameter is converted according to the above-mentioned conversion rule determined in advance, and the voice is output by resynthesizing. Converts the voice quality of the speaker to that of the speaker to be converted. Details of the voice quality conversion method based on the above-mentioned LPC analysis / synthesis method are described in, for example, DGCHILDERS and Ke WU, VOICE CONVERSION (Speech Communication 8 (1989) pp.147-158).
【0005】
[Problems to be Solved by the Invention]
By the way, since the conversion rules used in the above-mentioned conventional voice quality conversion method are experimentally or empirically determined from appropriate sample data, it is said that any input voice emitted by the input speaker can be appropriately converted. There is no guarantee.
【0006】
Further, in the voice actually emitted by the input speaker, there is a complicated correlation between the LPC parameter and the parameter representing the sound source pulse, and it is extremely difficult to determine the conversion rule in consideration of all of them. For this reason, when voice quality conversion is performed using a conventional voice quality conversion method, there is a problem that quality deterioration such as a change in phoneme may occur in the converted voice. The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a voice quality conversion method for controlling voice quality while maintaining voice quality.
【0007】
[Means for solving problems]
The voice quality conversion method according to the present invention is a spectrum analysis that spectrally analyzes the waveform of the input voice in the voice quality conversion method that converts the input voice by the input speaker into a voice having the voice quality of the speaker to be converted different from the input speaker. A vector quantization process in which the process and the analysis result obtained in the spectrum analysis process are vector-quantized based on a code book of an input speaker prepared in advance, and a code obtained in the vector quantization process. The conversion rule corresponding to the vector is selected from the spectral conversion rules in which the characteristics of the input voice and the characteristics of the voice of the speaker to be converted are associated with each other by a statistical method, and the conversion rules are used to describe the above. It comprises a spectrum conversion process for converting the spectrum of the waveform of the input voice obtained in the spectrum analysis process, and is characterized in that the sound corresponding to the spectrum converted in the spectrum conversion process is output.
【0008】
[Action]
According to the above method, the result of the spectrum analysis is vector-quantized based on the code book of the input speaker, and the conversion rule corresponding to the code vector obtained by this vector quantization is selected from the spectrum conversion rules. Applies to input audio spectra. The conversion rule associates the characteristics of the input voice with the characteristics of the voice of the speaker to be converted by using a statistical method, and is selectively selected with respect to the input voice. Therefore, it is possible to control the voice quality while maintaining the voice quality.
【0009】
[Example]
Hereinafter, an embodiment of the present invention will be described with reference to the drawings. FIG. 1A is a flowchart showing a partial procedure of the voice quality conversion method according to the embodiment of the present invention. In the procedure shown in this figure, in order to efficiently express a voice signal, a parameter indicating the characteristics of the voice signal (hereinafter referred to as a voice feature amount) is calculated, and the calculated voice feature amount is statistically classified. It is to create a classification table called a codebook. The voice features include LPC parameters by LPC analysis and spectral density by FFT (fast Fourier transform) analysis. Here, an example using LPC parameters will be described.
【0010】
In FIG. 1 (a), first, in step 11, the above-mentioned LPC analysis process is performed on the input speaker learning voice data 13 corresponding to the input voice generated by the input speaker, and the LPC parameter is calculated. Will be done. The LPC analysis is performed on a sufficiently large number of input speaker learning voice data 13 for statistical accuracy. Next, in step 12, clustering (classification) is performed on the collected LPC parameters. As a clustering method, there is an LBG (Linde-Buzo-Gray) algorithm, which is a typical method. Details of the LBG algorithm are described, for example, in "An algorithm for Vector Quantization Design" (IEEE COM-28 (1980-01)) by Linde et al.
【0011】
The input speaker codebook 14 is created through the procedure described above. FIG. 1B is a conceptual diagram showing the configuration of the input speaker codebook 14, and as shown in this figure, the input speaker codebook 14 is usually composed of code vectors 15 of about 256 to 512. In each code vector 15, 16 is a code vector number, for example, a natural number from 1 to 256 is assigned in order. Reference numeral 17 denotes a spectral feature corresponding to the input speaker learning voice data 13, which is composed of several LPC parameters.
【0012】
Next, the process of creating the mapping codebook 28 used in determining the spectral transformation rule will be described with reference to FIG. The mapping code book 28 statistically associates the voice signal of the input speaker with the voice signal of the speaker to be converted. First, in step 21, the conversion target speaker codebook 22 is created from the conversion target speaker learning voice data 23. Since this creation procedure is the same as the procedure shown in FIG. 1 (a), the description thereof will be omitted.
【0013】
Next, in steps 24 and 24, based on the input speaker and the converted speaker codebooks 14 and 22, the input speaker learning voice data 13 and the converted speaker learning voice data 23 are subjected to LPC analysis and LPC analysis, respectively. Vector quantization processing is applied. Here, in the vector quantization process, the code vector 15 having the spectral feature 17 most similar to the LPC parameter obtained by LPC analysis of each voice data 13 and 23 is extracted from each codebook 14 and 22. This is a process of outputting the spectral features 17 in the extracted code vector 15. Details of vector quantization are described, for example, in "Digital Speech Processing" by Sadaoki Furui.
【0014】
By the vector quantization process described above, the conversion target speaker code vector sequence 25 and the input speaker code vector sequence 26 are obtained. Next, in step 27, a mapping code vector that associates the input speaker code vector sequence 26 and the conversion target speaker code vector sequence 25 is generated. A plurality of mapping code vectors are generated, and a mapping code book 28 is created from these mapping code vectors.
【0015】
As a method for generating the mapping code vector, a known method is used in which a plurality of corresponding conversion target speaker code vector series 25 are aggregated for each input speaker code vector series 26 and generated by weighting averaging. Details of this method are described, for example, in "Voice Conversion through vector quantization" (JASJ (E) 11,2 (1990) pp.71-76) by Abe et al.
【0016】
The process of creating the spectral transformation rule 33 using the mapping code book 28 thus created will be described with reference to FIG. The spectrum conversion rule 33 is a rule for converting the formant frequency, which is one of the features related to the individuality of speech. In FIG. 3, first, in steps 31 and 31, formant analysis is performed on each code vector 15 in the input speaker codebook 14 and each mapping code vector in the mapping codebook 28, respectively. As a result, the formant frequency for each vector is obtained.
【0017】
There are many methods for analyzing formant frequencies, and for example, a method based on LPC pole extraction can be easily used. For details of the formant frequency analysis method, see, for example, Itakura et al., "Estimation of speech spectral density and formant frequency by statistical method" (Shingakuron, (1970), 53-A, 1, pp.35-42). Are listed.
【0018】
Next, in step 32, the spectral conversion rule 33 is obtained. Specifically, first, as shown in FIG. 4, the first to fourth formants F1 to F4 in the code vector 15 in the input speaker codebook 14 are obtained. Next, the mapping code vector corresponding to this code vector 15 is searched from the mapping code book 28, and the code vector corresponding to the speaker to be converted is extracted from the mapping code vector. Then, the first to fourth formants F'1 to F'4 in the extracted code vector are obtained, and they are associated with the first to fourth formants F1 to F4, respectively. The association between the two is done automatically or manually.
【0019】
Next, the frequencies ω1, ω2, ω3, ω4 corresponding to the first to fourth formants F1 to F4 and the frequencies ω'1, ω'corresponding to the first to fourth formants F'1 to F'4. Record 2, ω'3, ω'4 in spectrum conversion rule 33. Here, depending on the phonological type, the fourth formant may not exist, and in that case, the fourth formant is not recorded.
【0020】
In this way, the spectral transformation rule 33 is created. An example of the spectrum conversion rule 33 is shown in FIG. As shown in this figure, the spectrum conversion rule 33 is composed of a plurality of records, and each record is assigned a spectrum conversion rule number 34, which is a natural number from 1 to 256. This spectral conversion rule number 34 is assigned to correspond one-to-one with the code vector number 16 in the input speaker codebook 14.
【0021】
In addition, the corresponding frequencies are recorded in each record for each of the first to fourth formants. For example, in the record in which the spectrum conversion rule number is "1", the frequency ω1 (710) and the frequency ω'1 (815) are recorded in association with each other for the first formant.
【0022】
The process of converting an input voice signal into a converted voice signal having a different voice quality using the spectrum conversion rule 33 created through the above process will be described with reference to FIG. In FIG. 6, first, in step 41, a spectrum analysis process is performed on the input voice signal. The spectrum analysis process comprises an LPC analysis process and an FFT analysis process, and LPC parameters and FFT parameters (spectrums) corresponding to the input audio signal are obtained.
【0023】
Next, in step 42, the LPC parameters obtained in step 41 are vector-quantized based on the input speaker codebook 14 created in advance. As a result, the code vector corresponding to the input voice signal is obtained. Next, in step 43, the FFT parameter obtained in step 41 is transformed. This conversion process will be described below.
【0024】
Specifically, first, the record corresponding to the code vector obtained in step 42 is extracted from the spectrum conversion rule 33 created in advance. Then, the formant frequency of the FFT parameter (spectrum) obtained in step 41 is converted according to the conversion rule expressed in the extracted record. Details of the formant frequency conversion method are described in Mizuno et al., "Formant frequency conversion method with a high degree of control freedom" (Sound Lectures, pp.319-340). Stop.
【0025】
In the conversion method of this embodiment, the input voice signal is cut out in units of one pitch, and the formant of the input voice is extracted by LPC pole analysis. Then, when converting the frequency of a certain formant, the desired formant frequency was converted while suppressing the difference between the spectral density of the formant and the spectral density desired in the formant to a certain value or less by iterative processing. Determine omnipolar spectral characteristics. Next, an omnipolar filter having the omnipolar spectral characteristics thus obtained is constructed and repeatedly acted on the original sound until the desired formant frequency characteristics are obtained to convert the sound to the desired formant frequency. ..
【0026】
Next, in step 44, the audio signal is synthesized by IFFT from the FFT parameter (spectrum) obtained by the spectrum conversion in step 43, and the converted audio signal is output. This converted audio signal has the voice quality of the speaker to be converted.
【0027】
As explained above, the first to fourth formants F1 to F4 in the code vector 15 in the input speaker codebook 14 and the first to fourth formants F'1 to in the mapping code vector corresponding to this code vector 15. It is associated with F'4. Further, the mapping code vector is generated from the conversion target speaker codebook 22 which is weighted and averaged corresponding to each code vector 15 in the input speaker codebook 14. Therefore, by using the spectrum conversion rule 33, adaptive conversion can be performed on the input voice. This ensures that the converted audio signal is of high quality.
【0028】
[Effect of the invention]
As described above, according to the present invention, the result of the spectrum analysis is vector-quantized based on the code book of the input speaker, and the conversion rule corresponding to the code vector obtained by this vector quantization is the spectrum. It is selected from the conversion rules and applied to the spectrum of the input voice. The conversion rule associates the characteristics of the input voice with the characteristics of the voice of the speaker to be converted by using a statistical method, and is selectively selected with respect to the input voice. Therefore, there is an effect that the voice quality can be controlled while maintaining the voice quality.
[Simple explanation of drawings]
[Figure 1]
It is a figure for demonstrating the voice quality conversion method by one Example of this invention.
[Figure 2]
It is a figure which shows the creation process of the mapping code book 28.
[Fig. 3]
It is a figure which shows the process of making spectrum conversion rule 33.
[Fig. 4]
It is a figure for demonstrating spectrum transformation rule 33.
[Fig. 5]
It is a conceptual diagram which shows the structure of the spectrum transformation rule 33.
[Fig. 6]
It is a figure which shows the voice quality conversion process using the spectrum conversion rule 33.
[Explanation of symbols]
14 Input Speaker Codebook 22 Speaker codebook to be converted 28 Mapping Codebook 33 Spectral conversion rules
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US7379873B2 | Cited by | United States of America | Applicant |
| JP2008116534A | Cited by | Japan | Search report |
| US8099282B2 | Cited by | United States of America | Applicant |
| JP2020507819A | Cited by | Japan | Search report |
| JP2006330343A | Cited by | Japan | Search report |
| JP2001282267A | Cited by | Japan | Search report |
| WO2007063827A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US7228273B2 | Cited by | United States of America | Applicant |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 24718493 | Japan | A | |
| JP19930247184 | – | – | – |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS |
Numbers
- Publication
- 7-104792
- Publication, DOCDB
- H07104792
- Publication, EPODOC
- JPH07104792
- Application
- 5247184
- Application, DOCDB
- 24718493
- Application, EPODOC
- JP19930247184
Titles2
- Japanese
- 【発明の名称】声質変換方法
- English
- [Title of Invention] Voice Quality Conversion Method
Classification
- IPC, 2
- G10L13 00
- G10L19 00