Computational effectiveness enhancement of frequency domain pitch estimators
Abstract
Estimating a speech signal pitch frequency by determining a speech signal frame line spectrum including spectral lines having respective line amplitudes and frequencies, selecting a predefined number of spectral lines having highest amplitudes, fewer then the total number of the spectral lines, calculating a preliminary utility function over a pitch frequency range to provide a preliminary utility function value for each pitch frequency in the range measuring the compatibility of the selected spectral lines with the pitch frequency, identifying a predefined number of preliminary pitch frequency candidates at least partly responsive to the preliminary utility function, where each candidate is a local maximum of the preliminary utility function, calculating a final utility score for each of the candidates, and selecting any of the candidates to be an estimated pitch frequency of the speech signal at least partly responsive to any of the final utility scores.
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
31 claims: 18 independent, 13 dependent
- 1一種用於估算一語音訊號之一音調頻率之方法,其包含:判定一語音訊號之一訊框之一線頻譜,該頻譜包含具有個別線振幅及線頻率之複數個頻譜線;在該等頻譜線中選擇一預定數目之具有最高振幅的該等頻譜線,其中所選擇的頻譜線之數目少於該等複數個頻譜線之總數;在一音調頻率範圍上計算一初步效用函數,藉此爲該範圍中之每一音調頻率提供一能量測該等所選擇的頻譜線與該音調頻率之一相容性的初步效用函數值;至少部分地回應於該初步效用函數以識別一預定數目之初步音調頻率候選物,其中每一初步音調頻率候選物爲該初步效用函數之一局部最大值;爲該等初步音調頻率候選物中之每一個計算一最終效用得分;且至少部分地回應於該等最終效用得分中之任一個,選擇該等複數個初步音調頻率候選物中之任一個作爲該語音訊號之一估算的音調頻率。
- 2如申請專利範圍第1項之方法,其中該計算一初步效用函數步驟包含:計算一關於該等所選擇的頻譜線中之每一個的影響函數,其中該影響函數在該頻譜線之頻率與任何音調頻率之一比率中係呈周期性;且計算該等影響函數之一疊加。
- 3如申請專利範圍第2項之方法,其中該計算一影響函數步驟包含計算該比率之一函數,其在該比率之整數值處具有最大值,且在該比率之整數值之間具有最小值。
- 4如申請專利範圍第3項之方法,其中該計算一影響函數步驟包含計算一分段線性函數c(f)之值,其在一圍繞f=0之第一間隔中具有一最大值,在一圍繞f=1/2之第二間隔中具有一最小值,及在該等第一間隔與第二間隔之間的一過渡間隔中具有一呈分段線性變化的值。
- 5如申請專利範圍第2項之方法,其中該等影響函數爲分段線性函數,且其中該計算一疊加之步驟包含計算該等影響函數在其各斷點處之值,使得藉由該等斷點之間的內插以判定該初步效用函數。
- 6如申請專利範圍第5項之方法,其中該計算該影響函數步驟包含爲來自該等所選擇的頻譜線中之第一及第二頻譜線接連地計算至少第一及第二影響函數,且其中該計算一初步效用函數步驟包含:計算一包括該第一影響函數之部分效用函數;及藉由計算該第二影響函數在該初步效用函數之該等斷點處之該等值,並計算該初步效用函數在該第二影響函數之該等斷點處之該等值,將該第二影響函數添加至該初步效用函數。
- 7如申請專利範圍第6項之方法,其中該判定一音調頻率候選物步驟包含擇優地選擇該初步效用函數之一局部最大值,其頻率靠近該語音訊號之一先前訊框之一預先估算的音調頻率。
- 8如申請專利範圍第1項之方法,其中該計算一最終效用得分步驟包含:計算一關於該等頻譜線中之每一個的影響函數,其中該影響函數在該頻譜線之該頻率與任何音調頻率之一比率中係呈周期性;及計算該等影響函數之一和。
- 9如申請專利範圍第8項之方法,其中該計算一影響函數步驟包含計算該比率之一函數,其在該比率之整數值處具有最大值,且在該比率之整數值之間具有最小值。
- 10如申請專利範圍第9項之方法,其中該計算該比率之函數的步驟包含計算一分段線性函數c(f)之各值,其在一圍繞f=0之第一間隔中具有一最大值,在一圍繞f=1/2之第二間隔中具有一最小值,及在該等第一間隔與第二間隔之間的一過渡間隔中具有一呈分段線性變化的值。
- 11如申請專利範圍第1項之方法,其中該選擇一音調頻率步驟包含擇優地選擇該等初步音調頻率候選物的其中一個,其具有一高於該等初步音調頻率候選物中之另一個的最終效用得分。
- 12如申請專利範圍第1項之方法,其中該選擇一音調頻率步驟包含擇優地來選擇該等初步音調頻率候選物的其中一個,其具有一高於該等初步音調頻率候選物中另一個的頻率。
- 13如申請專利範圍第1項之方法,其中該選擇一音調頻率步驟包含擇優地選擇該等初步音調頻率候選物的其中一個,其頻率靠近該語音訊號之一先前訊框之一預先估算的音調頻率。
- 14如申請專利範圍第1項之方法,其進一步包含藉由將該所估算的音調頻率之該最終效用得分與一預定臨限值進行比較以判定該語音訊號是有聲還是無聲。
- 15如申請專利範圍第1項之方法,其進一步包含回應於該估算的音調頻率以對該語音訊號進行編碼。
- 16一種用於估算一語音訊號之一音調頻率的裝置,其包含:用於判定一語音訊號之一訊框之一線頻譜的構件,該頻譜包含具有個別線振幅及線頻率之複數個頻譜線;用於在該等頻譜線中選擇一預定數目之具有最高振幅的該等頻譜線的構件,其中所選擇的頻譜線之數目係少於該等複數個頻譜線之總數;用於在一音調頻率範圍上計算一初步效用函數的構件,藉此爲該範圍中之每一音調頻率提供一能量測該等所選擇的頻譜線與該音調頻率之一相容性的初步效用函數值;用於至少部分地回應於該初步效用函數以識別一預定數目之初步音調頻率候選物的構件,其中每一初步音調頻率候選物爲該初步效用函數之一局部最大值;用於爲該等初步音調頻率候選物中之每一個計算一最終效用得分的構件;及用於至少部分地回應於該等最終效用得分中之任一個,選擇該等複數個初步音調頻率候選物中之任一個可成爲該語音訊號之一估算的音調頻率的構件。
- 17如申請專利範圍第16項之裝置,其中可操作該用於計算一初步效用函數之構件,以:計算一關於該等所選擇的頻譜線中之每一個的影響函數,其中該影響函數在該頻譜線之該頻率與任何音調頻率之一比率中係呈周期性;及計算該等影響函數之一疊加。
- 18如申請專利範圍第17項之裝置,其中可操作該用於計算一影響函數之構件,以計算該比率之一函數,其在該比率之整數值處具有最大值,且在該比率之整數值之間具有最小值。
- 19如申請專利範圍第18項之裝置,其中可操作該用於計算一影響函數之構件,以計算一分段線性函數c(f)之各值,其在一圍繞f=0之第一間隔中具有一最大值,在一圍繞f=1/2之第二間隔中具有一最小值,及在該等第一間隔與第二間隔之間的一過渡間隔中具有一呈分段線性變化的值。
- 20如申請專利範圍第17項之裝置,其中該等影響函數爲分段線性函數,且其中可操作該用於計算一疊加之構件,以計算該等影響函數在其各斷點處之值,使得藉由該等斷點之間的內插以判定該初步效用函數。
- 21如申請專利範圍第20項之裝置,其中可操作該用於計算該影響函數之構件,以爲來自該等所選擇的頻譜線中之第一及第二頻譜線接連地計算至少第一及第二影響函數,且其中可操作該用於計算一初步效用函數之構件,以:計算一包括該第一影響函數之部分效用函數;及藉由計算該第二影響函數在該初步效用函數之該等斷點處之該等值,並計算該初步效用函數在該第二影響函數之該等斷點處之該等值,將該第二影響函數添加至該初步效用函數。
- 22如申請專利範圍第21項之裝置,其中可操作該用於判定一音調頻率候選物之構件,以擇優地選擇該初步效用函數之一局部最大值,其頻率靠近該語音訊號之一先前訊框之一預先估算的音調頻率。
- 23如申請專利範圍第16項之裝置,其中可操作該用於計算一最終效用得分之構件,以:計算一關於該等頻譜線中之每一個之影響函數,其中該影響函數在該頻譜線之頻率與任何音調頻率之一比率中係呈周期性;及計算該等影響函數之一和。
- 24如申請專利範圍第23項之裝置,其中可操作該用於計算一影響函數之構件,以計算該比率之一函數,其在該比率之整數值處具有最大值,且在該比率之整數值之間具有最小值。
- 25如申請專利範圍第24項之裝置,其中可操作該用於計算該比率之函數的構件,以計算一分段線性函數c(f)之各值,其在一圍繞f=0之第一間隔中具有一最大值,在一圍繞f=1/2之第二間隔中具有一最小值,及在該等第一間隔與第二間隔之間的一過渡間隔中具有一呈分段線性變化的值。
- 26如申請專利範圍第16項之裝置,其中可操作該用於選擇一音調頻率之構件,以擇優地選擇該等初步音調頻率候選物的其中一個,其具有一高於該等初步音調頻率候選物中另一個之最終效用得分。
- 27如申請專利範圍第16項之裝置,其中可操作該用於選擇一音調頻率之構件,以擇優地選擇該等初步音調頻率候選物的其中一個,其具有一高於該等初步音調頻率候選物中另一個之頻率。
- 28如申請專利範圍第16項之裝置,其中可操作該用於選擇一音調頻率之構件,以擇優地選擇該等初步音調頻率候選物的其中一個,其頻率靠近該語音訊號之一先前訊框之一預先估算的音調頻率。
- 29如申請專利範圍第16項之裝置,其進一步包含用於藉由將該估算的音調頻率之該最終效用得分與一預定臨限值進行比較以判定該語音訊號是有聲還是無聲之構件。
- 30如申請專利範圍第16項之裝置,其進一步包含用於回應於該估算的音調頻率以對該語音訊號進行編碼的構件。
- 31一電腦可讀取媒體,其上具有一電腦程式,該電腦程式包含:一第一程式碼區塊,可操作以判定一語音訊號之一訊框的一線頻譜,該頻譜包含具有個別線振幅及線頻率之複數個頻譜線;一第二程式碼區塊,可操作以在該等頻譜線中選擇一預定數目之具有最高振幅之該等頻譜線,其中所選擇的頻譜線之數目係少於該等複數個頻譜線之總數;一第三程式碼區塊,可操作以在一音調頻率範圍上計算一初步效用函數,藉此爲該範圍中之每一音調頻率提供一能量測該等所選擇的頻譜線與該音調頻率之一相容性的初步效用函數值;一第四程式碼區塊,可操作以至少部分地回應於該初步效用函數以識別一預定數目之初步音調頻率候選物,其中每一初步音調頻率候選物爲該初步效用函數之一局部最大值;一第五程式碼區塊,可操作以爲該等初步音調頻率候選物中之每一個來計算最終效用得分;及一第六程式碼區塊,可操作以至少部分地回應於該等最終效用得分中之任一個,選擇該等複數個初步音調頻率候選物中之任一個可成爲該語音訊號之一估算的音調頻率。
Independent claims31
94 paragraphs, as filed
The enhancement of the calculation efficiency of the frequency domain pitch estimator
The present invention generally relates to a method and apparatus for processing audio signals, and in particular, the present invention relates to a method for estimating the pitch of a voice signal.
By modulating the air flow in the voice track, voice sounds can be produced. The non-sound is derived from the disturbance noise caused by a certain compression position in the vocal tract, while the sound is excited in the larynx by the periodic vibration of the vocal cords. Roughly speaking, the variable period of the throat vibration causes the tone of the voice sound. The low bit rate speech coding mechanism usually separates the modulation from the speech source (voiced or unvoiced), and encodes the two elements separately. In order for the speech to be reconstructed properly, it is necessary to accurately estimate the pitch of the voiced part of the speech during encoding. Various techniques have been developed for this purpose, including time domain and frequency domain methods.
The Fourier transform of periodic signals such as voiced speech has the form of a series of pulses or peaks in the frequency domain. This burst corresponds to the line spectrum of the signal, which can be expressed as a sequence {(a<sub>i</sub>, θ<sub>i</sub>)}, where θ<sub>i</sub>Is the frequency of the peak, and a<sub>i</sub>Is the amplitude of the individual complex-valued line spectrum. In order to determine whether a given speech signal segment is sound or silent, and to calculate the pitch of the segment if it is sound, the time domain signal can be multiplied by a finite smooth window for the first time. Then the Fourier transform of the windowed signal is given by:<maths><img file="TWI282972B_D0001.tif" /></maths>Where W(θ) is the Fourier transform of the window.
Given any tone frequency, the line spectrum corresponding to that tone frequency can include all multiples of line spectrum components at that frequency. Therefore, it can be understood that any frequency that appears in the line spectrum can be a multiple of the pitch frequency of many different candidates. As a result, for any peak appearing in the transformed signal, there will be a sequence of candidate pitch frequencies that can cause that particular peak, where each candidate frequency is the integer divisor of the frequency of the peak. There is this ambiguity: whether the spectrum is analyzed in the frequency domain, or whether it is transformed back to the time domain for further analysis.
Frequency domain pitch estimation is generally based on, for example, analyzing the position and amplitude of the peak in the transformed signal X(θ) by associating the spectrum with the "teeth" of the prototype spectrum "comb". The pitch frequency is given by the comb frequency that maximizes the correlation between the comb function and the transformed speech signal.
Related mechanisms for pitch estimation are commonly known as "cepstrum" mechanisms, in which a log operation is applied to the frequency spectrum of the voice signal, and then the logarithmic spectrum is transformed back to the time domain to generate a cepstrum signal. The pitch frequency is the position of the first peak of the time-domain cepstrum signal. This precisely corresponds to maximizing the correlation between the logarithm of the amplitude corresponding to the line frequency z(i) and cos(ω(i)T) over the period T. For each guess of the pitch period T, the function cos(ωT) is a periodic function of ω. It has a peak at a frequency corresponding to multiple times the tone frequency 1/T. If their peaks happen to coincide with the line frequency, then 1/T is the tone frequency, or a good candidate for some multiple thereof.
A common method for time-domain pitch estimation uses related-type mechanisms that search for an audio period that maximizes the correlation between a signal section concentrated at time t and a signal section concentrated at time tT T. The tone frequency is the reciprocal of T.
Both time-domain and frequency-domain methods for pitch determination are subject to instability and error, and therefore accurate pitch determination is computationally enhanced. For example, in time domain analysis, high-frequency components in the line spectrum cause an oscillation period to be added to the correlation. When the frequency of the component is high, this period changes rapidly with the estimated pitch period T. In this case, even if there is only a slight deviation between T and the true pitch period, the correlation value will be substantially reduced, and may lead to rejection of correct estimation. High frequency components will also add a large number of peaks to the correlation, and these peaks complicate the search for the true maximum. In the frequency domain, a small error in the estimation of the candidate's pitch frequency will cause a large deviation in the estimated value of any spectral component that is an integer multiple of the candidate's frequency.
With currently known techniques, it is necessary to perform a thorough search with high resolution on all possible candidates and their multiples to avoid missing the best candidate pitch for a given input spectrum. Depending on the actual tone frequency, it is often necessary to search the sampled high frequency spectrum, such as above 1500 Hz. At the same time, the analysis interval or window must have sufficient time to capture at least a few cycles of each imaginable pitch candidate in the spectrum, resulting in additional complexity. Similarly, in the time domain, it is necessary to search for the best pitch period T in a wide range of time and with high resolution. Search in either case consumes substantial computing resources. The search criteria cannot be relaxed even during the interval that can be silent, because the interval can be judged as silent only after all candidate pitch frequencies or periods have been eliminated. Although the pitch value from the previous frame is usually used to guide the search of the current value, the search is not limited to the neighborhood of the previous pitch. Otherwise, the error in one interval will always exist in the following interval, and the interference of the voiced section will be silent.
An object of the present invention is to provide an improved method and device for determining the pitch of an audio signal, and especially the pitch of a voice signal.
In one aspect of the present invention, a method for estimating the pitch frequency of a voice signal is provided, which includes: discovering a line spectrum of the signal, the spectrum including spectral lines having individual line amplitudes and line frequencies; and a given pitch frequency The pitch frequency of each candidate in the range calculates a utility function of the compatibility of the spectrum with the pitch frequency of the candidate, the utility function is indicative; and the pitch frequency of the speech signal is estimated in response to the utility function.
In another aspect of the present invention, calculating the utility function includes calculating at least one influence function, the influence function being periodic in the ratio of the frequency of one of the spectral lines to the pitch frequency of the candidate. Calculating the at least one influence function also preferably includes calculating a function of the ratio, which has a maximum value at the integer values of the ratio and a minimum value between the integer values of the ratio. The function for calculating the ratio preferably also includes calculating the value of a piecewise linear function c(f), which has a maximum value in the first interval around f=0 and has a second interval around f=1/2 The minimum value, and has a linearly changing value in the transition interval between the first interval and the second interval.
In another aspect of the present invention, calculating at least one influence function includes calculating individual influence functions for multiple lines in the frequency spectrum, and calculating the utility function includes calculating a superposition of the influence functions. Preferably, the individual influence functions include piecewise linear functions with breakpoints, and calculating the superposition includes calculating the value of the influence function at the breakpoints, so that the interpolation between the breakpoints can be used to obtain Determine the utility function. Calculating individual influence functions also preferably includes successively calculating at least the first and second influence functions for the first and second lines in the spectrum, and calculating the utility function includes calculating a partial utility function including the first influence function, and then By calculating the value of the second influence function at the breakpoint of the partial utility function and calculating the value of the partial utility function at the breakpoint of the second influence function, the second influence function is added to the partial utility function .
In another aspect of the present invention, a method for estimating the pitch frequency of a voice signal is provided, which includes: determining the line spectrum of the frame of the voice signal, the spectrum including a plurality of frequency spectra having individual line amplitudes and line frequencies Line; select a predetermined number of spectrum lines with the highest amplitude among the spectrum lines, wherein the number of the selected spectrum lines is less than the total number of the plurality of spectrum lines; calculate a preliminary utility function on a pitch frequency range , Thereby providing a preliminary utility function value for each tone frequency in the range that can be used to measure the compatibility between the selected spectrum line and the tone frequency; at least in part in response to the preliminary utility function to identify a predetermined number Preliminary pitch frequency candidates, where each preliminary pitch frequency candidate is the local maximum of the preliminary utility function; calculate the final utility score for each preliminary pitch frequency candidate; and at least partially respond to the final utility scores For any one of them, selecting any one of the plurality of preliminary pitch frequency candidates can become one of the estimated pitch frequencies of the speech signal.
In another aspect of the present invention, the step of calculating the preliminary utility function includes: calculating an influence function for each selected spectral line, wherein the influence function is periodic in the ratio of the frequency of the spectral line to any tone frequencySex; and calculate the superposition of these influence functions.
In another aspect of the present invention, the step of calculating the influence function includes calculating a function of the ratio, which has a maximum value at the integer values of the ratio and a minimum value between the integer values of the ratio.
In another aspect of the present invention, the step of calculating the influence function includes calculating the value of a piecewise linear function c(f), which has a maximum value in a first interval around f=0, and a value around f= The second interval of 1/2 has a minimum value, and there is a piecewise linear change value in the transition interval between the first interval and the second interval.
In another aspect of the present invention, the influence functions are piecewise linear functions, and the step of calculating a superposition includes calculating the value of the influence function at its breakpoints, so that the Interpolate to determine the preliminary utility function.
In another aspect of the present invention, the step of calculating the influence function includes successively calculating at least the first and second influence functions from the first and second spectrum lines of the selected spectrum lines, and wherein the preliminary utility function is calculated The steps include: calculating a partial utility function including the first influence function; and calculating the value of the second influence function at the breakpoint of the preliminary utility function and calculating the preliminary utility function at the breakpoint of the second influence function The value of, the second influence function is added to the preliminary utility function.
In another aspect of the present invention, the step of determining the pitch frequency candidates includes preferentially selecting the local maximum of the preliminary utility function whose frequency is close to the pre-estimated pitch frequency of the previous frame of the speech signal.
In another aspect of the present invention, the step of calculating the final utility score includes: calculating an influence function for each spectral line, wherein the influence function is periodic in the ratio of the frequency of the spectral line to any tone frequency; and Calculate the sum of these influence functions.
In another aspect of the present invention, the step of calculating the influence function includes calculating a function of the ratio, which has a maximum value at the integer values of the ratio and a minimum value between the integer values of the ratio.
In another aspect of the present invention, the step of calculating the function of the ratio includes calculating the value of a piecewise linear function c(f), which has a maximum value in a first interval around f=0, and The second interval where f=1/2 has the minimum value, and the transition interval between the first interval and the second interval has a piecewise linear change value.
In another aspect of the present invention, the step of selecting the pitch frequency includes preferentially selecting one of the preliminary pitch frequency candidates that has a final utility score higher than the other one of the preliminary pitch frequency candidates.
In another aspect of the present invention, the step of selecting the pitch frequency includes preferentially selecting one of the preliminary pitch frequency candidates that has a higher frequency than another of the preliminary pitch frequency candidates.
In another aspect of the present invention, the step of selecting the pitch frequency includes preferentially selecting one of the preliminary pitch frequency candidates whose frequency is close to the pre-estimated pitch frequency of the previous frame of the speech signal.
In another aspect of the present invention, the method further includes determining whether the voice signal is voiced or unvoiced by comparing the final utility score of the estimated pitch frequency with a predetermined threshold.
In another aspect of the present invention, the method further includes encoding the speech signal in response to the estimated pitch frequency.
In another aspect of the present invention, a device for estimating the pitch frequency of a voice signal is provided, which includes: a component for determining the line spectrum of the frame of the voice signal, the spectrum including individual line amplitudes and line frequencies A plurality of spectrum lines; a component used to select a predetermined number of spectrum lines with the highest amplitude among the spectrum lines, wherein the number of selected spectrum lines is less than the total number of the plurality of spectrum lines; The component for calculating the preliminary utility function on the pitch frequency range, thereby providing an energy for each pitch frequency in the range to measure the compatibility of the selected spectrum line with the pitch frequency of the preliminary utility function value; used at least partially The component that responds to the preliminary utility function to identify a predetermined number of preliminary pitch frequency candidates, where each preliminary pitch frequency candidate is the local maximum of the preliminary utility function; used to calculate the final value for each preliminary pitch frequency candidate A component of the utility score; and for at least partially responding to any one of the final utility scores, selecting any one of the plurality of preliminary pitch frequency candidates can be a component for estimating the pitch frequency of one of the speech signals.
In another aspect of the present invention, the component for calculating the preliminary utility function is operated to calculate an influence function for each selected spectrum line and to calculate the superposition of the influence functions, wherein the influence function is on the spectrum line The ratio of the frequency to any tone frequency is periodic.
In another aspect of the present invention, the component for calculating the influence function is operated to calculate a function of the ratio, which has a maximum value at the integer value of the ratio and a minimum value between the integer values of the ratio value.
In another aspect of the present invention, the component used to calculate the influence function is operated to calculate the value of a piecewise linear function c(f), which has a maximum value in a first interval around f=0, in A second interval around f=1/2 has a minimum value, and has a piecewise linear change value in the transition interval between the first interval and the second interval.
In another aspect of the present invention, the influence functions are piecewise linear functions, and the operation is used to calculate a superimposed component to calculate the value of the influence function at its break point, so that the Interpolate between points to determine the preliminary utility function.
In another aspect of the present invention, the means for calculating the influence functions are operated to successively calculate at least the first and second influence functions for the first and second spectrum lines from the selected spectrum line, And the operation is used to calculate the component of the preliminary utility function to calculate a partial utility function including the first influence function, and by calculating the value of the second influence function at the breakpoint of the preliminary utility function and calculating the preliminary utility function in the first The value at the break point of the second influence function, and the second influence function is added to the preliminary utility function.
In another aspect of the present invention, the component for judging pitch frequency candidates is operated to preferentially select the local maximum of the preliminary utility function whose frequency is close to the pre-estimated pitch frequency of the previous frame of the speech signal.
In another aspect of the present invention, the component used to calculate the final utility score is operated to calculate an influence function on each spectrum line and calculate the sum of the influence functions, wherein the frequency of the influence function on the spectrum line is The ratio of any tone frequency is periodic.
In another aspect of the present invention, the component for calculating the influence function is operated to calculate a function of the ratio, which has a maximum value at the integer value of the ratio and a minimum value between the integer values of the ratio value.
In another aspect of the present invention, the means for calculating the function of the ratio is operated to calculate the value of a piecewise linear function c(f), which has a maximum value in a first interval around f=0 , Has a minimum value in a second interval around f=1/2, and has a piecewise linear change value in the transition interval between the first interval and the second interval.
In another aspect of the present invention, the means for selecting the pitch frequency is operated to preferentially select one of the preliminary pitch frequency candidates having a final utility score higher than the other one of the preliminary pitch frequency candidates one of.
In another aspect of the present invention, the means for selecting the pitch frequency is operated to preferentially select one of the preliminary pitch frequency candidates having a frequency higher than the other one of the preliminary pitch frequency candidates .
In another aspect of the present invention, the component for selecting the pitch frequency is operated to preferentially select one of the preliminary pitch frequency candidates whose frequency is close to the pre-estimated pitch frequency of the previous frame of the speech signal.
In another aspect of the present invention, the device further includes means for judging whether the voice signal is voiced or unvoiced by comparing the final utility score of the estimated pitch frequency with a predetermined threshold value.
In another aspect of the present invention, the device further includes means for encoding the speech signal in response to the estimated pitch frequency.
In another aspect of the present invention, a computer program embodied on a computer readable medium is provided. The computer program includes: a first code section, which is operated to determine the line spectrum of the frame of the voice signal, the The spectrum includes a plurality of spectrum lines with individual line amplitudes and line frequencies; the second code section is operated to select a predetermined number of spectrum lines with the highest amplitude among the spectrum lines, wherein the selected spectrum The number of lines is less than the total number of the plurality of spectrum lines; the third code section is operated to calculate the preliminary utility function on the pitch frequency range, thereby providing an energy for each pitch frequency in the range to measure the Wait for the preliminary utility function value of the compatibility of the selected spectrum line with the pitch frequency; the fourth code section is operated to at least partially respond to the preliminary utility function to identify a predetermined number of preliminary pitch frequency candidates, where Each preliminary pitch frequency candidate is the local maximum of the preliminary utility function; the fifth code section is operated to calculate the final utility score for each preliminary pitch frequency candidate; and the sixth code section is operated on Selecting any one of the plurality of preliminary pitch frequency candidates in response at least in part to any one of the final utility scores may become one of the estimated pitch frequencies of the speech signal.
FIG. 1 is a schematic diagram of a system 20 for analyzing and encoding voice signals according to a preferred embodiment of the present invention. The system includes an audio input device 22 such as a microphone, which is coupled to an audio processor 24. Alternatively, the audio input to the processor can be provided by a communication line in either analog or digital form or can be restored from a storage device. The processor 24 preferably includes a general-purpose computer programmed by suitable software for executing the functions described below. The software can be provided to the processor in electronic form via the network, for example, or it can be provided on a physical medium such as a CD-ROM or non-volatile memory. Alternatively or additionally, the processor 24 may include a digital signal processor (DSP) or hard-wired logic.
FIG. 2 is a flowchart schematically illustrating a method for processing voice signals by using the system 20 according to one of the preferred embodiments of the present invention. At the input step 30, the voice signal is input from the device 22 or from another source and digitized for further processing (if the signal is not already in digital form). Divide the digitized signal into frames of appropriate duration and relative offset generally of 25 ms and 10 ms for subsequent processing. In the pitch recognition step 32, the processor 24 extracts the approximate line spectrum of the signal for each frame. As described below, the spectrum can be extracted by analyzing the signal at multiple time intervals at the same time. Preferably, each frame uses two intervals: a short interval for extracting high-frequency pitch values; and a long interval for extracting low-frequency values. Alternatively, a larger number of intervals can be used. The low frequency and high frequency parts together better cover the entire range of possible pitch values. Based on the extracted spectrum, the pitch frequency of the candidate for the current frame can be identified.
In the pitch selection step 34, the best estimate for the pitch frequency of the current frame is selected from the candidate frequencies in all parts of the frequency spectrum. At the utterance decision step 36, based on the selected tone, the system 24 determines whether the current frame is actually voiced or unvoiced. At the output encoding step 38, the voiced/unvoiced decision and the selected pitch frequency are used to encode the current frame. Any suitable encoding method may be used, such as those described in U.S. Patent Application Nos. 09/410,085 and 09/432,081. Preferably, the encoded output includes the modulation characteristics of the audio stream, together with utterance and pitch information. The encoded output is generally transmitted via a communication link and/or stored in the memory 26 (FIG. 1). The method for pitch determination described herein can also be used in other audio processing applications, and then it may be coded or not coded.
Fig. 3 is a flowchart schematically illustrating the details of the pitch recognition step 32 according to a preferred embodiment of the present invention. In the transformation step 40, a dual-window short-time Fourier transform (STFT) is applied to each frame of the speech signal. The range of possible tonal frequencies for speech signals is generally 55 to 420 Hz. Preferably, this range can be divided into two regions: from 55 Hz to the intermediate frequency F<sub>b</sub>(Generally about 90 Hz) lower area; and from F<sub>b</sub>Up to the higher area of 420 Hz. As described below, a short time window can be defined for each frame for searching higher frequency regions, and a long time window can be defined for lower frequency regions. Alternatively, a larger number of adjacent windows can be used. STFT is applied to each time window to calculate the individual high-frequency and low-frequency spectrum of the voice signal.
The processing of the short-window and long-window spectra is preferably performed on separate parallel tracks. In the spectrum estimation steps 42 and 44, it has the form defined above {(a<sub>i</sub>, θ<sub>i</sub>)} The high-frequency and low-frequency line spectra are derived from individual STFT results. At steps 46 and 48 of candidate frequency discovery, the line spectrum is used to find individual sets of high-frequency and low-frequency candidate values of the tones. The pitch candidates are sent to step 34 (FIG. 2) for selecting the best pitch frequency estimation among the candidates. Referring to Figures 4, 5 and 6A-6D, the details of steps 40 to 48 are described below.
Fig. 4 is a block diagram schematically illustrating the details of the transformation step 40 according to a preferred embodiment of the present invention. The windowing block 50 applies a windowing function to the current frame of the speech signal. The windowing function is preferably a Hamming window with a duration of 25 ms known in the art. Depending on the sampling rate, the transform block 52 applies an appropriate frequency transform to the windowed frame. The frequency transform is preferably a fast Fourier transform (FFT) with a resolution of 256 or 512 frequency points. .
Preferably, the Dirichlet kernel function (Dirichlet kernel)<img file="TWI282972B_D0002.tif" />Applied to FFT output coefficient X<sup>d</sup>[k], so that the output of block 52 can be sent to the interpolation block 54, which is used to increase the resolution of the spectrum. The interpolated spectrum coefficients are given:<maths><img file="TWI282972B_D0003.tif" /></maths>
For effective interpolation, a small number of coefficients X<sup>d</sup>[k] is preferably used in the vicinity of each frequency θ. Generally, 16 coefficients are used, and the resolution of the spectrum is doubled in this way, so that the number of points in the interpolated spectrum is L=2N. The output of block 54 is given a short window transformation, which is passed to step 42 (Figure 3).
By changing the short window of the current frame to X<sup>s</sup>Change Y with the short window of the previous frame<sup>s</sup>Combine to calculate the long window transformation to be passed to step 44, which is blocked by the delay block 56. Before combining, at the multiplier 58, the coefficient from the previous frame is multiplied by a phase shift of 2πmk/L, where m is the number of samples in the frame. At the adder 60, the long window spectrum X is generated by adding the short window coefficients from the current frame and the previous frame (with appropriate phase shift)<sup>1</sup>,given:<i>x</i><sup>1</sup>(2<i>πk</i>/<i>L</i>)=<i>x</i><sup><i>S</i></sup>(2<i>πk</i>/<i>L</i>)+<i>Y</i><sup><i>S</i></sup>(2<i>πk</i>/<i>L</i>)exp(<i>j</i>2<i>πmk</i>/<i>L</i>) Formula 3
Here, k is an integer taken from a set of integers so that the frequency 2πk/L spans the entire range of the frequency. The method illustrated in FIG. 4 can therefore allow a little more computational effort to allow the spectrum to be derived for multiple, overlapping windows, requiring the method to perform STFT operations on a single window.
FIG. 5 is a flowchart schematically illustrating the details of the line spectrum estimation steps 42 and 44 according to a preferred embodiment of the present invention. The line spectrum estimation method described in this figure is applied to the long window and short window transformation X(θ) generated in step 40. The purpose of steps 42 and 44 is to determine the estimation of the absolute line spectrum of the current frame {(<i>â</i><sub><i>i</i>∣</sub>∣,<img file="TWI282972B_D0004.tif" />)}. Peak frequency {<img file="TWI282972B_D0005.tif" />The sequence of} is derived from the position of the local maximum of X(θ), and <i>â</i><sub><i>i</i></sub>∣=∣<i>x</i>(<img file="TWI282972B_D0006.tif" />). This estimation is based on the following assumption: Compared with the pitch frequency, the width of the transformed main lobe of the windowing function (block 50) in the frequency domain is small. Therefore, the interaction between adjacent windows in the spectrum is small.
The estimation of the line spectrum starts at the peak finding step 70, and the approximate frequency of the peak is found from the interpolated spectrum (each equation (2)). Generally, these frequencies are calculated to be accurate to integers. At the interpolation step 72, the peak frequency and amplitude are preferably calculated to be accurate to floating point by using quadratic interpolation based on the three nearest integer multiples of 2π/L.
In the distortion evaluation step 74, the array of peaks found in the previous steps is processed to estimate whether the distortion is present in the input speech signal, and if there is, try to correct the distortion. Preferably, the analyzed frequency range is divided into three equal regions, and the maximum value of all amplitudes in the region is calculated for each region. These areas completely cover the frequency range. If the maximum value of any one of the middle frequency and the high frequency range is too high compared to the maximum value in the low frequency range, the attenuation step 76 is performed to attenuate the value of the peak value in the middle and/or high range. It has been heuristically found that if the maximum value in the middle frequency range is greater than 65% of the maximum value in the low frequency range, or if the maximum value in the high frequency range is greater than 45% of the maximum value in the low frequency range, attenuation should be applied . Attenuating the peaks in this way can "regenerate" the spectrum into a more likely shape. Generally speaking, if the speech signal is not distorted at first, step 74 will not change its frequency spectrum.
At peak count step 78, the number of peaks found at step 72 is counted. In the significant peak evaluation step 80, the number of peaks is compared with a predetermined maximum number, where the predetermined maximum number is generally set to seven. If seven or fewer peaks are found, the process proceeds directly to step 46 or 48. Otherwise, at the classification step 82, the peaks are classified in the descending order of their amplitude values. Once a predetermined number of highest peaks have been found (generally equal to the maximum number of peaks used in step 80), a threshold value is set at threshold setting step 84 to be equal to the lowest peak in the group of highest peaks A certain fraction of the amplitude value. In the false peak rejection step 86, the peak value below the threshold is discarded. Or, if at a certain stage of the classification step 82, the sum of the classified peaks exceeds a predetermined score of the total sum of all the peak values found, generally 95%, then the classification process is stopped. Then at step 86, all remaining smaller peaks are discarded. The purpose of this step is to eliminate small spurious peaks that can interfere with the pitch determination at steps 34 and 36 (FIG. 2) or with the voiced/unvoiced decision.
FIG. 6A is a flowchart schematically showing the details of steps 46 and 48 (FIG. 3) of finding candidate pitch frequencies according to a preferred embodiment of the present invention. As shown and described above, these steps are applied to the short-window and long-window line spectrum output from steps 42 and 44, respectively {<i>â</i><sub><i>i</i></sub>∣,<img file="TWI282972B_D0007.tif" />)}. In step 46, a tone candidate whose frequency is higher than a certain threshold is generated, and its utility function is calculated by using the procedure outlined below based on the line spectrum generated in the short analysis interval. In step 48, the line spectrum generated in the long analysis interval also generates a list of tone candidates, and only the tone candidates whose frequencies are lower than the threshold value are used to calculate the utility function. For both the long and short windows, the line spectrum is normalized at the normalization step 90 to produce a normalized amplitude b<sub>i</sub>And frequency f<sub>i</sub>Line of, where b<sub>i</sub>And f<sub>i</sub>Given by the following formula:<maths><img file="TWI282972B_D0008.tif" /></maths><maths><img file="TWI282972B_D0009.tif" /></maths>
In the two equations 4 and 5, i is from 1 to K, where K is the number of spectral lines (peaks), and T<sub>s</sub>Is the sampling interval. In other words, 1/T<sub>S</sub>Is the sampling frequency of the original voice signal, and therefore f<sub>i</sub>Is the frequency in the sample of the spectrum line per second.
At step 92 of selecting the main line, a predetermined number of spectral lines with the highest amplitude value are selected. Next, at step 94, a preliminary utility function is calculated for each candidate pitch frequency in the given pitch frequency range, which can indicate the compatibility of the main spectrum line selected at step 92 with the candidate pitch frequency. Referring to Figures 7 and 8, the definition of the utility function according to a preferred embodiment of the present invention will be described in more detail below, while referring to Figure 6B, a preferred method for calculating the preliminary utility function will be described in more detail below. . Next, at step 96 of selecting preliminary candidates, a predetermined number of pitch frequency candidates are selected by using the preliminary utility function. Referring to FIG. 6C, a preferred method for selecting preliminary candidates will be described in more detail below. Next, at step 98 of calculating the final utility score for the preliminary candidates, a utility score is calculated for each preliminary candidate. Refer to Figure 6D below to describe in more detail a preferred method for calculating the final utility score.
According to a preferred embodiment of the present invention, the utility function is defined by an influence function, such as shown in FIG. 7, which is a curve showing a cycle of the influence function 120 identified as c(f). The influence function preferably has the following characteristics: 1.c(f+1)=c{f}, which means that the function is periodic and the period is 1.
2.0<img file="TWI282972B_D0010.tif" />c(f)<img file="TWI282972B_D0011.tif" />1。
3. c(0)=1.
4. c(f)=c(-f).
5. When r<img file="TWI282972B_D0012.tif" />f<img file="TWI282972B_D0013.tif" />When 1/2, c{f}=0, where r is the parameter <1/2.
6. In [0, r], c(f) is piecewise linear and non-increasing.
In the preferred embodiment shown in Fig. 7, the influence function is trapezoidal, and one period has the following form:<maths><img file="TWI282972B_D0014.tif" /></maths>
Alternatively, another periodic function may be used, preferably a piecewise linear function whose value is zero higher than a certain predetermined distance from the origin.
Fig. 8 is a utility function U(f<sub>p</sub>The curve of the component 130 of) can be generated for the candidate pitch frequency f by using the influence function c(f)<sub>p</sub>The utility function U(f<sub>p</sub>). Based on the line spectrum {(b<sub>i</sub>, f<sub>i</sub>)) to generate the utility function U(f<sub>p</sub>), as given by the following formula:<maths><img file="TWI282972B_D0015.tif" /></maths>
Then it will be used for a single spectrum line (b<sub>i</sub>, f<sub>i</sub>) The component U of this function<sub>i</sub>(f<sub>p</sub>) Is defined as:<maths><img file="TWI282972B_D0016.tif" /></maths>
Figure 8 shows this component, where f<sub>i</sub>=700 Hz, and the component is evaluated at the tone frequency in the range from 50 to 400 Hz. This component contains a plurality of petals 132, 134, 136, 138..., each of which defines the pitch frequency of the candidate in which the candidate can appear and appears in f<sub>i</sub>The area that can cause the frequency range of the spectrum line.
Because of the value b<sub>i</sub>Is standardized, and c(f)<img file="TWI282972B_D0017.tif" />1, so the utility function for any given candidate pitch frequency will be between zero and one. Since according to the definition c(f<sub>i</sub>/f<sub>p</sub>) At f<sub>i</sub>It is cyclical and has a period f<sub>p</sub>, Thus for a given pitch frequency f<sub>p</sub>The high value indicator sequence of the utility function {f<sub>i</sub>Most of the frequencies in} are close to a certain multiple of the tone frequency. Therefore, by calculating the utility function for all possible pitch frequencies in the appropriate frequency range with a specific resolution, and selecting the candidate pitch frequencies by the high-efficiency value, it can be found in a straightforward (but inefficient) way for the current The pitch frequency of the frame.
Now return to Fig. 6A, in the main line selection step 92, select the spectrum line related to the highest amplitude of M from the K lines {(b<sub>ij</sub>, f<sub>ij</sub>)}, j=1, 2, ..., the number M of M. In the preferred embodiment of the present invention, M is set to seven. The preliminary utility function calculated at step 94 mentioned above is given by the following formula:<maths><img file="TWI282972B_D0018.tif" /></maths>Only the M main lines selected at step 92 are used. By referring to FIG. 6B to use the fast method described below, the preliminary utility function can be calculated over the entire pitch frequency search range. Since the influence function c(f) is piecewise linear, Uij(f<sub>p</sub>) The value at any point will be defined by its value at the breakpoint of the function (that is, the point that is not continuous in the first derivative). The breakpoints can be points such as those shown in Figure 8. 140 and 142. Although U<sub>ij</sub>(f<sub>p</sub>) Itself is not piecewise linear, but it can be approximated as a linear function in all regions. UD(f<sub>p</sub>) Fast calculation method using component U<sub>ij</sub>(f<sub>p</sub>) To create a complete function UD(f<sub>p</sub>). Each component U<sub>ij</sub>(f<sub>p</sub>) Add its own breakpoints to the complete function, and at the same time, you can find the value of the utility function between these breakpoints by performing linear interpolation.
Used to establish UD(f<sub>p</sub>) Uses a series of partial utility functions PU<sub>j</sub>, Which by successively assigning each main spectrum line (b<sub>ij</sub>, f<sub>ij</sub>) To add (add in) ingredient U<sub>ij</sub>(f<sub>p</sub>) To produce:<maths><img file="TWI282972B_D0019.tif" /></maths>
Continuing to refer to Figure 6B, the influence function c(f) is repeatedly applied to each main line in the normalized line spectrum (b<sub>ij</sub>, f<sub>ij</sub>) To generate part of the utility function PU<sub>j</sub>The continuity. Since the first component U<sub>il</sub>(f<sub>p</sub>) Start the process. This component corresponds to the main spectral line (b<sub>il</sub>, f<sub>il</sub>). In the utility function component generation step 102, in the search f<sub>p</sub>To calculate U<sub>il</sub>(f<sub>p</sub>) The value at all its breakpoints. Part of the utility function PU at this stage<sub>l</sub>Just equal to U<sub>il</sub>. In subsequent iterations of this step, the new component U<sub>ij</sub>(f<sub>p</sub>) At its own breakpoint and in part of the utility function PU<sub>j-1</sub>(f<sub>p</sub>) Both are judged at all breakpoints. It is better to calculate U by interpolation<sub>ij</sub>(f<sub>p</sub>) In PU<sub>j-1</sub>(f<sub>p</sub>) The value at the breakpoint. Calculate PU in the same way<sub>j-1</sub>(f<sub>p</sub>) In U<sub>ij</sub>(f<sub>p</sub>) The value at the breakpoint. If U<sub>ij</sub>(f<sub>p</sub>) Contains very close to PU<sub>j-1</sub>Among the existing breakpoints, it is better to discard these new breakpoints as superfluous at the discarding step 103. The best way to discard the frequency and the frequency of the existing breakpoint does not exceed 0.0006*f in this way<sub>p</sub><sup>2</sup>The breakpoint. Then at the adding step 104, the U<sub>ij</sub>Add to PU at all remaining breakpoints<sub>j-1</sub>To produce PU<sub>j</sub>。
At the termination step 105, when the last main spectrum line (b<sub>iM</sub>, f<sub>iM</sub>) Component U<sub>iM</sub>When evaluating, complete the process and synthesize the utility function UD(f<sub>p</sub>) Is passed to the preliminary pitch candidate selection step 96. This function has the form of a set of frequency breakpoints and the value of the preliminary utility function at these breakpoints. Otherwise, if the other main spectrum lines are kept to be evaluated, the next main line is taken out at step 106, and the iterative process is continued from step 102 until all main spectrum lines have been evaluated.
It can be observed that the method of Fig. 6B can search all possible pitch frequencies in the search range, but its efficiency for optimization is also the same, because there are few spectral lines involved and only at specific breakpoints, not between pitch frequencies. Calculate the contribution of each line to the utility function over the entire search range.
Fig. 6C is a flowchart schematically illustrating the details of the preliminary pitch candidate selection step 96 (Fig. 6A) according to a preferred embodiment of the present invention. Select a predetermined number of m preliminary pitch candidates. In a preferred embodiment of the present invention, m is set to four. The selection of preliminary pitch frequency candidates is based on the preliminary utility function output from step 94, including all breakpoints that have been found. The breakpoints of the preliminary utility function are evaluated, and some breakpoints are selected as preliminary pitch candidates.
At step 110, they are found to represent the breakpoints of the local maximum of the preliminary utility function. Then, the m (usually four) highest local maxima are selected as the original set of preliminary candidates {(f<sub>1</sub>, UD(f<sub>1</sub>)), (f<sub>2</sub>, UD(f<sub>2</sub>)), ..., (f<sub>m</sub>, UD(f<sub>m</sub>))}. Make (f<sub>k</sub>, UD(f<sub>k</sub>)) is the lowest item of the group, which means that if ik, then UD(fk)<UD(f<sub>j</sub>)。
It is usually necessary to select a pitch close to the pitch of the previous frame for the current frame, as long as the pitch is stable in the previous frame. Therefore, in the previous frame estimation step 112, it is determined whether the pitch of the previous frame is stable. Preferably, if the determined continuity criterion is satisfied on the six previous frames, the pitch is considered stable. It may be required, for example, that the pitch change between consecutive frames is less than a predetermined value such as 22%, and the predetermined value of the utility function is maintained in all frames. If the pitch has been stabilized, at the closest maximum selection step 113, the alternative pitch frequency candidate that is closest to the previous pitch frequency related to the local maximum is selected<img file="TWI282972B_D0020.tif" />. Next, test the frequency of alternative candidates by evaluating the following conditions<img file="TWI282972B_D0021.tif" />And the previous tone frequency f<sub>prev</sub>Closeness between:<maths><img file="TWI282972B_D0022.tif" /></maths>Here, R is set to a predetermined value, such as 1.22. If this condition is met, then at the comparison step 114, relative to the lowest group item UD(f<sub>k</sub>) To evaluate the frequency of alternative candidates<img file="TWI282972B_D0023.tif" />The initial utility function of the place. If the difference between the utility function values at these two frequencies does not exceed the predetermined threshold value T<sub>1</sub>, Such as 0.06, in step 114 by (<img file="TWI282972B_D0024.tif" />) To replace the lowest group item (f<sub>k</sub>, UD(f<sub>k</sub>)). Otherwise, leave the original set of preliminary candidates unchanged. If the pitch of the previous frame is found to be unstable at step 112, and if no local maximum is found near the previous pitch at step 113, the original set of preliminary candidates is also selected.
FIG. 6D is a flowchart schematically illustrating the details of the calculation step 98 (FIG. 6A) of the final utility score related to the preliminary pitch frequency candidate f. It is preferable to apply the sequence of steps shown in FIG. 6D to each preliminary candidate pitch frequency found at step 96. Use all the spectrum lines to execute the final utility score by using Equation 7. At the initialization step 116, the score is set to zero and the first spectrum line (b<sub>1</sub>, f<sub>1</sub>). At step 117, Equation 6 is used to calculate the weighted influence function. This includes the ratio f<sub>1</sub>/f calculation; obtain the fractional part of the ratio to make it deviate from the main period of the influence function (-1, +1); apply formula 6 and multiply by b<sub>1</sub>. Add the obtained value to the score. It is preferable to repeat the steps of FIG. 6D for all spectral lines.
9A and 9B are flowcharts illustrating the details of the optimal pitch frequency selection step 34 (FIG. 2). The utility score of the preliminary pitch candidates calculated in step 98 is used to select the best pitch candidate from the preliminary pitch candidates. Generally, priority is given to high pitch frequencies to avoid mistaking the integer divisor of the pitch frequency (corresponding to an integer multiple of the pitch period) as a true pitch. Therefore, at the frequency classification step 152, the preliminary candidates<img file="TWI282972B_D0025.tif" />Classification such that:<maths><img file="TWI282972B_D0026.tif" /></maths>
In the initialization step 154, it is preferable to set the estimated pitch<img file="TWI282972B_D0027.tif" />Initially set equal to the highest frequency candidate<img file="TWI282972B_D0028.tif" />. Each remaining candidate is evaluated in order of decreasing frequency relative to the current value of the estimated pitch.
At the next frequency step 156, with the candidate pitch<img file="TWI282972B_D0029.tif" />Come and start the evaluation process. In the evaluation step 158, the value of the utility function U(<img file="TWI282972B_D0030.tif" />) And U(<img file="TWI282972B_D0031.tif" />)Compare. like<img file="TWI282972B_D0032.tif" />The utility function is compared with<img file="TWI282972B_D0033.tif" />The utility function is larger by at least a threshold difference T<sub>2</sub>, Or if<img file="TWI282972B_D0034.tif" />Close to<img file="TWI282972B_D0035.tif" />And has a larger utility function, then<img file="TWI282972B_D0036.tif" />Think it is better than the current<img file="TWI282972B_D0037.tif" />The pitch frequency estimation. Preferably, T<sub>2</sub>=0.06, and if 1.17<img file="TWI282972B_D0038.tif" />><img file="TWI282972B_D0039.tif" />, Then<img file="TWI282972B_D0040.tif" />Thought to be close to<img file="TWI282972B_D0041.tif" />. In this case, in the candidate setting step 160,<img file="TWI282972B_D0042.tif" />Set as the new candidate value<img file="TWI282972B_D0043.tif" />. For all preliminary candidates<img file="TWI282972B_D0044.tif" />Repeat steps 156 to 160 in sequence until the last frequency is reached at the last frequency step 162<img file="TWI282972B_D0045.tif" />。
It is usually necessary to select a pitch close to the pitch of the previous frame for the current frame, as long as the pitch is stable in the previous frame. Therefore, in FIG. 9B, a process similar to that used for the selection of preliminary candidates and shown in FIG. 6D can also be applied to the selection of the best pitch candidates. As described above, in the previous frame estimation step 170, it is determined whether the pitch of the previous frame has stabilized. If the pitch has stabilized, select the group {<img file="TWI282972B_D0046.tif" />} The alternate tone frequency closest to the previous tone frequency<img file="TWI282972B_D0047.tif" />. Next, the condition of Equation 11 is evaluated to determine whether the replacement candidate is sufficiently close to the previous pitch frequency. If the condition is met, then in the comparison step 174, the utility function U(<img file="TWI282972B_D0048.tif" />) To evaluate the substitution frequency U(<img file="TWI282972B_D0049.tif" />) At the utility function. If the difference between the utility function values at these two frequencies does not exceed the predetermined threshold value T<sub>2</sub>, Then at step 176 to select an alternative frequency for the current frame<img file="TWI282972B_D0050.tif" />, Making it the estimated pitch frequency<img file="TWI282972B_D0051.tif" />. Generally T<sub>2</sub>Set to 0.06. Otherwise, if the value of the utility function differs by more than T<sub>2</sub>, Then at the candidate frequency setting step 178, continue to use the currently estimated pitch frequency from step 162 in the current frame<img file="TWI282972B_D0052.tif" />As the selected tone frequency. If it is found in step 170 that the pitch of the previous frame is unstable, and if it is found in step 172 that there is no preliminary candidate near the previous pitch, the estimated value is also selected.
FIG. 10 is a flowchart schematically showing the details of the utterance decision step 36 according to a preferred embodiment of the present invention. The decision is based on the utility function U{ at the estimated pitch at the threshold value comparison step 180<img file="TWI282972B_D0053.tif" />} And the threshold T mentioned above<sub>uv</sub>Compare. Generally, T<sub>uv</sub>=0.75. If the utility function is higher than the threshold value, the current frame is classified as sound at the sound setting step 188.
However, during the transition period in the voice stream, the periodic structure of the voice signal can change, which often results in a low value of one of the utility functions even when the current frame should be considered to be voiced. Therefore, when the utility function applied to the current frame is lower than the threshold T<sub>uv</sub>At the time, in the previous frame checking step 182, the utility function of the previous frame is checked. If the estimated pitch of the previous frame has a high-efficiency value, generally at least 0.84, and the pitch of the current frame is found to be close to the pitch of the previous frame at the pitch checking step 184, and the difference is generally no more than 18%, then step 188 The current frame is classified as sound, regardless of its low utility value. Otherwise, in the silent setting step 186, the current frame is classified as silent.
It should be understood that one or more steps in any of the methods described herein can be omitted or performed in an order different from the order shown, without departing from the true spirit and scope of the present invention.
Although the methods and devices disclosed in this article have been described with or without reference to specific computer hardware or software, it should be understood that it is not difficult to implement the methods described in this article in computer hardware or software by using conventional technologies.Method and device. Method and device.
We will understand that the preferred embodiments described above can be cited by examples, and the present invention is not limited to the embodiments specifically shown and described above. On the contrary, the true spirit and scope of the present invention include both the combination and sub-combination of the various features described above, as well as their changes and modifications. These changes and modifications can enable those skilled in the art to read the foregoing description as soon as possible. It will be remembered and not disclosed in the prior art.
<p>20System</p><p>22Audio input equipment</p><p>24Audio processor</p><p>26Memory</p><p>50Window block</p><p>52Transform block</p><p>54Interpolation block</p><p>56Delayed block</p><p>58Multiplier</p><p>60Adder</p><p>120Influence function</p><p>130 Ingredients</p><p>132, 134, 136, 138 petals</p><p>140, 142 points</p>
From the above detailed description of the preferred embodiments of the present invention, together with the reference drawings, the present invention can be more fully understood, in which: Figure 1 is a method for speech analysis and coding according to one of the preferred embodiments of the present invention Schematic illustration of the system; Fig. 2 is a flowchart, which schematically illustrates a method for pitch determination and speech coding according to one of the preferred embodiments of the present invention; Fig. 3 is a flowchart, which It schematically illustrates a method for extracting a line spectrum for a voice signal and finding the pitch value of a candidate according to one of the preferred embodiments of the present invention; FIG. 4 is a block diagram that schematically illustrates the method according to the present invention. One of the preferred embodiments is a method for simultaneously extracting line spectrum through a long time interval and a short time interval; A method for finding peaks in an online spectrum; FIGS. 6A, 6B, 6C, and 6D are flowcharts, which schematically illustrate one of the preferred embodiments of the present invention for evaluating candidates based on an input line spectrum The method of pitch frequency; Fig. 7 is a cycle curve for evaluating the influence function of the pitch frequency of a candidate according to one of the methods of Figs. 6A-6D; Fig. 8 is one of the preferred embodiments according to the present invention by The curve of part of the utility function derived by applying the influence function of Fig. 7 to a component of the line spectrum. 9A and 9B are flowcharts, which schematically illustrate a method for selecting an estimated pitch frequency for a frame of speech from a plurality of candidate pitch frequencies according to one of the preferred embodiments of the present invention; and FIG. 10 is a flowchart schematically illustrating a method for determining whether a voice frame is voiced or unvoiced according to one of the preferred embodiments of the present invention.
6 members in 3 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 10373260 | United States of America | – | |
| 37326003 | United States of America | A | |
| 20030373260 | – | – | – |
| US20030373260 | – | – | – |
Members6
| Document | Office | Kind | |
|---|---|---|---|
| US2004167775A1 | United States of America | A1 | |
| CN1525435A | China | A | |
| TW200508581A | Taiwan Province of China | A | |
| CN1265351C | China | C | |
| TWI282972BThis record | Taiwan Province of China | B | |
| US7272551B2 | United States of America | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Annulment or lapse of patent due to non-payment of feesLapsedMM4A | MM4A |
Numbers
- Publication
- I282972
- Publication, DOCDB
- I282972
- Publication, EPODOC
- TWI282972B
- Application
- 93104139
- Application, DOCDB
- 93104139
- Application, EPODOC
- TW200493104139
Titles4
- Chinese
- 頻率域音調估算器之計算效率的強化
- English
- COMPUTATIONAL EFFECTIVENESS ENHANCEMENT OF FREQUENCY DOMAIN PITCH ESTIMATORS
- Unlabeled
- 頻率域音調估算器之計算效率的強化
- Unlabeled
- The enhancement of the calculation efficiency of the frequency domain pitch estimator
Classification
- CPC, 1
- G10L25/90
- IPC, 6
- G10L19 00
- G01L19 00
- G10L15 28
- G10L19 04
- G10L25 90
- G10L25 93