Device and method for eliminating echo, and sound reproducing device
Abstract
[Task] There was no echo elimination method for a system that outputs an output signal with a high sampling frequency due to band expansion and a signal with a wide audio frequency band while the input sampling frequency remains at 8 KHz.
Solution.The downsampling circuit 32 converts the sampling frequency 16KHz of the wideband audio signal to be the output signal into the sampling frequency 8KHz of the narrowband audio signal input from the input terminal 10. The adaptive filter 37 estimates an echo signal having echopath characteristics by the echopath filter 34 that wraps around from the speaker 28 to the microphone 35 using the wideband audio signal whose sampling frequency is downsampled to 8 KHz by the downsampling circuit 32. The subtraction circuit 36 subtracts the estimated echo signal estimated by the adaptive filter 37 from the microphone input signal.
Term
Term ended
Projected expiry passed 13 January 2019, 7.7 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
24 claims: 9 independent, 15 dependent
- 1【特許請求の範囲】 【請求項1】 第1のサンプリング周波数のマイクロホンから入ったマイク入力信号から、第2のサンプリング周波数のスピーカ出力音声信号に基づいて発生するエコー成分を消去するエコー消去装置であって、 上記スピーカ出力音声信号の第2のサンプリング周波数を上記マイク入力音声信号の第1のサンプリング周波数に変換するサンプリング周波数変換手段と、 上記サンプリング周波数変換手段によりサンプリング周波数が変換された上記出力音声信号を用いて上記スピーカから上記マイクロホンに回り込むエコー信号を推定するエコー推定手段と、 上記マイク入力信号から上記エコー推定手段により推定された推定エコー信号を減算する減算手段とを備えることを特徴とするエコー消去装置。
- 2【請求項2】 上記出力音声信号の周波数帯域を上記マイク入力音声信号の周波数帯域の範囲内に制限する帯域制限手段を備え、この帯域制限手段で帯域制限された出力信号のサンプリング周波数を上記サンプリング周波数変換手段が第1のサンプリング周波数に変換し、この変換した出力を用いて上記エコー推定手段がエコー信号を推定し、この推定エコー信号を減算手段が上記マイク入力信号から減算することを特徴とする請求項1記載のエコー消去装置。
- 3【請求項3】 上記マイク入力音声信号の第1のサンプリング周波数が上記出力音声信号の第2のサンプリング周波数の整数n分の一であるとき、上記サンプリング周波数変換手段は上記出力音声信号をn分の一に間引くことを特徴とする請求項1記載のエコー消去装置。
- 4【請求項4】 狭帯域音声信号を入力とし、広帯域音声信号を出力とする帯域幅の拡張を伴った音声再生装置に発生するエコーを消去するエコー消去装置であって、サンプリング周波数変換手段によりサンプリング周波数が変換された上記広帯域音声信号を用いて上記エコー推定手段が推定した推定エコー信号を上記減算手段により上記マイク入力信号から減算することを特徴とする請求項1記載のエコー消去装置。
- 5【請求項5】 第1のサンプリング周波数の入力音声信号を第2のサンプリング周波数の出力音声信号に変換してスピーカから発音するときに、第1のサンプリング周波数のマイクロホンから入ったマイク入力信号から、上記スピーカの出力音声信号に基づいて発生するエコー成分を消去するエコー消去装置であって、 上記第1のサンプリング周波数と等しく、かつ出力しない信号を用いて、上記スピーカから上記マイクロホンに回り込むエコー成分を推定するエコー推定手段と、 上記マイク入力信号から上記エコー推定手段により推定された推定エコー信号を減算する減算手段とを備えることを特徴とするエコー消去装置。
- 6【請求項6】 上記入力音声信号を狭帯域音声信号とし、上記出力音声信号を広帯域音声信号とする帯域幅の拡張を伴った音声再生装置に発生するエコーを消去するエコー消去装置であって、上記減算手段はサンプリング周波数を変換しない上記狭帯域音声信号を用いて上記エコー推定手段が推定した推定エコー信号を上記マイク入力信号から減算することを特徴とする請求項5記載のエコー消去装置。
- 7【請求項7】 第1のサンプリング周波数のマイクロホンから入ったマイク入力信号から、第2のサンプリング周波数のスピーカ出力音声信号に基づいて発生するエコー成分を消去するエコー消去装置であって、 上記マイク入力信号のサンプリング周波数を第2のサンプリング周波数に変換するサンプリング周波数変換手段と、 上記出力音声信号を用いて上記スピーカから上記マイクロホンに回り込むエコー信号を推定するエコー推定手段と、 上記エコー推定手段により推定された推定エコー信号を、上記周波数変換手段により周波数が変換されたマイク入力信号から減算する減算手段とを備えることを特徴とするエコー消去装置。
- 8【請求項8】 狭帯域音声信号を入力とし、広帯域音声信号を出力とする帯域幅の拡張を伴った音声再生装置に発生するエコーを消去するエコー消去装置であって、上記サンプリング周波数変換手段によりサンプリング周波数が変換された上記マイク入力信号から上記エコー推定手段が上記広帯域音声信号を用いて推定した推定エコー信号を上記減算手段で減算することを特徴とする請求項7記載のエコー消去装置。
- 9【請求項9】 第1のサンプリング周波数のマイクロホンから入ったマイク入力信号から、第2のサンプリング周波数のスピーカ出力音声信号に基づいて発生するエコー成分を消去するエコー消去方法であって、 上記出力音声信号の第2のサンプリング周波数を上記マイク入力音声信号の第1のサンプリング周波数に変換し、このサンプリング周波数が変換された上記出力音声信号を用いて上記スピーカから上記マイクロホンに回り込むエコー信号を推定し、上記マイク入力信号から上記推定エコー信号を減算することを特徴とするエコー消去方法。
- 10【請求項10】 上記出力音声信号の周波数帯域を上記マイク入力音声信号の周波数帯域の範囲内に制限し、帯域制限された出力音声信号のサンプリング周波数を第1のサンプリング周波数に変換し、この変換出力を用いて推定した推定エコー信号を上記マイク入力信号から減算することを特徴とする請求項9記載のエコー消去方法。
- 11【請求項11】 上記マイク入力音声信号の第1のサンプリング周波数が上記出力音声信号の第2のサンプリング周波数の整数n分の一であるとき、上記出力音声信号をn分の一に間引いてサンプリング周波数を変換することを特徴とする請求項9記載のエコー消去方法。
- 12【請求項12】 狭帯域音声信号を入力とし、広帯域音声信号を出力とする帯域幅の拡張を伴った音声再生装置に発生するエコーを消去するエコー消去方法であって、サンプリング周波数が変換された上記広帯域音声信号を用いて推定した推定エコー信号を上記マイク入力信号から減算することを特徴とする請求項9記載のエコー消去方法。
- 13【請求項13】 第1のサンプリング周波数の入力音声信号を第2のサンプリング周波数の出力音声信号に変換してスピーカから発音するときに、第1のサンプリング周波数のマイクロホンから入ったマイク入力信号から、上記スピーカの出力音声信号に基づいて発生するエコー成分を消去するエコー消去方法であって、 上記第1のサンプリング周波数と等しく、かつ出力しない信号を用いて上記スピーカから上記マイクロホンに回り込むエコー成分を推定し、上記マイク入力信号から上記推定エコー信号を減算することを特徴とするエコー消去方法。
- 14【請求項14】 上記入力音声信号を狭帯域音声信号とし、上記出力音声信号を広帯域音声信号とする帯域幅の拡張を伴った音声再生装置に発生するエコーを消去するエコー消去方法であって、サンプリング周波数を変換しない上記狭帯域音声信号を用いて推定した推定エコー信号を上記マイク入力信号から減算することを特徴とする請求項13記載のエコー消去方法。
- 15【請求項15】 第1のサンプリング周波数のマイクロホンから入ったマイク入力信号から、第2のサンプリング周波数のスピーカ出力音声信号に基づいて発生するエコー成分を消去するエコー消去方法であって、 上記マイク入力信号のサンプリング周波数を第2のサンプリング周波数に変換すると共に、上記出力音声信号を用いて上記スピーカから上記マイクロホンに回り込むエコー信号を推定し、上記推定エコー信号を、上記サンプリング周波数が変換されたマイク入力信号から減算することを特徴とするエコー消去方法。
- 16【請求項16】 狭帯域音声信号を入力とし、広帯域音声信号を出力とする帯域幅の拡張を伴った音声再生装置に発生するエコーを消去するエコー消去方法であって、上記サンプリング周波数が変換された上記マイク入力信号から、上記広帯域音声信号を用いて推定した推定エコー信号を減算することを特徴とする請求項15記載のエコー消去方法。
- 17【請求項17】 第1のサンプリング周波数の入力音声信号を第2のサンプリング周波数の出力音声信号に変換してスピーカから発音すると共に、マイクロホンに入った第1のサンプリング周波数のマイク入力信号を処理する音声再生装置において、 上記出力音声信号の第2のサンプリング周波数を上記マイク入力音声信号の第1のサンプリング周波数に変換するサンプリング周波数変換手段と、 上記サンプリング周波数変換手段によりサンプリング周波数が変換された上記出力音声信号を用いて上記スピーカから上記マイクロホンに回り込むエコー信号を推定するエコー推定手段と、 上記マイク入力信号から上記エコー推定手段により推定された推定エコー信号を減算する減算手段とを備えることを特徴とする音声再生装置。
- 18【請求項18】 上記出力音声信号の周波数帯域を上記マイク入力音声信号の周波数帯域の範囲内に制限する帯域制限手段を備え、この帯域制限手段で帯域制限された出力音声信号のサンプリング周波数を上記サンプリング周波数変換手段が第1のサンプリング周波数に変換し、この変換出力を用いて上記エコー推定手段がエコー信号を推定し、この推定エコー信号を上記減算手段が上記マイク入力信号から減算することを特徴とする請求項17記載の音声再生装置。
- 19【請求項19】 上記マイク入力音声信号の第1のサンプリング周波数が上記出力音声信号の第2のサンプリング周波数の整数n分の一であるとき、上記サンプリング周波数変換手段は上記出力音声信号をn分の一に間引くことを特徴とする請求項17記載の音声再生装置。
- 20【請求項20】 上記入力音声信号を狭帯域音声信号とし、上記出力音声信号を広帯域音声信号とする帯域幅の拡張を伴った音声再生装置であって、サンプリング周波数変換手段によりサンプリング周波数が変換された上記広帯域音声信号を用いて上記エコー推定手段が推定した推定エコー信号を上記減算手段により上記マイク入力信号から減算することを特徴とする請求項17記載の音声再生装置。
- 21【請求項21】 第1のサンプリング周波数の入力音声信号を第2のサンプリング周波数の出力音声信号に変換してスピーカから発音すると共に、マイクに入った第1のサンプリング周波数のマイク入力信号を処理する音声再生装置において、 上記第1のサンプリング周波数と等しく、かつ出力しない信号を用いて、上記スピーカから上記マイクロホンに回り込むエコー成分を推定するエコー推定手段と、 上記マイク入力信号から上記エコー推定手段により推定された推定エコー信号を減算する減算手段とを備えることを特徴とする音声再生装置。
- 22【請求項22】 上記入力音声信号を狭帯域音声信号とし、上記出力音声信号を広帯域音声信号とする帯域幅の拡張を伴った音声再生装置であって、上記減算手段はサンプリング周波数を変換しない上記狭帯域音声信号を用いて上記エコー推定手段が推定した推定エコー信号を上記マイク入力信号から減算することを特徴とする請求項21記載の音声再生装置。
- 23【請求項23】 第1のサンプリング周波数の入力音声信号を第2のサンプリング周波数の出力音声信号に変換してスピーカから発音すると共に、マイクに入った第1のサンプリング周波数のマイク入力信号を処理する音声再生装置において、 上記マイク入力信号の周波数を第2のサンプリング周波数に変換する周波数変換手段と、 上記出力音声信号を用いて上記スピーカから上記マイクロホンに回り込むエコー信号を推定するエコー推定手段と、 上記エコー推定手段により推定された推定エコー信号を、上記周波数変換手段により周波数が変換されたマイク入力信号から減算する減算手段とを備えることを特徴とする音声再生装置。
- 24【請求項24】 上記入力音声信号を狭帯域音声信号とし、上記出力音声信号を広帯域音声信号とする帯域幅の拡張を伴った音声再生装置であって、上記周波数変換手段により周波数が変換された上記マイク入力信号から上記エコー推定手段が上記広帯域音声信号を用いて推定した推定エコー信号を上記減算手段で減算することを特徴とする請求項23記載の音声再生装置。
Independent claims24
193 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
The present invention eliminates the echo component generated based on the speaker output audio signal from the microphone input signal when the microphone input audio signal of the first sampling frequency and the speaker output audio signal of the second sampling frequency are present. The present invention relates to an echo erasing device and a sound reproducing device.
【0002】
[Conventional technology]
In a voice terminal device such as a mobile phone device, there is a problem of echo due to voice wraparound from the terminal speaker to the microphone. To solve this problem, there is an echo cancellation device (canceller) that estimates the echo path characteristics with an adaptive filter and subtracts the estimated echo signal from the microphone input signal. This echo canceller has two signals: a signal obtained by A / D converting the microphone input (hereinafter, simply referred to as a microphone input signal) and a signal before D / A conversion for outputting to a speaker (hereinafter referred to as a speaker output signal). It is used as an input, and the signal with the echo eliminated is used as an output.
【0003】
On the other hand, bandwidth expansion technology that estimates out-of-band components from narrowband signals with a sampling frequency of 8KHz and a voice frequency band of about 300 to 3400Hz, and synthesizes a wideband signal with a sampling frequency of 16KHz and a voice frequency band of about 300 to 6000Hz. There is.
【0004】
[Problems to be Solved by the Invention]
By the way, while the sampling frequency of the microphone input signal remains at 8 KHz, there is no echo elimination method for a system that outputs an output signal with a high sampling frequency due to band expansion and a signal with a wide audio frequency band. .. This is because, in order to eliminate the echo, the estimated echo signal presumed to wrap around from the speaker output signal must be subtracted from the microphone input signal, and for this purpose, the sampling frequency must be adjusted.
【0005】
The present invention has been made in view of the above problems, and an object of the present invention is to provide an echo erasing device capable of erasing echo components even if the sampling rates of the microphone input and the speaker output are different.
【0006】
Another object of the present invention is to provide an audio reproduction device capable of erasing an echo component even if the sampling rates of the microphone input and the speaker output are different.
【0007】
[Means for solving problems]
In order to solve the above problems, the echo erasing device according to the present invention generates an echo component from a microphone input signal input from a microphone having a first sampling frequency based on a speaker output audio signal having a second sampling frequency. An echo erasing device for erasing, the sampling frequency converting means for converting the second sampling frequency of the output audio signal to the first sampling frequency of the microphone input audio signal, and the sampling frequency converting means for converting the sampling frequency. It is provided with an echo estimation means that estimates an echo signal that wraps around the microphone from the speaker using the output voice signal, and a subtraction means that subtracts the estimated echo signal estimated by the echo estimation means from the microphone input signal. ..
【0008】
Further, the echo erasing device includes a band limiting means for limiting the frequency band of the output audio signal within the frequency band of the microphone input audio signal, and the sampling frequency of the output signal band limited by the band limiting means. Is converted to the first sampling frequency by the sampling frequency conversion means, the echo estimation means estimates the echo signal using the converted output, and the subtraction means subtracts the estimated echo signal from the microphone input signal.
【0009】
Further, when the first sampling frequency of the microphone input audio signal is an integral nth of the second sampling frequency of the output audio signal, the sampling frequency conversion means reduces the output audio signal to one nth. Thin out.
【0010】
Further, the echo erasing device is an echo erasing device that erases an echo generated in an audio reproduction device with an expansion of a bandwidth that receives a narrow band audio signal as an input and outputs a wide band audio signal, and is used for sampling frequency conversion. The estimated echo signal estimated by the echo estimating means is subtracted from the microphone input signal by the subtracting means using the wideband audio signal whose sampling frequency is converted by the means.
【0011】
In order to solve the above problems, the echo erasing device according to the present invention converts an input audio signal of a first sampling frequency into an output audio signal of a second sampling frequency and produces a first sound from a speaker. An echo erasing device that erases the echo component generated based on the output audio signal of the speaker from the microphone input signal input from the microphone of the sampling frequency, and uses a signal that is equal to the first sampling frequency and does not output. Further, it includes an echo estimation means for estimating an echo component that wraps around the microphone from the speaker, and a subtraction means for subtracting the estimated echo signal estimated by the echo estimation means from the microphone input signal.
【0012】
In order to solve the above problems, the echo erasing device according to the present invention generates an echo component from a microphone input signal input from a microphone having a first sampling frequency based on a speaker output audio signal having a second sampling frequency. An echo erasing device for erasing, which estimates an echo signal that wraps around the microphone from the speaker using the sampling frequency conversion means that converts the sampling frequency of the microphone input signal to the second sampling frequency and the output audio signal. It includes an echo estimation means and a subtraction means for subtracting the estimated echo signal estimated by the echo estimation means from the microphone input signal whose frequency has been converted by the frequency conversion means.
【0013】
Further, in the echo erasing method according to the present invention, in order to solve the above problems, an echo generated from a microphone input signal input from a microphone having a first sampling frequency based on a speaker output audio signal having a second sampling frequency. This is an echo erasing method for erasing components, in which the second sampling frequency of the output audio signal is converted to the first sampling frequency of the microphone input audio signal, and the output audio signal to which this sampling frequency is converted is used. The echo signal that wraps around the microphone is estimated from the speaker, and the estimated echo signal is subtracted from the microphone input signal.
【0014】
In this echo erasing method, the frequency band of the output audio signal is limited to the range of the frequency band of the microphone input audio signal, the sampling frequency of the band-limited output audio signal is converted to the first sampling frequency, and this method is performed. The estimated echo signal estimated using the conversion output is subtracted from the microphone input signal.
【0015】
Further, when the first sampling frequency of the microphone input audio signal is an integer nth of the second sampling frequency of the output audio signal, the output audio signal is thinned to one nth to convert the sampling frequency. To do.
【0016】
Further, the echo erasing method is an echo erasing method in which a narrow band audio signal is input and an echo generated in an audio reproduction device with a bandwidth expansion that outputs a wide band audio signal is erased, and the sampling frequency is high. The estimated echo signal estimated using the converted wideband audio signal is subtracted from the microphone input signal.
【0017】
In order to solve the above problems, the echo erasing method according to the present invention has a first method of converting an input audio signal having a first sampling frequency into an output audio signal having a second sampling frequency and producing a sound from a speaker. This is an echo erasing method that eliminates the echo component generated based on the output audio signal of the speaker from the microphone input signal input from the microphone of the sampling frequency, and uses a signal that is equal to the first sampling frequency and does not output. The echo component that wraps around the microphone is estimated from the speaker, and the estimated echo signal is subtracted from the microphone input signal.
【0018】
In the echo erasing method according to the present invention, in order to solve the above problems, an echo component generated based on a speaker output audio signal of a second sampling frequency is generated from a microphone input signal input from a microphone of the first sampling frequency. This is an echo erasing method for erasing, in which the sampling frequency of the microphone input signal is converted to a second sampling frequency, and the echo signal that wraps around the microphone from the speaker is estimated using the output audio signal, and the estimated echo is used. The signal is subtracted from the microphone input signal to which the sampling frequency has been converted.
【0019】
In order to solve the above problems, the voice reproduction device according to the present invention converts the input voice signal of the first sampling frequency into the output voice signal of the second sampling frequency, emits the sound from the speaker, and enters the microphone. In an audio reproduction device that processes a microphone input signal having a first sampling frequency, a sampling frequency conversion means for converting a second sampling frequency of the output audio signal into a first sampling frequency of the microphone input audio signal, and the sampling. An echo estimation means that estimates an echo signal that wraps around the microphone from the speaker using the output audio signal whose sampling frequency has been converted by the frequency conversion means, and an estimated echo signal that is estimated by the echo estimation means from the microphone input signal. It is provided with a subtraction means for subtracting.
【0020】
In order to solve the above problems, the audio reproduction device according to the present invention converts the input audio signal of the first sampling frequency into the output audio signal of the second sampling frequency, emits the signal from the speaker, and enters the microphone. In an audio reproduction device that processes a microphone input signal of the first sampling frequency, an echo estimation means that estimates an echo component that wraps around the microphone from the speaker using a signal that is equal to the first sampling frequency and does not output. The microphone input signal is provided with a subtracting means for subtracting the estimated echo signal estimated by the echo estimating means.
【0021】
In order to solve the above problems, the audio reproduction device according to the present invention converts the input audio signal of the first sampling frequency into the output audio signal of the second sampling frequency, emits the signal from the speaker, and enters the microphone. In an audio reproduction device that processes a microphone input signal of a first sampling frequency, a frequency conversion means for converting the frequency of the microphone input signal into a second sampling frequency and an output audio signal are used to convert the speaker to the microphone. The echo estimating means for estimating the wraparound echo signal and the subtracting means for subtracting the estimated echo signal estimated by the echo estimating means from the microphone input signal whose frequency is converted by the frequency converting means are provided.
【0022】
As described above, in the present invention, the conventional echo canceller can be used as it is by converting the sampling frequency of the speaker output and matching it with the sampling frequency of the microphone input.
【0023】
As this conversion method, downsampling is performed when the sampling rate is converted to the lower one. Further, downsampling may be performed after limiting the band. Further, when the sampling rate is an integral multiple, a method of simply thinning out and converting may be used. If it is not an integral multiple, a linear interpolation filter or the like may be used without limiting the band.
【0024】
In addition, a method of matching with the microphone input sampling frequency by using an intermediate output having the same sampling rate as the microphone input instead of the output signal was also used. For example, in band expansion processing, a narrow band signal having the same sampling frequency as the microphone input is used.
【0025】
On the contrary, by adjusting the microphone input to the sampling frequency of the output, the echo canceller can be used. In this case, upsampling is also performed. Echo cancellation can be performed by any of the above methods.
【0026】
BEST MODE FOR CARRYING OUT THE INVENTION
Hereinafter, embodiments of the present invention will be described with reference to the drawings.
【0027】
An echo canceller (canceller) in the case where the sampling frequency is the same for both the microphone input voice and the output voice, for example, 8 KHz, has been conventionally known and is represented as shown in FIG. This echo canceller estimates the echo path characteristics by the echo path (reverberation path) impulse response filter 3 by the adaptive filter 5, and subtracts the echo signal estimated by the adaptive filter 5 from the microphone input signal of the microphone 4 by the subtractor 6. If the initial value of the tap coefficient B (z) of the adaptive filter 5 is given by a constant, and the microphone input and speaker 2 output of the same sampling frequency are given, the tap coefficient B (z) is updated and the sampling frequency is the same. The signal after echo cancellation is obtained and can be supplied to the output terminal 7.
【0028】
However, when the band is expanded, the sampling frequencies of the microphone input and the speaker output are different, so this echo canceller cannot be used as it is. However, if the two are made equal by converting the sampling frequency, this echo canceller can be used.
【0029】
FIG. 2 shows a voice band expansion device that requires an echo canceller to which the present invention can be applied. PSI-CELP (Pitch Synchronus Innovation --CELP: Pitch Synchronus Innovation-CELP) or VSELP (Pitch Synchronus Innovation-CELP), which is one of the personal digital cellular (PDC) codecs with a sampling frequency of fs = 8KHz and a voice frequency band of 300 to 3400Hz. This is an example of expanding the band of Vector Sum Excited Linear Prediction and applying it to the sampling frequency fs = 16KHz and the voice frequency band 300 to 6000Hz.
【0030】
This voice bandwidth expansion device expands the voice bandwidth by using, for example, the coding parameters sent from the voice encoder on the transmitting side of the digital mobile phone device.
【0031】
The above coding parameters are decoded by the voice decoder in the previous stage. If the coding method in the voice encoder on the transmitting side is based on the PSI-CELP coding method, the decoding method in this voice decoder is also based on PSI-CELP. Further, if the coding method in the voice encoder is based on the VSELP coding method, the decoding method in the voice decoder in the previous stage of this voice band expansion device is also based on VSELP.
【0032】
The parameters ExcN1, ExcN2 or ExcN related to the excitation source, which is the first coding parameter among the above coding parameters decoded by the voice decoder in the previous stage, are excited source switching & expansion from the input terminals 12, 13 or 14. It is supplied to the unit 20 or the excitation source expansion unit 21. The excitation source expansion output from the excitation source switching & expansion unit 20 or the excitation source expansion unit 21 is switched by the switch 22 depending on whether the coding method is VSELP or PSI-CELP, and is supplied to the LPC synthesis unit 19. ..
【0033】
Further, the linear prediction coefficient αN, which is the second coding parameter among the above coding parameters, is supplied from the input terminal 11 to the αN rN (linear prediction coefficient autocorrelation) conversion circuit 15.
【0034】
In addition, the voice band expansion device includes a wide band codebook (rwCB) 17 created in advance using autocorrelation parameters extracted from the wide band sound. Then, using the wideband codebook 17, the autocorrelation r is expanded by the wideband section 16 to be rw, converted into αw again by the conversion section 18, and then supplied to the LPC synthesis section 19. The LPC synthesizer 19 synthesizes wideband speech based on the wideband linear prediction coefficient αw from the rw αw converter 18 and the extended excitation source from the changeover switch 22.
【0035】
In addition, this audio band expansion device has an upsample circuit 26 that oversamples the sampling frequency of the narrow band audio signal (decoded audio) input from the input terminal 10 from 8 kHz to 16 kHz after being decoded by the audio decoder in the previous stage. , The signal component of the frequency band 300Hz to 3400Hz of the input narrow band audio signal is removed from the composite output from the LPC synthesis circuit 19, the signal component of 3400Hz or higher is extracted, and the high frequency component is suppressed according to the user's preference. High-frequency extraction & suppression filter 23, a multiplier 24 that multiplies the filter output from this filter 23 by the gain adjusted by the gain adjuster from the terminal 25, and an upsample circuit to the output obtained by multiplying the gain by the multiplier 24. It is equipped with an adder 27 that adds the basic narrowband audio signal components of the frequency band 300Hz to 3400Hz with a sampling frequency of 16kHz from 26.
【0036】
Then, a digital audio signal having a frequency band of 300 to 7000 Hz and a sampling frequency of 16 kHz is output from the output terminal 28.
【0037】
This voice band expansion device operates as follows as a whole. This voice band expansion device estimates the wide band parameter from the narrow band parameter and performs wide band LPC synthesis. After that, the low frequency side, which is the frequency band of the original voice, is replaced with an upsampled version of the original voice to 16 KHz. That is, a high frequency pass filter is applied to leave only the high frequency, the high frequency component among the high frequency components is suppressed, the gain is further adjusted, and then the original sound is added to the upsampled one.
【0038】
Here, the estimation of the wideband parameter requires two widening of α and a widening of the excitation source. Further, in order to widen the bandwidth of α, it is necessary to prepare in advance a codebook based on the autocorrelation r, which is a parameter that can be mutually converted with α. The autocorrelation r is widened by the quantization and inverse quantization by this codebook.
【0039】
First, widening the bandwidth of α will be described. Focusing on the fact that α is a filter coefficient representing the spectral envelope, it is once converted to the autocorrelation r, which is a parameter representing another spectral envelope that makes it easy to estimate the high frequency side, and this is widened, and then the wideband autocorrelation rw. Is converted back to αw. Vector quantization is used for expansion. The narrowband autocorrelation rn can be vector-quantized and the corresponding rw can be obtained from the index.
【0040】
Since a certain relationship holds between narrow-band autocorrelation and wide-band autocorrelation as described later, it is only necessary to prepare a codebook based on wide-band autocorrelation, and narrow-band autocorrelation can be vector-quantized by this, and inverse quantum. Wideband autocorrelation can be obtained by the conversion.
【0041】
Assuming that the narrowband signal is band-limited by the wideband signal, the wideband autocorrelation and the narrowband autocorrelation have the relationship shown in Eq. (1) below.
【0042】
[Number 1]
<img file="JP2000200099A_D0001.tif" />【0043】
Here, φ is an autocorrelation, xn is a narrowband signal, xw is a wideband signal, and h is the impulse response of the band limiting filter.
【0044】
Furthermore, the following equation (2) can be obtained from the relationship between the autocorrelation and the power spectrum.
【0045】
[Number 2]
<img file="JP2000200099A_D0002.tif" />【0046】
Considering another band limiting filter having a frequency characteristic equal to the power characteristic of this band limiting filter, and letting this be H', the above equation (2) becomes the following equation (3).
【0047】
[Number 3]
<img file="JP2000200099A_D0003.tif" />【0048】
The passing area and blocking area of this new filter are the same as the original band limiting filter, and the attenuation characteristics are squared. Therefore, this new filter can also be said to be a bandwidth limiting filter. With this in mind, narrowband autocorrelation is simplified to the convolution of wideband autocorrelation with the impulse response of a band-limited filter, i.e., band-limited wideband autocorrelation. That is, the following equation (4) is obtained.
【0049】
[Number 4]
<img file="JP2000200099A_D0004.tif" />【0050】
From the above, when vector-quantizing narrowband autocorrelation, if only a wideband codebook is prepared, the narrowband vector required for quantization can be created by calculation, and the codebook can be prepared in advance from narrowband autocorrelation. You don't have to prepare it.
【0051】
Furthermore, since each rw code vector has a curve that decreases monotonically or gradually increases or decreases, there is no significant change even if it is passed through the low frequency range by H', and rn quantization can be performed directly in the rw codebook. However, since the sampling frequency is 1/2, it is necessary to compare every other order.
【0052】
By dividing the expansion of α into voiced sound (V) and unvoiced sound (UV), more accurate expansion is possible, so this is also done. Along with this, two codebooks, one for V and one for UV, are used.
【0053】
Next, the expansion of the excitation source will be described. In PSI-CELP, the excitation source in a narrow band is upsampled by inserting a zero value at the zero packing part 12, and the one that generates aliasing distortion is used. Although this method is very simple, it can be said that it is of sufficient quality as an excitation source because the difference in the power of the original voice and the tuning structure is preserved.
【0054】
Then, LPC synthesis is performed by the LPC synthesis circuit 19 using the wideband α and the wideband excitation source obtained above.
【0055】
In addition, since the quality of the voice synthesized by the wideband LPC is poor as it is, the low frequency side is replaced with the original voice SNDN of the codec output. For this purpose, 4KHz or more of the synthesized sound is extracted, while the codec output is upsampled to fs = 16KHz, and these are added.
【0056】
At this time, the gain to be multiplied on the high frequency side by the multiplier 24 can be adjusted by the gain adjuster according to the user's preference. This value is variable because there are large individual differences for each user. The value of the gain on the high frequency side is set in advance by the input from the user, and the multiplication is performed with reference to this value.
【0057】
In addition, before addition, the high frequency side is filtered by the high frequency extraction & suppression filter 23 to slightly suppress the components of about 6 KHz or higher to make the sound easy to hear. This filter coefficient can be selected, and by performing processing with a filter selected in advance, the frequency band on the high frequency side can be selected according to preference. The selection of this filter is also set by the input of the user.
【0058】
However, since the processing using this filter 23 does not affect the power characteristics on the low frequency side, it may be performed after addition. Alternatively, it is also possible to dare to apply a filter that also affects the low frequency side after addition. Wideband audio can be obtained from the above.
【0059】
Here, the creation of a codebook used in the voice band expansion device will be described.
【0060】
The codebook is created by the well-known GLA (Generalized Lloyd Algorithm) method. Broadband audio is divided into frames for a certain period of time, for example, every 20 msec, and the autocorrelation up to the constant order, for example, the sixth order is obtained for each frame. This autocorrelation for each frame is used as training data to create a 6-dimensional codebook. At this time, the voiced sound and the unvoiced sound may be distinguished, and the autocorrelation of the voiced sound and the autocorrelation of the unvoiced sound may be collected separately to create a codebook for each. In this case, the codebook is referred to when α is expanded during the band expansion process. At this time as well, voiced and unvoiced sounds are discriminated and the corresponding codebook is used.
【0061】
The voice band expansion device uses a wideband voiced sound codebook and a wideband unvoiced sound codebook, and the creation thereof will be described in detail with reference to FIGS. 3 and 4.
【0062】
First, a wideband audio signal is prepared for learning, and framed to 20 msec per frame in step S31. Next, in step S32, in each frame, for example, the voiced sound (V) or the unvoiced sound (UV) is classified by examining the frame energy, the value of zero cross, and the like.
【0063】
Then, in step S33, in the wideband voiced sound frame, for example, the autocorrelation parameter r up to the sixth order is calculated. Further, in step S34, the autocorrelation parameter r up to the sixth order, for example, in the wideband unvoiced frame is obtained.
【0064】
From the 6th-order autocorrelation parameters of each frame, wideband parameters are extracted in step S41 of FIG. 4, and a dimension 6 wideband V (UV) codebook is created in step S42 by GLA.
【0065】
As described above, in the audio bandwidth expansion device using the decoding method by PSI-CELP, the wide band audio that the user likes can be provided by making the high frequency gain and the high frequency suppression filter variable.
【0066】
However, in VSELP, the vowels of the original voice are turbid. If the above method is applied to this as it is, annoying noise remains in the high frequency range. To improve this, perform the following processing.
【0067】
PSI-CELP does the codec itself, especially V, to make it audibly smooth, but VSELP doesn't, and it sounds a bit noisy when the bandwidth is extended. Therefore, when creating a wideband excitation source, the following processing is performed by the excitation source switching & expansion unit 20.
【0068】
The excitation source of VSELP is based on the parameters β (long-term prediction coefficient), bL [i] (long-term filter state), γ (gain), c1 [i] (excitation code vector) used in the codec. β * bL [i] + γ * c1 [i] Of these, the former represents the pitch component and the latter represents the noise component, so this is divided into β * bL [i] and γ * c1 [i], and the energy of the former is large in a certain time range. In this case, it is considered that the pitch is a strong voiced sound, so the excitation source is a pulse train, and the part without the pitch component is suppressed to 0. If the energy is not large, the conventional method is used, and the narrow-band excitation source thus created is filled with 0 by the zero-packed portion like PSI-CELP and upsampled to obtain a wide-band excitation source. This has improved the audible quality of voiced sounds in VSELP.
【0069】
Then, the original audio SNDN upsampled by the upsample circuit 26 and the adder 27 add the original audio SNDN. At this time, the high-frequency side is filtered by the high-frequency suppression filter 23 that slightly suppresses the components of about 6 KHz or higher to make the sound easy to hear. This filter coefficient can be selected.
【0070】
Further, the multiplier 24 is used to adjust the gain on the high frequency side according to the user's preference.
【0071】
When the voice bandwidth expansion processing by the above voice bandwidth expansion device and the echo canceller coexist, it is necessary to match the sampling frequencies of the microphone input and the speaker output.
【0072】
As a first specific example for this, the obtained wideband audio signal is downsampled. At this time, in order to suppress the occurrence of aliasing, band limitation is performed, and then downsampling is performed. This first specific example will be described with reference to FIG. The first specific example is a voice band expansion device equipped with an echo canceller 30.
【0073】
The echo canceller 30 uses the microphone input signal that enters the microphone 35 when the voice band expansion device converts a narrow band signal having a sampling frequency of 8 KHz into a wide band signal having a sampling frequency of 16 KHz and emits the signal from the speaker 28. A device that eliminates the echo component generated based on the output, with a downsampling circuit 32 that downsamples the sampling frequency of the wideband signal from 16KHz to 8KHz, and an adaptive filter 37 that estimates the echo signal that wraps around the microphone 35 from the speaker 28. A subtraction circuit 36 for subtracting an estimated echo signal from the microphone input signal is provided.
【0074】
First, the downsampling circuit 32 converts the sampling frequency 16 KHz of the wideband audio signal to be the output audio signal into the sampling frequency 8 KHz of the narrow band audio signal input from the input terminal 10.
【0075】
The adaptive filter 37 can estimate the echo path characteristics of the echo path filter 34, and uses the wideband audio signal whose sampling frequency is downsampled to 8 KHz by the downsampling circuit 32 to wrap around from the speaker 28 to the microphone 35. Estimate the signal.
【0076】
The subtraction circuit 36 subtracts the estimated echo signal estimated by the adaptive filter 37 from the microphone input signal.
【0077】
Further, the echo canceller 30 passes the wideband audio signal through the LPF 31 before supplying it to the downsampling circuit 32 to suppress the occurrence of aliasing. However, in the case of this first specific example, the obtained fs = 8KHz signal is affected by the LPF31 with respect to the speaker 28 output, which should be the input of the echo canceller 30, and especially when the phase characteristic becomes a problem. There is. Therefore, it is preferable to use it when a filter whose phase characteristics are not so problematic can be used.
【0078】
Next, as a second specific example, the wideband audio signal, which is effective when the above phase characteristic becomes a problem and the amount of calculation of the band limiting filter becomes a problem, is downsampled as it is by thinning out, as shown in FIG. The echo canceller 40 shown is listed. The echo canceller 40 is different in that the LPF 31 is removed from the echo canceller 30 shown in FIG. 5 and the downsampling circuit 32 is subjected to the thinning process. However, in this case, aliasing occurs due to thinning out, so it is preferable to use it when aliasing does not pose a problem in terms of echo characteristics and frequency characteristics of output audio.
【0079】
As a third specific example, as shown in FIG. 7, the audio band expansion provided with the echo canceller 45 that stores the data before the upsampling of the narrowband audio supplied from the input terminal 10 and uses it. It is a device. Storage memory is required, but downsampling operations are not required. However, unlike the speaker 28 output, this signal is not affected by the post-process at the time of upsampling or after that, so the echo canceller 45 is the echo path that includes this change in the original echo path. Therefore, it should be used when the effects of these changes are less of a problem.
【0080】
As a fourth specific example, as shown in FIG. 8, the microphone input sound entering the microphone 35 is upsampled by the upsampling circuit 51, and the sampling frequency of the wideband audio signal is adjusted to 16 KHz. In this case, it is assumed that the echo canceller 50 also operates at the sampling frequency of the output audio, that is, 16KHz in this case, and the number of target data is doubled as in the case of 8KHz, and the B (z) of the adaptive filter 37. It is also necessary to change the initial value of. Then, the signal after echo cancellation is downsampled by the down dumple circuit 52 to 8 KHz. The echo path is the original echo path including changes due to upsampling. It should be used when these effects are more advantageous than other methods.
【0081】
As described above, in these band expansion devices, a means for adjusting the sampling frequency is provided inside the echo canceller. In this case, the appearance of the echo canceller uses two systems with different sampling frequencies as inputs, but since the sampling rate is converted internally, it is substantially the same as the conventional one.
【0082】
The sampling frequency ratio is not limited to an integral multiple except for the thinning method. Moreover, it is not limited to the application to the bandwidth expansion technology.
【0083】
Further, the echo cancellation method is suitable for software of a program that functions inside a central processing unit (CPU) or a digital signal processing unit (DSP).
【0084】
[Effect of the invention]
According to the present invention, the echo canceller can be used and operated even when the sampling frequency of the microphone input and the sampling frequency of the speaker output are different, especially when the frequency is doubled as in the case of voice band expansion. , The echo-erased voice is obtained.
[Simple explanation of drawings]
[Figure 1]
It is a block diagram of an echo canceller when the sampling frequency is the same for both input and output, for example, 8 KHz.
[Figure 2]
It is a block diagram of the voice band expansion apparatus which requires an echo canceller to which this invention can apply.
[Fig. 3]
It is a flowchart for demonstrating the training data generation processing used in the wideband voiced sound codebook and the wideband unvoiced sound codebook used in the said voice band expansion apparatus.
[Fig. 4]
It is a flowchart for demonstrating the generation of the said codebook.
[Fig. 5]
It is a block diagram of the voice band expansion apparatus which becomes the 1st specific example.
[Fig. 6]
It is a block diagram of the voice band expansion apparatus which becomes the 2nd specific example.
[Fig. 7]
It is a block diagram of the voice band expansion apparatus which becomes the 3rd specific example.
[Fig. 8]
It is a block diagram of the voice band expansion apparatus which becomes the 4th specific example.
[Explanation of symbols]
28 Speakers, 30 Echo Cancellers, 31 LPF, 32 Downsampling Circuits, 34 Echo Pass Filters, 35 Microphones, 36 Subtractors, 37 Adaptive Filters
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2010224321A | Cited by | Japan | Examiner |
| JPWO2014103099A1 | Cited by | Japan | Search report |
| WO2022024188A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO2010137203A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US7486719B2 | Cited by | United States of America | Applicant |
| WO2004040552A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US9130526B2 | Cited by | United States of America | Applicant |
| JPWO2014103099A1 | Cited by | Japan | Search report |
| WO2023188661A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US9792902B2 | Cited by | United States of America | Applicant |
| WO2015093013A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO2014103099A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2015118307A | Cited by | Japan | Search report |
| US10262653B2 | Cited by | United States of America | Applicant |
| US10127910B2 | Cited by | United States of America | Applicant |
| US8515085B2 | Cited by | United States of America | Applicant |
| JP2011055405A | Cited by | Japan | Examiner |
| JP2005130475A | Cited by | Japan | Examiner |
| JP2012181561A | Cited by | Japan | Examiner |
| JP2015118307A | Cited by | Japan | Search report |
3 members in 2 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 10304303 | Japan | – | |
| 30430398 | Japan | A | |
| 30430398 | Japan | A | |
| 700499 | Japan | A | |
| 304303 | – | – | – |
| JP19980304303 | – | – | – |
| JP19990007004 | – | – | – |
Members3
| Document | Office | Kind | |
|---|---|---|---|
| JP2000200099AThis record | Japan | A | |
| US6694018B1 | United States of America | B1 | |
| JP4296622B2 | Japan | B2 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2000-200099
- Publication, DOCDB
- 2000200099
- Publication, EPODOC
- JP2000200099
- Application
- 11007004
- Application, DOCDB
- 700499
- Application, EPODOC
- JP19990007004
Titles2
- Japanese
- 【発明の名称】エコ―消去装置及び方法、並びに音声再生装置
- English
- INDUSTRIAL APPLICABILITY: Eco-erasing device and method, and voice reproducing device
Classification
- CPC, 1
- H04M9/082
- IPC, 6
- G10L19 00
- G10L15 20
- G10L21 0208
- G10L21 0224
- H04B3 23
- H04M9 08