Systems, methods, apparatus, and computer program products for wideband speech coding
Abstract
An audio coding method is described in which an excitation signal for a first frequency band of an audio signal is used to calculate an excitation signal for a second frequency band separated from the first frequency band of the audio signal.

Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
49 claims: 35 independent, 14 dependent
- 1A method for processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separated from the low-frequency sub-band, the method comprising:filtering the audio signal to obtain a narrow Frequency signal and an ultra high frequency signal;calculate an encoded narrow frequency excitation signal based on the information from the narrow frequency signal;calculate an ultra high frequency excitation signal based on the information from the encoded narrow frequency excitation signal;The information calculation of the high-frequency signal characterizes a plurality of filter parameters of the spectrum envelope of the high-frequency sub-band;and by evaluating the relationship between a signal based on the UHF signal and a signal based on the UHF excitation signal To calculate a plurality of gain factors, wherein the narrow-frequency signal is based on the frequency component in the low-frequency sub-band, and wherein the UHF signal is based on the frequency component in the high-frequency sub-band, and A width of the low-frequency sub-band is at least three kilohertz, and the low-frequency sub-band is separated from the high-frequency sub-band by a distance that is at least equal to half of the width of the low-frequency sub-band. 一種處理一具有在一低頻率次頻帶中及在一與該低頻率次頻帶分開之高頻率次頻帶中之頻率成分的音訊信號的方法,該方法包含:對該音訊信號進行濾波以獲得一窄頻信號及一超高頻信號;基於來自該窄頻信號之資訊計算一經編碼之窄頻激勵信號;基於來自該經編碼之窄頻激勵信號之資訊計算一超高頻激勵信號;基於來自該超高頻信號之資訊計算特性化該高頻率次頻帶之一頻譜包絡的複數個濾波器參數;及藉由評估一基於該超高頻信號之信號與一基於該超高頻激勵信號之信號之間的一時變關係來計算複數個增益因子,其中該窄頻信號係基於該低頻率次頻帶中之該頻率成分,且其中該超高頻信號係基於該高頻率次頻帶中之該頻率成分,且其中該低頻率次頻帶之一寬度為至少三千赫茲,且其中該低頻率次頻帶與該高頻率次頻帶以一距離分開,該距離至少等於該低頻率次頻帶之該寬度之一半。
- 4Such as the method of any one of claims 1 to 3, wherein the plurality of filter parameters include complex FCH filter coefficients that characterize a spectral envelope of a frame of the high-frequency sub-band, and wherein the method includes calculating Characterize the complex FCL filter coefficients of one of the low-frequency sub-bands corresponding to the spectral envelope of a frame, and where FCH is smaller than FCL. 如請求項1至3中任一項之方法,其中該複數個濾波器參數包括特性化該高頻率次頻帶之一訊框之一頻譜包絡的複數FCH個濾波器係數,且其中該方法包括計算特性化該低頻率次頻帶之一對應訊框之一頻譜包絡的複數FCL個濾波器係數,且其中FCH小於FCL。
- 5The method of any one of claims 1 to 4, wherein the filtering the audio signal comprises:re-sampling a signal based on the frequency component in the high-frequency sub-band to obtain a re-sampled signal;and A spectrum inversion operation is performed based on the signal of the resampled signal to obtain a spectrum inverted signal, wherein the UHF signal is based on the spectrum inverted signal. 如請求項1至4中任一項之方法,其中該對該音訊信號進行濾波包括:重取樣一基於該高頻率次頻帶中之該頻率成分之信號以獲得一經重取樣之信號;及對一基於該經重取樣之信號之信號執行一頻譜反轉操作以獲得一經頻譜反轉之信號,其中該超高頻信號係基於該經頻譜反轉之信號。
- 6The method of any one of claims 1 to 5, wherein the calculating the UHF excitation signal includes:raising a signal based on the information from the encoded narrowband excitation signal to a sampling frequency to generate an interpolated And extending a spectrum of a signal based on the interpolated signal to generate a spectrum-spread signal, and wherein the UHF excitation signal is based on the spectrum-spread signal. 如請求項1至5中任一項之方法,其中該計算該超高頻激勵信號包括:將一基於來自該經編碼之窄頻激勵信號之該資訊的信號升高取樣頻率以產生一經內插之信號;及擴展一基於該經內插之信號之信號的頻譜以產生一經頻譜擴展之信號,且其中該超高頻激勵信號係基於該經頻譜擴展之信號。
- 11Such as the method of any one of claims 1 to 10, wherein the high-frequency sub-band includes a frequency range from eight kilohertz (8 kHz) to eight thousand five hundred hertz (8500 Hz), and wherein the high-frequency sub-band includes The frequency range from 13 kilohertz (13 kHz) to 13.5 kilohertz (13,500 Hz). 如請求項1至10中任一項之方法,其中該高頻率次頻帶包括自八千赫茲(8 kHz)至八千五百赫茲(8500 Hz)之頻率範圍,且其中該高頻率次頻帶包括自十三千赫茲(13 kHz)至十三點五千赫茲(13,500 Hz)之頻率範圍。
- 12The method of any one of claims 1 to 11, wherein the audio signal has a frequency component in a frequency sub-band that is different from the frequency sub-band in the low-frequency sub-band, and wherein the filtering the audio signal includes obtaining a frequency component based on The high-frequency signal of the frequency component in the mid-frequency sub-band, and the method includes:calculating a high-frequency excitation signal based on information from the encoded narrow-frequency excitation signal;and calculating characteristics based on the information from the high-frequency signal Transform a plurality of filter parameters of a spectral envelope of the mid-frequency sub-band;and calculate a second complex number by evaluating a time-varying relationship between a signal based on the high-frequency signal and a signal based on the high-frequency excitation signal Gain factor. 如請求項1至11中任一項之方法,其中該音訊信號具有在一不同於該低頻率次頻帶之中頻率次頻帶中之頻率成分,且其中該對該音訊信號進行濾波包括獲得一基於該中頻率次頻帶中之該頻率成分之高頻信號,且其中該方法包括:基於來自該經編碼之窄頻激勵信號之資訊計算一高頻激勵信號;基於來自該高頻信號之資訊計算特性化該中頻率次頻帶之一頻譜包絡的複數個濾波器參數;及藉由評估一基於該高頻信號之信號與一基於該高頻激勵信號之信號之間的一時變關係來計算第二複數個增益因子。
- 14The method of any one of claims 12 and 13, wherein the calculating the UHF excitation signal includes expanding the frequency spectrum of the encoded narrow-frequency excitation signal into a frequency range occupied by the high-frequency subband, and Wherein the calculating the high-frequency excitation signal includes expanding the frequency spectrum of the encoded narrow-frequency excitation signal into a frequency range occupied by a mid-frequency band. 如請求項12及13中任一項之方法,其中該計算該超高頻激勵信號包括將該經編碼之窄頻激勵信號之頻譜擴展至一由該高頻率次頻帶佔據之頻率範圍中,且其中該計算該高頻激勵信號包括將該經編碼之窄頻激勵信號之該頻譜擴展至一由中頻率頻帶佔據之頻率範圍中。
- 15The method of any one of claims 12 to 14, wherein the middle frequency subband includes frequencies between five kilohertz and six kilohertz, and wherein the high frequency subband includes between ten kilohertz and eleven kilohertz Frequency of. 如請求項12至14中任一項之方法,其中該中頻率次頻帶包括五千赫茲與六千赫茲之間的頻率,且其中該高頻率次頻帶包括十千赫茲與十一千赫茲之間的頻率。
- 18The method of any one of claims 12 to 17, wherein the plurality of filter parameters that characterize a spectral envelope of the high-frequency subband includes complex numbers that characterize a spectral envelope of a frame of the high-frequency subband FCH filter coefficients, and the plurality of filter parameters characterizing the spectral envelope of one of the intermediate frequency subbands include complex FCM filter coefficients characterizing one of the intermediate frequency subbands corresponding to a spectral envelope of the frame , And where FCM is less than FCH. 如請求項12至17中任一項之方法,其中特性化該高頻率次頻帶之一頻譜包絡的該複數個濾波器參數包括特性化該高頻率次頻帶之一訊框之一頻譜包絡的複數FCH個濾波器係數,且其中特性化該中頻率次頻帶之一頻譜包絡的該複數個濾波器參數包括特性化該中頻率次頻帶之一對應訊框之一頻譜包絡的複數FCM個濾波器係數,且其中FCM小於FCH。
- 19A device for processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separated from the low-frequency sub-band, the device comprising:for filtering the audio signal A component for obtaining a narrowband signal and a UHF signal;a component for calculating an encoded narrowband excitation signal based on information from the narrowband signal;a component for calculating an encoded narrowband excitation signal based on information from the encoded narrowband excitation signal Means for calculating an ultra-high frequency excitation signal;means for calculating and characterizing a plurality of filter parameters of a spectral envelope of the high-frequency sub-band based on information from the ultra-high frequency signal;and for evaluating a number of filter parameters based on A component for calculating a plurality of gain factors based on a time-varying relationship between a signal of the UHF signal and a signal of the UHF excitation signal, wherein the narrowband signal is based on the frequency component in the low frequency subband , And wherein the UHF signal is based on the frequency component in the high-frequency sub-band, and wherein a width of the low-frequency sub-band is at least three kilohertz, and wherein the low-frequency sub-band and the high-frequency sub-band are Separated by a distance that is at least equal to half of the width of the low-frequency subband. 一種用於處理一具有在一低頻率次頻帶中及在一與該低頻率次頻帶分開之高頻率次頻帶中之頻率成分的音訊信號的裝置,該裝置包含:用於對該音訊信號進行濾波以獲得一窄頻信號及一超高頻信號的構件;用於基於來自該窄頻信號之資訊計算一經編碼之窄頻激勵信號的構件;用於基於來自該經編碼之窄頻激勵信號之資訊計算一超高頻激勵信號的構件;用於基於來自該超高頻信號之資訊計算特性化該高頻率次頻帶之一頻譜包絡的複數個濾波器參數的構件;及用於藉由評估一基於該超高頻信號之信號與一基於該超高頻激勵信號之信號之間的一時變關係來計算複數個增益因子的構件,其中該窄頻信號係基於該低頻率次頻帶中之該頻率成分,且其中該超高頻信號係基於該高頻率次頻帶中之該頻率成分,且其中該低頻率次頻帶之一寬度為至少三千赫茲,且其中該低頻率次頻帶與該高頻率次頻帶以一距離分開,該距離至少等於該低頻率次頻帶之該寬度之一半。
- 22Such as the device of any one of claims 19 to 21, wherein the plurality of filter parameters include complex FCH filter coefficients that characterize a spectral envelope of a frame of the high-frequency sub-band, and wherein the device includes A component of complex FCL filter coefficients of one of the low-frequency sub-bands corresponding to a spectral envelope of the frame is calculated and characterized, and the FCH is smaller than the FCL. 如請求項19至21中任一項之裝置,其中該複數個濾波器參數包括特性化該高頻率次頻帶之一訊框之一頻譜包絡的複數FCH個濾波器係數,且其中該裝置包括用於計算特性化該低頻率次頻帶之一對應訊框之一頻譜包絡的複數FCL個濾波器係數的構件,且其中FCH小於FCL。
- 23The device of any one of claims 19 to 22, wherein the means for filtering the audio signal includes:for re-sampling a signal based on the frequency component in the high-frequency sub-band to obtain a re-sampled A signal component;and a component for performing a spectrum inversion operation on a signal based on the resampled signal to obtain a spectrum inverted signal, wherein the UHF signal is based on the spectrum inverted Signal. 如請求項19至22中任一項之裝置,其中該用於對該音訊信號進行濾波的構件包括:用於重取樣一基於該高頻率次頻帶中之該頻率成分之信號以獲得一經重取樣之信號的構件;及用於對一基於該經重取樣之信號之信號執行一頻譜反轉操作以獲得一經頻譜反轉之信號的構件,其中該超高頻信號係基於該經頻譜反轉之信號。
- 24The device of any one of claims 19 to 23, wherein the means for calculating the UHF excitation signal comprises:for up-sampling a signal based on the information from the encoded narrowband excitation signal Frequency to generate an interpolated signal;and a member for expanding the spectrum of a signal based on the interpolated signal to generate a spectrum-spread signal, and wherein the UHF excitation signal is based on the frequency spectrum The signal of expansion. 如請求項19至23中任一項之裝置,其中該用於計算該超高頻激勵信號的構件包括:用於將一基於來自該經編碼之窄頻激勵信號之該資訊的信號升高取樣頻率以產生一經內插之信號的構件;及用於擴展一基於該經內插之信號之信號的頻譜以產生一經頻譜擴展之信號的構件,且其中該超高頻激勵信號係基於該經頻譜擴展之信號。
- 30The device of any one of claims 19 to 29, wherein the audio signal has a frequency component in a frequency sub-band different from that of the low-frequency sub-band, and wherein the means for filtering the audio signal It includes means for obtaining a high-frequency signal based on the frequency component in the middle-frequency sub-band, and wherein the device includes:a means for calculating a high-frequency excitation signal based on information from the encoded narrow-frequency excitation signal Means;means for calculating and characterizing a plurality of filter parameters of a spectral envelope of the mid-frequency sub-band based on information from the high-frequency signal;and for evaluating a signal based on the high-frequency signal and a component based on A time-varying relationship between the signals of the high-frequency excitation signal is used to calculate the second plurality of gain factors. 如請求項19至29中任一項之裝置,其中該音訊信號具有在一不同於該低頻率次頻帶之中頻率次頻帶中之頻率成分,且其中該用於對該音訊信號進行濾波的構件包括用於獲得一基於該中頻率次頻帶中之該頻率成分之高頻信號的構件,且其中該裝置包括:用於基於來自該經編碼之窄頻激勵信號之資訊計算一高頻激勵信號的構件;用於基於來自該高頻信號之資訊計算特性化該中頻率次頻帶之一頻譜包絡的複數個濾波器參數的構件;及用於藉由評估一基於該高頻信號之信號與一基於該高頻激勵信號之信號之間的一時變關係來計算第二複數個增益因子的構件。
- 32The device of any one of claims 30 and 31, wherein the means for calculating the UHF excitation signal includes spreading the frequency spectrum of the encoded narrow-frequency excitation signal to a frequency occupied by the high-frequency sub-band In the range, and wherein the means for calculating the high-frequency excitation signal includes expanding the frequency spectrum of the encoded narrow-frequency excitation signal into a frequency range occupied by a mid-frequency band. 如請求項30及31中任一項之裝置,其中該用於計算該超高頻激勵信號的構件包括將該經編碼之窄頻激勵信號之頻譜擴展至一由該高頻率次頻帶佔據之頻率範圍中,且其中該用於計算該高頻激勵信號的構件包括將該經編碼之窄頻激勵信號之該頻譜擴展至一由中頻率頻帶佔據之頻率範圍中。
- 33The device of any one of claims 30 to 32, wherein the medium frequency sub-band includes frequencies between five kilohertz and six kilohertz, and wherein the high-frequency sub-band includes frequencies between ten kilohertz and eleven kilohertz Frequency of. 如請求項30至32中任一項之裝置,其中該中頻率次頻帶包括五千赫茲與六千赫茲之間的頻率,且其中該高頻率次頻帶包括十千赫茲與十一千赫茲之間的頻率。
- 36The device of any one of claims 30 to 35, wherein the plurality of filter parameters that characterize a spectrum envelope of the high-frequency sub-band includes complex numbers that characterize a spectrum envelope of a frame of the high-frequency sub-band FCH filter coefficients, and the plurality of filter parameters characterizing the spectral envelope of one of the intermediate frequency subbands include complex FCM filter coefficients characterizing one of the intermediate frequency subbands corresponding to a spectral envelope of the frame , And where FCM is less than FCH. 如請求項30至35中任一項之裝置,其中特性化該高頻率次頻帶之一頻譜包絡的該複數個濾波器參數包括特性化該高頻率次頻帶之一訊框之一頻譜包絡的複數FCH個濾波器係數,且其中特性化該中頻率次頻帶之一頻譜包絡的該複數個濾波器參數包括特性化該中頻率次頻帶之一對應訊框之一頻譜包絡的複數FCM個濾波器係數,且其中FCM小於FCH。
- 37A device for processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separated from the low-frequency sub-band, the device comprising:a filter bank, the filter The group is configured to filter the audio signal to obtain a narrowband signal and a UHF signal;a narrowband encoder configured to calculate an encoded based on information from the narrowband signal The narrow-band excitation signal;and an ultra-high-frequency encoder configured to: (A) calculate an ultra-high-frequency excitation signal based on the information from the encoded narrow-frequency excitation signal, (B ) Based on the information from the UHF signal, calculate and characterize a plurality of filter parameters of the spectral envelope of the high frequency subband, and (C) by evaluating a signal based on the UHF signal and a signal based on the UHF signal A time-varying relationship between the signals of the high-frequency excitation signal is used to calculate a plurality of gain factors, wherein the narrow-frequency signal is based on the frequency component in the low-frequency sub-band, and wherein the UHF signal is based on the high-frequency sub-band The frequency component in a frequency band, and one of the low-frequency sub-bands has a width of at least three kilohertz, and wherein the low-frequency sub-band and the high-frequency sub-band are separated by a distance that is at least equal to the low-frequency sub-band It is half of the width. 一種用於處理一具有在一低頻率次頻帶中及在一與該低頻率次頻帶分開之高頻率次頻帶中之頻率成分的音訊信號的裝置,該裝置包含:一濾波器組,該濾波器組經組態以對該音訊信號進行濾波以獲得一窄頻信號及一超高頻信號;一窄頻編碼器,該窄頻編碼器經組態以基於來自該窄頻信號之資訊計算一經編碼之窄頻激勵信號;及一超高頻編碼器,該超高頻編碼器經組態以:(A)基於來自該經編碼之窄頻激勵信號之資訊計算一超高頻激勵信號,(B)基於來自該超高頻信號之資訊計算特性化該高頻率次頻帶之一頻譜包絡的複數個濾波器參數,及(C)藉由評估一基於該超高頻信號之信號與一基於該超高頻激勵信號之信號之間的一時變關係來計算複數個增益因子,其中該窄頻信號係基於該低頻率次頻帶中之該頻率成分,且其中該超高頻信號係基於該高頻率次頻帶中之該頻率成分,且其中該低頻率次頻帶之一寬度為至少三千赫茲,且其中該低頻率次頻帶與該高頻率次頻帶以一距離分開,該距離至少等於該低頻率次頻帶之該寬度之一半。
- 40Such as the device of any one of claims 37 to 39, wherein the plurality of filter parameters include complex FCH filter coefficients that characterize a spectral envelope of a frame of the high-frequency sub-band, and wherein the narrowband coding The filter is configured to calculate complex FCL filter coefficients that characterize one of the low-frequency sub-bands corresponding to the spectral envelope of a frame, and where FCH is less than FCL. 如請求項37至39中任一項之裝置,其中該複數個濾波器參數包括特性化該高頻率次頻帶之一訊框之一頻譜包絡的複數FCH個濾波器係數,且其中該窄頻編碼器經組態以計算特性化該低頻率次頻帶之一對應訊框之一頻譜包絡的複數FCL個濾波器係數,且其中FCH小於FCL。
- 41The device of any one of claims 37 to 40, wherein the filter bank includes:a resampler configured to resample a signal based on the frequency component in the high-frequency subband to obtain A resampled signal;and a spectrum inversion module configured to perform a spectrum inversion operation on a signal based on the resampled signal to obtain a spectrum inverted signal, The UHF signal is based on the frequency-inverted signal. 如請求項37至40中任一項之裝置,其中該濾波器組包括:一重取樣器,該重取樣器經組態以重取樣一基於該高頻率次頻帶中之該頻率成分之信號以獲得一經重取樣之信號;及一頻譜反轉模組,該頻譜反轉模組經組態以對一基於該經重取樣之信號之信號執行一頻譜反轉操作以獲得一經頻譜反轉之信號,其中該超高頻信號係基於該經頻譜反轉之信號。
- 42Such as the device of any one of claims 37 to 41, wherein the UHF encoder includes:an up-sampling frequency sampler, the up-sampling frequency sampler is configured to base a narrow The information signal of the frequency excitation signal raises the sampling frequency to produce an interpolated signal;and a spectrum expander configured to expand the spectrum of a signal based on the interpolated signal to produce an interpolated signal Spectrum-spread signal, and wherein the UHF excitation signal is based on the spectrum-spread signal. 如請求項37至41中任一項之裝置,其中該超高頻編碼器包括:一升高取樣頻率取樣器,該升高取樣頻率取樣器經組態以將一基於來自該經編碼之窄頻激勵信號之該資訊的信號升高取樣頻率以產生一經內插之信號;及一頻譜擴展器,該頻譜擴展器經組態以擴展一基於該經內插之信號之信號的頻譜以產生一經頻譜擴展之信號,且其中該超高頻激勵信號係基於該經頻譜擴展之信號。
- 43The device of any one of claims 37 to 43, wherein the narrowband signal has a first sampling rate, and wherein the width of the high-frequency subband is greater than 50% of the first sampling rate. 如請求項37至43中任一項之裝置,其中該窄頻信號具有一第一取樣率,且其中該高頻率次頻帶之寬度大於該第一取樣率之百分之五十。
- 45The device of any one of claims 37 to 45, wherein the width of the high-frequency subband is at least six kilohertz. 如請求項37至45中任一項之裝置,其中該高頻率次頻帶之該寬度為至少六千赫茲。
- 46Such as the device of any one of claims 37 to 46, wherein the high frequency sub-band includes a frequency range from eight kilohertz (8 kHz) to eight thousand five hundred hertz (8500 Hz), and wherein the high-frequency sub-band includes The frequency range from 13 kilohertz (13 kHz) to 13.5 kilohertz (13,500 Hz). 如請求項37至46中任一項之裝置,其中該高頻率次頻帶包括自八千赫茲(8 kHz)至八千五百赫茲(8500 Hz)之頻率範圍,且其中該高頻率次頻帶包括自十三千赫茲(13 kHz)至十三點五千赫茲(13,500 Hz)之頻率範圍。
- 47The device of any one of claims 37 to 47, wherein the audio signal has a frequency component in a frequency sub-band different from the middle frequency sub-band of the low-frequency sub-band, and wherein the filter bank is configured to obtain a The high-frequency signal of the frequency component in the mid-frequency sub-band, and wherein the device includes:a high-frequency encoder configured to: (A) based on the encoded narrow-frequency excitation signal Calculate a high-frequency excitation signal based on the information from the high-frequency signal, (B) calculate and characterize a plurality of filter parameters of a spectral envelope of the mid-frequency sub-band based on the information from the high-frequency signal, and (C) evaluate a filter based on the high A time-varying relationship between the signal of the high-frequency signal and a signal based on the high-frequency excitation signal is used to calculate the second plurality of gain factors. 如請求項37至47中任一項之裝置,其中該音訊信號具有在一不同於該低頻率次頻帶之中頻率次頻帶中之頻率成分,且其中該濾波器組經組態以獲得一基於該中頻率次頻帶中之該頻率成分之高頻信號,且其中該裝置包括:一高頻編碼器,該高頻編碼器經組態以:(A)基於來自該經編碼之窄頻激勵信號之資訊計算一高頻激勵信號,(B)基於來自該高頻信號之資訊計算特性化該中頻率次頻帶之一頻譜包絡的複數個濾波器參數,及(C)藉由評估一基於該高頻信號之信號與一基於該高頻激勵信號之信號之間的一時變關係來計算第二複數個增益因子。
Independent claims35
231 paragraphs, as filed
System, method, device and computer program product for broadband speech coding
The present invention relates to speech processing.
<b>Claim of priority based on 35 USC §119</b>
This patent application claims the title of "Systems, Methods, Devices and Computer Program Products for Broadband Speech Coding (SYSTEMS, METHODS, APPARATUS, AND COMPUTER PROGRAM FOR WIDEBAND SPEECH CODING)" filed on June 1, 2010. The priority of provisional application No. 61/350,425 (agent case No. 092086P1) has been assigned to its assignee.
For example, the traditional wireless voice service of the public switched telephone network (PSTN) is based on narrowband audio between 300 Hz and 3400 Hz. This quality is being challenged by the increasing focus on wideband (WB) high definition (HD) voice systems, which are designed to reproduce voice frequencies between 50 Hz and 7 kHz or 8 kHz. Increasing the bandwidth more than twice in this way can lead to significant improvements in perceived quality and intelligibility. Broadband desktop phones in the enterprise and VoIP clients based on personal computers (PC) (for example, Skype) (these clients provide communication with other clients of the same type) Zhongzheng is facing resistance.
Since broadband conversational voice is beginning to meet resistance, codec developers consider the next development step in the audio bandwidth for conversational voice. The new ultra-wideband (SWB) voice codec is now a trend, which reproduces frequencies from 50 Hz to 14 kHz.
Extending the bandwidth used for voice to 14 kHz will bring a new conversational audio experience to cellular calls. By covering almost the entire audio frequency spectrum, the added bandwidth can contribute to an improved sense of presence. Voiced speech usually decays at a rate of about six decibels per octave. Therefore, above 14 kHz, there is almost no energy.
A method of processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separated from the low-frequency sub-band according to a general configuration includes filtering the audio signal to obtain A narrowband signal and a UHF signal. The method includes calculating an encoded narrowband excitation signal based on information from the narrowband signal, and calculating an ultra-high frequency excitation signal based on information from the encoded narrowband excitation signal. The method includes: calculating and characterizing a plurality of filter parameters of a spectral envelope of the high-frequency sub-band based on information from the UHF signal, and evaluating a signal based on the UHF signal and a signal based on the UHF signal. A time-varying relationship between the signals of the high-frequency excitation signal is used to calculate a plurality of gain factors. In this method, the narrow-frequency signal is based on the frequency component in the low-frequency sub-band, and the UHF signal is based on the frequency component in the high-frequency sub-band. In this method, a width of the low-frequency sub-band is at least three kilohertz, and the low-frequency sub-band is separated from the high-frequency sub-band by a distance that is at least equal to half of the width of the low-frequency sub-band .
An apparatus for processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separated from the low-frequency sub-band according to another general configuration includes: The audio signal is filtered to obtain a narrowband signal and a UHF signal; a component for calculating an encoded narrowband excitation signal based on information from the narrowband signal; and a component for calculating an encoded narrowband excitation signal based on the encoded narrowband signal The information of the high frequency excitation signal is used to calculate a UHF excitation signal. This device also includes: means for calculating and characterizing a plurality of filter parameters of a spectral envelope of the high frequency sub-band based on information from the UHF signal, and for evaluating a filter parameter based on the UHF signal A component for calculating a plurality of gain factors based on a time-varying relationship between the signal of the UHF excitation signal. In this device, the narrow-frequency signal is based on the frequency component in the low-frequency sub-band, and the UHF signal is based on the frequency component in the high-frequency sub-band. In this device, a width of the low-frequency sub-band is at least three kilohertz, and the low-frequency sub-band and the high-frequency sub-band are separated by a distance that is at least equal to half of the width of the low-frequency sub-band .
An apparatus for processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separated from the low-frequency sub-band according to another general configuration includes: a filter bank , The filter bank is configured to filter the audio signal to obtain a narrowband signal and an ultra-high frequency signal; and a narrowband encoder configured to be based on the narrowband signal The information is calculated as an encoded narrowband excitation signal. This device also includes an UHF encoder configured to: (A) calculate an UHF excitation signal based on information from the encoded narrowband excitation signal; (B) based on The information calculation of the UHF signal characterizes a plurality of filter parameters of a spectral envelope of the high frequency subband; and (C) by evaluating a signal based on the UHF signal and a signal based on the UHF excitation A time-varying relationship between signals to calculate multiple gain factors. In this device, the narrow-frequency signal is based on the frequency component in the low-frequency sub-band, and the UHF signal is based on the frequency component in the high-frequency sub-band. In this device, a width of the low-frequency sub-band is at least three kilohertz, and the low-frequency sub-band and the high-frequency sub-band are separated by a distance that is at least equal to half of the width of the low-frequency sub-band .
Conventional narrow-band (NB) speech codecs usually reproduce signals with a frequency range from 300 Hz to 3400 Hz. The wideband speech codec extends this coverage to 50 Hz to 7000 Hz. The SWB speech codec as described herein can be used to reproduce a much wider frequency range, such as from 50 Hz to 14 kHz. The expanded bandwidth can provide listeners with a more natural sounding experience with greater presence.
The proposed spectrum-efficient SWB speech codec provides a new speech coding and decoding technology, so that the processed speech has a much wider bandwidth than the traditional speech codec can provide. Compared with other existing speech codecs, which are usually narrowband (0 kHz to 3.5 kHz) or wideband (0 kHz to 7 kHz), the SWB speech codec gives mobile end users a much more practical and clearer experience .
Unless clearly limited by the context, the term "signal" is used herein to indicate any of its ordinary meanings, including memory locations (or memory locations, as expressed on wires, buses, or other transmission media). Collection). Unless clearly limited by the context, the term "produced" is used herein to indicate any of its ordinary meanings, such as calculating or otherwise producing. Unless clearly limited by the context, the term "calculation" is used herein to indicate any of its ordinary meanings, such as calculating, evaluating, estimating, and/or selecting from a plurality of values. Unless clearly limited by the context, the term "obtained" is used to indicate any of its ordinary meanings, such as computing, deriving, receiving (for example, from an external device), and/or retrieving (for example, from a storage element array) . Unless clearly limited by the context, the term "choice" is used to indicate any of its ordinary meanings, including identifying, indicating, applying, and/or using at least one and less of the set of two or more. For all. Where the term "comprising" is used in this description and the scope of the patent application, it does not exclude other elements or operations. The term "based on" (as in "A is based on B") is used to indicate any of its ordinary meanings, including the following conditions: (i) "derived from" (for example, "B is the predecessor of A"Object"); (ii) "at least based on" (e.g., "A is based on at least B"); and if appropriate in a particular context, (iii) "equal to" (e.g., "A equals B" or "A Same as B"). Similarly, the term "response to" is used to indicate any of its ordinary meanings, including "respond to at least."
Unless otherwise indicated, the term "series" is used to indicate a sequence of two or more items. The term "logarithm" is used to indicate a logarithm with a base of ten, but it is also within the scope of the present invention to extend this operation to other bases. The term "frequency component" is used to indicate one of a set of frequencies or frequency bands of a signal, such as a sample (or "frequency grid") of a signal's frequency domain representation (for example, as produced by a fast Fourier transform) or a sub-frequency band of the signal (For example, Bark scale or mel scale subbands).
Unless otherwise indicated, any disclosure of the operation of a device with specific characteristics is also expressly intended to disclose a method with similar characteristics (and vice versa), and any disclosure of the operation of a device with a specific configuration is also clear It is intended to reveal methods based on similar configurations (and vice versa). The term "configuration" can be used with respect to methods, devices, and/or systems as dictated by its specific context. Unless otherwise indicated by a particular context, the terms "method", "processing procedure", "procedure" and "technology" are used generically and interchangeably. Unless otherwise indicated by a particular context, the terms "device" and "device" are also used generically and interchangeably. The terms "component" and "module" are usually used to indicate part of a larger configuration. Unless clearly limited by the context, the term "system" is used herein to indicate any of its ordinary meanings, including "a group of elements that interact to achieve a common purpose." Any incorporation of a part of a document by reference shall also be understood as incorporating the definitions of the terms or variables cited in that part (where these definitions appear elsewhere in the document), and in the incorporated part Any diagram referenced in.
The terms "encoder", "codec" and "encoding system" are used interchangeably to mean a frame that includes a frame configured to receive and encode an audio signal (which may be used in one or other filtering operations such as perceptual weighting and/or other filtering operations). After multiple preprocessing operations) at least one encoder and a corresponding decoder system configured to generate a decoded representation of the frame. The encoder and decoder are usually deployed at opposite ends of the communication link. In order to support full-duplex communication, examples of both encoder and decoder are usually deployed at each end of the link.
Unless otherwise indicated by a specific context, the term "narrow frequency" refers to having less than 6 kHz (for example, from 0 Hz, 50 Hz, or 300 Hz to 2000 Hz, 2500 Hz, 3000 Hz, 3400 Hz, 3500 Hz, or 4000 Hz) The term "wideband" refers to a signal with a bandwidth in the range from 6 kHz to 10 kHz (for example, from 0 Hz, 50 Hz, or 300 Hz to 7000 Hz or 8000 Hz); and the term "Ultra-wideband" refers to a signal with a bandwidth greater than 10 kHz (for example, from 0 Hz, 50 Hz, or 300 Hz to 12 kHz, 14 kHz, or 16 kHz). Generally speaking, the terms "low frequency", "high frequency" and "ultra high frequency" are used in a relative sense, so that the frequency range of the low frequency signal is lower than the frequency range of the corresponding high frequency signal and the frequency range of the high frequency signal is higher In the frequency range of the low-frequency signal, the frequency range of the high-frequency signal is lower than the frequency range of the corresponding UHF signal and the frequency range of the UHF signal is higher than the frequency range of the high-frequency signal.
Several session codecs that support ultra-wide bandwidth have been standardized in ITU-T (International Telecommunication Union (Geneva, CH)-Telecommunication Standardization Department) such as G.719 and G.722.1C. Speex (available online at www.speex.org) is another SWB codec, which is already available as part of the GNU Project (www.gnu.org). However, these codecs may not be suitable for constrained applications such as cellular communication networks. Using this codec in this network to deliver reasonable communication quality to end users will usually require unacceptably high bit rates, while conversion-based speech codecs such as G.722.1C can be used in lower bits. Provide unsatisfactory voice quality at a low rate.
Methods for encoding and decoding general audio signals include transformation-based methods, such as the AAC (Advanced Audio Coding) series of codecs (for example, the European Telecommunications Standards Association TS 102005, the International Organization for Standardization (ISO)/International Electrotechnical Commission ( IEC) 14496-3:2009), which is intended for streaming audio content. These codecs have several features (for example, longer delay and higher bit rate) that may be problematic when the codecs are directly applied to speech signals for conversational speech on capacity-sensitive wireless networks. The third generation partnership project (3GPP) standard enhanced adaptive multi-rate-wideband (AMR-WB+) is another codec intended for streaming audio content, which is usually capable of operating at low rates (for example, as low as 10.4 It encodes high-quality SWB voice under thousands of bits/sec), but it may not be suitable for conversational purposes due to high calculation delay.
Existing wideband speech codecs include model-based sub-band methods, such as the Third Generation Partnership Project 2 (3GPP2, Arlington, VA) standard enhanced variable rate codec-wideband (EVRC-WB) codec (in Available online at www.3gpp2.org) and G.729.1 codec. This codec can implement a two-band model, which uses information from low-frequency sub-bands to reconstruct signal content in high-frequency sub-bands. For example, the EVRC-WB codec uses spectrum expansion for the excitation of the low-frequency part of the signal (50 Hz to 4000 Hz) to simulate high-frequency excitation.
In EVRC-WB, a spectrally efficient bandwidth extension model is used to reconstruct the high-frequency part (4 kHz to 7 kHz) of the speech signal. LP analysis is still performed on the HB signal to obtain spectral envelope information. However, the audible HB excitation signal is no longer the actual residue of the HB LPC analysis. The reality is that the excitation signal of the NB part is processed through the nonlinear model to generate the HB excitation for the voiced speech.
This method can be used to generate high frequency excitation with a wider bandwidth. After using an appropriate envelope and energy level to modulate a wider excitation, the SWB speech signal can be reconstructed. However, it is not unimportant to extend this method to include a wider frequency range for SWB speech coding, and it is not clear whether this model-based method can effectively handle SWB speech signals with ideal quality and reasonable delay. coding. Although this SWB speech coding method is applicable to some conversational applications on the Internet, the proposed method can provide quality advantages.
The proposed SWB codec appropriately and effectively handles the extra bandwidth by introducing a multi-band method to synthesize SWB speech signals. Regarding the proposed SWB speech codec described in this article, a multi-band technology has been designed to effectively expand the bandwidth coverage so that the codec can reproduce twice or even larger bandwidth. The proposed method of synthesizing the SWB speech signal using a multi-band model-based method represents the ultra-high frequency (SHB) part with high spectral efficiency in order to recover the widest frequency components of the SWB speech signal. Due to its model-based nature, this method avoids the higher latency associated with transformation-based methods. Due to the additional SHB signal, the output voice is more natural and provides a greater sense of presence, and therefore provides a much better conversation experience to the end user. Multi-band technology also provides embedded scalability from WB to SWB, which may not be available in a two-band method.
In a typical example, the proposed codec is implemented using a three-band sub-band method, in which the input speech signal is divided into three frequency bands: low frequency (LB), high frequency (HB), and ultra high frequency (SHB). Since the energy in human speech attenuates as the frequency increases, and human hearing is less sensitive as the frequency increases higher than narrowband speech, more aggressive modeling can be used for higher frequency bands, and the results are perceptually satisfactory .
In the proposed codec, similar to the high-frequency excitation extension of EVRC-WB, the non-linear extension of LB excitation is used to model the SHB excitation signal instead of using the actual SHB excitation signal. Since the non-linear extension is less computationally complex than the actual excitation calculation and encoding, less power and less delay are involved at the encoder and decoder in this part of the processing procedure.
The proposed method uses SHB excitation signal, SHB spectral envelope and SHB time gain parameters to reconstruct the SHB component. The spectral envelope information of SHB can be obtained by calculating linear predictive coding (LPC) coefficients based on the original SHB signal. The SHB time gain parameter can be estimated by comparing the energy of the original SHB signal with the energy of the estimated SHB signal. The appropriate choice of the LPC order and the number of time gains per frame may be important to the quality obtained using this method, and it may be necessary to achieve a reproduced speech quality and the number of bits required to represent the SHB envelope and time gain parameters. Properly balanced.
The proposed SWB codec can be implemented to include an extension that is configured to encode the SHB part (7 kHz to 14 kHz) of the speech signal using a method similar to the encoding of the HB part in EVRC-WB. In one such example as shown in Figure 10, a non-linear function is used to blindly extend the LPC residue of LB (50 Hz to 4000 Hz) up to 7 kHz to 14 kHz to generate the SHB excitation signal XS10. The LPC filter parameter CPS10a (for example, obtained by the eighth-order LPC analysis) represents the spectral envelope of the SHB, and represents the ten sub-frame gains of the difference between the gain envelope (for example, energy) of the original SHB signal and the synthesized SHB signal And a frame gain contains the time envelope of the SHB signal.
Figure 1 shows a high-level block diagram of SWB encoder SWE100 (which can also be configured to perform quantization of spectrum and time envelope parameters) including this SHB encoder. The corresponding SWB and SHB decoders (which can also be configured to perform inverse quantization of spectrum and time envelope parameters) are illustrated in FIGS. 3 and 21, respectively.
The proposed method can be implemented to encode using the same technology used in the EVRC-B narrowband speech codec standardized by 3GPP2 as Service Option 68 (SO68) (and available online at www.3gpp2.org) The low frequency (LB) of the SWB signal (for example, 50 Hz to 4000 Hz). Regarding active voiced speech, EVRC-B uses a compression technique based on Code Excited Linear Prediction (CELP) to encode low frequencies. The basic idea behind this technology is the source filter speech generation model, which describes speech as the result of linear filtering of quasi-periodic excitation (source). This filter shapes the spectral envelope of the original input speech. The LPC coefficients can be used to approximate the spectral envelope of the input signal. The LPC coefficients describe each sample as a linear combination of previous samples. The excitation is modeled using adaptive and fixed codebook items, which are selected to best match the residue of the LPC analysis. Although extremely high quality is possible, quality can be compromised due to bit rates below about 8 kbps. Regarding active silent speech, EVRC-B uses a compression technique based on Noise Excited Linear Prediction (NELP) to encode low frequencies.
Theoretically, the SHB model can be applied to any LB and HB coding technology. Any traditional vocoder can be used to process the LB signal. The traditional vocoder performs the analysis and synthesis of the excitation signal and the shaping of the signal's spectrum envelope. The HB part can be encoded and decoded by any codec that can reproduce HB frequency components. Obviously, it is not necessary for HB to use model-based methods (for example, CELP). For example, transform-based techniques can be used to encode HB. However, the use of model-based methods to encode HB usually entails lower bit rate requirements and less encoding delay.
The proposed method can also be implemented to use the same high-frequency modeling method as the EVRC-WB codec standardized by 3GPP2 as Service Option 70 (SO 70) (and available on www.3gpp2.org) Encode the high frequency (HB) part (4 kHz to 7 kHz) of the signal of the SWB codec. In this case, HB is the blind extension of the LB linear prediction residual by the low-rate coding of the spectral envelope added by the nonlinear function, five sub-frame gains (for example, as shown in FIG. 23A), and one frame gain.
It may be necessary to implement the proposed codec so that most of the bits are allocated to high-quality coding in the lowest frequency band. For example, EVRC-WB allocates 155 bits to encode LB and 16 bits to encode HB, resulting in a total allocation of 171 bits per 20 millisecond frame. The proposed SWB codec allocates an additional 19 bits to encode SHB, resulting in a total allocation of 190 bits per 20 millisecond frame. Therefore, the proposed SWB codec doubles the bandwidth of WB, while the increase in bit rate is less than 12%. An alternative implementation of the proposed SWB codec allocates an additional 24 bits to encode SHB (getting a total allocation of 195 bits per twenty millisecond frame). Another alternative implementation of the proposed SWB codec allocates an additional 38 bits to encode SHB (getting a total allocation of 209 bits per twenty millisecond frame).
A version of the proposed encoder transmits the following three sets of high-frequency parameters to the decoder for reconstruction of the SHB signal: LSF parameters, sub-frame gain, and frame gain. The LSF parameter and sub-frame gain of each frame are multi-dimensional, and the frame gain is a scalar. Regarding the quantization of multi-dimensional parameters, it may be necessary to minimize the number of bits required to use vector quantization (VQ). Since the vector dimensions of high-frequency LSF parameters and sub-frame gains are often relatively high, split VQ can be used. To achieve a specific quantization quality, the VQ codebook may be larger. For the situation where single vector VQ has been selected, multi-stage VQ can be used to reduce memory requirements and reduce the complexity of codebook search.
Figure 1 shows the block diagram of the ultra-wideband encoder SWE100 according to the general configuration. The filter bank FB100 is configured to filter the ultra-wideband signal SISW10 to generate a narrow-band signal SIL10, a high-frequency signal SIH10, and an ultra-high-frequency signal SIS30. The narrowband encoder EN100 is configured to encode a narrowband signal SIL10 to generate narrowband (NB) filter parameters FPN10 and an encoded NB excitation signal XL10. As described in more detail herein, the narrowband encoder EN100 is generally configured to generate narrowband filter parameters FPN10 as a codebook index or in another quantized form and an encoded narrowband excitation signal XL10. The high frequency encoder EH100 is configured to encode the high frequency signal SIH10 according to the information XL10a from the encoded narrowband excitation signal XL10 to generate the high frequency encoding parameter CPH10. As described in more detail herein, the high-frequency encoder EH100 is generally configured to generate the high-frequency encoding parameter CPH10 as a codebook index or in another quantized form. The UHF encoder ES100 is configured to encode the UHF signal SIS10 according to the information XL10b from the encoded narrowband excitation signal XL10 to generate the UHF encoding parameter CPS10. As described in more detail herein, the UHF encoder ES100 is generally configured to generate the UHF encoding parameter CPS10 as a codebook index or in another quantized form.
A specific example of the ultra-wideband encoder SWE100 is configured to encode the ultra-wideband signal SISW10 at a rate of about 9.75 kbps (kilobits per second), of which about 7.75 kbps is used for the narrowband filter parameter FPN10 and the encoded narrow The frequency excitation signal XL10, about 0.8 kbps is used for the high frequency coding parameter CPH10, and about 0.95 kbps is used for the UHF coding parameter CPS10. Another specific example of the ultra-wideband encoder SWE100 is configured to encode the ultra-wideband signal SISW10 at a rate of about 9.75 kbps, of which about 7.75 kbps is used for the narrowband filter parameter FPN10 and the encoded narrowband excitation signal XL10, approximately 0.8 kbps is used for the high frequency coding parameter CPH10, and about 1.2 kbps is used for the UHF coding parameter CPS10. Another specific example of the ultra-wideband encoder SWE100 is configured to encode the ultra-wideband signal SISW10 at a rate of about 10.45 kbps, of which about 7.75 kbps is used for the narrowband filter parameter FPN10 and the encoded narrowband excitation signal XL10, approximately 0.8 kbps is used for the high frequency coding parameter CPH10, and about 1.9 kbps is used for the UHF coding parameter CPS10.
It may be necessary to combine the encoded narrowband, high frequency, and UHF signals into a single bit stream. For example, the encoded signals may need to be multiplexed together for transmission (for example, via wired, optical, or wireless transmission channels) or storage as encoded ultra-wideband signals. Figure 2 shows a block diagram of the implementation of SWE110 of the ultra-wideband encoder SWE100. The implementation of SWE110 includes the configuration of the narrowband filter parameter FPN10, the encoded narrowband excitation signal XL10, the high frequency encoding parameter CPH10, and the UHF encoding The parameter CPS10 is combined into a multiplexer MPX100 (for example, a bit packer) of the multiplexed signal SM10.
The device including the encoder SWE110 may also include a circuit configured to transmit the multiplexed signal SM10 to a transmission channel such as a wired, optical or wireless channel. The device can also be configured to perform one or more channel encoding operations on the signal (such as error correction encoding (e.g., rate compatible cyclotron encoding) and/or error detection encoding (e.g., cyclic redundancy encoding)), and /Or one or more layers of network protocol encoding (for example, Ethernet, TCP/IP, cdma2000).
It may be necessary to configure the multiplexer MPX100 to embed the encoded narrowband signal (including the narrowband filter parameters FPN10 and the encoded narrowband excitation signal XL10) as a separable sub-stream of the multiplex signal SM10, so that The encoded narrowband signal can be recovered and decoded independently of another part of the multiplex signal SM10 (such as a high frequency signal, an ultra high frequency signal and/or a low frequency signal). For example, the multiplex signal SM10 can be configured such that the encoded narrowband signal can be recovered by removing the high frequency encoding parameter CPH10 and the UHF encoding parameter CPS10. One of the potential advantages of this feature is that it avoids the need to pass the encoded ultra-wideband signal to a system that supports the decoding of narrow-band signals but does not support the decoding of high-frequency or ultra-high-frequency parts. The need for transcoding.
Alternatively or in addition, the multiplexer MPX100 may need to be configured to embed the encoded wideband signal (including the narrowband filter parameter FPN10, the encoded narrowband excitation signal XL10, and the high frequency encoding parameter CPH10) into a multiplexed signal The separable sub-stream of SM10 allows the encoded narrowband signal to be recovered and decoded independently of another part of the multiplexed signal SM10 (such as ultra-high frequency signal and/or low frequency signal). For example, the multiplexed signal SM10 can be configured so that the encoded broadband signal can be recovered by removing the UHF encoding parameter CPS10. One potential advantage of this feature is that it avoids the need to transcode the encoded UWB signal before passing it to a system that supports the decoding of the broadband signal but does not support the decoding of the UHF part. .
Figure 3 is a block diagram of an ultra-wideband decoder SWD100 based on a general configuration. The narrowband decoder DN100 is configured to decode the narrowband filter parameter FPN10 and the encoded narrowband excitation signal XL10 to generate the decoded narrowband signal SDL10. The high frequency decoder DH100 is configured to generate the decoded high frequency signal SDH10 based on the high frequency encoding parameter CPH10 and the information XL10a from the encoded excitation signal XL10. The UHF decoder DS100 is configured to generate the decoded UHF signal SDS10 based on the UHF encoding parameter CPS10 and the information XL10b from the encoded excitation signal XL10. The filter bank FB200 is configured to combine the decoded narrowband signal SDL10, the decoded high frequency signal SDH10 and the decoded UHF signal SDS10 to generate the ultra wideband output signal SOSW10.
Figure 4 is a block diagram of an implementation of SWD110 of the ultra-wideband decoder SWD100. The implementation of SWD110 includes a demultiplexer DMX100 (e.g., bit Meta decapsulator). The device including the decoder SWE110 may include a circuit configured to receive the multiplexed signal SM10 from a transmission channel such as a wired, optical or wireless channel. The device can also be configured to perform one or more channel decoding operations on the signal (such as error correction decoding (e.g., rate compatible convolutional decoding) and/or error detection decoding (e.g., cyclic redundancy decoding)), and /Or one or more layers of network protocol decoding (for example, Ethernet, TCP/IP, cdma2000).
*****(1: filter bank)
The filter bank FB100 is configured to filter the input signal according to the sub-band scheme to generate a plurality of sub-band signals with limited bandwidth, each of which contains the frequency components of the corresponding sub-band of the input signal. Depending on the design criteria of a particular application, the output sub-band signals may have equal or unequal bandwidth and may be overlapping or non-overlapping. The configuration of the filter bank FB100 that generates more than three sub-band signals is also possible. For example, this filter bank can be configured to generate one or more low-frequency signals, the one or more low-frequency signals included in a frequency range lower than the frequency range of the narrow-band signal SIL10 (such as from 0 Hz, 20 Hz or 50 Hz to 200 Hz, 300 Hz or 500 Hz). It is also possible to configure this filter bank to generate one or more UHF signals, the one or more UHF signals included in a frequency range higher than the frequency range of the UHF signal SIH10 (such as, 14 kHz to 20 kHz, 16 kHz to 20 kHz, or 16 kHz to 32 kHz range). In this situation, the ultra-wideband encoder SWE100 can be implemented to separately encode this signal or these signals, and the multiplexer MPX100 can be configured to include the or these additional encoded signals in the multiplexed signal SM10 (For example, as a separable part).
The filter bank FB100 is configured to receive an ultra-wideband signal SISW10 having a low frequency subband, a middle frequency subband, and a high frequency subband. Figure 5A shows a block diagram of an implementation of FB110 of the filter bank FB100, which is configured to generate three sub-band signals with reduced sampling rates (narrowband signal SIL10, high frequency signal SIH10, and ultra high frequency signal SIS10 ). The filter bank FB110 includes a broadband analysis processing path PAW10 configured to receive the ultra-wideband signal SISW10 and generate a broadband signal SIW10, and an ultra-high frequency analysis processing path configured to receive the ultra-wideband signal SISW10 and generate a UHF signal SIS30 PAS10. The filter bank FB110 also includes a narrowband analysis processing path PAN10 configured to receive the wideband signal SIW10 and generate a narrowband signal SIL10, and a high frequency analysis processing path configured to receive the wideband speech signal SIW10 and generate a high frequency signal SIH10 PAH10. The narrow-band signal SIL10 contains frequency components in the low-frequency sub-band, the high-frequency signal SIH10 contains the frequency components in the middle-frequency sub-band, and the wide-band signal SIW10 contains the frequency components in the low-frequency sub-band and the mid-frequency sub-band. The signal SIS10 contains frequency components of high frequency sub-bands.
Because the sub-band signal has a narrower bandwidth than the ultra-wideband signal SISW10, the sampling rate of the sub-band signal can be reduced to some extent (for example, to reduce the computational complexity without losing information). FIG. 6A shows a block diagram of the implementation of FB112 of the filter bank FB110, in which the wideband analysis processing path PAW10 is implemented by an integer downsampler (decimator) DW10 and the narrowband analysis processing path PAN10 is implemented by an integer downsampler DN10. The filter bank FB112 also includes: the implementation PAH12 of the high-frequency analysis processing path PAH10, which has a spectrum inversion module RHA10 and an integer down sampler DH10; and the implementation PAS12 of the ultra-high frequency analysis processing path PAS10, which has a spectrum inversion Module RSA10 and integer down sampler DS10.
Each of the integer downsamplers DW10, DN10, DH10, and DS10 can be implemented as a low-pass filter (for example, to prevent frequency overlap) followed by a downsampler. For example, FIG. 8A shows a block diagram of this implementation DS12 of an integer down sampler DS10 that is configured to downsample the input signal by a factor of two. Under these conditions, the low-pass filter can be implemented with a cutoff frequency of<i>f</i><sub><i>s</i></sub><i>(2k</i><sub><i>d</i></sub><i>)</i>Finite impulse response (FIR) or infinite impulse response (IIR) filter, where<i>f</i><sub><i>s</i></sub>Is the sampling rate of the input signal and<i>k</i><sub><i>d</i></sub>The downsampling factor is an integer multiple, and the downsampling frequency can be performed by removing samples of the signal and/or replacing the samples with an average value.
Alternatively, one or more (possibly all) of the integer down-samplers DW10, DN10, DH10, and DS10 can be implemented as filters that integrate low-pass filtering and down-sampling frequency operations. One such example of an integer downsampler is configured to perform integer downsampling by a factor of 2 by using a three-stage polyphase implementation, so that for even numbers<i>n</i><img file="TW201214419A_D0001.tif" />0, input signal to be downsampled by integer multiples<i>S</i><sub><i>in</i></sub>[<i>n</i>] The samples are filtered by the all-pass filter given by the transfer function:
<maths><img file="TW201214419A_D0002.tif" /></maths>
And for odd numbers<i>n</i><img file="TW201214419A_D0003.tif" />0, input signal<i>S</i><sub><i>in</i></sub>[<i>n</i>] The samples are filtered by the all-pass filter given by the transfer function:
<maths><img file="TW201214419A_D0004.tif" /></maths>
Add the output of these two polyphase components (for example, averaging) to obtain an output signal that has been downsampled by an integer multiple<i>S</i><sub><i>out</i></sub><i>[n]</i>. In a particular instance, the value (<i>a</i><sub><i>down2, 0, 0</i></sub>,<i>adown2,0,.1,adown2,0,2,adown2,1,0,adown2,1,1,adown2,1,2</i>) Is equal to (0.06056541924291,0.42943401549235, 0.80873048306552, 0.22063024829630, 0.63593943961708, 0.94151583095682). This implementation may allow the reuse of functional blocks of logic and/or code. For example, it is clearly noted that any of the downsampling operations described herein by integer multiples of 2 can be performed in this manner (and may be performed by the same module at different times). In a specific example, this three-stage multi-phase implementation is used to implement integer downsamplers DH10 and DS10.
Alternatively or additionally, one or more (possibly all) of the integer downsamplers DW10, DN10, DH10, and DS10 are configured to use a polyphase implementation to perform integer downsampling by a factor of 2 such that the integer multiple The down-sampled input signal is divided into sub-sequences of odd time index and even time index filtered by a respective 13th order FIR filter. In other words, for even-numbered sample indexes<i>n</i><img file="TW201214419A_D0005.tif" />0, input signal to be downsampled by integer multiples<i>S</i><sub><i>in</i></sub><i>[n]</i>The sample is passed through the first 13th order FIR filter<i>H</i><sub><i>dec1</i></sub><i>(z)</i>To filter, and for odd numbers<i>n</i><img file="TW201214419A_D0006.tif" />0, input signal<i>S</i><sub><i>i</i></sub><sub><i>n</i></sub><i>(n)</i>The samples are passed through the second 13th order FIR filter<i>H</i><sub><i>dec2</i></sub><i>(</i><i>z)</i>To filter. Add the output of these two polyphase components (for example, averaging) to obtain an output signal that has been downsampled by an integer multiple<i>S</i><sub><i>out</i></sub><i>[n]</i>. In a specific example, the coefficients of the filter<i>H</i><sub><i>dec1</i></sub><i>(z)</i>and<i>H</i><sub><i>dec2</i></sub><i>(z)</i>Shown in the table below:
<tables><img file="TW201214419A_D0007.tif" /></tables>
This implementation may allow the reuse of functional blocks of logic and/or code. For example, it is obvious to note that any of the 2 integer downsampling operations described herein can be performed in this manner (and may be performed by the same module at different times). In a specific example, this FIR polyphase implementation is used to implement integer downsamplers DW10 and DN10.
In the high-frequency analysis processing path PAH12, the spectrum inversion module RHA10 inverts the spectrum of the wideband signal SIW10 (for example, by making the signal and the function<i>e</i><sup><i>jn</i></sup><sup>π</sup>Or sequence(-1)<sup><i>n</i></sup>Multiply, sequence (-1)<sup><i>n</i></sup>The value of is alternated between +1 and -1), and the integer down-sampler DH10 reduces the sampling rate of the spectrum-inverted signal according to the desired integer down-sampling factor to generate the high-frequency signal SIH10. In the UHF processing path PAS12, the spectrum inversion module RSA10 inverts the spectrum of the UHF signal SISW10 (for example, by making the signal and the function<i>e</i><sup><i>jn</i></sup><sup>π</sup>Or sequence(-1)<sup><i>n</i></sup>Multiply), and the integer down-sampler DS10 reduces the sampling rate of the spectrum-inverted signal according to the desired integer down-sampling factor to generate the UHF signal SIS10. It also covers the configuration of the filter bank FB112 that generates more than three passband signals.
The filter bank FB200 is configured to filter the passband signal with low frequency components, the passband signal with middle frequency components, and the passband signal with high frequency components according to the sub-band scheme to generate output signals, where the limited bandwidth Each of the sub-band signals contains frequency components of the corresponding sub-band of the output signal. Depending on the design criteria of a particular application, the output sub-band signals may have equal or unequal bandwidth and may be overlapping or non-overlapping. Figure 5B shows a block diagram of an implementation FB210 of the filter bank FB200, which is configured to receive three passband signals with reduced sampling rates (decoded narrowband signal SDL10, decoded high frequency signal SDH10 , And the decoded UHF signal SDS10) and combine the frequency components of the passband signals to generate an ultra-wideband output signal SOSW10.
The filter bank FB210 includes a narrowband synthesis processing path PSN10 configured to receive a narrowband signal SDL10 (for example, a decoded version of the narrowband signal SIL10) and generate a narrowband output signal SOL10, and a narrowband synthesis processing path PSN10 configured to receive a high frequency signal SDH10 (For example, the decoded version of the high-frequency signal SIH10) and the high-frequency synthesis processing path PSH10 that generates the high-frequency output signal SOH10. The filter bank FB210 also includes an adder ADD10 configured to generate the decoded wideband signal SDW10 (for example, a decoded version of the wideband signal SIW10) as the sum of the passband signals SOL10 and SOH10. The adder ADD10 may also be implemented to generate the decoded wideband signal SDW10 as a weighted sum of the two passband signals SOL10 and SOH10 according to one or more weights received and/or calculated by the ultra-high frequency decoder SWD100. In one such example, the adder ADD10 is configured to generate the decoded wideband signal SDW10 according to the expression SDW10[n]=SOL10[n]+0.9*SOH10[n].
The filter bank FB210 also includes a broadband synthesis processing path PSW10 configured to receive the decoded broadband signal SDW10 and generate a broadband output signal SOW10, and a broadband synthesis processing path PSW10 configured to receive the ultra-high frequency signal SDS10 (for example, the ultra-high frequency signal SIS10 Decoded version) and UHF synthesis processing path PSS10 that generates UHF output signal SOS10. The filter bank FB210 also includes an adder ADD20 configured to generate the ultra-wideband output signal SOSW10 (for example, a decoded version of the ultra-wideband signal SISW10) as the sum of the signals SOW10 and SOS10. The adder ADD20 may also be implemented to generate the ultra-wideband output signal SOSW10 as a weighted sum of the two passband signals SOW10 and SOS10 according to one or more weights received and/or calculated by the ultra-high frequency decoder SWD100. In one such example, the filter bank FB210 is configured to follow the expression SOSW10[n]=SO W10[n]+0.9*SOS10[n] to generate the ultra-wideband output signal SOSW10. The narrowband signals SDL10 and SOL10 contain the frequency components of the low frequency subband of the signal SOSW10, the high frequency signals SDH10 and SOH10 contain the frequency components of the middle frequency subband of the signal SOSW10, and the wideband signals SDW10 and SOW10 contain the low frequency subband of the signal SOSW10. Frequency components and frequency components of the mid-frequency sub-band, and the ultra-high frequency signals SDS10 and SOS10 contain frequency components of the high-frequency sub-band of the signal SOSW10.
The configuration of the filter bank FB210 combining three or more sub-band signals is also possible. For example, the filter bank can be configured to generate an output signal having frequency components from one or more low-frequency signals, the one or more low-frequency signals included in a frequency range lower than the frequency range of the narrow-band signal SDL10 (For example, from 0 Hz, 20 Hz, or 50 Hz to 200 Hz, 300 Hz, or 500 Hz). It is also possible to configure this filter bank to generate an output signal with frequency components from one or more UHF signals. The component of the frequency range of the frequency range (such as the range of 14 kHz to 20 kHz, 16 kHz to 20 kHz, or 16 kHz to 32 kHz). In this situation, the ultra-wideband decoder SWD100 can be implemented to decode this signal or these signals separately, and the demultiplexer DMX100 can be configured to extract the or these additional encoded signals from the multiplexed signal SM10 (For example, as a separable part).
Since the sub-band signal has a narrower bandwidth than the ultra-wideband output signal SOSW10, the sampling rate of the sub-band signal can be lower than the sampling rate of the signal SOSW10. 6B shows a block diagram of the implementation of FB212 of the filter bank FB210, in which the narrowband synthesis processing path PSN10 is implemented by the interpolator IN10 and the wideband synthesis processing path PSW10 is implemented by the interpolator IW10. The filter bank FB212 also includes: a high frequency synthesis processing path PSH10 implementation PSH12, which has an interpolator IH10 and a spectrum inversion module RHD10; and an ultra high frequency synthesis processing path PSS10 implementation PSS12, which has an interpolator IS10 and Spectrum inversion module RSD10.
Each of the interpolators IW10, IN10, IH10, and IS10 may be implemented as an upsampler followed by a low-pass filter (for example, to prevent frequency overlap). For example, FIG. 8B shows a block diagram of this implementation IS12 of an interpolator IS10 configured to interpolate an input signal by a factor of 2. Under these conditions, the low-pass filter can be implemented with a cutoff frequency of<i>f</i><sub><i>s</i></sub><i>(2k</i><sub><i>d</i></sub><i>)</i>Finite impulse response (FIR) or infinite impulse response (IIR) filter, where<i>f</i><sub><i>s</i></sub>Is the sampling rate of the input signal and<i>k</i><sub><i>d</i></sub>It is an interpolation factor, and the sampling frequency can be increased by zero padding and/or by copying samples.
Alternatively, one or more (possibly all) of the interpolators IW10, IN10, IH10, and IS10 may be implemented as filters that integrate up-sampling frequency and low-pass filtering operations. One such instance of the interpolator is configured to perform interpolation by a factor of 2 by using a three-stage polyphase implementation so that for even numbers<i>n</i><img file="TW201214419A_D0008.tif" />0, interpolated signal<i>S</i><sub><i>out</i></sub>[<i>n</i>] Is obtained by applying the transfer function to the input signal by the all-pass filter given by<i>S</i><sub><i>in</i></sub>[<i>n</i>/2] Obtained by filtering:
<maths><img file="TW201214419A_D0009.tif" /></maths>
And for odd numbers<i>n</i><img file="TW201214419A_D0010.tif" />0, interpolated signal<i>S</i><sub><i>out</i></sub>[<i>n</i>] Is obtained by applying the transfer function to the input signal by the all-pass filter given by<i>S</i><sub><i>in</i></sub>[(<i>n</i>-<i>1</i>)/<i>2</i>] Obtained by filtering:
<maths><img file="TW201214419A_D0011.tif" /></maths>
In a particular instance, the value (<i>a</i><sub><i>u</i></sub><sub><i>p</i></sub><sub>2,0,0</sub>,<i>a</i><sub><i>u</i></sub><sub><i>p</i></sub><sub>2,0,1</sub>,<i>a</i><sub><i>up</i></sub><sub>2,0,2</sub>) Is equal to (0.22063024829630, 0.63593943961708, 0.94151583095682) and the value (<i>a</i><sub><i>up</i></sub><sub>2,1,0</sub>,<i>aup</i>2,1,1,<i>aup</i>2,1,2) is equal to (0.06056541924291,0.42943401549235,0.80873048306552). This implementation may allow the reuse of functional blocks of logic and/or code. For example, it is clearly noted that any of the 2-by-2 interpolation operations described herein can be performed in this manner (and may be performed by the same module at different times). In a specific example, this three-stage multi-phase implementation is used to implement interpolators IH10 and IS10.
Alternatively or in addition, one or more (possibly all) of the interpolators IW10, IN10, IH10, and IS10 are configured to use a polyphase implementation to perform interpolation by a factor of 2 so that the input signal to be interpolated is Filtered by two different 15th order FIR filters to generate odd time index and even time index subsequences of the interpolated signal. In other words, for even-numbered sample indexes<i>n</i><img file="TW201214419A_D0012.tif" />0, interpolated signal<i>S</i><sub><i>out</i></sub>[<i>n</i>] Sample is by using the first 15th order FIR filter<i>H</i><sub><i>int1</i></sub><i>(z)</i>Input signal to be inserted<i>S</i><sub><i>in</i></sub>[<i>n/2</i>] Is filtered and generated, and for odd numbers<i>n</i><img file="TW201214419A_D0013.tif" />0, interpolated signal<i>S</i><sub><i>out</i></sub>[<i>n</i>] Sample is by using the second 15th order FIR filter<i>H</i><sub><i>int2</i></sub><i>(z)</i>Sample input signal<i>S</i><sub><i>in</i></sub>[(<i>n</i>-<i>1</i>)/2] is generated by filtering. In a specific example, the coefficients of the filter<i>H</i><sub><i>int1</i></sub><i>(z)</i>and<i>H</i><sub><i>int2</i></sub><i>(z)</i>Shown in the table below:
<tables><img file="TW201214419A_D0014.tif" /></tables>
This implementation may allow the reuse of functional blocks of logic and/or code. For example, it is clearly noted that any of the downsampling operations described herein by integer multiples of 2 can be performed in this manner (and may be performed by the same module at different times). In a specific example, this FIR multiphase implementation is used to implement interpolators IN10 and IW10.
In the high-frequency synthesis processing path PSH12, the interpolator IH10 increases the sampling rate of the decoded high-frequency signal SDH10 according to the required interpolation factor, and the spectrum inversion module RHD10 inverts the frequency spectrum of the signal with the increased sampling frequency ( For example, by making the signal and function<i>e</i><sup><i>jn</i></sup><sup>π</sup>Or sequence(-1)<sup><i>n</i></sup>Multiply) to generate a high-frequency output signal SOH10. Then the two passband signals SOL10 and SOH10 are summed to form a decoded wideband signal SDW10. The filter bank FB212 may also be implemented to generate the decoded wideband signal SDW10 as a weighted sum of the two passband signals SOL10 and SOH10 according to one or more weights received and/or calculated by the ultra-high frequency decoder SWD100. In one such example, the filter bank FB212 is configured to generate the decoded wideband signal SDW10 according to the expression SDW10[n]=SOL10[n]+0.9*SOH10[n].
In the UHF synthesis processing path PSS12, the interpolator IS10 increases the sampling rate of the decoded UHF signal SDS10 according to the required interpolation factor, and the spectrum inversion module RSD10 inverts the signal with the increased sampling frequency Spectrum (for example, by making the signal and the function<i>e</i><sup><i>jnπ</i></sup>Or sequence(-1)<sup><i>n</i></sup>Multiply) to generate the ultra-high frequency output signal SOS10. Then, the two passband signals SOW10 and SOS10 are summed to form an ultra-wideband output signal SOSW10. The filter bank FB212 may also be implemented to generate the ultra-wideband output signal SOSW10 as a weighted sum of the two passband signals SOW10 and SOS10 according to one or more weights received and/or calculated by the ultra-high frequency decoder SWD100. In one such example, the filter bank FB212 is configured to generate the ultra-wideband output signal SOSW10 according to the expression SOSW10[n]=SOW10[n]+0.9*SOS10[n]. It also covers the configuration of the filter bank FB212 combining three or more decoded passband signals.
In a typical example, the narrow-band signal SIL10 contains frequency components in the low-frequency sub-band. The low-frequency sub-band includes the limited PSTN range from 300 Hz to 3400 Hz (for example, the frequency band from 0 kHz to 4 kHz), but in In other examples, the low frequency subband may be narrow (for example, 0 Hz, 50 Hz, or 300 Hz to 2000 Hz, 2500 Hz, or 3000 Hz). 7A, 7B, and 7C show the relative bandwidths of the narrowband signal SIL10, the high frequency signal SIH10, and the UHF signal SIS10 in three different implementation examples. In all these specific examples, the ultra-wideband signal SISW10 has a sampling rate of 32 kHz (representing frequency components in the range of 0 kHz to 16 kHz), and the narrow-band signal SIL10 has a sampling rate of 8 kHz (representing the sampling rate at 0 kHz). Frequency components within the range of 4 kHz), and each of FIGS. 7A to 7C shows the part of the frequency components of the ultra-wideband signal SISW10 contained in each of the signals generated by the filter bank Instance.
The term "frequency component" is used herein to refer to the energy present at the specified frequency of the signal, or the energy distribution across the specified frequency band of the signal. The narrow-band signal SIL10 contains frequency components in the low-frequency sub-band, the high-frequency signal SIH10 contains the frequency components in the middle-frequency sub-band, and the wide-band signal SIW10 contains the frequency components in the low-frequency sub-band and the frequency components in the middle-frequency sub-band. The signal SIS10 contains frequency components of high frequency sub-bands. The width of the sub-band is defined as the distance between minus 20 decibel points in the frequency response of the filter bank path that selects the frequency components of the sub-band. Similarly, the overlap of two sub-bands can be defined as the point at which the frequency response of the filter bank path that selects the frequency components of the higher frequency sub-band drops to minus 20 decibels until the frequency component of the lower frequency sub-band is selected The distance between the points where the frequency response of the filter bank path drops to minus 20 decibels.
In the example of Figure 7A, there is no significant overlap between the three sub-bands. The implementation of the high-frequency analysis processing path PAH10 with a passband of 4 kHz to 8 kHz can be used to obtain the high-frequency signal SIH10 shown in this example. In this situation, it may be necessary for the processing path PAH10 to reduce the sampling rate to 8 kHz by down-sampling the signal by an integer factor of 2. This operation (which can be expected to significantly reduce the computational complexity of further processing operations on the signal) reduces the frequency components of the frequency sub-bands from 4 kHz to 8 kHz to the range of 0 kHz to 4 kHz without loss of information.
Similarly, the implementation of the UHF analysis processing path PAS10 with a passband of 8 kHz to 16 kHz can be used to obtain the UHF signal SIS10 shown in this example. In this situation, it may be necessary for the processing path PAS10 to reduce the sampling rate to 16 kHz by down-sampling the signal by an integer factor of 2. This operation (which can be expected to significantly reduce the computational complexity of further processing operations on the signal) reduces the frequency components of the high frequency sub-band from 8 kHz to 16 kHz to the range of 0 kHz to 8 kHz without losing information.
In the alternative example of FIG. 7B, the low-frequency sub-band and the middle-frequency sub-band have a significant overlap, so that the narrow-frequency signal SIL10 and the high-frequency signal SIH10 both describe the 3.5 kHz to 4 kHz region. The implementation of the high-frequency analysis processing path PAH10 with a passband of 3.5 kHz to 7 kHz can be used to obtain the high-frequency signal SIH10 shown in this example. In this situation, it may be necessary for the processing path PAH10 to reduce the sampling rate to 7 kHz by downsampling the signal by an integer multiple of 16/7. This operation (which can be expected to significantly reduce the computational complexity of further processing operations on the signal) reduces the frequency components of the frequency sub-bands from 3.5 kHz to 7 kHz to the range of 0 kHz to 3.5 kHz without losing information. Other specific examples of the high frequency analysis processing path PAH10 have passbands of 3.5 kHz to 7.5 kHz and 3.5 kHz to 8 kHz.
Figure 7B also shows an example of the high-frequency sub-band extending from 7 kHz to 14 kHz. The implementation of the UHF analysis processing path PAS10 with a passband of 7 kHz to 14 kHz can be used to obtain the UHF signal SIS10 shown in this example. In this situation, it may be necessary for the processing path PAS10 to reduce the sampling rate from 32 kHz to 7 kHz by downsampling the signal by integer multiples by a factor of 32/7. This operation (which can be expected to significantly reduce the computational complexity of further processing operations on the signal) reduces the frequency components of the high-frequency sub-band from 7 kHz to 14 kHz to the range of 0 kHz to 7 kHz without losing information.
FIG. 8C shows a block diagram of an implementation FB120 of the filter bank FB112, which can be used for the application shown in FIG. 7B. The filter bank FB120 is configured to receive a sampling rate<i>f</i><sub><i>s</i></sub>(For example, 32 kHz) ultra-wideband signal SISW10. The filter bank FB120 includes: an implementation DW20 of an integer down-sampler DW10, which is configured to perform integer down-sampling on the signal SISW10 by a factor of 2 to obtain a sampling rate<i>f</i><sub><i>SW</i></sub>For example, the wideband signal SIW10 of 16 kHz); and the implementation of the integer down-sampler DN10 DN20, which is configured to perform integer down-sampling of the signal SIW10 by a factor of 2 to obtain a sampling rate<i>f</i><sub><i>SN</i></sub>(For example, 8 kHz) narrowband signal SIL10.
The filter bank FB120 also includes the implementation PAH20 of the high-frequency analysis processing path PAH12, which is configured to be based on a non-integer factor<i>f</i><sub><i>SH</i></sub>/<i>f</i><sub><i>SW</i></sub>Integer downsampling the wideband signal SIW10, where<i>f</i><sub><i>SH</i></sub>It is the sampling rate of the high frequency signal SIH10 (for example, 7 kHz). Path PAH20 includes: Interpolation block IAH10, which is configured to interpolate signal SIW10 by a factor of 2 to achieve the sampling rate<i>f</i><sub><i>SW</i></sub>×2 (for example, to 32 kHz); resample block, which is configured to resample the interpolated signal to the sampling rate<i>f</i><sub><i>SH</i></sub>×4 (for example, by a factor of 7/8 to reach 28 kHz); and an integer downsampling block DH30, which is configured to perform integer downsampling on the resampled signal by a factor of 2 to achieve Sampling rate<i>f</i><sub><i>SH</i></sub>×2 (for example, up to 14 kHz). The integer downsampling block DH30 can be implemented according to any of the examples of this operation as described herein (for example, the three-stage multi-phase example described herein). Path PAH20 also includes spectrum inversion block and integer down sampler DH10. Decimate-by-two implementation of DH20 by 2 integer downsampling. The spectrum inversion block and implementation DH20 can respectively refer to the path above The PAH12 module RHA10 and the integer down sampler DH10 are implemented as described.
In this particular example, the path PAH20 also includes an optional spectrum shaping block FAH10. The optional spectrum shaping block FAH10 can be implemented as a low-pass filter configured to shape the signal to obtain the desired overall filter response Device. In a specific example, the spectrum shaping block FAH10 is implemented as a first-order IIR filter with the following transfer function:
<maths><img file="TW201214419A_D0015.tif" /></maths>
The interpolation block IAH10 of the path PAH20 can be implemented according to any of the examples of this operation as described herein (for example, the three-stage multi-phase example described herein). One such instance of an interpolator is configured to perform interpolation by a factor of 2 by using a two-stage polyphase implementation so that for even numbers<i>n</i><img file="TW201214419A_D0016.tif" />0, interpolated signal<i>S</i><sub><i>out</i></sub>[<i>n</i>] Is obtained by applying the transfer function to the input signal sub-sequence using the all-pass filter given by<i>S</i><sub><i>in</i></sub>[<i>n</i>/2] Obtained by filtering:
<maths><img file="TW201214419A_D0017.tif" /></maths>
And for odd numbers<i>n</i><img file="TW201214419A_D0018.tif" />0, interpolated signal<i>S</i><sub><i>out</i></sub>[<i>n</i>] Is obtained by applying the transfer function to the input signal sub-sequence using the all-pass filter given by<i>S</i><sub><i>in</i></sub>[(<i>n</i>-1)/2] is filtered to obtain:
<maths><img file="TW201214419A_D0019.tif" /></maths>
In a particular instance, the value (<i>a</i><sub><i>up</i></sub><sub>2,0,0</sub>,<i>a</i><sub><i>up</i></sub><sub>2,0,1</sub>,<i>a</i><sub><i>up</i></sub><sub>2,1,0</sub>,<i>a</i><sub><i>up</i></sub><sub>2,1,1</sub>,) is equal to (0.06262441299567, 0.49326511845632, 0.23754715248027, 0.80890715711734).
The 7/8 resample block of the path PAH20 can be implemented to use multiphase interpolation to resample the input signal with a sampling rate of 32 kHz<i>S</i><sub><i>in</i></sub>To generate an output signal with a sampling rate of 28 kHz<i>S</i><sub><i>out</i></sub>. This interpolation can be (for example) based on things such as<i>s</i><sub><i>out</i></sub>(7<i>n</i>+<i>j</i>)=<img file="TW201214419A_D0020.tif" /><i>h</i><sub><i>32 to 28</i></sub>(<i>j,k</i>)<i>s</i><sub><i>in</i></sub>(<i>8n</i>+<i>j</i>)(n=0,1,2,...,(320/8)-1 and j=0,1,2,...,6) expressions, where<i>h</i><sub>32 to 28</sub>It is a 7×10 matrix. matrix<i>h</i><sub>32 to 28</sub>The value of the left half is shown in the table below:
<tables><img file="TW201214419A_D0021.tif" /></tables>
Flip this half matrix horizontally and vertically to obtain a matrix<i>h</i><sub>32 to 28</sub>The value of the right half of (that is, the element at column r and row c has the same value as the element at column (8-r) and row (11-c)).
The filter bank FB120 also includes the implementation PAS20 of the UHF analysis processing path PAS12, which is configured to be based on a non-integer factor<i>f</i><sub><i>s</i></sub>/<i>f</i><sub><i>ss</i></sub>Integer down-sampling the ultra-wideband signal SISW10, where<i>f</i><sub><i>ss</i></sub>It is the sampling rate of the UHF signal SIS10 (for example, 14 kHz). Path PAS20 includes: Interpolation block IAS10, which is configured to interpolate the signal SISW10 by a factor of 2 to achieve the sampling rate<i>f</i><sub><i>s</i></sub>×2 (for example, to 64 kHz); resample block, which is configured to resample the interpolated signal to the sampling rate<i>f</i><sub><i>ss</i></sub>×4 (for example, by a factor of 7/8 to reach 56 kHz); and the integer downsampling block DS30, which is configured to perform integer downsampling on the resampled signal by a factor of 2 to achieve Sampling rate<i>f</i><sub><i>ss</i></sub>×2 (for example, up to 28 kHz). The interpolation block IAS10 may be implemented according to any of the examples of this operation as described herein (for example, the two-stage multi-phase example described herein). The integer downsampling block DS30 can be implemented according to any of the examples of this operation as described herein (for example, the three-stage multi-phase example described herein). Path PAS20 also includes a spectrum inversion block and an integer downsampler DS10, which implements DS20 by downsampling by 2 integers. The spectrum inversion block and the implementation DH20 can be as described above for the module RSA10 and integer multiples of path PAS12, respectively. Downsampler DS10 described to implement.
It may be necessary to apply the UHF analysis processing path PAS20 to extract the UHF signal SIS10 from the input UHF signal SISW10 with a sampling rate of 32 kHz. The UHF signal SIS10 has a sampling rate of 14 kHz and a high frequency of 7 kHz to 14 kHz The frequency components of the sub-band. 9A to 9F show step-by-step examples of the frequency spectrum of the signal processed in this application of path PAS20 (at each of the corresponding points labeled A to F in FIG. 8C). In Figures 9A to 9F, the shaded area indicates the frequency components of the high-frequency sub-band from 7 kHz to 14 kHz, and the vertical axis indicates the magnitude. Figure 9A shows the representative spectrum of the 32 kHz ultra-wideband signal SISW10. Fig. 9B shows the spectrum after raising the sampling frequency of the signal SISW10 to a sampling rate of 64 kHz. Figure 9C shows the spectrum after resampling the signal with the up-sampling frequency by a factor of 7/8 to reach a sampling rate of 56 kHz. Figure 9D shows the frequency spectrum after the resampled signal is downsampled by integer times to a sampling rate of 28 kHz. Figure 9E shows the frequency spectrum after inverting the frequency spectrum of the integer downsampled signal. FIG. 9F shows the spectrum after integer downsampling of the spectrum-inverted signal to generate the UHF signal SIS10 with a sampling rate of 14 kHz.
The interpolation block IAS10 and the integer downsampling block DS30 of the path PAS20 can be implemented according to any of the examples of such operations as described herein (for example, the multi-stage multi-phase example described herein). The 7/8 resample block of the path PAS20 can be implemented to use a multiphase implementation to resample the input signal with a sampling rate of 64 kHz<i>S</i><sub><i>in</i></sub>To generate an output signal with a sampling rate of 56 kHz<i>S</i><sub><i>out</i></sub>. This resampling can, for example, be based on<img file="TW201214419A_D0022.tif" />And j=0,1,2,...,6), where<i>h</i><sub>64 to 56</sub>It is a 7×10 matrix. matrix<i>h</i><sub>64 to 56</sub>The values of the left half of the specific implementation are shown in the table below:
<maths><img file="TW201214419A_D0023.tif" /></maths>
Flip this half matrix horizontally and vertically to obtain the matrix<i>h</i><sub>64 to 56</sub>The value of the right half of this particular implementation (ie, the elements in column r and row c have the same values as the elements in column (8-r) and row (11-c)).
Fig. 7C shows another example in which the mid-frequency sub-band extends from 3.5 kHz to 7.5 kHz, so that the narrow-band signal SIL10 and the high-frequency signal SIH10 both describe the region from 3.5 kHz to 4 kHz and the high-frequency signal SIH10 and the ultra-high frequency signal Both SIS10 describe the area from 7 kHz to 7.5 kHz.
In some implementations, providing the overlap between the sub-bands as in the example of FIG. 7B and FIG. 7C allows the use of processing paths with smooth attenuation in the overlap region. These filters are generally easier to design, have lower computational complexity, and/or introduce less delay than filters with sharper or "brick-wall" responses. Filters with sharp transition regions tend to have higher side lobes (side lobes can cause frequency overlap) than filters of similar order with smooth attenuation. Filters with sharp transition regions can also have long impulse responses, which can cause ringing artifacts. For filter bank implementations with one or more IIR filters, allowing smooth attenuation in the overlap region can enable the use of one or more filters with each pole farther from the unit circle. This pair ensures stable fixation Point implementation is very important.
The overlap of the sub-bands allows for smooth mixing of the sub-bands, which can cause less audio artifacts, reduced frequency overlap, and/or less noticeable transitions from one sub-band to another. For implementations where two or more of the narrowband encoder EN100, the high frequency encoder EH100, and the UHF encoder ES100 operate according to different encoding methods, one or more of these features may be particularly ideal. For example, different encoding techniques can produce signals that sound very different. An encoder that encodes the spectrum envelope in the form of a codebook index can generate a signal whose sound is different from the sound of the signal produced by an encoder that encodes the amplitude spectrum. The time domain encoder (for example, pulse code modulation or PCM encoder) can generate a signal whose sound is different from the sound of the signal produced by the frequency domain encoder. An encoder that encodes a signal with the representation of the spectral envelope and the corresponding residual signal can generate a signal whose sound is different from the signal produced by an encoder that only encodes the signal with the representation of the spectral envelope (for example, a transform-based encoder) the sound of. An encoder that encodes a signal as a representation of its waveform can produce an output whose sound is different from the sound from a sinusoidal encoder. Under these conditions, the use of filters with sharp transition regions to define non-overlapping sub-bands can result in sudden and clearly perceptible transitions between the sub-bands in the synthesized ultra-wideband signal.
In addition, the coding efficiency of an encoder (for example, a waveform encoder) may decrease as the frequency increases. At low bit rates, especially in the presence of background noise, the encoding quality may be reduced. Under these conditions, providing overlap of sub-bands can increase the quality of the reproduced frequency components in the overlap region.
The overlap of two sub-bands (for example, the overlap of the low-frequency sub-band and the middle-frequency sub-band, or the overlap of the middle-frequency sub-band and the high-frequency sub-band) is defined as the decrease in the frequency response of the path that generates the higher-frequency sub-band The distance from the -20 dB point to the point where the frequency response of the path that produces the lower frequency subband drops to -20 dB. In each instance of the filter bank FB100 and/or FB200, this overlap is in the range of about 200 Hz to about 1 kHz. The range of about 400 Hz to about 600 Hz may represent an ideal compromise between coding efficiency and perceptual smoothness. In the specific example shown in Figures 7B and 7C, each overlap is about 500 Hz.
It should be noted that as a result of the spectrum inversion operation in the processing paths PAH12 and PAS12, the frequency components of the high-frequency signal SIH10 and the ultra-high-frequency signal SIS10 are inverted. The subsequent operations in the encoder and corresponding decoder can be configured accordingly. For example, the high-frequency excitation generator GXH100 as described herein can be configured to generate the high-frequency excitation signal SXH10 that also has a form of spectrum inversion.
Figure 10 shows a block diagram of an implementation FB220 of the filter bank FB212, which can be used for the application shown in Figure 7B. The filter bank FB220 includes an implementation PSN20 of the narrowband synthesis processing path PSN10, which is configured to receive a sampling rate<i>f</i><sub><i>SN</i></sub>(For example, 8 kHz) narrowband signal SDL10 and perform interpolation by 2 to generate a sampling rate<i>f</i><sub><i>SW</i></sub>(For example, 16 kHz) narrowband output signal SOL10. In this example, the path PSN20 includes an implementation IN20 of the interpolator IN10 (e.g., an FIR polyphase implementation as described herein) and an optional shaping filter FSL10 (e.g., a first-order pole-zero filter). In a specific example, the shaping filter FSL10 is implemented as a second-order IIR filter with the following transfer function:
<maths><img file="TW201214419A_D0024.tif" /></maths>
The filter bank FB220 also includes the implementation PSH20 of the high frequency synthesis processing path PSH12, which is configured to be based on a non-integer factor<i>f</i><sub><i>SW</i></sub>/<i>f</i><sub><i>SH</i></sub>Sampling rate<i>f</i><sub><i>SH</i></sub>(For example, 7 kHz) high-frequency signal SDH10. Path PSH20 includes: implementation IH20 of interpolator IH10, which is configured to interpolate signal SDH10 by a factor of 2 to achieve the sampling rate<i>f</i><sub><i>SH</i></sub>×2 (for example, up to 14 kHz); spectrum inversion block, which can be implemented as described above with reference to module RHS10 of path PSH12; interpolation block IH30, which is configured to interpolate by a factor of 2 The signal whose spectrum is inverted to achieve the sampling rate<i>f</i><sub><i>SH</i></sub>×4 (for example, up to 28 kHz); and a resample block, which is configured to resample (for example, by a factor of 4/7) the interpolated signal to reach the sampling rate<i>f</i><sub><i>SW</i></sub>. In this particular example, the path PSH20 also includes an optional spectrum shaping filter FSW10. The optional spectrum shaping filter FSW10 can be implemented as a low-pass filter that is configured to shape the signal to obtain the desired overall filter response. And/or implemented as a notch filter configured to attenuate the components of the signal at 7100 Hz. In a specific example, the shaping filter FSW10 is implemented as a notch filter, which has the following transfer function:
<maths><img file="TW201214419A_D0025.tif" /></maths>
Or the following transfer function:
<maths><img file="TW201214419A_D0026.tif" /></maths>
The interpolation block IH30 of the path PSH20 can be implemented according to any of the examples of this operation as described herein (for example, the three-stage multi-phase example described herein). The 4/7 resample block of the path PSH20 can be implemented to use a multiphase implementation to resample the input signal with a sampling rate of 28 kHz<i>S</i><sub><i>in</i></sub>To generate an output signal with a sampling rate of 16 kHz<i>S</i><sub><i>out</i></sub>. This resampling can, for example, be based on<i>s</i><sub><i>out</i></sub>(4<i>n</i>+<i>j</i>)=<img file="TW201214419A_D0027.tif" />(<i>j</i>,<i>k</i>)<i>s</i><sub><i>in</i></sub>(7<i>n</i>+<i>j</i>)(n=0,1,2,...and j=0,1,2,3) expression, where<i>h</i><sub>28 to 16</sub>It is a 4×10 matrix. matrix<i>h</i><sub>28 to 128 to 16</sub>The values of the left half of the specific implementation are shown in the table below:
<tables><img file="TW201214419A_D0028.tif" /></tables>
matrix<i>h</i><sub>28 to 16</sub>The values of the right half of this particular implementation are shown in the table below:
<tables><img file="TW201214419A_D0029.tif" /></tables>
<tables><img file="TW201214419A_D0030.tif" /></tables>
The filter bank FB220 also includes an implementation PSW20 of the broadband synthesis processing path PSW12, which is configured to receive a sampling rate<i>f</i><sub><i>sw</i></sub>(For example, 16 kHz) wideband signal SDW10 and perform interpolation by 2 to generate a sampling rate<i>f</i><sub><i>s</i></sub>(For example, 32 kHz) wideband output signal SOW10. In this example, the path PSW20 includes an implementation IW20 of the interpolator IW10 (e.g., an FIR polyphase implementation as described herein) and an optional shaping filter (e.g., a second-order pole-zero filter).
The filter bank FB220 also includes the implementation PSS20 of the UHF synthesis processing path PSS12, which is configured to be based on a non-integer factor<i>f</i><sub><i>s</i></sub>/<i>f</i><sub><i>ss</i></sub>Sampling rate<i>f</i><sub><i>ss</i></sub>(For example, 14 kHz) UHF signal SDS10, where<i>f</i><sub><i>s</i></sub>It is the sampling rate of the ultra-wideband signal SOSW10 (for example, 32 kHz). The filter bank FB220 includes: the implementation IS20 of the interpolator IS10, which is configured to interpolate the signal SDS10 by a factor of 2 to achieve the sampling rate<i>f</i><sub><i>ss</i></sub>×2 (for example, up to 28 kHz); spectrum inversion block, which can be implemented as described above with reference to module RHD10 of path PSS12; interpolation block IS30, which is configured to interpolate by a factor of 2 The signal whose spectrum is inverted to achieve the sampling rate<i>f</i><sub><i>ss</i></sub>×4 (for example, up to 56 kHz); resample block, which is configured to resample (for example, by a factor of 8/7) the interpolated signal to reach the sampling rate<i>f</i><sub><i>s</i></sub>×2; and the integer down-sampling block DSS10, which is configured to perform integer down-sampling on the re-sampled signal by a factor of 2 to achieve the sampling rate<i>f</i><sub><i>s</i></sub>(For example, up to 32 kHz). In this particular example, the path PSS20 also includes an optional spectrum shaping block, which can be implemented as a filter configured to shape the signal to obtain the desired overall filter response (e.g., 30th order FIR filter).
It may be necessary to apply the UHF synthesis processing path PSS20 to extract the UHF signal SOS10 from the decoded input UHF signal SDS10 with a sampling rate of 14 kHz. The UHF signal SOS10 has a sampling rate of 32 kHz and 7 kHz to The frequency component of the high frequency sub-band of 14 kHz. 11A to 11F show step-by-step examples of the frequency spectrum of the signal processed in this application of path PSS20 (at each of the corresponding points labeled A to F in FIG. 10). In Figures 11A to 11F, the shaded area indicates the frequency components of the high-frequency subband from 7 kHz to 14 kHz, and the vertical axis indicates the magnitude. Figure 11A shows the representative spectrum of the 14 kHz UHF signal SDS10, which contains the frequency components of the high frequency subband from 7 kHz to 14 kHz, which have been inverted. Figure 11B shows the spectrum after interpolating the signal SDS10 to reach a sampling rate of 28 kHz. Figure 11C shows the frequency spectrum after inverting the frequency spectrum of the interpolated signal. Figure 11D shows the spectrum after interpolating the spectrum-inverted signal to reach a sampling rate of 56 kHz. Figure 11E shows the spectrum after resampling the interpolated signal by a factor of 8/7 to reach a sampling rate of 64 kHz. FIG. 11F shows the frequency spectrum after integer downsampling of the resampled signal to generate an ultra-high frequency signal SOS10 with a sampling rate of 32 kHz.
The integer downsampling block DSS10 of the path PSS20 can be implemented according to any of the examples of this operation as described herein (for example, the three-stage polyphase example described herein). The interpolators IH20, IH30, IS20, and IS30 of the paths PSH20 and PSS20 can be implemented according to any of the examples of this operation as described herein. In a specific example, each of the interpolators IH20, IH30, IS20, and IS30 are implemented according to the three-stage polyphase example described herein.
The 8/7 resample block of path PSS20 can be implemented to use multiphase interpolation to resample the input signal with a sampling rate of 56 kHz<i>S</i><sub><i>in</i></sub>To generate an output signal with a sampling rate of 64 kHz<i>S</i><sub><i>out</i></sub>. In one instance, the basis of<img file="TW201214419A_D0031.tif" />And j=0,1,2,...,6) multi-phase interpolation to perform this resampling, where<i>h</i><sub>56 to 64</sub>It is an 8×5 matrix. matrix<i>h</i><sub>56 to 64</sub>The values of the specific implementation are shown in the table below:
<maths><img file="TW201214419A_D0032.tif" /></maths>
The narrowband encoder EN100 is implemented based on the source filter model. The source filter model encodes the input speech signal as: (A) describe a set of parameters of the filter; and (B) drive the described filter to generate the input speech signal The synthetic reproduced excitation signal. Figure 12A shows an example of the spectral envelope of a speech signal. The peak that characterizes this spectral envelope represents the resonance of the vocal tract and is called the formant. Most speech encoders encode at least this rough spectrum structure into a set of parameters, such as filter coefficients.
Figure 12B shows an example of a basic source filter configuration as applied to the encoding of the spectral envelope of the narrowband signal SIL10. The analysis module calculates a set of parameters, which characterizes the filter corresponding to the speech sound within a time period (usually ten milliseconds or twenty milliseconds). A whitening filter (also called an analysis or prediction error filter) configured according to their filter parameters removes the spectrum envelope to flatten the signal on the spectrum. The resulting whitened signal (also called residual) has less energy, and therefore has less variation and is easier to encode than the original speech signal. The errors generated by the encoding of the residual signal can also be spread more evenly on the frequency spectrum. Filter parameters and residuals are usually quantized to obtain efficient transmission on the channel. At the decoder, the synthesis filter configured according to the filter parameters is excited by the signal based on the residue to produce a synthesized version of the original speech sound. The synthesis filter is usually configured to have a transfer function that is the inverse function of the transfer function of the whitening filter.
Figure 13 shows a block diagram of the basic implementation EN110 of the narrowband encoder EN100. In this example, the linear predictive coding (LPC) analysis module LPN10 encodes the spectral envelope of the narrowband signal SIL10 into a set of linear predictive (LP) coefficients (for example, the coefficients of the all-pole filter 1/A(z)). The analysis module usually processes the input signal into a series of non-overlapping frames, in which a new set of coefficients is calculated for each frame. The frame period is generally the period when the signal is expected to be locally stable; a common example is 20 milliseconds (equivalent to 160 samples at a sampling rate of 8 kHz). In one example, the LPC analysis module LPN10 is configured to calculate a set of ten LP filter coefficients to characterize the formant structure of each twenty millisecond frame. It is also possible to implement an analysis module to process the input signal into a series of overlapping frames.
The analysis module can be configured to directly analyze the samples of each frame, or can first weight the samples according to a windowing function (eg, Hamming window). The analysis of the frame can also be performed in a window larger than the frame (such as a 30 millisecond window). This window can be symmetric (for example, 5-20-5, so that it includes 5 ms immediately before and after the 20 ms frame) or asymmetric (for example, 10-20, so that it includes the last 10 of the previous frame). millisecond). The LPC analysis module is usually configured to use Levinson-Durbin recursion or Leroux-Gueguen algorithm to calculate LP filter coefficients. In another implementation, the analysis module can be configured to calculate a set of cepstral coefficients instead of a set of LP filter coefficients for each frame.
By quantizing these filter parameters, the output rate of the encoder EN110 can be significantly reduced, which has relatively little impact on the reproduction quality. Linear prediction filter coefficients are difficult to effectively quantize and are usually mapped to another representation, such as line spectrum pair (LSP) or line spectrum frequency (LSF), for quantization and/or entropy coding. In the example of FIG. 13, the LP filter coefficient to LSF transform XLN10 transforms the set of LP filter coefficients into a set of corresponding LSF. Other one-to-one representations of LP filter coefficients include: partial autocorrelation coefficient; logarithmic area ratio; immittance spectrum pair (IPS); and immittance spectrum frequency (ISF), all of which are used in GSM (Global System for Mobile Communications) AMR -WB (Adaptive Multi-rate Broadband) codec. Generally, the transformation between a set of LP filter coefficients and a set of corresponding LSFs is reversible, but the embodiment also includes the implementation of the encoder EN110 in which the transformation is not error-free and reversible.
The quantizer QLN10 is configured to quantize the set of narrowband LSF (or other coefficient representation), and the narrowband encoder EN110 is configured to output the result of this quantization as the narrowband filter parameter FPN10. This quantizer usually includes a vector quantizer that encodes the input vector as an index to the corresponding vector item in the table or codebook.
May need quantizer QLN10 and time noise shaping. Figure 14 shows a block diagram of this implementation QLN20 of the quantizer QLN10. For each frame, calculate the LSF quantization error vector and multiply the LSF quantization error vector with a scale factor V40 whose value is less than one. In the next frame, add this scaled quantization error to LSF before quantization. The value of the scale factor V40 can be dynamically adjusted depending on the amount of fluctuations that already exist in the unquantized LSF vector. For example, when the difference between the current LSF vector and the previous LSF vector is large, the value of the scale factor V40 is close to zero, so that noise shaping is hardly performed. When there is a small difference between the current LSF vector and the previous LSF vector, the value of the scale factor V40 is close to one. It can be expected that the resulting LSF quantization minimizes spectral distortion when the speech signal changes, and minimizes spectral fluctuations when the speech signal is relatively constant from one frame to another.
FIG. 15 shows a block diagram of another noise shaping implementation QLN30 of the quantizer QLN10. Additional description of temporal noise shaping in vector quantization can be found in U.S. Published Patent Application No. 2006/0271356 (Vos et al.) published on November 30, 2006.
As shown in FIG. 13, the narrowband encoder EN110 can be configured to pass the narrowband signal SIL10 through a whitening filter WF10 (also called analysis or prediction error filter) configured according to the set of filter coefficients. Generate residual signal. In this particular example, the whitening filter WF10 is implemented as an FIR filter, but it can also be implemented using IIR. This residual signal will usually contain perceptually important information (such as the long-term structure related to the pitch) of the speech frame that is not represented in the narrowband filter parameter FPN10. QXN10 was quantizer configured to calculate a residual signal of this quantization table shown XL10 excitation signal to output to the encoded narrowband. This quantizer usually includes a vector quantizer that encodes the input vector as an index to the corresponding vector item in the table or codebook. Alternatively, the quantizer can be configured to send one or more parameters, and a vector can be dynamically generated at the decoder based on the one or more parameters instead of fetching from the memory as in the sparse codebook method. This method is used in coding schemes such as Algebraic CELP (Codebook Excited Linear Prediction) and in codecs such as 3GPP2 (3rd Generation Partnership Project 2) EVRC (Enhanced Variable Rate Codec).
The narrowband encoder EN110 may be required to generate the encoded narrowband excitation signal according to the same filter parameter values that will be available to the corresponding narrowband decoder. In this way, the resulting encoded narrowband excitation signal may have resolved the non-idealities of their parameter values, such as quantization errors, to some extent. Accordingly, it may be necessary to configure the whitening filter with the same coefficient values that will be available at the decoder. In the basic example of the encoder EN110 as shown in FIG. 13, the dequantizer IQN10 dequantizes the narrowband coding parameter FPN10, and the LSF to LP filter coefficient transform IXN10 maps the obtained values back to a set of corresponding LPs. Filter coefficients, and this set of coefficients is used to configure the whitening filter WF10 to generate the residual signal quantized by the quantizer QXN10. Some implementations of the narrowband encoder EN100 are configured to calculate the encoded narrowband excitation signal XL10 by identifying one of the set of codebook vectors that best matches the residual signal. However, it is noted that the narrowband encoder EN100 can also be implemented to calculate the quantized representation of the residual signal without actually generating the residual signal. For example, the narrowband encoder EN100 can also be configured to: use multiple codebook vectors to generate the corresponding composite signal (for example, according to a set of current filter parameters), and select the best one in the perceptually weighted domain Ground matches the codebook vector associated with the generated signal of the original narrowband signal SIL10.
Figure 16 shows a block diagram of the narrowband decoder DN100 implementing DN110. The dequantizer IQXN10 dequantizes the narrowband filter parameters FPN10 (in this case, dequantizes into a set of LSF), and the LSF to LP filter coefficient transform IXN20 transforms the LSF into a set of filter coefficients (for example, see above The inverse quantizer IQN10 and transform IXN10 of the narrowband encoder EN110 are described). The dequantizer IQLN10 dequantizes the encoded narrowband excitation signal XL10 to generate a decoded narrowband excitation signal XLD10. Based on the filter coefficients and the narrowband excitation signal XLD10, the narrowband synthesis filter FNS10 synthesizes the narrowband signal SDL10. In other words, the narrowband synthesis filter FNS10 is configured to perform spectrum shaping on the narrowband excitation signal XLD10 according to the inverse quantized filter coefficients to generate the narrowband signal SDL10. The narrowband decoder DN110 also provides the narrowband excitation signal XL10a to the high frequency encoder DH100. The high frequency encoder DH100 uses the narrowband excitation signal XL10a to derive the high frequency excitation signal XHD10 as described in this article, and the narrowband decoder DN110 The narrowband excitation signal XL10b is provided to the SHB encoder DS100, and the SHB encoder DS100 uses the narrowband excitation signal XL10b to derive the SHB excitation signal XSD10 as described herein. In some implementations as described below, the narrowband decoder DN110 can be configured to provide additional information related to the narrowband signal (such as spectral tilt, pitch gain, and delay and/or voice mode) to the high frequency Decoder DH100 and/or to SHB decoder DS100.
The system of the narrowband encoder EN110 and the narrowband decoder DN110 is a basic example of an analysis-by-synthesis speech codec. Codebook Excited Linear Prediction (CELP) coding is a popular coding for analysis by synthesis, and the implementation of these encoders can perform residual waveform coding, including operations such as: selecting each item from a fixed and adaptive codebook , Error minimization operation, and/or perceptual weighting operation. Other implementations of coding with synthesis for analysis include Mixed Excited Linear Prediction (MELP), Algebraic CELP (ACELP), Relaxed CELP (RCELP), Regular Pulse Excitation (RPE), Multi-pulse CELP (MPE), and Vector Sum Excited Linear Prediction ( VSELP) encoding. Related coding methods include multi-band excitation (MBE) and prototype waveform interpolation (PWI) coding. Examples of standardized speech codecs that use synthesis for analysis include: ETSI (European Telecommunications Standards Institute)-GSM full-rate codec (GSM 06.10), which uses residual excitation linear prediction (RELP); GSM enhanced full-rate coding Decoder (ETSI-GSM 06.60); ITU (International Telecommunications Union) standard 11.8 kb/s G.729 Annex E encoder; IS (temporary standard)-641 codec for IS-136 (a time-sharing multiple access mechanism); GSM adaptive multi-rate (GSM-AMR) codec; and 4GV<sup>TM</sup>(Fourth generation Vocoder<sup>TM</sup>) Codec (QUALCOMM Incorporated, San Diego, CA). The narrowband encoder EN110 and the corresponding decoder DN110 can be based on any of these technologies, or the speech signal can be expressed as (A) describing a set of parameters of the filter and (B) used to drive the described filter Any other speech coding technology (whether known or to be developed) that reproduces the excitation signal of the speech signal is implemented.
Even after the whitening filter has removed the coarse spectrum envelope from the narrowband signal SIL10, a considerable amount of fine harmonic structure may still be retained, especially for voiced speech. FIG. 17A shows a spectrogram of an example of a residual signal that can be generated by a whitening filter for a voiced signal such as a vowel. The periodic structure seen in this example is related to pitch, and different voiced sounds spoken by the same speaker may have different formant structures but similar pitch structures. Figure 17B shows a time-domain diagram of an example of this residual signal, which shows the sequence of pitch pulses over time.
By using one or more parameter values to encode the characteristics of the pitch structure, coding efficiency and/or speech quality can be increased. One of the important characteristics of the pitch structure is the frequency of the first harmonic (also called the fundamental frequency), which is usually in the range of 60 Hz to 400 Hz. This characteristic is usually encoded as the reciprocal of the fundamental frequency (also known as pitch lag). The pitch lag indicates the number of samples in one pitch period, and can be encoded as an offset from the minimum or maximum pitch lag value and/or encoded as one or more codebook indexes. Voice signals from male speakers tend to have greater pitch lag than voice signals from female speakers. Another signal characteristic related to the pitch structure is periodicity, which indicates the strength of the harmonic structure, or in other words, the degree to which the signal is harmonic or non-harmonic. Two typical indicators of periodicity are zero crossing and normalized autocorrelation function (NACF). Periodicity can also be indicated by pitch gain, which is usually encoded as codebook gain (for example, quantized adaptive codebook gain).
The narrowband encoder EN100 may include one or more modules configured to encode the long-term harmonic structure of the narrowband signal SIL10. As shown in Figure 17C, a typical CELP example that can be used includes an open loop LPC analysis module that encodes short-term characteristics or coarse spectrum envelopes, followed by a closed loop long-term predictive analysis stage that encodes fine pitch or harmonic structure. The short-term characteristics are encoded as filter coefficients, and the long-term characteristics are encoded as parameter values such as pitch lag and pitch gain.
For example, the LPC residue coded by the CELP coding technique usually includes a fixed codebook part and an adaptive codebook part. For example, the narrowband encoder EN100 can be configured to output an encoded narrowband excitation signal XL10, which includes one or more fixed codebook indexes and corresponding gain values, and one or more adaptive codebook gains The form of value. The calculation of this quantized representation of the narrowband residual signal (for example, by the quantizer QXN10) may include selecting these indexes and calculating these gain values.
The structure retained after the residual long-term predictive analysis can be encoded as one or more indexes in the fixed codebook and one or more corresponding fixed codebook gains. The quantization of the fixed codebook can be performed using pulse coding techniques such as factor or combined pulse coding. The encoding of the pitch structure may also include interpolation of the pitch prototype waveform. This operation may include calculating the difference between successive pitch pulses. For frames corresponding to silent speech (which is usually noise-like and unstructured), the modeling of the long-term structure can be disabled. Alternatively, a modified discrete cosine transform (MDCT) technique or other transform-based techniques can be used to encode the LPC residue, especially for popular audio or non-speech applications (e.g., music).
The implementation of the narrowband decoder DN110 according to the example shown in Figure 17C can be configured to: after the long-term structure (pitch or harmonic structure) has been restored, the narrowband excitation signal XL10a is output to the high-frequency decoder DH100, and/or output the narrowband excitation signal XL10b to the SHB decoder DS100. For example, the decoder can be configured to output the narrowband excitation signal XL10a and/or XL10b as an inverse quantized version of the encoded narrowband excitation signal XL10. Of course, it is also possible to implement the narrowband decoder DN100, so that the high frequency decoder DH100 performs inverse quantization of the encoded narrowband excitation signal XL10 to obtain the narrowband excitation signal XL10a, and/or makes the SHB decoder DS100 perform the encoded The narrowband excitation signal XL10 is dequantized to obtain the narrowband excitation signal XL10b.
In the implementation of the ultra-wideband speech encoder SWE100 according to the example shown in FIG. 17, the high-frequency encoder EH100 and/or the SHB encoder ES100 can be configured to receive a narrow frequency such as that produced by a short-term analysis or a whitening filter. Frequency excitation signal. In other words, the narrowband encoder EN100 can be configured to output the narrowband excitation signal XL10a to the high frequency encoder EH100 and/or output the narrowband excitation signal XL10b to the SHB encoder ES100 before encoding the long-term structure. However, it may be necessary for the high-frequency encoder EH100 to receive the same encoding information from the narrow-band channel that will be received by the high-frequency decoder DH100, so that the encoding parameters generated by the high-frequency encoder EH100 may have been considered to some extent. Ideality. Therefore, it may be better for the high frequency encoder EH100 to reconstruct the high frequency excitation signal XH10 from the same parameterized and/or quantized encoded narrow frequency excitation signal XL10 output by the SWB encoder SWE100. For example, the narrowband encoder EN100 can be configured to output the narrowband excitation signal XL10a as an inverse quantized version of the encoded narrowband excitation signal XL10. One of the potential advantages of this method is to more accurately calculate the high frequency gain factor CPH10b (described below).
Similarly, it may be necessary for the SHB encoder ES100 to receive the same encoding information from the narrowband channel that will be received by the SHB decoder DS100, so that the encoding parameters generated by the SHB encoder ES100 may have taken into account the non-idealities in the information to some extent . Therefore, it is preferable that the SHB encoder ES100 reconstruct the SHB excitation signal XS10 from the same parameterized and/or quantized encoded narrowband excitation signal XL10 output by the SWB encoder SWE100. For example, the narrowband encoder EN100 can be configured to output the narrowband excitation signal XL10b as an inverse quantized version of the encoded narrowband excitation signal XL10. One of the potential advantages of this method is to more accurately calculate the SHB gain factor CPS10b (described below).
In addition to the parameters that characterize the short-term and/or long-term structure of the narrow-band signal SIL10, the narrow-band encoder EN100 can also generate parameter values related to other characteristics of the narrow-band signal SIL10. This equivalent value (which can be appropriately quantized so as to be output by the SWB speech encoder SWE100) can be included in the narrowband filter parameter FPN10 or output separately. The high-frequency encoder EH100 can also be configured to calculate the high-frequency encoding parameter CPH10 based on one or more of these additional parameters (for example, after inverse quantization). At the SWB decoder SWD100, the high frequency decoder DH100 can be configured to receive the parameter values via the narrowband decoder DN100 (for example, after dequantization). Alternatively, the high frequency decoder DH100 can be configured to directly receive (and possibly inverse quantize) the parameter values. Likewise, the SHB encoder ES100 may be configured to calculate the SHB encoding parameter CPS10 according to one or more of these additional parameters (for example, after inverse quantization). At the SWB decoder SWD100, the SHB decoder DS100 can be configured to receive the parameter values via the narrowband decoder DN100 (for example, after dequantization). Alternatively, the SHB decoder DS100 can be configured to directly receive (and possibly dequantize) the parameter values.
In an example of the additional narrowband coding parameters, the narrowband encoder EN100 generates spectral tilt values and speech mode parameters for each frame. The spectral tilt is related to the shape of the spectral envelope on the passband, and is usually represented by the quantized first reflection coefficient. For most vocal sounds, the spectral energy decreases as the frequency increases, so that the first reflection coefficient is negative and can be close to -1. Most silent sounds have a flat frequency spectrum, so that the first reflection coefficient is close to zero, or have a spectrum with greater energy at high frequencies, so that the first reflection coefficient is positive and can be close to +1.
The voice mode (also called the voice mode) indicates that the current frame represents voiced voice or unvoiced voice. This parameter can have a binary value, which is based on one or more measurements of the periodicity of the frame (such as zero crossing, NACF, pitch gain) and/or voice activity, such as this measurement and threshold The relationship between. In other implementations, the speech mode parameter has one or more other states to indicate modes such as: quiet or background noise, or transition between quiet and voiced speech.
It is not an unimportant task to determine the order of the LPC analysis of the SHB signal SIS10. Generally speaking, because the SHB signal SIS10 has a relatively large bandwidth (for example, 7 kHz), a relatively higher order of LPC coefficients may be required to support the reconstruction of the SWB signal SISW10, and the perception result is satisfactory. An example of this implementation uses traditional linear predictive coding (LPC) analysis to obtain eight spectrum parameters to describe the spectrum envelope of the SHB signal SIS10, and uses similar analysis to obtain six spectrum parameters to describe the spectrum envelope of the high-frequency signal SIH10. In order to obtain efficient coding, these prediction coefficients are converted into line spectrum frequencies (LSF) and then quantized using a vector quantizer as described herein (for example, using a temporal noise shaping vector quantizer).
FIG. 18 shows a block diagram of an implementation of EH110 of the high frequency encoder EH100, and FIG. 19 shows a block diagram of an implementation of ES110 of the SHB encoder ES100. The high frequency encoder EH100 and the SHB encoder ES100 can be configured to have an LPC analysis path similar to the LPC analysis path in the narrowband encoder EN110. For example, the narrowband encoder EN110 includes the LPC analysis path (including quantization and dequantization) LPN10-XLN10-QLN10-IQN10-IXN10, while the high-frequency encoder EH110 includes the similar path LPH10-XFH10-QLH10-IQH10-IXH10, and The SHB encoder EH110 includes similar paths LPS10-XFS10-QLS10-IQS10-IXS10. Therefore, two or more of the encoders EN100, EH100 and ES100 can be configured to use the same LPC analysis processing path (which may include quantization and may also include inverse quantization) at different times and in different configurations. . The high-frequency encoder EH110 includes a synthesis filter FSH10. The synthesis filter FSH10 is configured to generate a synthesized high-frequency signal SYH10 according to the high-frequency excitation signal XH10 and the LPC parameters generated by the transformation IXH10, and the SHB encoder ES110 includes a synthesis The filter FSS10 and the synthesis filter FSS10 are configured to generate the synthesized SHB signal SYS10 according to the SHB excitation signal XS10 and the LPC parameters generated by the transformation IXS10.
For different types of voice frames, different numbers of bits can be allocated in the high-frequency and SHB quantization processing procedures. Since the quiet period often does not contain many high-frequency or SHB components, not sending high-frequency or SHB information during the quiet period can save the total bit rate requirement. It is also possible to process the audio frame and the silent frame in different ways during the VQ training and encoding process. Generally speaking, when there are not many constraints on the codebook size and codebook search complexity, the single-stage large codebook VQ can be used by the high-frequency encoder EH100 and/or the SHB encoder ES100. On the other hand, if there are strict constraints on the complexity of the memory and the quantization process, the multi-stage and/or split VQ can be adopted by the high-frequency encoder EH100 and/or the SHB encoder ES100.
As shown in FIG. 19, the SHB encoder ES110 includes an SHB excitation generator XGS10 configured to generate an SHB excitation signal XS10 from a narrowband excitation signal XL10b. As shown in FIG. 21, the SHB decoder DS110 also includes an example of the SHB excitation generator XGS10 configured to generate the SHB excitation signal XS10 from the narrowband excitation signal XL10b. Figure 22A shows a block diagram of an implementation XGS20 of the SHB excitation generator XGS10, which implementation XGS20 is configured to generate an SHB excitation signal XS10 from a narrowband excitation signal XL10b. The generator XGS20 includes the spectrum expander SX10, the SHB analysis filter bank FBS10 and the adaptive whitening filter AW10.
The spectrum expander SX10 is configured to expand the frequency spectrum of the narrowband excitation signal XL10b into the frequency range occupied by the SHB signal SIS10. The spectrum expander SX10 can be configured to apply memoryless non-linear functions to the narrow-frequency excitation signal XL10b, such as absolute value functions (also known as full-wave rectification), half-wave rectification, squaring, cube-finding, or cutting . The spectrum expander SX10 can be configured to increase the sampling frequency of the narrowband excitation signal XL10b before applying the nonlinear function (for example, to a sampling rate of 32 kHz, or to a sampling rate equal to or closer to the sampling rate of the SHB signal SIS10 Rate). Then the analysis filter bank FBS10 (which can be the same high-frequency analysis filter bank used to generate the high-frequency excitation signal (for example, the HB analysis processing path PAH10, PAH12, or PAH20)) is applied to the spectrum-spread signal to generate The desired sampling rate (e.g.,<i>f</i><sub><i>ss</i></sub>Or 14 kHz) signal.
The spectrum-spread signal is likely to have a significant decrease in amplitude as the frequency increases. The whitening filter WF20 (for example, an adaptive sixth-order linear prediction filter) can be used to flatten the harmonically expanded result on the spectrum to generate the SHB excitation signal XS10. Another implementation of the SHB excitation generator XGS20 can be configured to mix the harmonically expanded signal and the noise signal, which can be time-modulated according to the time-domain envelope of the narrow-band signal SIL10 or the narrow-band excitation signal XL10b.
It should be noted that SHB excitation is generated at both the encoder and the decoder. In order to make the decoding process consistent with the encoding process, the encoder and decoder may need to generate the same SHB excitation. This result can be achieved by using information from the encoded narrowband excitation signal XL10 available to both the encoder and the decoder to generate SHB excitation at both the encoder and the decoder. For example, the dequantized narrowband excitation signal can be used as input XL10b to the SHB excitation generator XGS10 at both the encoder and the decoder. When a sparse codebook (a codebook whose items are mostly zero values) has been used to calculate the quantized representation of the residue, artifacts may appear in the synthesized speech signal. Especially when the narrowband excitation signal has been coded at a low bit rate, codebook sparseness may occur. Artifacts caused by codebook sparseness are usually quasi-periodic in time, and mostly occur above 3 kHz. Because the human ear has better time resolution at higher frequencies, these artifacts may be more noticeable at high frequencies and/or UHF.
The embodiment includes the implementation of a high frequency excitation generator XGS10 configured to perform anti-sparse filtering. 22B shows a block diagram of an implementation XGS30 of the SHB excitation generator XGS20, which implementation XGS30 includes an anti-sparse filter ASF10 configured to filter the narrowband excitation signal XL10b. In an example, the anti-sparse filter ASF10 is implemented with<img file="TW201214419A_D0033.tif" /><img file="TW201214419A_D0034.tif" />All-pass filter in the form of.
The anti-sparse filter ASF10 can be configured to change the phase of its input signal. For example, it may be necessary for the anti-sparse filter ASF10 to be configured and configured to randomize the phase of the SHB excitation signal XS10, or otherwise distribute it more evenly over time. It may also be necessary that the response of the anti-sparse filter ASF10 be flat in the frequency spectrum so that the magnitude spectrum of the filtered signal does not change significantly. In an example, the anti-sparse filter ASF10 is implemented as an all-pass filter with a transfer function according to the following expression:
<maths><img file="TW201214419A_D0035.tif" /></maths>
One function of this filter can be to spread the energy of the input signal so that the energy is no longer concentrated in only a few samples.
Artifacts caused by codebook sparseness are usually more pronounced for noise-like signals, where the residual includes less pitch information, and is also more pronounced for speech in background noise. In the case where the excitation has a long-term structure, sparseness usually causes less artifacts, and in fact phase modification can cause noise in the audible signal. Therefore, it may be necessary to configure the anti-sparse filter ASF10 to filter the unvoiced signal and pass at least some of the voiced signal without change. The ASF filter ASF10 can be selected for use based on factors such as vocalization, periodicity, and/or spectral tilt. A silent signal is characterized by a low pitch gain (for example, quantized narrowband adaptive codebook gain) and a spectral tilt close to zero or a positive number (for example, the quantized first reflection coefficient), which is close to zero or A positive spectral tilt indicates that the spectral envelope is flat or tilts upward as the frequency increases. A typical implementation of the anti-sparse filter ASF10 is configured to filter silent sounds (for example, as indicated by the value of the spectral tilt) so that when the pitch gain is below the threshold (or not greater than the threshold) Filter the voiced sound, and in other cases make the signal pass without change.
Additional implementations of the anti-sparse filter ASF10 include two or more filters that are configured to have different maximum phase modification angles (e.g., at most 180 degrees). In this case, the anti-sparse filter ASF10 can be configured to select among these constituent filters according to the value of the pitch gain (for example, quantized adaptive codebook or LTP gain), so that the larger The maximum phase modification angle is used for frames with lower pitch gain values. One implementation of the anti-sparse filter ASF10 may also include different composition filters configured to modify the phase in a larger or smaller range of the spectrum, so that it is configured to modify the phase in a wider frequency range of the input signal The filter of is used for frames with lower pitch gain values.
As shown in FIG. 18, the high frequency encoder EH110 includes a high frequency excitation generator XGH10 configured to generate a high frequency excitation signal XH10 from a narrowband excitation signal XL10a. As shown in FIG. 20, the high frequency decoder DH110 also includes an example of the high frequency excitation generator XGH10 configured to generate the high frequency excitation signal XH10 from the narrowband excitation signal XL10a. The high frequency excitation generator XGH10 can be implemented in the same way as the SHB excitation generator XGS20 or XGS30 as described herein, where the spectrum expander SX10 is configured to increase the sampling frequency to 16 kHz instead of 32 kHz. The additional description of the high-frequency excitation generator XGH10 can be found in (for example) the October 2010 document 3GPP2 C.S0014-D, v3.0, section 4.3.3.3 (pages 4.21 to 4.22) "Enhanced Variable Rate Codec, Speech Service Options 3,68,70,73 for Wideband Spread Spectrum Digital Systems" (available online at www.3gpp2.org).
In order to accurately reproduce the encoded speech signal, the ratio between the high frequency part and the narrow frequency part of the synthesized SWB signal SOSW10 may be similar to the level of the high frequency part and the narrow frequency part of the original SWB signal SISW10 The ratio between. In addition to the spectrum envelope as represented by the SHB encoding parameter CPS10, the SHB encoder ES100 can also be configured to characterize the SHB signal SIS10 by a predetermined time or gain envelope. As shown in FIG. 19, the SHB encoder ES110 includes an SHB gain factor calculator GCS10, which is configured and configured to be based on the relationship between the SHB signal SIS10 and the synthesized SHB signal SYS10 (such as the two The difference or ratio between the energy of a signal in a frame or a certain part of the frame) is used to calculate one or more gain factors. In other implementations of the SHB encoder ES110, the SHB gain calculator GCS10 can be configured the same but instead is configured to calculate based on the time-varying relationship between the SHB signal SIS10 and the narrowband excitation signal XL10b or the SHB excitation signal XS10 Gain envelope.
The time envelopes of the narrowband excitation signal XL10b and the SHB signal SIS10 are likely to be similar. Therefore, the gain envelope encoding based on the relationship between the SHB signal SIS10 and the narrowband excitation signal XL10b (or a signal derived therefrom, such as the SHB excitation signal XS10 or the synthesized SHB signal SYS10) will generally be better than encoding based only on the SHB signal SIS10 The gain envelope is more effective. In a typical implementation, the quantizer QGS10 of the SHB encoder ES110 is configured to output a quantization index (for example, with 8, 10, 12, 14, 16, 18, or 20 bits) and a normalization factor as The SHB gain factor CPS10b of each frame. The quantization index specifies ten sub-frame gain factors (for example, for each of the ten sub-frames as shown in FIG. 23B).
The SHB gain factor calculator GCS10 can be configured to perform the gain factor calculation by calculating the gain value of the corresponding sub-frame based on the relative energy of the SHB signal SHB10 and the synthesized SHB signal SYS10. The calculator GCS10 can be configured to calculate the energy of the corresponding sub-frames of the respective signals (for example, the energy is calculated as the sum of the squares of the samples of the respective sub-frames). The calculator GCS10 can then be configured to calculate the gain factor of the sub-frame as the square root of the ratio of their energy (for example, the gain factor is calculated as the energy of the SHB signal SIS10 on the sub-frame and the synthesized SHB The square root of the ratio of the energy of the signal SYS10).
It may be necessary for the SHB gain factor calculator GCS10 to be configured to calculate the sub-frame energy according to the windowing function. For example, the calculator GCS10 can be configured to apply the same windowing function to the SHB signal SIS10 and the synthesized SHB signal SYS10, calculate the energy of each window, and calculate the gain factor of the sub-frame as these The square root of the ratio of energy. Once the sub-frame gain factors of the frame have been calculated, the calculator GCS10 may be required to calculate a normalization factor for the frame and normalize the sub-frame gain factors according to the normalization factor.
It may be necessary to apply a windowing function that overlaps with adjacent sub-frames. For example, a windowing function that generates gain factors that can be applied in an overlap-addition manner can help reduce or avoid discontinuities between sub-frames. In one example, the SHB gain factor calculator GCS10 is configured to apply the trapezoidal windowing function as shown in Figure 23C, where the window overlaps each of the two adjacent subframes for one millisecond. Other implementations of the SHB gain factor calculator GCS10 can be configured to apply windowing functions with different overlap periods and/or different window shapes (eg, rectangular, Hamming), and the window shapes can be symmetrical or asymmetrical. The implementation of the SHB gain factor calculator GCS10 may also be configured to apply different windowing functions to different sub-frames in a frame, and/or a frame may also include sub-frames with different lengths.
The SHB encoder can be configured to determine side information about the gain factor by comparing the synthesized SHB signal with the original SHB signal. The decoder then uses these gains to scale the synthesized SHB signal appropriately.
Although the higher-order SHB LPC coefficients can be expected to model the fine structure of the spectrum with sufficient detail, it may also be necessary to use a relatively high time-domain resolution to reproduce a good SWB signal. In an implementation as described above, for each 20 millisecond frame of the input speech signal, ten time gain parameters are calculated, each of which represents a scale factor for the corresponding two millisecond sub-frame (for example, as Shown in Figure 23B). The gain parameter can be calculated by comparing the energy in each sub-frame of the input SHB signal with the energy in the corresponding sub-frame of the unscaled synthesized SHB excitation signal. A rectangular window in time that selects only the samples of a specific subframe, or a windowing function extended to the previous and/or next subframe (for example, as shown in Figure 23C) can be used to execute each subframe. Calculation of frame gain. It may also be necessary to calculate the frame gain of each frame to adjust the total speech energy level. In order to improve the subsequent quantization process, each sub-frame gain vector can be normalized by the corresponding frame gain value. The frame gain value can also be adjusted to compensate for the normalization of the sub-frame gain.
It may be necessary to configure the SHB gain factor calculator GCS10 to perform attenuation of the gain factor in response to a large change in the gain factor over time, which may indicate that the synthesized signal is greatly different from the original signal. Alternatively or in addition, it may be necessary to configure the SHB gain factor calculator GCS10 to perform time smoothing of the gain factor (for example, to reduce changes that may cause audio artifacts).
Similarly, the time envelopes of the narrow-band excitation signal XL10a and the high-frequency signal SIH10 are likely to be similar. As shown in FIG. 18, the high frequency encoder EH100 can be implemented to include a high frequency gain factor calculator GCH10, which is configured and configured to be based on the high frequency signal SIH10 and the narrowband excitation signal XL10a ( Or calculate one or more gain factors based on the relationship between its signals, such as the synthesized high-frequency signal SYH10 or the high-frequency excitation signal XH10. The calculator GCH10 can be implemented in the same way as the calculator GCS10, except that the calculator GCH10 may be required to calculate the gain factor for fewer subframes per frame compared to the calculator GCS10. In a typical implementation, the quantizer QGH10 of the high frequency encoder EH110 is configured to output a quantization index (for example, with eight to twelve bits) and a normalization factor as the high frequency of each frame The gain factor CPH10b, the quantization index specifies five sub-frame gain factors (for example, for each of the five sub-frames as shown in FIG. 23A).
Figure 20 shows a block diagram of the implementation of DH110 of the high frequency decoder DH100. The high frequency decoder DH110 includes an example of the high frequency excitation generator XGH10 as described herein, which is configured to generate the high frequency excitation signal XH10 based on the narrow frequency excitation signal XL10a. The decoder DH110 includes an inverse quantizer IQH20, the inverse quantizer IQH20 is configured to inversely quantize the high-frequency filter parameters CPH10a (in this example, inverse quantization into a set of LSF), and the LSF to LP filter coefficient conversion IXH20 is configured To transform the LSF into a set of filter coefficients (for example, as described above with reference to the inverse quantizer IQXN10 and the transform IXN20 of the narrowband decoder DN110). As mentioned above, in other implementations, different sets of coefficients (e.g., cepstral coefficients) and/or coefficient representations (e.g., ISP) may be used. The high frequency synthesis module FSH20 is configured to generate a synthesized high frequency signal according to the high frequency excitation signal XH10 and the set of filter coefficients. For systems where the high-frequency encoder includes a synthesis filter (for example, as in the example of the encoder EH110 described above), it may be necessary to implement the high-frequency synthesis module FSH20 to have the same response as the synthesis filter ( For example, the same transfer function).
The high-frequency decoder DH110 also includes: an inverse quantizer IQGH10, which is configured to inversely quantize the high-frequency gain factor CPH10b; and a gain control element GH10 (for example, a multiplier or amplifier), which is configured and configured to reverse the The quantized gain factor is applied to the synthesized high-frequency signal to generate the high-frequency signal SDH10. For the situation where the gain envelope of a frame is specified by more than one gain factor, the gain control element GH10 may include a gain calculator (for example, the high-frequency gain calculator GCH10 ) The logic of applying the gain factor to each sub-frame with the same or different windowing function applied. Similarly, the gain control element GH10 may include logic configured to apply the normalization factor to the gain factor before applying the gain factor to the signal. In other implementations of the high frequency decoder DH110, the gain control element GH10 is similarly configured but instead configured to apply the inverse quantized gain factor to the narrow frequency excitation signal XL10a or to the high frequency excitation signal XH10.
As mentioned above, it may be necessary to obtain the same state in the high-frequency encoder and the high-frequency decoder (for example, by using the inverse-quantized value during encoding). Therefore, it may be necessary to ensure the same state of the corresponding noise generators in the high-frequency excitation generators of the encoder and the decoder in the encoding system according to this implementation. For example, the high-frequency excitation generator of this implementation can be configured so that the state of the noise generator is the information that has been encoded in the same frame (for example, narrowband filter parameter FPN10 or part thereof, and/or The deterministic function of the coded narrowband excitation signal XL10 or its part).
Figure 21 shows the block diagram of the DS110 implementation of the SHB decoder DS100. The SHB decoder DS110 includes an example of the SHB excitation generator XGS10 as described herein, which is configured to generate the SHB excitation signal XS10 based on the narrowband excitation signal XL10b. The decoder DS110 includes an inverse quantizer IQS20, the inverse quantizer IQS20 is configured to inversely quantize the SHB filter parameter CPS10a (in this example, inversely quantized into a set of LSF), and the LSF to LP filter coefficient transform IXS20 is configured to The LSF is transformed into a set of filter coefficients (for example, as described above with reference to the inverse quantizer IQXN10 and the transform IXN20 of the narrowband decoder DN110). As mentioned above, in other implementations, different sets of coefficients (e.g., cepstral coefficients) and/or coefficient representations (e.g., ISP) may be used. The SHB synthesis module FSS20 is configured to generate a synthesized SHB signal according to the SHB excitation signal XS10 and the set of filter coefficients. For systems where the SHB encoder includes a synthesis filter (for example, as in the example of the encoder ES110 described above), it may be necessary to implement the SHB synthesis module FSS20 to have the same response as the synthesis filter (for example, The same transfer function).
The SHB decoder DS110 also includes: an inverse quantizer IQGS10, which is configured to inversely quantize the SHB gain factor CPS10b; and a gain control element GS10 (for example, a multiplier or amplifier), which is configured and configured to inversely quantize the The gain factor is applied to the synthesized SHB signal to generate the SHB signal SDS10. For the situation where the gain envelope of a frame is specified by more than one gain factor, the gain control element GS10 may include a gain control element GS10 configured to be able to be matched with the gain calculator (for example, SHB gain calculator GCS10) of the corresponding SHB encoder. Apply the same or different windowing function to the logic of applying the gain factor to each sub-frame. Similarly, the gain control element GS10 may include logic configured to apply the normalization factor to the gain factor before applying the gain factor to the signal. In other implementations of the SHB decoder DS110, the gain control element GS10 is similarly configured but instead configured to apply the dequantized gain factor to the narrowband excitation signal XL10b or to the SHB excitation signal XS10.
As mentioned above, it may be necessary to obtain the same state in the SHB encoder and the SHB decoder (for example, by using the inverse quantized value during encoding). Therefore, it may be necessary to ensure the same state of the corresponding noise generators in the SHB excitation generators of the encoder and the decoder in the encoding system according to this implementation. For example, the SHB excitation generator of this implementation can be configured so that the state of the noise generator is the information that has been encoded in the same frame (for example, the narrowband filter parameter FPN10 or its part and/or encoded The narrow-band excitation signal XL10 or its part) deterministic function. One or more of the quantizers of the elements described herein (eg, quantizers QLN10, QLH10, QLS10, QGH10, or QGS10) may be configured to perform classification vector quantization. For example, this quantizer can be configured to select one of a set of codebooks based on information that has been encoded in the same frame in the narrow frequency channel and/or the high frequency channel. This technique usually increases coding efficiency at the cost of additional codebook storage.
The encoded narrowband excitation signal XL10 can describe a signal that is distorted in time (for example, by relaxing CELP or other pitch regularization techniques). For example, it may be necessary to time warp the narrow-band signal SIL10 or the signal based on the narrow-band residual according to the model of the pitch structure of the low-frequency sub-band. In this case, it may be necessary to configure the high frequency encoder EH100 to be based on the time distortion described in the encoded narrowband excitation signal (for example, as applied to narrowband signals or to residuals) and also based on low frequency subbands The difference between the sampling rate of the high-frequency signal SIH10 and the high-frequency signal SIH10 offsets the high-frequency signal SIH10 before the calculation of the gain factor. Similarly, it may be necessary to configure the SHB encoder ES100 to be based on the time distortion described in the encoded narrowband excitation signal (for example, as applied to the narrowband signal or to the residual) and also based on the low frequency subband and the SHB signal The difference in the sampling rate of SIS10 shifts the SHB signal SIS10 before the gain factor calculation. This time warping may include a different time offset for each of at least two consecutive subframes of the time-warped signal, and/or may include rounding the calculated time offset to an integer sample value. The time warping of the signal SIH10 or SIS10 can be performed upstream or downstream of the corresponding LPC analysis of the signal.
The encoded signal will likely be carried on a packet-switched network. For circuit-switched operation, the codec may be required to implement discontinuous transmission (DTX) during quiet periods to reduce bandwidth.
The method according to the first general configuration includes calculating a first excitation signal (for example, narrowband excitation signal XL10) based on information from the first frequency band of the speech signal. The method also includes calculating a second excitation signal (for example, SHB excitation signal XS10) for the second frequency band of the speech signal based on the information from the first excitation signal. In this method, the first frequency band is separated from the second frequency band by a distance that is at least half the width of the first frequency band. In an example, the excitation signal includes a component having a frequency of at least 3000 Hz, and the second excitation signal includes a component having a frequency not greater than 8 kHz. In another example, the first frequency band is separated from the second frequency band by at least 2500 Hz. In the implementation as described herein, the first frequency band extends from 50 Hz to 3500 Hz, and the second frequency band extends from 7 kHz to 14 kHz.
The method according to the second general configuration includes calculating a first excitation signal (eg, narrowband excitation signal XL10) based on information from the first frequency band of the speech signal. The method also includes calculating a second excitation signal (for example, SHB excitation signal XS10) for the second frequency band of the speech signal based on the information from the first excitation signal. In this method, the second excitation signal includes energy at each of the first frequency component and the second frequency component, and these equal components are separated by a distance that is at least the sampling rate of the first excitation signal Fifty percent. In another example, the second excitation signal includes energy in the range of 8000 Hz to 8500 Hz and 13,000 Hz to 13,500 Hz. In the implementation as described herein, the sampling rate of the first excitation signal is 8 kHz, and the second excitation signal includes energy at components in the range of 7 kHz (for example, from 7 kHz to 14 kHz).
The method according to the third general configuration includes calculating the first excitation signal (eg, narrowband excitation signal XL10) based on information from the first frequency band of the speech signal. The method also includes: calculating a second excitation signal (for example, a high-frequency excitation signal) for the second frequency band of the speech signal based on the information from the first excitation signal, and calculating the second excitation signal for the speech signal based on the information from the first excitation signal The third excitation signal of the third frequency band (for example, SHB excitation signal XS10). In this method, the second frequency band is different from the first frequency band (but can overlap), the third frequency band is different from the second frequency band (but can overlap), and the third frequency band is separated from the first frequency band. In one example, calculating the second excitation signal includes expanding the spectrum of the first excitation signal into the second frequency band, and calculating the third excitation signal includes expanding the spectrum of the first excitation signal into the third frequency band. In another example, the second frequency band includes frequencies between 5 kHz and 6 kHz, and the third frequency band includes frequencies between 10 kHz and 11 kHz. In the implementation as described herein, the second excitation signal extends from 3500 Hz to 7 kHz, and the third excitation signal extends from 7 kHz to 14 kHz.
The method according to the fourth general configuration includes calculating a first excitation signal (eg, narrowband excitation signal XL10) based on information from the first frequency band of the speech signal. The method also includes: calculating a second excitation signal (for example, a high-frequency excitation signal) for the second frequency band of the speech signal based on the information from the first excitation signal, and calculating the second excitation signal for the speech signal based on the information from the first excitation signal The third excitation signal of the third frequency band (for example, SHB excitation signal XS10). In this method, the second frequency band is different from the first frequency band (but can overlap), the third frequency band is different from the second frequency band (but can overlap), and the third frequency band is separated from the first frequency band.
The method includes calculating a first complex number m gain factors that describe (A) a frame of a signal based on information from a first frequency band and (B) a frame based on information from a second excitation signal One of the signals corresponds to a relationship between the frames. The method also includes calculating a second plural n gain factors, the second plural n gain factors describing (A) the frame of the signal based on the information from the first frequency band and (B) the frame based on the information from the third excitation signal One of the signals corresponds to a relationship between the frames, where n is greater than m.
In an example, each of the first plurality of m gain factors corresponds to one of the m subframes, and each of the second plurality of n gain factors corresponds to one of the n subframes By. In another example, calculating the first complex m gain factors includes normalizing the first complex m gain factors according to the first gain frame value, and calculating the second complex n gain factors includes normalizing the second gain frame value. Calculate the second complex number n gain factors. In the implementation as described herein, m is equal to five and n is equal to ten.
24A shows a flowchart of a method M100 for processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separate from the low-frequency sub-band according to a general configuration. The method M100 includes the task T100 of filtering the audio signal to obtain a narrowband signal and an ultra-high frequency signal (for example, as described with reference to the filter bank FB100 in this article), and calculating a process based on the information from the narrowband signal The task T200 of the encoded narrowband excitation signal (for example, as described herein with reference to the narrowband encoder EN100), and the task T300 of calculating an UHF excitation signal based on the information from the encoded narrowband excitation signal (for example, , As described herein with reference to the SHB encoder ES100). The method M100 also includes the task T400 of calculating and characterizing a plurality of filter parameters of a spectral envelope of the high frequency subband based on the information from the UHF signal (for example, as described herein with reference to the SHB gain factor calculator GCS100) . In this method, the narrow-frequency signal is based on the frequency component in the low-frequency sub-band, and the UHF signal is based on the frequency component in the high-frequency sub-band. In this method, a width of the low-frequency sub-band is at least two kilohertz, and the low-frequency sub-band is separated from the high-frequency sub-band by a distance that is at least equal to half of the width of the low-frequency sub-band . The method M100 may also include the task of calculating a plurality of gain factors by evaluating a time-varying relationship between a signal based on the UHF signal and a signal based on the UHF excitation signal.
FIG. 24B shows a block diagram of an apparatus MF100 for processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separate from the low-frequency sub-band according to a general configuration. The device MF100 includes: a component F100 for filtering the audio signal to obtain a narrowband signal and an ultra-high frequency signal (for example, as described herein with reference to the filter bank FB100), for filtering the audio signal from the narrowband signal The information calculates an encoded narrowband excitation signal component F200 (for example, as described herein with reference to the narrowband encoder EN100), and is used to calculate an UHF excitation based on the information from the encoded narrowband excitation signal The signal component F300 (for example, as described herein with reference to the SHB encoder ES100). The device MF100 also includes a component F400 for calculating and characterizing a plurality of filter parameters of a spectral envelope of the high-frequency sub-band based on the information from the UHF signal (for example, as described herein with reference to the SHB gain factor calculator GCS100 describe). In this device, the narrow-frequency signal is based on the frequency component in the low-frequency sub-band, and the UHF signal is based on the frequency component in the high-frequency sub-band. In this device, a width of the low-frequency sub-band is at least two kilohertz, and the low-frequency sub-band is separated from the high-frequency sub-band by a distance that is at least equal to half of the width of the low-frequency sub-band . The device MF100 may also include means for calculating a plurality of gain factors by evaluating a time-varying relationship between a signal based on the UHF signal and a signal based on the UHF excitation signal.
The methods and devices disclosed herein can generally be applied to any transceiving and/or audio sensing applications, especially mobile applications or other portable examples of these applications. For example, the scope of the configuration disclosed herein includes communication devices residing in a wireless telephone communication system configured to use a code division multiple access (CDMA) air interface. However, those skilled in the art will understand that the methods and devices having the characteristics described herein can reside in any of various communication systems using a wide range of technologies known to those skilled in the art, such as A system using Voice over Internet Protocol (VoIP) via a wired and/or wireless (for example, CDMA, TDMA, FDMA, and/or TD-SCDMA) transmission channel.
Explicitly covered and hereby disclosed, the communication devices disclosed herein can be adapted for use in packet-switched networks (for example, wired and/or wireless networks configured to carry audio transmission according to protocols such as VoIP) and/ Or in a circuit-switched network. It is also expressly covered and hereby disclosed that the communication device disclosed herein can be adapted for use in a narrow-band coding system (for example, a system that encodes an audio frequency range of approximately four kilohertz or five kilohertz) and/or for broadband In coding systems (for example, systems that encode audio frequencies greater than five kilohertz), these systems include full-band bandwidth coding systems and frequency-divided bandwidth coding systems.
The presentation of the configuration described in this article is provided to enable anyone familiar with the art to manufacture or use the methods and other structures disclosed in this article. The flowcharts, block diagrams, and other structures shown and described herein are only examples, and other variations of these structures are also within the scope of the present invention. Various modifications to these configurations are possible, and the general principles presented in this article can also be applied to other configurations. Therefore, the present invention is not intended to be limited to the configuration shown above, but conforms to the scope of this document (including the scope of the attached application, which forms a part of the original disclosure). The broadest category that is consistent with the principles and novel features disclosed in any way.
Those familiar with this technology will understand that any of a variety of different technologies and techniques can be used to represent information and signals. For example, the data, instructions, commands, information, signals, bits and symbols that may be mentioned in the above description can be made by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles or any combination thereof. Express.
Important design requirements for the implementation of the configuration as disclosed herein may include minimizing processing delay and/or computational complexity (usually measured in millions of instructions per second or MIPS), Especially for computationally intensive applications, such as compressed audio or audiovisual information (e.g., files or streams encoded according to a compressed format, such as one of the examples identified in this article) playback or for broadband communications (e.g., The application of voice communications at a sampling rate higher than eight kilohertz (such as 12 kHz, 16 kHz, 44.1 kHz, 48 kHz, or 192 kHz). The goals of a multi-microphone processing system as described in this article may include: achieving a total noise reduction of 10 dB to 12 dB, maintaining the voice level and tone during the movement of the desired speaker, and obtaining the noise that has been moved to the background Instead of active noise removal, speech dereverberation, and/or enable post-processing options (for example, spectrum masking and/or another spectrum modification operation based on noise estimation, such as spectrum subtraction or text Wiener filtering to obtain a more aggressive noise reduction.
The various processing elements (for example, the components of the encoder SWE100 and the decoder SWD100 and the encoder SWE100 and the decoder SWD100) of the implementation of the device as disclosed herein can be embodied in hardware, software, and/or firmware deemed suitable for the intended application. In any combination of the body. For example, these components can be manufactured as electronic and/or optical devices that reside on, for example, the same chip or two or more chips in a chip set. An example of such a device is a fixed or programmable array of logic elements (such as transistors or logic gates), or any of these elements can be implemented as one or more of these arrays. Any two or more or even all of these elements can be implemented in the same array or arrays. This array or these arrays may be implemented in one or more chips (e.g., in a chip set including two or more chips).
One or more elements of various implementations of the devices disclosed herein (for example, elements of encoder SWE100 and decoder SWD100, and encoder SWE100 and decoder SWD100) may also be implemented in whole or in part as one or more instructions The one or more instruction sets are configured to be executed on one or more fixed or programmable logic element arrays, such as microprocessors, embedded processors, IP cores, digital signal processors, FPGAs Programmable gate array), ASSP (standard product for special application) and ASIC (integrated circuit for special application). Any of the various elements implemented by one of the devices disclosed herein may also be embodied as one or more computers (e.g., including one or more sets of instructions or sequences of instructions programmed to execute Array machines are also called "processors"), and any two or more or even all of these components can be implemented in the same computer or computers.
The processor or other components used for processing as disclosed herein can be manufactured as one or more electronics and/or residing on, for example, the same chip or two or more chips in a chipset. optical instrument. An example of such a device is a fixed or programmable array of logic elements (such as transistors or logic gates), or any of these elements can be implemented as one or more of these arrays. This array or these arrays may be implemented in one or more chips (e.g., in a chip set including two or more chips). Examples of such arrays include fixed or programmable logic element arrays, such as microprocessors, embedded processors, IP cores, DSPs, FPGAs, ASSPs, and ASICs. The processor or other components used for processing as disclosed herein may also be embodied as one or more computers (for example, including a machine programmed to execute one or more arrays of one or more instruction sets or instruction sequences ) Or other processors. It is possible to use the processor as described herein to perform tasks that are not directly related to the implementation of the method M100 (or another method as disclosed with reference to the operation of the device or device described herein), or to perform tasks that are not directly related to the method. Other instruction sets directly related to the program implemented by the M100, such as tasks related to another operation of the device or system (for example, a voice communication device) embedded with the processor. It is also possible that the processor of the audio sensor device executes a part of the method as disclosed herein and executes another part of the method under the control of one or more other processors.
Those familiar with this technology will understand that various descriptive modules, logic blocks, circuits, tests, and other operations described in conjunction with the configuration disclosed in this article can be implemented as electronic hardware, computer software, or a combination of both. General-purpose processors, digital signal processors (DSP), ASICs or ASSPs, FPGAs or other programmable logic devices, discrete gates or transistor logic, discrete hardware components or their designs can be used to generate groups as disclosed herein Any combination of states to implement or execute these modules, logic blocks, circuits, and operations. For example, this configuration can be implemented at least partially as a hard-wired circuit, as a circuit configuration manufactured in a special application integrated circuit, or as a firmware program loaded into a non-volatile memory or As a software program loaded from a data storage medium or loaded into a data storage medium as a machine-readable program code, the program code is an instruction that can be executed by a logic element array (such as a general-purpose processor or other digital signal processing unit). The general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with DSP cores, or any other such configuration. Software modules can reside in RAM (random access memory), ROM (read-only memory), non-volatile RAM (NVRAM) (such as flash RAM), erasable programmable ROM (EPROM), electronic Erasable programmable ROM (EEPROM), temporary memory, hard disk, removable disk or CD-ROM non-transitory storage medium; or reside in any other form of storage medium known in the art . The illustrative storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. In the alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in the ASIC. The ASIC can reside in the user terminal. In the alternative, the processor and the storage medium may reside as discrete components in the user terminal.
It should be noted that various methods disclosed herein (for example, method M100 and other methods disclosed with reference to the operation of various devices described herein) can be executed by an array of logic elements such as a processor, and can be implemented as described herein. The various elements of the described device are partially implemented as modules designed to execute on this array. As used herein, the term "module" or "sub-module" can refer to any method, device, device, unit, or computer instruction (for example, logical expression) in the form of software, hardware, or firmware. Computer-readable data storage media. It should be understood that multiple modules or systems can be combined into one module or system, and one module or system can be divided into multiple modules or systems to perform the same function. When implemented by software or other computer executable instructions, the elements of the processing program are basically code fragments used to perform related tasks, such as routines, programs, objects, components, data structures, and the like. The term "software" should be understood to include source code, assembly language code, machine code, binary code, firmware, macro code, microcode, any one or more instruction sets or instruction sequences that can be executed by an array of logic elements, and Any combination of these instances. The program or code fragments can be stored in a processor-readable storage medium or transmitted by a computer data signal embodied in a carrier wave through a transmission medium or a communication link.
The implementation of the methods, solutions, and techniques disclosed herein may also be tangibly embodied (for example, in the tangible computer-readable features of one or more computer-readable storage media as listed herein) as may include logic elements The machine of the array (e.g., processor, microprocessor, microcontroller, or other finite state machine) executes one or more sets of instructions. The term "computer-readable medium" can include any medium that can store or transmit information, including volatile, non-volatile, removable, and non-removable storage media. Examples of computer-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks or other magnetic storage, CD-ROM/DVD or other optical storage, hard disk Disk, or any other media that can be used to store the desired information, optical fiber media, radio frequency (RF) links, or any other media that can be used to carry the desired information and can be accessed. Computer data signals can include any signals that can be propagated via transmission media such as electronic network channels, optical fibers, air, electromagnetic, and RF links. The code snippets can be downloaded via a computer network such as the Internet or corporate intranet. In any case, the scope of the present invention should not be construed as being limited by these embodiments.
Each of the tasks of the methods described herein can be directly embodied in hardware, in a software module executed by a processor, or in a combination of the two. In a typical application implemented as one of the methods disclosed herein, an array of logic elements (eg, logic gates) is configured to perform one, more than one, or even all of the various tasks of the method. One or more of these tasks (possibly all) can also be implemented as embodied in computer program products (for example, one or more data storage media, such as disks, flash memory cards or other non-volatile memory cards, Program code (for example, one or more instruction sets) in a semiconductor memory chip, etc., which can be used by a machine that includes an array of logic elements (for example, a processor, a microprocessor, a microcontroller, or other finite state machine) (For example, computer) to read and/or execute. The tasks performed by one of the methods disclosed herein can also be performed by more than one such array or machine. In these or other implementations, these tasks can be performed in a device for wireless communication (such as a cellular phone) or other device with such communication capabilities. This device can be configured to communicate with a circuit-switched network and/or a packet-switched network (for example, using one or more protocols such as VoIP). For example, the device may include an RF circuit configured to receive and/or transmit the encoded frame. It is clearly disclosed that the various methods disclosed herein can be executed by a portable communication device such as a mobile phone, a headset, or a portable digital assistant (PDA), and the various devices described herein can be included in this device. A typical real-time (for example, online) application is a phone call using this mobile device.
In one or more exemplary embodiments, the operations described herein can be implemented by hardware, software, firmware, or any combination thereof. If implemented in software, these operations can be stored as one or more instructions or program codes on a computer-readable medium or transmitted via the computer-readable medium. The term "computer-readable medium" includes both computer-readable storage media and communication (eg, transmission) media. As an example and not limitation, a computer-readable storage medium may include an array of storage elements, such as semiconductor memory (which may include (but is not limited to) dynamic or static RAM, ROM, EEPROM, and/or flash RAM), or ferroelectric , Magnetoresistive, bidirectional, polymer or phase change memory; CD-ROM or other optical disk storage; and/or magnetic disk storage or other magnetic storage devices. Such storage media can store information in the form of commands or data structures that can be accessed by a computer. The communication medium may include any medium that can be used to carry the desired program code in the form of a command or a data structure that can be accessed by a computer, including any medium that facilitates the transfer of a computer program from one place to another. Also, any connection is appropriately referred to as a computer-readable medium. For example, if you use coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and/or microwave) to transmit from a website, server or other remote source For software, coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology (such as infrared, radio and/or microwave) are included in the definition of media. As used in this article, disks and optical discs include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVD), flexible discs and Blu-ray Discs<sup>TM</sup>(Blu-Ray Disc Association, Universal City, CA), where disks usually reproduce data magnetically, and optical disks reproduce data optically through lasers. Combinations of the above should also be included in the category of computer-readable media.
The acoustic signal processing apparatus as described herein can be incorporated into an electronic device (such as a communication device) that accepts voice input in order to control certain operations or can otherwise benefit from the separation of the desired noise from the background noise. Many applications can benefit from enhancing a clear desired sound or separating a clear desired sound from background sounds originating from multiple directions. These applications may include man-machine interfaces in electronic or computing devices that incorporate capabilities such as voice recognition and detection, voice enhancement and separation, voice-activated control, and the like. It may be necessary to implement this acoustic signal processing device to be suitable for devices that only provide limited processing capabilities.
The components of the various implementations of the modules, components, and devices described herein can be manufactured as electronic and/or optical devices that reside on, for example, the same chip or among two or more chips in a chipset. An example of such a device is a fixed or programmable array of logic elements such as transistors or gates. One or more elements of the various implementations of the devices described herein may also be implemented in whole or in part as one or more instruction sets configured to execute on one or more fixed or programmable On an array of logical logic elements, such as microprocessors, embedded processors, IP cores, digital signal processors, FPGAs, ASSPs and ASICs.
It is possible to use one or more elements implemented by one of the devices described herein to perform tasks that are not directly related to the operation of the device, or to perform other instruction sets that are not directly related to the operation of the device, such as those with embedded A task related to another operation of the device or system of the device. One or more components implemented by one of the devices may also have a common structure (for example, processors corresponding to parts of different components that are used to execute code at different times, and are executed to execute at different times corresponding to The instruction set of tasks of different components, or the configuration of electronic and/or optical devices that perform operations of different components at different times).
<p>ADD10. . . Adder</p><p>ADD20. . . Adder</p><p>ASF10. . . Anti-sparse filter</p><p>CPH10. . . High frequency coding parameters</p><p>CPH10a. . . High frequency filter parameters</p><p>CPH10b. . . High frequency gain factor</p><p>CPS10. . . UHF encoding parameters</p><p>CPS10a. . . UHF filter parameters/linear prediction coding filter parameters</p><p>CPS10b. . . UHF gain factor</p><p>DH10. . . Integer downsampler</p><p>DH100. . . High frequency decoder</p><p>DH110. . . High frequency decoder</p><p>DH20. . . Integer downsampler implements 2 integer downsampling</p><p>DH30. . . Integer downsampling block</p><p>DMX100. . . Demultiplexer</p><p>DN10. . . Integer downsampler</p><p>DN110. . . Narrowband decoder</p><p>DN20. . . Integer downsampler</p><p>DS10. . . Integer downsampler</p><p>DS100. . . UHF decoder</p><p>DS110. . . UHF decoder</p><p>DS12. . . Integer downsampler</p><p>DS20. . . Integer downsampler implements 2 integer downsampling</p><p>DS30. . . Integer downsampling block</p><p>DSS10. . . Integer downsampling block</p><p>DW10. . . Integer downsampler</p><p>DW20. . . Integer downsampler</p><p>EH100. . . High frequency encoder</p><p>EH110. . . High frequency encoder</p><p>EN100. . . Narrowband encoder</p><p>EN110. . . Narrowband encoder</p><p>ES100. . . UHF encoder</p><p>ES110. . . UHF encoder</p><p>F100. . . A component for filtering the audio signal to obtain a narrowband signal and a UHF signal</p><p>F200. . . Means for calculating an encoded narrowband excitation signal based on information from the narrowband signal</p><p>F300. . . Means for calculating a UHF excitation signal based on information from the encoded narrowband excitation signal</p><p>F400. . . A means for calculating and characterizing a plurality of filter parameters of a spectral envelope of the high frequency sub-band based on the information from the UHF signal</p><p>FAH10. . . Spectrum shaping block</p><p>FB100. . . Filter bank</p><p>FB110. . . Filter bank</p><p>FB112. . . Filter bank</p><p>FB120. . . Filter bank</p><p>FB200. . . Filter bank</p><p>FB210. . . Filter bank</p><p>FB220. . . Filter bank</p><p>FBS10. . . UHF analysis filter bank</p><p>FNS10. . . Narrowband synthesis filter</p><p>FPN10. . . Narrowband filter parameters</p><p>FPN40. . . Coded signal</p><p>FSH10. . . Synthesis filter</p><p>FSH20. . . High frequency synthesis module</p><p>FSL10. . . Shaping filter</p><p>FSS10. . . Synthesis filter</p><p>FSS20. . . UHF synthesis module</p><p>FSW10. . . Spectrum Shaping Filter</p><p>GCH10. . . High frequency gain factor calculator</p><p>GCS10. . . UHF Gain Factor Calculator</p><p>GH10. . . Gain control element</p><p>GS10. . . Gain control element</p><p>IAH10. . . Interpolated block</p><p>IAS10. . . Interpolated block</p><p>IH20. . . Interpolator</p><p>IH30. . . Interpolator</p><p>IN20. . . Interpolator</p><p>IQGH10. . . Inverse quantizer</p><p>IQGS10. . . Inverse quantizer</p><p>IQH20. . . Inverse quantizer</p><p>IQLN10. . . Inverse quantizer</p><p>IQN10. . . Inverse quantizer</p><p>IQS20. . . Inverse quantizer</p><p>IQXN10. . . Inverse quantizer</p><p>IS12. . . Interpolator</p><p>IS20. . . Interpolator</p><p>IS30. . . Interpolator</p><p>IW20. . . Interpolator</p><p>IXH10. . . Transform</p><p>IXN10. . . Linear spectrum frequency to linear prediction filter coefficient transformation</p><p>IXN20. . . Linear spectrum frequency to linear prediction filter coefficient transformation</p><p>IXS10. . . Transform</p><p>IXS20. . . Linear spectrum frequency to linear prediction filter coefficient transformation</p><p>LPH10. . . Analysis module</p><p>LPN10. . . Linear predictive coding analysis module</p><p>LPS10. . . Analysis module</p><p>M100. . . Method for processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separate from the low-frequency sub-band</p><p>MF100. . . Device for processing an audio signal having frequency components in a low-frequency sub-band and a high-frequency sub-band separate from the low-frequency sub-band</p><p>MPX100. . . Multiplexer</p><p>PAH10. . . High-frequency analysis processing path</p><p>PAH12. . . High-frequency analysis processing path</p><p>PAH20. . . High-frequency analysis processing path</p><p>PAN10. . . Narrowband analysis processing path</p><p>PAS10. . . UHF analysis processing path</p><p>PAS12. . . UHF analysis processing path</p><p>PAS20. . . UHF analysis processing path</p><p>PAW10. . . Broadband analysis processing path</p><p>PSH10. . . High frequency synthesis processing path</p><p>PSH20. . . High frequency synthesis processing path</p><p>PSN10. . . Narrowband synthesis processing path</p><p>PSN20. . . Narrowband synthesis processing path</p><p>PSS10. . . UHF synthesis processing path</p><p>PSS20. . . UHF synthesis processing path</p><p>PSW10. . . Broadband synthesis processing path</p><p>PSW20. . . Broadband synthesis processing path</p><p>QGH10. . . Quantizer</p><p>QGS10. . . Quantizer</p><p>QLH10. . . Quantizer</p><p>QLN10. . . Quantizer</p><p>QLN20. . . Quantizer</p><p>QLN30. . . Quantizer</p><p>QLS10. . . Quantizer</p><p>QXN10. . . Quantizer</p><p>RHA10. . . Spectrum Inversion Module</p><p>RSA10. . . Spectrum Inversion Module</p><p>SDH10. . . High frequency signal</p><p>SDL10. . . Narrowband signal</p><p>SDS10. . . UHF signal</p><p>SDW10. . . Decoded broadband signal</p><p>SIH10. . . High frequency signal</p><p>SIL10. . . Narrowband signal</p><p>SIS10. . . UHF signal</p><p>SISW10. . . Ultra-wideband signal</p><p>SIW10. . . Broadband signal</p><p>SM10. . . Multiplexed signal</p><p>SOH10. . . High frequency output signal/passband signal</p><p>SOL10. . . Narrowband output signal/passband signal</p><p>SOS10. . . UHF output signal/passband signal</p><p>SOSW10. . . Ultra-wideband output signal</p><p>SOW10. . . Wideband output signal/passband signal</p><p>SWD100. . . Ultra-wideband decoder</p><p>SWD110. . . Ultra-wideband decoder</p><p>SWE100. . . Ultra-wideband encoder</p><p>SWE110. . . Ultra-wideband encoder</p><p>SX10. . . Spectrum spreader</p><p>SYH10. . . Synthesized high frequency signal</p><p>SYS10. . . Synthesized UHF signal</p><p>T100. . . The task of filtering the audio signal to obtain a narrowband signal and a UHF signal</p><p>T200. . . The task of calculating an encoded narrowband excitation signal based on the information from the narrowband signal</p><p>T300. . . The task of calculating a UHF excitation signal based on the information from the encoded narrowband excitation signal</p><p>T400. . . The task of calculating and characterizing a plurality of filter parameters of a spectral envelope of the high-frequency sub-band based on the information from the UHF signal</p><p>V40. . . Scale Factor</p><p>WF10. . . Whitening filter</p><p>WF20. . . Whitening filter</p><p>XFH10. . . Linear prediction filter coefficients to line spectrum frequency conversion</p><p>XGH10. . . High frequency excitation generator</p><p>XGS10. . . UHF excitation generator</p><p>XGS20. . . UHF excitation generator</p><p>XGS30. . . UHF excitation generator</p><p>XH10. . . High frequency excitation signal</p><p>XL10. . . Encoded narrowband excitation signal</p><p>XL10a. . . Narrowband excitation signal</p><p>XL10b. . . Narrowband excitation signal</p><p>XLD10. . . Narrowband excitation signal</p><p>XLN10. . . Linear prediction filter coefficients to line spectrum frequency conversion</p><p>XS10. . . UHF excitation signal</p>
Figure 1 shows the block diagram of the ultra-wideband encoder SWE100 according to the general configuration;
Figure 2 shows the block diagram of the implementation of SWE110 of the ultra-wideband encoder SWE100;
Figure 3 is a block diagram of an ultra-wideband decoder SWD100 according to a general configuration;
Figure 4 is a block diagram of the implementation of SWD110 of the ultra-wideband decoder SWD100;
FIG. 5A shows a block diagram of the implementation of FB110 of the filter bank FB100;
FIG. 5B shows a block diagram of the implementation of FB210 of the filter bank FB200;
FIG. 6A shows a block diagram of the implementation of FB112 of the filter bank FB100;
FIG. 6B shows a block diagram of the implementation of FB212 of the filter bank FB210;
7A, 7B, and 7C show the relative bandwidth of the narrowband signal SIL10, the high frequency signal SIH10, and the ultra high frequency signal SIS10 in three different implementation examples;
FIG. 8A shows a block diagram of the DS12 implementation of the integer down sampler DS10;
FIG. 8B shows a block diagram of the implementation of IS12 of the interpolator IS10;
FIG. 8C shows a block diagram of the implementation of FB120 of the filter bank FB112;
9A to 9F show a step-by-step example of the frequency spectrum of the signal processed in the application of the path PAS20;
Figure 10 shows a block diagram of the implementation of FB220 of the filter bank FB212;
11A to 11F show a step-by-step example of the frequency spectrum of the signal processed in the application of the path PSS20;
Figure 12A shows an example of a graph of logarithmic amplitude and frequency of a speech signal;
Figure 12B shows a block diagram of a basic linear predictive coding system;
Figure 13 shows a block diagram of the narrowband encoder EN100 implementing EN110;
Figure 14 shows a block diagram of the quantizer QLN10 implementing QLN20;
Figure 15 shows a block diagram of the quantizer QLN10 implementing QLN30;
Figure 16 shows a block diagram of the narrowband decoder DN100 implementing DN110;
Figure 17A shows an example of a graph of logarithmic amplitude and frequency of the residual signal of voiced speech;
Figure 17B shows an example of a graph of logarithmic amplitude and time of the residual signal of voiced speech;
Figure 17C shows a block diagram of a basic linear predictive coding system that also performs long-term prediction;
Figure 18 shows a block diagram of the high frequency encoder EH100 implementation EH110;
Figure 19 shows a block diagram of the ES110 implementation of the UHF encoder ES100;
Figure 20 shows a block diagram of the implementation of the high frequency decoder DH100 DH110;
Figure 21 shows a block diagram of the DS110 implementation of the UHF decoder DS100;
Figure 22A shows a block diagram of the XGS20 implementation of the UHF excitation generator XGS10;
Figure 22B shows a block diagram of the XGS30 implementation of the UHF excitation generator XGS20;
FIG. 23A shows an example of dividing a frame into five sub-frames;
FIG. 23B shows an example of dividing a frame into ten sub-frames;
Figure 23C shows an example of a windowing function used for sub-frame gain calculation;
Figure 24A shows a flow chart of the method M100 according to the general configuration; and
FIG. 24B shows a block diagram of the device MF100 according to the general configuration.
2 sheets
Sheet 1 Sheet 2
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9818420B2 | Cited by | United States of America | Applicant |
| TWI461705B | Cited by | Taiwan Province of China | Examiner |
| TWI513246B | Cited by | Taiwan Province of China | Examiner |
| US10354666B2 | Cited by | United States of America | Applicant |
| TWI571867B | Cited by | Taiwan Province of China | Examiner |
| US10720172B2 | Cited by | United States of America | Applicant |
| US10229693B2 | Cited by | United States of America | Applicant |
10 priority claims, no other members on record
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 35042510 | United States of America | P | |
| 35042510 | United States of America | P | |
| 61350425 | United States of America | – | |
| 13149874 | United States of America | – | |
| 201113149874 | United States of America | A | |
| 201113149874 | United States of America | A | |
| 20100350425P | – | – | – |
| 201113149874 | – | – | – |
| US20100350425P | – | – | – |
| US201113149874 | – | – | – |
Numbers
- Publication
- 201214419
- Publication, DOCDB
- 201214419
- Publication, EPODOC
- TW201214419
- Application
- 100119283
- Application, DOCDB
- 100119283
- Application, EPODOC
- TW20110119283
Titles5
- Chinese
- 用於寬頻語音編碼之系統、方法、裝置及電腦程式產品
- English
- SYSTEMS, METHODS, APPARATUS, AND COMPUTER PROGRAM PRODUCTS FOR WIDEBAND SPEECH CODING
- English
- System, method, device and computer program product for broadband speech coding
- Unlabeled
- 用於寬頻語音編碼之系統、方法、裝置及電腦程式產品
- Unlabeled
- System, method, device and computer program product for broadband speech coding
Classification
- CPC, 3
- G10L21/038
- G10L21/02
- G10L19/06
- IPC, 1
- G10L21 02