Individual channel shaping for bcc schemes and the like
Abstract
At an audio encoder, cue codes are generated for one or more audio channels, wherein an envelope cue code is generated by characterizing a temporal envelope in an audio channel. At an audio decoder, E transmitted audio channel (s) are decoded to generate C playback audio channels, where C>E1. Received cue codes include an envelope cue code corresponding to a characterized temporal envelope of an audio channel corresponding to the transmitted channel (s). One or more transmitted channel (s) are upmixed to generate one or more upmixed channels. One or more playback channels are synthesized by applying the cue codes to the one or more upmixed channels, wherein the envelope cue code is applied to an upmixed channel or s synthesized signal to adjust a temporal envelope of the synthesized signal based on the characterized temporal envelope such that the adjusted temporal envelope substantially matches the characterized temporal envelope.
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
48 claims: 39 independent, 9 dependent
- 1一種用以編碼聲道之方法,該方法包括:對一個或多個聲道產生一個或多個提示碼,其中至少一個提示碼為在一個或多個聲道之一者中藉特性化一時序包封所產生之一包封提示碼,其中該一個或多個提示碼進而包括一個或多個頻道間相關碼(ICC)、頻道間位準差異(ICLD)碼與頻道間時間差異(ICTD)碼,其中聯合該包封提示碼之第一次解析度,為較細緻於聯合其他提示碼之第二次解析度;及傳送該一個或多個提示碼。
- 2如申請專利範圍第1項之方法,進而包括傳送對應一個或多個聲道之E經傳送聲道,其中E≧1。
- 3如申請專利範圍第2項之方法,其中:該一個或多個聲道包括C輸入聲道,其中C>E;及該C輸入頻道被降混以產生E經傳送頻道。
- 4如申請專利範圍第1項之方法,其中該一個或多個提示碼被傳送,以啟動一解碼器,以在基於一個或多個提示碼之E經傳送聲道之解碼期間實施包封成形,其中E經傳送聲道對應該一個或多個頻道,其中E≧1。
- 5如申請專利範圍第4項之方法,其中該包封成形調整由解碼器產生之合成訊號之時序包封以實質匹配該特性化時序包封。
- 6如申請專利範圍第1項之方法,其中該時序包封僅於該對應聲道之特定頻率被特性化。
- 7如申請專利範圍第6項之方法,其中該時序包封僅於該對應頻道之頻率高於一特定截止頻率時被特性化。
- 8如申請專利範圍第1項之方法,其中將時序包封特性化用於頻域中該對應聲道。
- 9如申請專利範圍第8項之方法,其中將時序包封分別特性化,用於該對應聲道中之不同訊號次頻帶。
- 10如申請專利範圍第8項之方法,其中該頻域對應一快速富立葉轉換(FFT)。
- 11如申請專利範圍第8項之方法,其中該頻域對應一正交鏡相濾波器(QMF)。
- 12如申請專利範圍第1項之方法,其中將該時序包封特性化用於時域中該對應聲道。
- 13如申請專利範圍第1項之方法,其中進而包括決定是否啟動或關閉該特性化。
- 14如申請專利範圍第13項之方法,進而包括產生與傳送一啟動/關閉旗標,其係基於該決定以教示一解碼器是否於對應該一個或多個聲道之E經傳送頻道的解碼期間,實施包封成形,其中E≧1。
- 15如申請專利範圍第13項之方法,其中該決定係基於分析一聲道以偵測聲道中之暫態,使得假如一暫態之出現被偵側到,該特性化被啟動。
- 16一種用於編碼聲道之設備,該設備包括:用於對一個或多個聲道產生一個或多個提示碼之裝置,其中至少一個提示碼為藉一個或多個聲道中,特性化一時序包封所產生之包封提示碼,其中該一個或多個提示碼進而包括一個或多個頻道間相關碼(ICC)、頻道間位準差異(ICLD)碼與頻道間時間差異(ICTD)碼;其中聯合該包封提示碼之第一次解析度,為較細緻於聯合其他提示碼之第二次解析度;及用於傳送該一個或多個提示碼之裝置。
- 17一種用於編碼C輸入聲道之設備以產生E經傳送聲道之設備,該設備包括:一包封分析器適以特性化至少該C輸入頻道中之一者之一輸入時序包封;一碼估計器適以對該C輸入頻道之二者或多者產生提示碼,及;一降混器適以降混該C輸入頻道,以產生該E經傳送頻道,其中C>E≧1,其中該設備被適以傳送關於該提示碼之資訊,及該經特性化輸入時序包封以啟動一解碼器,以於E經傳送頻道之解碼期間執行合成與包封成形。
- 18如申請專利範圍第17項之設備,其中該設備為選自由一數位錄影機、一數位聲音錄音機、一電腦、一衛星發射器、一有線發射器、一陸上廣播發射器、一家用娛樂系統與一電影劇院系統組成之群之一系統;及該系統包括包封分析器、碼估計器與降混器。
- 19一種機器可讀取式媒體,具有編碼於其上的程式碼,其中當該程式碼以一機器執行時,該機器實行一方法用以編碼聲道,該方法包括:對一個或多個聲道產生一個或多個提示碼,其中至少一提示碼為藉於該一個或多個聲道之一者中,特性化一時序包封所產生之一包封提示碼,其中該一個或多個提示碼進而包括一個或多個頻道間相關碼(ICC)、頻道間位準差異(ICLD)碼與頻道間時間差異(ICTD)碼;其中聯合該包封提示碼之第一次解析度,為較細緻於聯合其他提示碼之第二次解析度;及傳送該一個或多個提示碼。
- 20一種儲存有經編碼聲音位元流之儲存裝置,其中該經編碼聲音位元流係由編碼聲道所產生,其中對一個或多個聲道產生一個或多個提示碼,其中至少一提示碼為於該一個或多個聲道之一者中,藉特性化一時序包封所產生之一包封提示碼;及該一個或多個提示碼及對應該一個或多個聲道之E經傳送聲道,其中E≧1,被編進該經編碼的聲音位元流。
- 21一種儲存有經編碼聲音位元流之儲存裝置,其中該經編碼聲音位元流包括一個或多個提示碼與E經傳送聲道,其中該一個或多個聲音碼被產生用於一個或多個聲道,其中至少一提示碼為於該一個或多個聲道之一者中,藉特性化一時序包封所產生之一包封提示碼,其中該一個或多個提示碼進而包括一個或多個頻道間相關碼(ICC)、頻道間位準差異(ICLD)碼與頻道間時間差異(ICTD)碼;其中聯合該包封提示碼之第一次解析度,為較細緻於聯合其他提示碼之第二次解析度;及該E經傳送聲道對應該一個或多個聲道。
- 22一種用於解碼E經傳送聲道以產生C放音聲道之方法,其中C>E≧1,該方法包括:接收對應至該E經傳送聲道之提示碼,其中該提示碼包括一包封提示碼,該包封提示碼對應E經傳送頻道之一聲道之一特性化時序包封;升混一個或多個該E經傳送頻道,以產生一個或多個升混頻道;及藉將該提示碼施加至該一個或多個升混頻道,合成一個或多個C放音頻道,其中該包封提示碼被施加至一升混頻道或一合成訊號,以調整基於該特性化時序包封之該合成訊號之一時序包封,使得該經調整的時序包封實質匹配該特性化時序包封。
- 23如申請專利範圍第22項之方法,其中該包封提示碼對應使用於產生該E經傳送頻道之原始輸入頻道中一特性化時序包封。
- 24如申請專利範圍第22項之方法,其中該提示碼進而包括一個或多個ICC、ICLD與ICTD碼。
- 25如申請專利範圍第24項之方法,其中聯合該包封提示碼之第一次解析度為較細緻於聯合其他提示碼之第二次解析度。
- 26如申請專利範圍第24項之方法,其中該合成包括遲回聲ICC合成。
- 27如申請專利範圍第24項之方法,其中該經合成訊號之時序包封先於ICLD合成被調整。
- 28如申請專利範圍第22項之方法,其中該經合成訊號的時序包封被特性化;及該經合成訊號的時序包封基於對應包封提示碼之特性化時序包封,及該經合成訊號之特性化時序包封二者被調整。
- 29如申請專利範圍第28項之方法,其中一比例化功能基於對應該包封提示碼,及該經合成訊號之特性化時序包封的特性化時序包封被產生;及該比例化功能被施加至該經合成訊號。
- 30如申請專利範圍第22項之方法,進而包括調整一基於該特性化時序包封之一經傳送頻道以產生一平坦化頻道,其中該升混與合成被施加至該平坦化頻道,以產生一對應的放音頻道。
- 31如申請專利範圍第22項之方法,其進而包括調整一基於該特性化時序包封之一升混頻道,以產生一平坦化頻道,其中該合成被施加至該平坦化頻道,以產生一對應的放音頻道。
- 32如申請專利範圍第22項之方法,其中將該經合成訊號之時序包封僅對特定頻率調整。
- 33如申請專利範圍第32項之方法,其中將該經合成訊號之時序包封僅對一特定截止頻率以上之頻率調整。
- 34如申請專利範圍第22項之方法,其中將該經合成訊號之時序包封於一頻域中調整。
- 35如申請專利範圍第34項之方法,其中時序包封被個別調整用於該經合成訊號中不同訊號次頻帶。
- 36如申請專利範圍第34項之方法,其中該頻域對應FFT。
- 37如申請專利範圍第34項之方法,其中該頻域對應QMT。
- 38如申請專利範圍第22項之方法,其中將該經合成訊號之時序包封於時域中調整。
- 39如申請專利範圍第22項之方法,進而包括決定是否啟動或關閉該經合成訊號之時序包封調整。
- 40如申請專利範圍第39項之方法,其中該決定係基於由產生該E經傳送頻道之一聲音編碼器所產生之一啟動/關閉旗標。
- 41如申請專利範圍第39項之方法,其中該決定係基於分析該E經傳送頻道以偵測暫態,使得該調整被啟動,假如一暫態出現被偵測到的話。
- 42如申請專利範圍第22項之方法,進而包括:特性化一經傳送頻道之一時序包封;及決定是否使用(1)對應該包封提示碼之特性化時序包封或(2)經傳送頻道之特性化時序包封,以調整該經合成訊號之時序包封。
- 43如申請專利範圍第22項之方法,其中在調整該時序包封後,該經合成訊號之一特定窗內之功率為實質相等於調整前該經合成訊號之一對應窗內之功率。
- 44如申請專利範圍第43項之方法,其中該特定窗對應聯合一個或多個非包封提示碼之一合成窗。
- 45一種用於解碼E經傳送聲道以產生C放音聲道之設備,其中C>E≧1,該設備包括:用於接收對應該E經傳送頻道之提示碼的裝置,其中提示碼包括對應該E經傳送頻道之一聲道之一特性化時序包封之一包封提示碼;用於升混一個或多個E經傳送頻道之裝置,以產生一個或多個升混頻道;及用於合成一個或多個C放音頻道之裝置,其係藉將該提示碼施加至一個或多個升混頻道,其中該包封提示碼被施加至一升混頻道或一合成訊號,以調整基於該特性化時序包封之該合成訊號之一時序包封,使得該經調整時序包封實質匹配該特性化時序包封。
- 46一種用於解碼E經傳送聲道以產生C放音聲道之設備,其中C>E≧1,該設備包括:一接收機適以接收對應該E經傳送頻道之提示碼,其中該提示碼包括對應該E經傳送頻道之一聲道之一特性化時序包封之包封提示碼;一升混器適以升混一個或多個E經傳送頻道,以產生一個或多個升混頻道;一合成器適以藉施加提示碼至該一個或多個升混頻道,以合成一個或多個C放音頻道,其中該包封提示碼被施加至一升混頻道或一合成訊號,以調整基於該特性化時序包封之該合成訊號之一時序包封,使得該經調整的時序包封實質匹配該經特性化時序包封。
- 47如申請專利範圍第46項之設備,其中該設備為選自由一數位放影機、一數位聲音放音機、一電腦、一衛星接收器、一有線接收器、一陸上廣播接收器、一家用娛樂系統與一電影劇院系統組成之群之一系統;及該系統包括接收器、升混器、合成器與包封調整器。
- 48一種機器可讀式媒體,具有編碼其上之程式碼,其中,當程式碼以一機器被執行,該機器實行一方法用以解碼E經傳送聲道以產生C放音聲道,其中,C>E≧1,該方法包括:接收對應該E經傳送頻道之提示碼,其中該提示碼包括一對應E經傳送頻道之一聲道之特性化時序包封之包封提示碼;升混一個或多個該E經傳送頻道,以產生一個或多個升混頻道;及藉施加該提示碼至該一個或多個升混頻道,合成一個或多個C放音頻道,其中該包封提示碼被施加至一升混頻道或一合成訊號,以調整基於該特性化時序包封之合成訊號之一時序包封,使得該經調整時序包封實質匹配該特性化時序包封。
Independent claims48
150 paragraphs, as filed
Individual channels used for the formation of two-channel prompt coding (BBC) schemes and similar schemes
The present invention relates to the encoding of sound signals and the subsequent synthesis of sound scenes from encoded sound data
When a person hears a sound signal (ie, sound) produced by a specific sound source, the sound signal usually arrives at the persons left and right ears at two different times and has two different sound levels (for example, decibels). These different times and levels are a function of the difference in the path. Through the path, the sound signal travels to the left ear and the right ear. The brain of the person interprets these differences in time and level and gives the person a sense of the received The sound signal is generated by a sound source located at a specific position (for example, direction and distance) relative to the person. A sound scene is that a person hears one or more sources located at one or more different positions relative to the person at the same time. The net effect of the sound signal produced by two different sound sources.
The existence of this processing by the brain can be used to synthesize sound scenarios, in which sound signals from one or more different sound sources can be purposefully modified to generate left and right signals which are given to different sound sources to be different from the listener Perception of location.
Figure 1 shows a high-level block diagram of the conventional two-channel signal synthesizer 100, which converts a single audio signal (for example, a single audio signal) into the left and right audio signals of a two-channel signal, among which the two-channel signal It is defined as the two signals received at the eardrum of a listener. In addition to the audio source signal, the synthesizer 100 receives a set of spatial cue signals corresponding to the desired position of the audio source relative to the listener. In normal implementation , This group of spatial prompt signals includes a channel intermediate level difference (ICLD) value (which identifies the difference in sound level between the left and right sound signals such as those received by the left and right ears respectively), and the time difference in a channel (ICTD) value (which recognizes the difference in arrival time between left and right sound signals, such as those received in the left and right ears respectively). In addition or as an option, some synthesis techniques involve the patterning of a direction-dependent transfer function for The sound from the signal source to the tympanic membrane is also referred to as the head-related transfer function (HRTF), see, for example, J. Blauert, The Psychophysics of Human Sound Localization, MIT Press, 1983, the teachings of which are incorporated in this submission refer to.
Using the two-channel signal synthesizer 100 in Figure 1, the single audio signal generated by a single audio source can be processed so that when listening through headphones, the audio source applies an appropriate spatial prompt signal set (for example, ICLD, ICTD, and / Or HRTF) is placed spatially to generate a sound signal for each ear, see, for example, D., R. Begault, 3-D Sound for Virtual Reality and Multimedia, Academic Press, Cambridge, MA, 1994.
The two-channel signal synthesizer 100 in Figure 1 generates the simplest type of sound scenes. They have a single sound source configuration relative to the listener. More complex sound scenes include two or more sound sources located at different positions relative to the listener. It is produced using a sound scene synthesizer. The sound scene synthesizer is essentially implemented using multiple two-channel signal synthesizer examples, where each two-channel signal synthesizer example generates a two-element signal corresponding to a different sound source Because each different sound source has a different position relative to the listener, a different spatial prompt signal group is used to generate binaural sound signals for each different sound source.
According to one embodiment, the present invention is a method, device, and machine-readable medium for encoding audio channels. One or more prompt codes are generated and transmitted for one or more audio channels, wherein at least one prompt code is An encapsulation prompt code generated by characterizing a timing encapsulation in one of the one or more channels.
According to another embodiment of the present invention, the present invention is a device for encoding the C input channel to generate the E transmitted channel. The device includes an envelope analyzer, a code estimator, and a downmixer (downmixer), the envelope analyzer characterizes at least one of the input timing envelopes of one of the C input channels, the code estimator generates prompt codes for two or more of the C input channels, and the downmixer reduces Mix the C input channel to generate the E transmitted channel, where C>E1, where the device transmits information about the prompt code and the characteristic input timing envelope to activate a decoder to transmit on E Perform synthesis and encapsulation during channel decoding.
According to another embodiment, the present invention is an encoded audio bit stream generated from an encoded channel, wherein one or more prompt codes are generated for one or more channels, and at least one prompt code is Or one of the multiple channels is characterized by a time-sequence encapsulation generated by an encapsulated prompt code, the one or more prompt codes and the E corresponding to one or more channels are transmitted through the channel, where E 1, is compiled into the encoded audio bit stream.
According to another embodiment, the present invention is an encoded audio bit stream including one or more prompt codes and E transmitted channels, the one or more audio codes are generated for one or more channels, at least one of which is The prompt code is an encapsulated prompt code generated by characterizing a timing encapsulation in one of the one or more channels, and the E corresponds to the one or more channels through the transmission channel.
According to another embodiment, the present invention is a method, device, and machine-readable medium for decoding E transmitted channel to generate C playback channel, where C>E1, receiving corresponds to E The prompt code of the transmitted channel, where the prompt code includes an encapsulated prompt code, and the encapsulated prompt code corresponds to one of the characteristic timing envelopes of one channel of the transmitted channel, and upmixes one or more of the E transmits the channel to generate one or more upmix channels, and synthesizes one or more C playback channels by applying the prompt code to the one or more upmix channels, wherein the enveloped prompt code is applied to an upmix channel Channel or a synthesized signal to adjust a timing envelope of the synthesized signal based on the characterization timing envelope, so that the adjusted timing envelope substantially matches the characterization timing envelope.
In two-channel cue coding (BCC), an encoder encodes the C input channel to generate E transmitted channels, where C>E1, especially two or more C input channels are provided in the frequency domain , And one or more prompt codes are generated for each of one or more different frequency bands of two or more input channels in the frequency domain. In addition, the C input channel is downmixed to generate an E transmitted channel, in In some downmixing implementations, at least one of the E-transmitted channels is based on two or more of the C input channels, and at least one of the E-transmitted channels is based on only one of the C input channels.
In one embodiment, a BCC encoder has two or more filter banks, a code estimator, and a downmixer, and the two or more filter banks input the C into two or more channels Converted from the time domain to the frequency domain, the code estimator generates one or more prompt codes for each of one or more different frequency bands in the two or more converted input channels, and the downmixer downmixes the C input Pass to to generate E through the transmission channel, where C>E1.
In BCC decoding, E transmitted channel is decoded to produce C playback channel, especially for each of one or more frequency bands, one or more E transmitted channels are upmixed in the frequency domain to be Two or more C playback channels are generated in the frequency domain, where C>E1, one or more prompt codes are applied to the one or more different frequency bands of the two or more playback channels in the frequency domain Each of them generates two or more modified channels, and the two or more modified channels are converted from the frequency domain to the time domain. In some upmixing implementations, at least one of the C audio channels is The one based on at least one of the E-transmission channels and at least one prompt code, and the one based on at least one of the C-play audio channels is based on only one of the E-transmitted channels and has nothing to do with any prompt code.
In one embodiment, a BCC decoder has an upmixer, a synthesizer, and one or more inverse filter banks. For each of one or more different frequency bands, the upmixer upscales in the frequency domain. Mix one or more E through the transmission channel to produce two or more C playback audio channels in the frequency domain, where C>E1, the synthesizer applies one or more prompt codes to the two or more in the frequency domain Each of the one or more different frequency bands in the multiple playback audio channels generates two or more modified channels, and the one or more inverse filter banks remove the two or more modified channels from the frequency domain Converted to time domain.
Depending on the specific implementation, a given audio channel can be based on a single transmitted channel, rather than a combination of two or more transmitted channels. For example, when there is only one transmitted channel, each of the C audio channels is Based on the transmitted channel, in these cases, the upmix corresponds to the copy of the transmitted channel. Thus, for applications with only one transmitted channel, the upmixer can use a replicator to copy it for each playback The channel is implemented via the transmission channel.
BCC encoders and/or decoders can be combined into some systems or applications, such as digital video recorders/players, digital recorders/players, computers, satellite transmitters/receivers, and a wired transmitter/receiver , Cable transmitter/receiver, terrestrial broadcast transmitter/receiver, home entertainment system and movie theater system.
(General BCC processing)
FIG. 2 is a block diagram of a conventional two-channel prompt coding (BCC) sound processing system 200, which includes an encoder 202 and a decoder 204. The encoder 202 includes a downmixer 206 and a BCC estimator 208.
Downmixer 206 converts the C input channel<i>x</i><sub><i>i</i></sub>(<i>n</i>) Into E through the transmission channel<i>y</i><sub><i>i</i></sub>(<i>n</i>), where C>E1. In this manual, the signal represented by the variable n is the time domain signal, and the signal represented by the variable k is the frequency domain signal. Depending on the specific implementation, downmixing can be implemented in In either the time domain or the frequency domain, the BCC estimator 208 generates BCC codes from the C input channel and transmits these BCC codes as either in-band or out-of-band side information relative to the E transmitted channel. The usual BCC codes include One or more of the inter-channel time difference code (ICTD), the inter-channel level difference code (ICLD) and the estimated inter-channel correlation code (ICC) data between certain pairs of input channels as a function of frequency and time, This special implementation will indicate that the BCC code is estimated between specific pairs of input channels.
The ICC data corresponds to the consistency of the two-channel signal, which is related to the width of the perception of the sound source. The wider the sound source, the lower the consistency between the left and right channels of the generated two-channel signal. For example, the correspondence is transmitted through an auditorium. The consistency of the two-channel signal of the orchestra of the podium is usually lower than the consistency of the two-channel signal corresponding to a single violin solo. Generally, a sound signal with lower consistency is usually perceived as propagating in the sound space Farther, so, ICC data is usually about the apparent sound source width and degree of the listener's environment, see, for example, J. Blauert, The Psychology of Human Sound Localization, MIT press, 1983.
Depending on the specific application, the E transmitted channel and the corresponding BCC code can be directly transmitted to the decoder 204 or stored in some suitable type of storage device for subsequent access by the decoder, depending on the situation, terminology "Transmission" can be referred to as either directly transmitting to a decoder or storing for subsequent supply to a decoder. In either case, the decoder 204 receives the transmitted channel and side information and performs upmixing and the use of BCC codes. BCC synthesis to convert E through the transmission channel into a playback channel exceeding E (usually, but not necessarily, C)<img file="TWI318079B_D0001.tif" />(n) For sound playback, depending on the specific implementation, upmixing can be performed either in the time domain or in the frequency domain.
In addition to the BCC processing shown in Figure 2, a conventional BCC sound processing system may contain additional encoding and decoding stages to further compress the sound signal in the encoder and then decompress the sound signal in the decoder. These sound codes can be Based on conventional sound compression/decompression techniques such as those based on Pulse Code Modulation (PCM), Differential PCM (DPCM) or Adaptive DPCM (ADPCM).
When the downmixer 206 generates a single sum signal (ie, E=1), the BCC code can represent a multi-channel audio signal at a bit rate that is only slightly higher than the signal required to represent a single sound. It is estimated that the ICTD, ICLD, and ICC data contain approximately two orders of magnitude less information than a sound waveform.
Not only the low bit rate of BCC encoding, but also the backward compatibility point of view is also interested. A single transmitted sum signal corresponds to the single-tone downmix of the original stereo or multi-channel signal. It does not support stereo or multi-channel for the receiver. Sound reproduction, listening to the transmitted sum signal is one of the correct ways to present the sound material on low-profile monophonic reproduction equipment. BCC encoding can therefore also be used to improve the transmission of monophonic sound material to multi-channel sound. Existing services, for example, the existing monophonic sound wireless broadcasting system can be upgraded for stereo or multi-channel playback. If the BCC side information can be embedded in the existing transmission channel, the similar ability exists when downmixing multi-channel sound to the corresponding stereo The second sum signal.
BCC processes sound signals with a certain time and frequency resolution. The frequency resolution used is mainly caused by the frequency resolution of the human auditory system. Psychiatrists suggest that spatial perception is most likely based on one of the thresholds of the sound input signal. (Critical) frequency band means that this frequency resolution is achieved by using a reversible filter library (for example, based on a fast Fourier transform (FFT) or A quadrature mirror filter (QMF)) is considered.
(General downmix)
In a preferred implementation, the transmitted sum signal includes all the signal components of the input audio signal, and the goal is that each signal component is completely preserved. The simple sum of the audio input channels often produces amplification or attenuation of the signal components, in other words In other words, the power of the signal components in a "simple" sum is often greater or less than the sum of the power of the corresponding signal components in each channel. A downmixing technique can be used, which equalizes the sum signal so that the sum signal is The power of the signal component in it is approximately equal to the corresponding power in all input channels.
Figure 3 shows a block diagram of a downmixer 300 which can be used in the downmixer 206 of Figure 2 according to the equivalent implementation of the BCC system 200. The downmixer 300 has a filter bank (FB) 302 for each One input channel<i>x</i><sub><i>i</i></sub>(<i>n</i>), a downmix block 304, an optional scaling/delay block 306 and an inverse FB (IFB) 308 for each encoded channel<i>y</i><sub><i>i</i></sub>(<i>n</i>)。
Each filter bank 302 takes a corresponding number in the time domain as the input channel<i>x</i><sub><i>i</i></sub>(<i>n</i>) Each frame (for example, 20msec) is converted into a set of input coefficients in the frequency domain<img file="TWI318079B_D0002.tif" />(<i>k</i>), the downmix block 305 downmixes each frequency band corresponding to the input coefficient of C into E. One of the frequency domain coefficients of the downmix is the corresponding subband. Equation (1) represents the input coefficient<img file="TWI318079B_D0003.tif" />(<i>k</i>), <img file="TWI318079B_D0004.tif" />(<i>k</i>),..., <img file="TWI318079B_D0005.tif" />(<i>k</i>) <img file="TWI318079B_D0006.tif" />Downmixing of the kth sub-band to generate the downmix coefficient<img file="TWI318079B_D0007.tif" />(<i>k</i>), <img file="TWI318079B_D0008.tif" />(<i>k</i>),..., <img file="TWI318079B_D0009.tif" />(<i>k</i>) <img file="TWI318079B_D0010.tif" />The kth sub-bands are as follows:<maths><img file="TWI318079B_D0011.tif" /></maths>in<i>D</i><sub><i>CE</i></sub>It is a real-valued C-by-E downmix matrix.
The optional scaling/delay block 306 includes a set of multipliers 310, each of which uses a scaling factor<i>e</i><sub><i>i</i></sub>(<i>k</i>) Multiplied by a corresponding downmix coefficient<img file="TWI318079B_D0012.tif" />(<i>k</i>) To generate a corresponding scaling factor<img file="TWI318079B_D0013.tif" />(<i>k</i>), the motivation for the scaling operation is equal to the generalized equalization used for downmixing with arbitrary weighting factors for each channel. If the input channels are independent, then the power of the downmixed signal in each frequency band<img file="TWI318079B_D0014.tif" />The equation (2) is as follows:<maths><img file="TWI318079B_D0015.tif" /></maths>,in<img file="TWI318079B_D0016.tif" />To borrow C-by-E downmix matrix<i>D</i><sub><i>CE</i></sub>Is obtained by squaring each matrix element, and<img file="TWI318079B_D0017.tif" />Is the power of the sub-band k of the input channel i.
If the sub-bands are not independent, then the power value of the downmixed signal<img file="TWI318079B_D0018.tif" />Make larger or smaller ones than those calculated using formula (2). Since the signal is amplified or cancelled when the signal components are in phase or out of phase, to avoid this, the downmixing operation of formula (1) is followed by a multiplier The scaling operation of 30 is applied to the sub-band, the scaling factor<i>e</i><sub><i>i</i></sub>(<i>k</i>)(1iE) can be obtained from the formula (3) as follows:<maths><img file="TWI318079B_D0019.tif" /></maths>,in,<img file="TWI318079B_D0020.tif" />Is the sub-band power as calculated by formula (2), and<img file="TWI318079B_D0021.tif" />Corresponding to the downmixed sub-band signal<img file="TWI318079B_D0022.tif" />(<i>k</i>) Of the power.
In addition to or without optional scaling, the scaling/delay block 306 can optionally apply a delay to the signal.
Each inverse filter bank 308 corresponds to a set of scaled coefficients in the frequency domain<img file="TWI318079B_D0023.tif" />(<i>k</i>) Is converted into a corresponding digital and transmitted channel<i>y</i><sub><i>i</i></sub>(<i>n</i>) A frame.
Although Figure 3 shows that all C of the input channels are converted to the frequency domain for subsequent downmixing, in an alternative implementation, one or more of the C input channels (but less than C-1) can avoid Figure 3 Some or all of the operations shown in and can be transmitted as an equal number of uncorrected channels, depending on the particular implementation, these uncorrected channels may or may not be generated by the BCC estimator 208 in Figure 2 Used by transmitting BCC code.
In the implementation of the downmixer 300, it generates a single sum signal y(n), E=1, and a signal of each frequency band of each input channel c<img file="TWI318079B_D0024.tif" />(<i>k</i>) Is added and then a factor<i>e</i>(<i>k</i>) Multiply, according to the formula (4) as follows:<maths><img file="TWI318079B_D0025.tif" /></maths>,factor<i>e</i>(<i>k</i>) Is obtained by formula (5) as follows:<maths><img file="TWI318079B_D0026.tif" /></maths>,in<img file="TWI318079B_D0027.tif" />(<i>k</i>) Is at time index k<img file="TWI318079B_D0028.tif" />(<i>k</i>) One of the power is estimated for a short time, and<img file="TWI318079B_D0029.tif" />(<i>k</i>) Is the power<img file="TWI318079B_D0030.tif" />In short-term estimation, the equalized sub-band is converted back to the time domain that generates the sum signal sent to the BCC decoder.
(General BCC synthesis)
Figure 4 shows a block diagram of a BCC synthesizer 100, which can be used in the decoder 204 of Figure 2 according to some implementations of the BCC system 200. The BCC synthesizer 400 has a filter bank 402 for each transmission Channel<i>y</i><sub><i>i</i></sub>(<i>n</i>), an upmix block 404, a delay 406, a multiplier 408, a related block 410 and an inverse filter bank 412 for each audio channel<img file="TWI318079B_D0031.tif" />(<i>n</i>)。
Each filter bank corresponds to a digital, transmitted channel in the time domain<i>y</i><sub><i>i</i></sub>(<i>n</i>) Each frame is converted into a set of input coefficients in the frequency domain<img file="TWI318079B_D0032.tif" />(<i>k</i>), the up-mixing block 404 up-mixes each sub-band corresponding to the transmitted channel coefficient of E into one of the corresponding sub-bands of the C up-mixed frequency domain coefficient. Equation (4) represents the transmitted channel coefficient<img file="TWI318079B_D0033.tif" />(<i>k</i>), <img file="TWI318079B_D0034.tif" />(<i>k</i>),..., <img file="TWI318079B_D0035.tif" />(<i>k</i>) <img file="TWI318079B_D0036.tif" />Upmixing of the kth sub-band to generate upmixing coefficients<img file="TWI318079B_D0037.tif" />(<i>k</i>), <img file="TWI318079B_D0038.tif" />(<i>k</i>),..., <img file="TWI318079B_D0039.tif" />(<i>k</i>) <img file="TWI318079B_D0040.tif" />The kth sub-bands are as follows:<maths><img file="TWI318079B_D0041.tif" /></maths>,in<i>U</i><sub><i>EC</i></sub>It is a real-valued E-by-C up-mixing matrix that performs up-mixing in the frequency domain so that the up-mixing can be individually applied to each different sub-band.
Each delay 406 applies a delay value based on a corresponding BCC code used for the ICTD data<i>d</i><sub><i>i</i></sub>(<i>k</i>) To ensure that the desired ICTD value appears in certain pairs of audio channels, each multiplier 408 applies a scaling factor based on a corresponding BCC code for the ICLD data<i>a</i><sub><i>i</i></sub>(<i>k</i>) To ensure that the desired ICLD value appears in some pairs of the audio channel, the relevant block 410 executes one of the corresponding BCC codes for the ICC data to de-correlate operation A to ensure that the desired ICC value appears in For some centering of audio channels, a further description of the operation of related blocks can be found in US Patent Application No. 10/155,437 such as Baumgarte 2-10 filed on May 24, 2002.
The synthesis of ICLD values can be less troublesome than the synthesis of ICLD and ICC values, because ICLD synthesis only involves the scaling of sub-band signals, because the ICLD prompt signal is the most commonly used directional prompt signal, and the ICLD value is similar to that of the original sound signal These values are usually more important. In this way, ICLD data can be estimated between all channel pairs. The scaling factor for each frequency band<i>a</i><sub><i>i</i></sub>(<i>k</i>), (1iC) should be selected so that the sub-band power of each audio channel is similar to the corresponding power of the original input channel.
A target can apply relatively little signal correction to synthesize the ICTD and ICC values. In this way, the BCC value may not include the ICTD and ICC values for all channel pairs. In this case, the BCC synthesizer 400 will only work on certain channels. Combine ICTD and ICC values between pairs.
Each inverse filter bank 412 combines a set of corresponding coefficients in the frequency domain<img file="TWI318079B_D0042.tif" />(<i>k</i>) Into a corresponding digital, audio channel<img file="TWI318079B_D0043.tif" />(<i>n</i>) One frame.
Although Figure 4 shows that all E transmitted channels are converted into frequency domain for subsequent upmixing and BCC processing, in another implementation, one or more (but not all) of the E transmitted channels can avoid Figure 4 Some or all of the processing shown, for example, one or more of the transmitted channels may be unmodified channels which do not accept any upmixing. Except as one or more C playback channels, these unmodified channels alternately , Can be but not necessarily used as a reference channel, and its BCC processing is applied to synthesize one or more other playback channels. In any case, these uncorrected channels can be delayed to compensate for the time involved in upmixing And/or the BCC operation used to generate the remaining audio channels.
Note that although Figure 4 shows that the audio channel C is synthesized from E via the transmission channel, where C is also the number of original input channels, BCC synthesis is not limited to the number of audio channels, usually, the number of audio channels It can be any number of channels, including the number greater than or less than C and possibly even when the number of audio channels is equal to or less than the number of transmitted channels.
("Relative difference in perception" between channels)
Assuming a single total signal, BCC synthesizes a stereo or multi-channel audio signal so that ICTD, ICLD and ICC are similar to the corresponding prompt signals of the original audio signal. Below, the role of ICTD, ICLD and ICC in the special value of audio space image will be discussed .
Stereo and multi-channel audio signals usually include a complex mix of synchronized active source signals that are superimposed on reflected signal components generated from recording in the surrounding space, or used by recording engineers to artificially generate a spatial impression Added, different audio signals and their reflection occupy different areas in the time-frequency plane, which are reflected by ICTD, ICLD and ICC, which change as a function of time and frequency. In this case, instantaneous ICTD , The relationship between ICLD and ICC and the direction of sound events and spatial impression is not obvious.
The strategy of some BCC embodiments is to unintentionally synthesize these prompt signals so that they approximate the corresponding prompt signals of the original sound signal.
A filter bank with a sub-band bandwidth equal to twice the equal rectangular bandwidth (ERB) is used, and informal listening will leak that the sound quality of BCC is not significantly improved when a higher frequency resolution is selected, a lower frequency resolution It may be required because it results in fewer ICTD, ICLD, and ICC values, which need to be transmitted to the decoder and therefore at a lower bit rate.
Regarding time resolution, ICTD, ICLD, and ICC are usually considered at a fixed time interval. When ICTD, ICLD, and ICC are considered every 4 to 16 ms, high performance can be obtained. Note that unless the prompt signal is in A very short time interval is considered, and the previous effect is not directly considered. Assume that one of the sound stimuli is classically leading-lagging. If the lead and lag are within a time interval, only one set of cue signals are synthesized, and then the part of the lead However, the sound quality achieved by BCC is reflected in an average MUSHRA score of about 87 (ie, "excellent" sound quality) on average, and almost 100 for some sound signals.
The often-obtained small perceptual difference between the reference signal and the synthesized signal implies that the cue signal about a wide range of sound spatial image characteristics is implicitly considered by the synthesized ICTD, ICLD, and ICC at a fixed time interval. Below, some The argument is about how ICTD, ICLD, and ICC can be related to a range of audio-spatial image characteristics.
(Estimation of space cue signal)
In the following, we will describe how ICTD, ICLD, and ICC are estimated. The bit rate used for the transmission of these (quantized and coded) spatial cue signals can be just a few kb/s. Therefore, using BCC, it may be in bits. The rate is close to that required for a single channel to transmit stereo and multi-channel audio signals.
Fig. 5 shows a block diagram of the BCC estimator 208 in Fig. 2 according to the present invention. The BCC estimator 208 includes a filter bank (FB) 502, which can be the same as the filter bank 302 in Fig. 3, and is similar to the estimation area Block 504 generates ICTD, ICLD and ICC spatial cue signals for each of the different frequencies generated by the filter library 502.
(Estimation of ICTD, ICLD and ICC for stereo signal)
The following measurement is used for ICTD, ICLD and ICC to correspond to the sub-band signal of two (for example, stereo) channels<img file="TWI318079B_D0044.tif" />(<i>k</i>)and<img file="TWI318079B_D0045.tif" />(<i>k</i>)。
ICTD[Example]<maths><img file="TWI318079B_D0046.tif" /></maths>It has a short-term estimate of one of the standardized cross-correlation functions obtained from the following formula (8).
<maths><img file="TWI318079B_D0047.tif" /></maths>in<i>d</i><sub>1</sub>=max{-<i>d</i>,0} (9)<i>d</i><sub>2</sub>=max{<i>d</i>,0} and,<img file="TWI318079B_D0048.tif" />(<i>d</i>,<i>k</i>)for<img file="TWI318079B_D0049.tif" />(<i>k</i>-<i>d</i><sub>1</sub>) <img file="TWI318079B_D0050.tif" />(<i>k</i>-<i>d</i><sub>2</sub>) A short-term estimate of one of the averages.
ICLD[dB]<maths><img file="TWI318079B_D0051.tif" /></maths>
ICC<i>c</i><sub>1</sub><sub>2</sub>(<i>k</i>)=max|Φ<sub>1</sub><sub>2</sub>(<i>d</i>,<i>k</i>)| (11) Note that the absolute value of the standardized cross-correlation is considered and<i>c</i><sub>1</sub><sub>2</sub>(<i>k</i>) Has a range [0,1].
(Estimation of ICTD, ICLD and ICC for multi-channel audio signal)
When there are more than two input channels, it is usually sufficient to define ICTD and ICLD (for example, channel number 1) and other channels between a reference channel, as illustrated in Figure 6 for the case of C=5 channel, where τ<sub>1</sub><sub><i>c</i></sub>(<i>k</i>) And <i>L</i><sub>1</sub><sub>2</sub>(<i>k</i>) Indicate ICTD and ICLD between reference channel 1 and channel c, respectively.
Contrary to ICTD and ICLD, ICC usually has more degrees of freedom. The defined ICC has different values among all possible input channel pairs. For channel C, it has the possibility of C(C-1)/2 Channel pairs, for example, for 5 channels, there will be 10 channel pairs as illustrated in Figure 7(a). However, these methods need to estimate each frequency band at each time index and transmit C(C-1) /2 ICC values, resulting in high computational complexity and high bit rate.
Alternatively, for each frequency band ICTD and ICLD determine the direction of the sound event of the corresponding signal component in the sub-band, a single ICC parameter in each frequency band can then be used to describe the overall consistency between all channels, and good results can be achieved by only It is obtained by estimating and transmitting the ICC signal between the two channels with the most energy in each frequency band of each time index. This example is shown in Figure 7(b), where the time instant k-1 and k are the channel pair (3,4) and (1,2) are the strongest respectively. Heurislic rules can be used to determine ICC among other channel pairs.
(Synthesis of Space Cue Signal)
Figure 8 shows a block diagram of an implementation of the BCC synthesizer 400 of Figure 4, which can be used in a BCC decoder to generate a single transmitted sum signal s(n) plus a spatial cue signal. Stereo or multi-channel audio signal, the sum signal s(n) is decomposed into sub-bands, where<img file="TWI318079B_D0052.tif" />(<i>k</i>) Indicates one of these sub-bands, in order to generate the corresponding sub-bands of each output channel, delay<i>d</i><sub><i>c</i></sub>Scaling factor<i>a</i><sub><i>c</i></sub>With filter<i>h</i><sub><i>c</i></sub>It is applied to the corresponding sub-band of the sum signal (for simplicity of labeling, the time index k is omitted in the delay, scaling factor and filter), ICTD by adding delay, ICTD by scaling and ICC by applying a decorrelation filter After being synthesized, the processing shown in Figure 8 is applied to each frequency band independently.
(ICTD synthesis)
delay<i>d</i><sub><i>c</i></sub>From ICTDs τ<sub>1</sub><sub><i>c</i></sub>(<i>k</i>) Is determined based on formula (12) as follows:<maths><img file="TWI318079B_D0053.tif" /></maths>Delay for reference channel<i>d</i><sub>1</sub>Be calculated to delay<i>d</i><sub><i>c</i></sub>The maximum size is minimized, the fewer sub-band signals are corrected, and the less man-made hazards are generated. If the sub-band sampling rate does not provide high enough time resolution for ICTD synthesis, the delay can be increased by using a suitable all-pass filter. Accurately added to it.
(ICLD synthesis)
To make the output sub-band signal have the required ICLDs on channel c and reference channel 1 <i>L</i><sub>1</sub><sub>2</sub>(<i>k</i>), gain factor<i>a</i><sub><i>c</i></sub>The formula (13) should be satisfied as follows:<maths><img file="TWI318079B_D0054.tif" /></maths>
In addition, the output sub-band should be standardized so that the total power of all non-output channels is equal to the power of the input total signal, because all the original signal power in each sub-band is stored in the total signal, which is the absolute sub-band power. The normalized result approximates the corresponding power of the original encoder sound signal for each output channel. Under these constraints, the scaling factor<i>a</i><sub><i>c</i></sub>It is obtained by the following formula (14).
<maths><img file="TWI318079B_D0055.tif" /></maths>
(ICC synthesis)
In some embodiments, the goal of ICC synthesis is to reduce the correlation between the sub-bands after the delay and the scaling has been applied without affecting ICTD and ICLD. This can be achieved by designing the filter in Figure 8.<i>h</i><sub><i>c</i></sub>To achieve, ICTD and ICLD are effectively changed as the same frequency function, so that the average variation in each frequency band (sound critical band) is 0.
Figure 9 illustrates how ICTD and ICLD are changed in the primary frequency band as the same frequency function. The amplitude of ICTD and ICLD variation determines the degree of decorrelation and is controlled as the same ICC function. Note that ICTD is changed gently (as in Section 9(a) ) Figure), while the ICLD is arbitrarily changed (as in Figure 9(b)), the ICLD can be changed as smoothly as the ICTD, but this will result in more sound signals.
Another method for synthesizing ICC, particularly suitable for multi-channel ICC synthesis, is described in more detail in C. Faller, "Parametric multi-channel ICC audio coding: Synthesis of coherence cues," IEEE Trans.On Speech and Audio Proc. , 2003, its teaching is incorporated here for reference. As a function of time and frequency, a specific amount of artificial late-reverberation is added to each output channel to obtain a desired ICC. In addition, , Spectrum correction can be applied so that the spectrum envelope of the generated signal is close to the spectrum envelope of the original sound signal.
Other related and unrelated ICC synthesis techniques for stereo signals (or channel pairs) have been published in E. Schuijers, W. Oomen, B.den Brinker, and J. Breebaart, "Advances in parametric coding for high-quality audio,"in Preprint 114<sup>t</sup><sup>h</sup>Conv.Aud.Eng.Soc.,Mar.2003,and J.Engdegard,H.Purnhagen,j.Roden,and L.Liljeryd,"Synthetic ambience in parametric stereo coding,"in Preprint 117<sup>t</sup><sup>h</sup>Cov.Aud.Eng.Soc., May 2004, the teachings of the two are incorporated here for reference.
(C-to-E BCC)
As previously described, BCC can be implemented on more than one transmission channel. A variant of BCC has been described, which means that the C channel is not a single (transmitted) channel, but the E channel, which is labeled C-to-E BCC. There are at least two motivations for C-to-E BCC: BCC with a transmission channel provides a backwards compatible path to upgrade the existing mono system for stereo or multi-channel sound playback, the upgraded The system transmits the BCC downmix sum signal via the existing monophonic structure, and the C-to-E BCC can be applied to the E-channel backward compatible C-channel sound coding.
C-to-E BCC introduces proportionality with varying degrees of reduction in the number of transmitted channels. It can be expected that more channels will be transmitted and there will be better sound quality.
The signal processing details of C-to-E BCC, such as how to define ICTD, ICLD, and ICC prompt signals, are described in US Patent No. 10/762,100 (Faller 13-1) filed on January 20, 2004.
(Individual channel shaping)
In some embodiments, both BCC and C-to-E BCC with one transmission channel involve an algorithm for ICLD ICTD and/or ICC synthesis. Generally, about every 4 to 30 ms is sufficient to synthesize ICTD, ICLD and/ Or ICC, but the perceptual phenomenon of the precedence effect implies that when the human auditory system evaluates the prompt signal with a higher time resolution (for example, every 1 to 10 ms), there is a specific time instant.
A single static filter bank usually cannot provide a high enough frequency resolution at a time instant, which is suitable for most time instants, and at the same time, when the previous effect becomes effective, it provides a high enough time resolution at the time instant.
Certain embodiments of the present invention are directed to use relatively low time resolution ICTD, ICLD, and/or ICC synthesis, while adding additional processing to emphasize the time instant when higher time resolution is required. In addition, in some In an embodiment, the system eliminates the need for signal adaptive window switching technology, which is usually difficult to integrate into a system architecture. In some embodiments, one or more of the original encoder input channels The timing envelope is estimated, which can be done, for example, directly by analyzing the signal time frame or by checking the autocorrelation of the signal spectrum in time. The two methods will be further explained in subsequent embodiments, including the information contained in these envelopes It is sent to the decoder (such as an encapsulated prompt code), if it is perceptually needed and advantageous.
In some embodiments, the decoder applies certain processing to impose these desired timing encapsulation on its output channels.
This can be achieved by TP processing, such as the signal encapsulation operation in which the time-domain samples of the signal are multiplied by a time-varying amplitude correction function. A similar processing can be applied to the spatial/sub-band samples, if the sub-band The time resolution is sufficiently high (at the cost of coarse frequency resolution).
Additionally, the signal spectrum in frequency indicates that a convolution/filtering can be used in a manner similar to that used in previous techniques, for quantized noise shaping of a low-bit-rate voice encoder or for intensity enhancement Stereo coded signal, which is better if the filter bank has a high frequency resolution and therefore a relatively low time resolution, the deconvolution/filtering method: the encapsulation shaping method is extended from intensity stereo to C-to -E multi-channel coding.
The technique includes a setting where the encapsulation is shaped to be controlled by parameter information (eg, binary flags) generated by the encoder, but is actually performed using the filter coefficient set derived from the decoder.
In another arrangement, the filter coefficient set is transmitted from the encoder, for example only when it is perceptually necessary and/or beneficial.
The same approach to the time domain/subband domain is also true. Therefore, thresholds (for example, transient detection and a speech estimation) can be introduced to additionally control the transmission of encapsulated information.
When it is advantageous to shut down the TP process to avoid potential artifacts, for safety reasons, a good strategy is to shut down the transient process without execution (ie, BCC will operate according to a conventional BCC method), The additional processing is only turned off when it is expected to be improved by the higher transient resolution of the channel, for example, when it is expected that the previous effect becomes active.
As described earlier, this enabling/disabling control can be achieved by transient detection, that is, if a transient is detected and then TP processing is activated, the previous effect is the most effective for the transient Yes, transient detection can be used with a look-ahead to effectively shape not only a single transient state but also signal components before or shortly after the transient state. Possible ways to detect transients include: Observe BCC If there is a sudden increase in the power of the encoder input signal or the timing envelope of the transmitted BCC sum signal, then a transient occurs.
Check that the linear predictive coding (LPC) gain is as estimated in the encoder or decoder. If the LPC predictive gain exceeds a critical value, it can be assumed that the signal is transient or highly fluctuating. The LPC analysis is in the spectrum The autocorrelation is calculated.
In addition, in order to avoid possible artifacts in the tone signal, the TP treatment should not be applied when the tone of the transmitted total signal is high.
According to some embodiments of the present invention, the timing envelope of individual original channels is estimated by a BCC encoder to activate a BCC decoder to generate timing envelopes with timing envelopes similar to (or perceptually similar) to the original channels Output channel. Some embodiments of the present invention focus on the previous effect phenomenon. Some embodiments of the present invention involve the transmission of encapsulation prompt codes in addition to other BCC codes, such as ICLD, ICTD, and/or ICC, as the BCC side Information part.
In some embodiments of the present invention, the time resolution of the timing envelope prompt signal is finer than the time resolution of other BCC codes (for example, ICLD, ICTD, ICC), which makes the envelope shape more precise The synthesis window is executed in the time period provided by the synthesis window, and the synthesis window corresponds to a block length of an input channel, which is derived from other BCC codes of the input channel.
(Example)
Figure 10 shows a block diagram of the time-domain processing of encoder 202 added to a BCC encoder in accordance with an embodiment of the present invention. As shown in Figure 10(a), each temporary The state processing analyzer (TPA) 1002 estimates the timing envelope of a different original input channel, although usually any one or more of the input channels can be analyzed.
Figure 10(b) shows a block diagram of a possible implementation of TPA1002 on a time-domain basis, where the input signal samples are squared (1006) and then low-pass filtered (1008) to characterize the timing envelope of the input signal. In another embodiment, the timing envelope is estimated using an auto correlation/LPC method or using other methods such as Hilbert transform.
Block 1004 in Figure 10(a) parameterizes, quantizes, and encodes the estimated timing envelope before transmission as transient processing (TP) information (ie, encapsulation hint code), which is included in Figure 2 Side information.
In one embodiment, one of the detectors (not shown) in the block 1004 determines whether the TP processing in the decoder will improve the sound quality, so that the block 1004 is only at these times when the sound quality will be improved by the TP processing. Send TP side information when it is improved.
Figure 11 illustrates an exemplary time-domain application of the TP processing of the BCC synthesizer 400 in Figure 4. In this embodiment, there is a single transmitted sum signal s(n), and the C base signal is copied by copying the sum signal Generated, and the envelope shaping is individually applied to different composite channels. In another embodiment, the order of delay, scaling, and other processing may be different, and, in another embodiment, the envelope shaping is not limited to Each channel is processed independently. This implementation based on convolution/filtering is particularly true. It seeks the consistency of the entire frequency band to derive information on the signal transient fine structure.
In Figure 11(a), the decoding block 1102 receives the transmitted TP side information from the BCC encoder and returns the timing encapsulation information a to each input channel, and each TP block 1104 applies the corresponding encapsulation information to Shape the envelope of the output channel.
Figure 11(b) shows a block diagram of a possible time-domain implementation of TP1104, where the synthesized signal samples are squared (1106), and then low-pass filtered (1108) to characterize the synthesized channel For the timing envelope, a scaling factor (for example, sqrt(a/b)) is generated (1110) and then applied (1112) to the synthesized channel to generate an output signal with a timing envelope, the output signal It is roughly equal to the output signal corresponding to the original input channel.
In another implementation of TPA1002 in Figure 10 and TP1004 in Figure 11, timing encapsulation is characterized by magnitude operations instead of squaring signal samples. In these implementations, the a/b ratio can be used It is a scale factor without having to use the square root operation.
Although the scaling operation in Figure 11(c) corresponds to a time-domain basic implementation of TP processing, TP processing (and TPA and inverse TP (ITP) processing) can also be implemented using frequency-domain signals, as shown in Figures 16-17 The embodiment (described below), so, for the purpose of this description, the term "scale function" should be interpreted to cover operations that are not time-domain or frequency-domain, such as the filtering operations in Figures 17(b) and (c).
Generally, each TP1004 should be designed so that it does not modify the signal power (ie, energy). Depending on the specific implementation, this signal power can be a short-term average signal power in each channel, for example, in a composite The total signal power per channel in the time domain defined by the window or some other suitable power measurement, so that the scaling for ICLD synthesis (for example, using the multiplier 408) can be applied before and after the envelope is formed.
Because the full-band scaling of the BCC output signal can produce artifacts, the encapsulation can be applied only at a specific frequency, for example, the frequency is greater than a comparable cut-off frequency<i>f</i><sub><i>TP</i></sub>(For example, 500 Hz). Note that the frequency range used for analysis (TPA) can be different from the frequency range used for synthesis (TP).
Figures 12(a) and (b) show the possible implementation of TPA1002 in Figure 10 and TP1104 in Figure 11, where the encapsulation can only be higher than the cut-off frequency<i>f</i><sub><i>TP</i></sub>The frequency is applied. In particular, Figure 12(a) shows the addition of the high-pass filter 1202, which filters lower than<i>f</i><sub><i>TP</i></sub>The frequency, Figure 12(b) shows that there is a cut-off frequency between the secondary frequency bands<i>f</i><sub><i>TP</i></sub>In addition to the dual-band filter bank 1204, where only the high frequency part is temporarily shaped, the dual-band inverted filter bank 1206 then combines the low frequency part and the transiently shaped high frequency part to generate an output channel.
Figure 13 shows a block diagram of frequency domain processing according to an embodiment of the present invention, which is added to a BCC encoder, as shown in Figure 2 for encoder 202 as shown in Figure 13(a), each The processing of TPA (1302) is applied to different sub-bands respectively, where each filter bank (FB) is the same as the one corresponding to FB302 in Figure 3, and block 1304 is similar to block 1004 in Figure 10. In another implementation, the sub-band used for TPA processing can be different from the BCC sub-band. As shown in Figure 13(b), TPA1302 can be implemented similar to TPA1002 in Figure 10.
Figure 14 illustrates an example of frequency domain application of TP processing in the BCC synthesizer 400 of Figure 4, the decoding block 1402 is similar to the decoding block 1102 of Figure 11, and each TP1404 is similar to that of each TP1104 of Figure 11 The primary frequency band is implemented as shown in Figure 14(b).
Figure 15 shows a block diagram of the frequency domain processing of encoder 202 attached to a BCC encoder according to another embodiment of the present invention. This method has the following setup for each input channel The encapsulated information is derived by calculation across frequency (1502), parameterization (1504), and quantization (1506), and is encoded into a bit stream (1508) by an encoder. Figure 17(a) illustrates Figure 15 In an embodiment of TPA1502, the side information to be sent to the multi-channel synthesizer (decoder) can be the reflection coefficient or line spectrum equivalent generated by the LPC filter coefficients calculated by an autocorrelation method, or to maintain the The data rate of the side information is small. For example, the LPC prediction gain is a parameter derived from the binary flag of "temporary presence/absence".
Figure 16 illustrates another exemplary frequency domain application of the TP processing of the BCC synthesizer 400 in Figure 4. The encoding process of Figure 15 and the decoding process of Figure 16 can be implemented to form an encoder/decoder profile. A matching pair, the decoding block 1602 is similar to the decoding block in Figure 14, and each TP1604 is similar to each TP1404 in Figure 14 in this multi-channel synthesizer, and the TP side information is transmitted to be decoded and used to Control the envelope shaping of individual channels. In addition, however, the synthesizer includes an envelope characterizer stage (TPA) 1606 for the analysis of the transmitted sum signal, and an inverse TP (ITP) 1608 for flattening each Timing encapsulation of the basic signal, in which the encapsulation adjuster (TP) 1604 applies a modified encapsulation to each output channel, depending on the specific implementation, ITP can be applied either before or after upmixing, In detail, this is done using a convolution/filtering method, in which the encapsulation shaping is performed by applying an LPC basic filter to the cross-frequency spectrum such as 17(a), (b), (c) used for TPA, ITP, and TP processing. ) Is achieved as illustrated in the figure. In figure 16, the control block 1610 decides whether the encapsulation shaping will be implemented, and if so, whether it will be based on (1) the TP side information is transmitted or (2) the TPA1606 Partially characterize encapsulated data.
Figures 18(a) and (b) illustrate the two exemplary modes of the operation control block 1610 of Figure 16. In the implementation of Figure 18(a), a set of filter coefficients are transmitted to the decoder and converted The convolution/filtering encapsulation shaping is done based on the transmitted coefficients. If the transient shaping is detected as not beneficial by the encoder, then no filter data is sent out and the filter is turned off (in section 18(a)) The picture is switched to a single filter coefficient group "[1,0...]".
In the implementation of Figure 18(b), only one "transient/non-transient flag" is transmitted for each channel and this flag is used to activate or make the decoder based on the transmitted downmix signal The calculated filter coefficient is invalid.
(Further another embodiment)
Although the present invention has been described in terms of BCC encoding, which has a single sum information, the present invention can also be implemented in terms of BCC encoding with two or more sum signals. In this case, it is used for each different "base". The timing envelope of the sum signal can be estimated before BCC synthesis is applied, and different BCC output channels can be generated based on different timing envelopes, depending on which sum signal is used to synthesize different output channels, an output channel from The synthesis of two or more sum channels can be generated based on an effective timing envelope that takes into account the relative effects of the constituent sum channels (for example, via weighted average).
Although the present invention has been described in terms of BCC codes involving ICTD, ICLD and ICC codes, the present invention can also be implemented in terms of BCC codes involving only one or two of these three types of codes (for example, ICTD, ICC instead of ICLD) And/or one or more of the additional code types, and the sequence of the BCC synthesis process and the encapsulation shaping can be changed in different implementations, for example, when the encapsulation shaping is applied to the frequency domain signal, as shown in Figures 14 and 16, Encapsulation shaping can be performed after ICTD synthesis (in those embodiments using ICTD synthesis) but before ICLD synthesis. In other embodiments, encapsulation shaping can be applied to the rise before any other BCC synthesis is applied. Mixed signal.
Although the present invention has been described in terms of a BCC encoder that generates a package reminder code from an original input channel, the package reminder code can be generated from a downmix channel corresponding to the original input channel, which will activate a processor (e.g., a The implementation of separate package prompt encoder), which can (1) receive the output of an encoder BCC that produces downmix channels and certain BCC codes (for example, ICLD, ICTD and/or ICC), and (2) characterization The timing envelope of one or more of the downmix channels is used to add the envelope prompt code to the BCC side information.
Although the present invention has been described in terms of BCC coding schemes, in which the envelope reminder code is transmitted in one or more channels and other BCC codes, in another embodiment, the envelope reminder code can be either alone or in combination with other BCC codes. The code is transmitted to a place (for example, a decoder or a storage device) which already has the transmitted channel and possibly other BCC codes.
Although the present invention has been described in terms of BCC coding schemes, the present invention can be implemented in other sound processing aspects, in which the sound signal is de-correlated or other sound processing that requires the de-correlation of the signal.
Although the present invention has been described in terms of implementation, the encoder receives the input audio signal in the time domain and generates the transmitted audio signal in the time domain, and the decoder receives the transmitted audio signal in the time domain, and in the time domain The present invention is not limited to this. For example, in other implementations, any one or more of the input, transmitted and played audio signals can be represented in the frequency domain.
BCC encoders and/or decoders can be connected to or incorporated into a variety of different applications or systems, including systems for television or electronic music distribution, movie theaters, broadcasting, streaming, and/or reception. These include The system is used for encoding/decoding transmission via, for example, terrestrial, satellite, cable TV, Internet, internet, or physical media (for example, floppy disk, digital floppy disk, semiconductor chip, hard disk, memory card and similarObject), BCC encoders and/or decoders can also be used in games and gaming systems, including, for example, interactive software products intended to interact with a user for entertainment and/or can be published for multiple machines , Platform or media education, and then BCC encoders and/or decoders can be incorporated into PC software applications. It is a combination of digital decoding (e.g., player, decoder) and a combination of digital encoding capabilities of software applications (e.g., encoding Recorder, recorder, jukebox).
The present invention can be implemented on a circuit-based manufacturing process, including the possible implementation as a single integrated circuit (eg, ASIC or FPGA), a multi-chip module, a single card, or a multi-card circuit set, which is useful for those skilled in the industry. It will be obvious that the various functions of the components can also be implemented as the processing steps of a software program. Such software can also be used in, for example, a digital signal processor, a microcontroller or a general computer.
The present invention can also be embodied in the types of methods and equipment used to implement these methods. The present invention can also be embodied in the types of code contained in physical media, such as floppy disks, CD-ROMs, hard disks or Any other machine-readable storage medium, where when the code is loaded and executed by a machine such as a computer, the machine becomes a device for implementing the present invention. The present invention can also be embodied in the form of code, for example , Whether it is stored in a storage medium, loaded or executed by a machine, or transmitted through some transmission medium or carrier, such as wire or cable, via optical fiber, or via electromagnetic radiation, where, when the program code is carried by a machine such as a computer Enter and execute, the machine becomes a device for implementing the present invention. When implemented on a general processor, the code segment combines with the processor to provide a special device that operates similar to a specific logic circuit.
It will further understand the various changes in the details, materials, and component configurations that have been described and explained to explain the essence of the present invention. Those skilled in the art can achieve this without departing from the scope of the present invention shown in the following patent applications.
Although the steps in the scope of the following method application, if any, can be described in a specific order and corresponding labels, unless the detailed description of the scope of the application also implies a specific order for implementing some or all of these steps, these steps It is not necessarily intended to be limited to being implemented in this particular order.
<p>1. . . Left</p><p>2. . . right</p><p>3. . . central</p><p>4. . . Back left</p><p>5. . . Back right</p><p>100. . . BCC synthesizer</p><p>200. . . BCC system</p><p>202. . . Encoder</p><p>204. . . decoder</p><p>206. . . Downmixer</p><p>208. . . BCC estimator</p><p>300. . . Downmixer</p><p>302. . . Filter library</p><p>304. . . Downmix block</p><p>306. . . Scaled/Delayed Block</p><p>308. . . Reverse FB (IFB)</p><p>310. . . Multiplier</p><p>400. . . BCC synthesizer</p><p>402. . . Filter library</p><p>404. . . Upmix block</p><p>406. . . Retarder</p><p>408. . . Multiplier</p><p>410. . . Related blocks</p><p>412. . . Inverse filter library</p><p>502. . . Filter library</p><p>504. . . Estimated block</p><p>1002. . . TPA</p><p>1004. . . Block</p><p>1006. . . square</p><p>1008. . . Low pass filter</p><p>1102. . . Decode block</p><p>1104. . . TP block</p><p>1106. . . square</p><p>1108. . . Low pass filter</p><p>1110. . . produce</p><p>1112. . . Impose</p><p>1202. . . High pass filter</p><p>1204. . . Dual-band filter library</p><p>1206. . . Dual-band inverse filter library</p><p>1302. . . TPA</p><p>1304. . . Block</p><p>1402. . . Decoded block</p><p>1404. . . TP</p><p>1502. . . LPC across frequency</p><p>1504. . . to parameterize</p><p>1506. . . Quantify</p><p>1508. . . Bit stream</p><p>1602. . . Decoded block</p><p>1604. . . Encapsulation adjuster (TP)</p><p>1606. . . Encapsulation Characteristic Stage (TPA)</p><p>1608. . . Reverse TP (ITP)</p><p>1610. . . Control block</p>
Other viewpoints, features and advantages of the present invention will become more fully apparent from the following detailed description, the scope of the attached patent application and the accompanying drawings, in which the same reference numbers in the accompanying drawings are regarded as similar or identical elements.
Figure 1 shows a high-level block diagram of a conventional two-channel signal synthesizer.
Figure 2 is a block diagram of a general two-channel cue coding (BCC) sound processing system.
Figure 3 shows a block diagram of the downmixer that can be used in Figure 2.
Figure 4 shows a block diagram of a BCC synthesizer that can be used in Figure 2.
Fig. 5 shows a block diagram of the BCC estimator in Fig. 2 according to an embodiment of the present invention.
Figure 6 illustrates the generation of ICTD and ICLD data for five-channel sound.
Figure 7 illustrates the generation of ICC data for five-channel sound.
Figure 8 shows a block diagram of an implementation of the BCC synthesizer of Figure 4, which can be used in a BCC decoder to generate a stereo sound under a single transmitted sum signal s(n) plus a spatial cue signal. Multi-channel audio signal.
Figure 9 illustrates how ICTD and ICLD, as a frequency function, are changed in the primary frequency band.
Fig. 10 shows a block diagram of time-domain processing according to an embodiment of the present invention, which is added to a BCC encoder such as the encoder of Fig. 2.
Fig. 11 illustrates an example of time domain application of TP processing in the BCC synthesizer in Fig. 4.
Figures 12(a) and (b) respectively show the possible implementation of the TPA in Figure 10 and the TP in Figure 11, where the encapsulation shaping is only applied when the frequency is higher than the cut-off frequency fTP.
Fig. 13 shows a block diagram of frequency domain processing added to a BCC encoder such as the encoder of Fig. 2 according to another embodiment of the present invention.
Figure 14 illustrates an example of the frequency domain application of TP processing in the BCC synthesizer of Figure 4.
Figure 15 shows a block diagram of the frequency domain processing added to a BCC encoder such as the encoder of Figure 2 according to another embodiment of the present invention.
Figure 16 illustrates that in the BCC synthesizer of Figure 4, one of the TP processes exemplifies frequency domain applications.
Figures 17(a)-(c) show the block diagrams of possible implementations of ITP and TP in Figure 16 and TPA in Figures 15 and 16.
Figures 18(a) and (b) illustrate the second exemplary mode of operation of the control block in Figure 16.
35 members in 21 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 60620480 | United States of America | – | |
| 62048004 | United States of America | P | |
| 62048004 | United States of America | P | |
| 11006482 | United States of America | – | |
| 648204 | United States of America | A | |
| 648204 | United States of America | A | |
| 11006482 | – | – | – |
| 60620480 | – | – | – |
| US20040006482 | – | – | – |
| US20040620480P | – | – | – |
Members35
| Document | Office | Kind | |
|---|---|---|---|
| US2006083385A1 | United States of America | A1 | |
| AU2005299068A1 | Australia | A1 | |
| CA2582485A1 | Canada | A1 | |
| WO2006045371A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200628001A | Taiwan Province of China | A | |
| NO20071493L | Norway | L | |
| KR20070061872A | Republic of Korea | A | |
| EP1803117A1 | European Patent Office (EPO) | A1 | |
| MX2007004726A | Mexico | A | |
| IL182236A0 | Israel | A0 | |
| CN101044551A | China | A | |
| HK1106861A1 | Hong Kong, China | A1 | |
| JP2008517333A | Japan | A | |
| BRPI0516405A | Brazil | A | |
| BRPI0516405A | Brazil | A | |
| AU2005299068B2 | Australia | B2 | |
| RU2339088C1 | Russian Federation | C1 | |
| EP1803117B1 | European Patent Office (EPO) | B1 | |
| AT424606T | Austria | T | |
| ATE424606T1 | Austria | T1 | |
| DE602005013103D1 | Germany | D1 | |
| PT1803117E | Portugal | E | |
| DK1803117T3 | Denmark | T3 | |
| ES2323275T3 | Spain | T3 | |
| PL1803117T3 | Poland | T3 | |
| KR100924576B1 | Republic of Korea | B1 | |
| TWI318079BThis record | Taiwan Province of China | B | |
| US7720230B2 | United States of America | B2 | |
| JP4664371B2 | Japan | B2 | |
| IL182236A | Israel | A | |
| CN101044551B | China | B | |
| CA2582485C | Canada | C | |
| NO338919B1 | Norway | B1 | |
| BRPI0516405A8 | Brazil | A8 | |
| BRPI0516405B1 | Brazil | B1 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Expiration of patent term of an invention patentMK4A | MK4A |
Numbers
- Publication
- I318079
- Publication, DOCDB
- I318079
- Publication, EPODOC
- TWI318079B
- Application
- 94136500
- Application, DOCDB
- 94136500
- Application, EPODOC
- TW200594136500
Titles4
- Chinese
- 用於雙聲道提示編碼(BBC)方案與類似方案成形之個別頻道
- English
- INDIVIDUAL CHANNEL SHAPING FOR BCC SCHEMES AND THE LIKE
- Unlabeled
- 用於雙聲道提示編碼(BBC)方案與類似方案成形之個別頻道
- Unlabeled
- Individual channels used for the formation of two-channel prompt coding (BBC) schemes and similar schemes
Classification
- CPC, 3
- G10L19/008
- G10L19/02
- H03M7/30
- IPC, 1
- H04S3 00