Audio signal generation
Abstract
An output audio signal (L, R) is generated based on an input audio signal, the input audio signal comprising a plurality of input subband signals (N). The input subband signals are delayed in a plurality of delay units ( 76 ) to obtain a plurality of delayed subband signals, wherein at least one input subband signal is delayed more than a further input subband signal of higher frequency, and wherein the output audio signal is derived ( 77 ) from a combination of the input audio signal and the plurality of delayed subband signals.
Term
Term ended
Projected expiry passed 14 April 2024, 2.4 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
18 claims: 9 independent, 9 dependent
- 1Patent claims Zastrzeżenia patentowe 1. A device for generating the audio output signal (L, R) based on w (^ j; ch ^ i (^ \ w ^ m audio signal containing a set of input subband signals (N), a device comprising:1. Urządzenie do generowania wyjściowego sygnału audio (L, R) bazującego na w(^j;śc^i(^\w^m sygnale audio zawierającym zbiór wejściowych sygnałów podpasmowych (N), urządzenie zawierające: - a set of delay units (76;501 ... 504) for delaying at least a portion of the input subband signals to obtain a set of delayed subband signals, wherein the at least one input subband signal is delayed more than the next higher frequency subband signal, - zbiór jednostek opóźniających (76;501... 504) do opóźniania przynajmniej części wejściowych sygnałów podpasmowych w celu uzyskania zbioru opóźnionych sygnałów podpasmowych, przy czym co najmniej jeden wejściowy sygnał podpasmowy jest opóźniony bardziej niż kolejny sygnał podpasmowy o wyższej częstotliwości, - a connecting unit (77) for providing the output audio signal based on a combination of the input audio signals and a set of delayed subband signals;- jednostkę łączącą (77) do dostarczania wyjściowego sygnału audio na podstawie kombinacji wejściowych sygnałów audio oraz zbioru opóźnionych sygnałów podpasmowych;- przy czym jednostki opóźniające (76;501 . .. 504) są dostosowane do opóźniania co najmniej części wejściowych sygnałów podpasmowych z opóźnieniem wynoszącym całkowitą liczbę próbek podpasmowych i do opóźniania co najmniej jednego z wejściowych sygnałów podpasmowych bardziej niż pozostałych wejściowych sygnałów podpasmowych wyższych częstotliwości, urządzenie ponadto zawiera: - wherein the delay units (76;501 ... 504) are adapted to delay at least a portion of the input subband signals with a delay of the total number of subband samples and to delay at least one of the input subband signals more than the other higher frequency input subband signals, the device further includes: - a partial delay unit (110) delaying at least a portion of the input subband signals with a delay that is a fraction of the time period between two consecutive subband samples, the partial delay may be constant for all or at least part of the input subband signals. - jednostkę opóźnienia cząstkowego (110) opóźniającą przynajmniej część wejściowych sygnałów podpasmowych z opóźnieniem które stanowi ułamek okresu czasu pomiędzy dwoma kolejnymi próbkami podpasmowymi, przy czym opóźnienie cząstkowe może być stałe dla wszystkich lub przynajmniej części wejściowych sygnałów podpasmowych.
- 8Office zniesttz.7, characterized in that bbnk. filtrppppspsowy (77.77) is a bbnk of subband filters with complex transformation. 8. Urząddznieweeług zzsttz.7, znamienne tym, że bbnk. filtrówppodpsmowyyh (77,77) to bbnk filtrów podpasmowych z transformacją zespoloną.
- 10Urządzenie według któregokolwiek z wcześniejszych zastrzeżeń, znamienne tym, że liczba jednostek opóźniających (501...504) jest mniejsza niż liczba wejściowych sygnałów podpasmowych oraz tym, że wejściowe sygnały podpasmowe są dzielone na grupy przypadające na kilka jednostek opóźniających. Ten. A device according to any one of the preceding claims, characterized in that the number of delay units (501 ... 504) is less than the number of input subband signals and in that the input subband signals are divided into groups for several delay units.
- 12Device according to any one of the preceding claims, characterized in that the delay units (501. and. 504) introduce delays that increase monotonically from high to low frequencies. 12. Urządzenie według któregokolwiek z wcześniejszych zastrzeżeń, znamienne tym, że jednostki opóźniające (501. i. 504) wprowadzają opóźnienia, które wzrastają monotonicznie od częstotliwości wysokich do częstotliwości niskich.
- 13A device according to any one of the preceding claims, characterized in that it is adapted to receive a mono audio input signal and to generate a stereo audio output signal. 13. Urządzenie według któregokolwiek z wcześniejszych zastrzeżeń, znamienne tym, że jest dostosowane do odbierania monofonicznego wejściowego sygnału audio oraz generowania stereofonicznego wyjściowego sygnału audio.
- 14A device according to any one of the preceding claims, further comprising an input unit (70) obtaining correlation parameters indicating the desired correlation between the first channel (L) and the second channel (R) of the audio output signal (L, R), the connecting unit (77) being adapted for obtaining the first channel (L) and the second channel (R) by a combination of the input audio signal and a set of delayed subband signals, depending on the correlation parameter. 14. Urządzenie według któregokolwiek z wcześniejszych zastrzeżeń, ponadto zawierające jednostkę wejściową (70) uzyskującą parametry korelacji wskazujące żądaną korelację pomiędzy pierwszym kanałem (L) oraz drugim kanałem (R) wyjściowego sygnału audio (L, R), przy czym jednostka łącząca (77) jest dostosowana do uzyskiwania pierwszego kanału (L) oraz drugiego kanału (R) przez kombinację wejściowego sygnału audio oraz zbioru opóźnionych sygnałów podpasmowych, w zależności od parametru korelacji.
- 16Device according to any one of the preceding claims, further comprising:16. Urządzenie według któregokolwiek z wcześniejszych zastrzeżeń patentowych, ponadto zawierające: - M podpasmowy bank filtrów analizy (72) do generowania M przefiltrowanych sygnałów podpasmowych na podstawie właściwego sygnału audio w dziedzinie czasu, oraz - M subband analysis filter bank (72) for generating M filtered subband signals based on the correct time domain audio signal, and - Generator wysokich częstotiiwości (73p 74) do generowania składowej wysokoczęstotliwościowej wydzielonej z M przefiltrowanych sygnałów podpasmowych, przy czym składowa wysokoczęstotliwościowa zawiera N-M sygnałów podpasmowych, przy czym N>M, N-M sygnałów podpasmowych zawiera sygnały podpasmowe o wyższej częstotliwości niż którekolwiek z podpasm z M podpasm, przy czym M przefiltrowanych sygnałów podpasmowych oraz N-M podpasm tworzy wspólnie zbiór wejściowych sygnałów podpasmowych (N). - High frequency generator (73p 74) for generating the high frequency component separated from M filtered subband signals, the high frequency component containing NM subband signals, where N> M, NM subband signals contain subband signals with a higher frequency than any subband with M subbands , where M filtered subband signals and NM subbands together form a set of input subband signals (N).
- 17A circuit (700) for providing an audio output signal comprising:17. Układ (700) do dostarczania wyjściowego sygnału audio, zawierający: - an input unit (70) obtaining an encoded audio signal, - jednostkę wejściową (70) uzyskującą zakodowany sygnał audio, EP 1 621 047 B1 EP 1 621 047 B1 - a decoder (71) for decoding an encoded audio signal to obtain a decoded audio signal comprising a set of subband signals, - dekoder (71) do dekodowania zakodowanego sygnału audio w celu uzyskania odkodowanego sygnału audio zawierającego zbiór sygnałów podpasmowych, - a device (76, 77, 110) according to any one of the preceding claims for obtaining an audio output signal based on the decoded signal, and - urządzenie (76, 77, 110) według któregokolwiek z wcześniejszych zastrzeżeń do uzyskiwania wyjściowego sygnału audio bazującego na sygnale odkodowanym, oraz - at least one output unit (78, 79) providing an audio output signal. - przynajmniej jedną jednostkę wyjściową (78, 79) dostarczającą wyjściowy sygnał audio.
- 18The method of obtaining audio output signal (L. R) based on the audio input signal, where the input audio signal contains a set of input subband signals (N), the method includes stages in which:18. Sposób uzzskiwania wyjścioweeo syynału audio (L. R) na podstawie wejścioweeo syynału audio, przy czym wejściowy sygnał audio zawiera zbiór wejściowych sygnałów podpasmowych (N), sposób zawiera etapy, w których: - at least a portion of the input subband signals are delayed (501 ... 504) to obtain a set of delayed subband signals, wherein at least one input subband signal is delayed more than the next higher frequency subband signal, - opóźnia się (501 ... 504) przynajmniej część wejściowych sygnałów podpasmowych w celu uzyskania zbioru opóźnionych sygnałów podpasmowych, przy czym przynajmniej jeden wejściowy sygnał podpasmowy jest opóźniany bardziej niż kolejny wejściowy sygnał podpasmowy o wyższej częstotliwości, - at least a portion (110) of the input subband signals is delayed, which is a fraction of the time period between two consecutive subband samples, such partial delay may be constant for all or at least part of the input subband signals, and - opóźnia się (110) przynajmniej część wejściowych sygnałów podpasmowych z opóźnieniem, które stanowi ułamek okresu czasu pomiędzy dwoma kolejnymi próbkami podpasmowymi, przy czym takie opóźnienie cząstkowe może być stałe dla wszystkich lub przynajmniej części wejściowych sygnałów podpasmowych, oraz - określa się (77) wyjściowy sygnał audio na podstawie kombinacji wejściowego sygnału audio oraz zbioru opóźnionych sygnałów podpasmowych. - the (77) audio output signal is determined based on a combination of the input audio signal and a set of delayed subband signals. ΕΡ 1 621 047 Β1 ΕΡ 1 621 047 Β1 Inputs Wejścia PCM fs. PCM fs. FIG.1 FIG.1 Q QMF band filter bank with transf. Μ pasmowy bank filtrów QMF z transf. complex analysis θ £ What zespoloną analizy θ £ c o T © T © I > ·/ O o c T c η ra > I> · / O oc T c η ra> fs/ Nfs/ Nfs / Nfs / N30 fs/Nfs/Nfs/Nfs/N30 FIG.2 FIG.2 M band QMF filter with transf. M pasmowy bankfiltrów QMF z transf. complex synthesis zespoloną syntezy FIG.3 FIG.3 Exit Wyjście PCM PCM ΕΡ 1 621 047 Β1 'VI ο ΕΡ 1 621 047 Β1 'VI ο (Λ (Λ Oh oh Ν Ν 50 100 150 200 50 100 150 200 Czas Time FIG.4 FIG.4 501 50 501 50 Ν input subband signals Ν wejściowych sygnałów podpasmowych Ν de-correlated subband signals Ν dekorelowanych sygnałów podpasmowych FIG. 5 FIG. 5 ΕΡ 1 621 047 Β1 ΕΡ 1 621 047 Β1 FIG. 6 SBR parameters FIG. 6 parametry SBR Kanał lewy Kanał prawy PCM PCM Left channel Right channel PCM PCM L R LR FIG. 7 FIG. 7 ΕΡ 1 621 047 Β1 ΕΡ 1 621 047 Β1 FIG. 8 beginning of transient FIG. 8 początek stanu nieustalonego FIG. 9 FIG. 9 FIG. 10 FIG. 10 EP 1 621 047 B1 EP 1 621 047 B1 Parametry SBR d SBR parameters d S - I S-—I AT- U-
Independent claims9
42 paragraphs, as filed
The object of the invention is to generate an output audio signal based on the input audio signal, and in particular a device providing the output audio signal.
Erik Schuijers, Werner Oomen, Bert den Brinker and Jeroen Breebaart, in the article "Advances in Paramertric Coding for High-Quality Audio", Preprint 5852, 114th AES Convention, Amsterdam The Netherlands, 22-25 March 2003 revealed a parametric coding scheme using effective parametric representation stereo image. Two input signals are combined into a single mono signal. However, relevant to the perception of sensual perception information about spatial distribution is precisely modeled. The combined signal is coded using a mono parametric encoder. Stereo parameters such as inter-channel intensity differences (IID), inter-channel time differences (ITD) and inter-channel cross-correlations (ICC) are quantized, encoded and multiplexed with a bit stream along with a quantized, encoded monophonic audio signal. On the decoder side, the bit stream is demultiplexed to the encoded mono signal and stereo parameters. The encoded monophonic signal is decoded to obtain a decoded monophonic signal m '(see Fig. 1). Based on the monophonic time-domain signal, the de-correlation signal is calculated using a filter D 10 generating a perceptible decor signal. Both signals, the monophonic signal in the time domain m 'and the decorating signal d are transformed into the frequency domain. Then, using the IIC, ITD, and ICC parameters, respectively, using scaling, phase modification and mixing in the parametric processing unit 11, a frequency domain stereo signal is processed to obtain a 1 'and r' decoded stereo pair. The obtained frequency domain representations are again transformed into the time domain.
In the MPEG-4 standard (ISO / IEC 14496-3: 2002) Proposed Changes (PDAM) 2, Section 5.4.6, such de-correlated signals are obtained by convolution / filtering of a monophonic signal with a predetermined impulse response.
Unpublished European Patent Application No. 02077863.5 (reference number PHNL 020639) describes a pass filter, for example a comb filter, comprising a frequency-dependent delay used to determine the decorrelated signal. At high frequencies, a slight real delay is used, giving a higher frequency resolution. At lower frequencies, a large delay is introduced to give a denser comb filter. Filtering can be combined with a band-constraining filter, so that one or more frequency bands are de-correlated.
Patent application US-A-4 039 755 discloses the division of the input signal into frequency bands, which are then subjected to various delays.
The object of the invention is to advantageously generate an output audio signal based on the input audio signal. Thus, the invention introduces a device, method and system as defined in the independent claims. Preferred embodiments are defined in the dependent claims.
According to a first aspect of the invention, the output audio signal is generated based on the input audio signal, wherein the input audio signal comprises a set of input subband signals, at least some of the input subband signals are delayed to obtain a set of delayed subband signals, and at least one input subband signal is delayed more than the next higher frequency subband signal, while the output audio signal is determined based on the combination of the input audio signal and the set of delayed subband signals. By introducing such frequency-dependent delays in the subband frequency domain, it is preferable to implement a parametric stereo signal in those audio decoders in which the main decoder already contains a subband filter bank. Filter banks are commonly used in the context of audio coding, for example MPEG -1/2 Layer I, II and III all of these formats use 32 band critically sampled subband filters. The set of delayed subband signals may be used as an equivalent in the field of subband signals of a de-correlated signal such as that described above. Under ideal conditions, the correlation between a set of delayed subband signals and an input audio signal is zero. However, in practical embodiments, the correlation can be up to 40% at
EP 1 621 047 B1 maintaining an acceptable audio signal quality, and up to 10% for medium or high quality signal and up to 2 or 3% for high quality signals.
In an embodiment of the invention, the output audio signal comprises a set of output subband signals. Combining delayed subband signals and input subband signals in the subband frequency domain to obtain a set of output subband signals is then relatively simple to implement. In practical examples, the time domain audio output is synthesized based on the set of output subband signals in the synthesis subband filter bank.
To achieve effective implementation, a set of delay units is introduced, where the number of delay units is less than the number of input subband signals, and in addition the input subband signals are divided into groups for which a set of delays has been assigned.
The highest quality of the audio signal is obtained in embodiments where the delays in the set of delay units increase monotonically from high frequencies towards low frequencies.
In a preferred embodiment of the invention, a complex transform filter bank is used which is efficiently re-sampled with a factor of two, because for each real input sample a composite output sample is generated that contains the actual two values; real component and imaginary component. Thus, the components resulting from distortion caused by overlapping spectra, which are a problem in critically sampled filter banks used in MPEG-1 and MPEG-2 formats, are eliminated.
In a preferred example of the implementation of the output audio signal generation, a bank of mirror quadrature filters (QMF) was used. Such filter banks are known as such from the article Per Ekstrand, "Bandwidth extension of audio signals by spectral band replication", Proc. 1st IEEE Benelux Workshop on Model based Processing and Coding of Audio (MPCA-2002), pp. 53-58, Leuven , Belgium, November 15, 2002. Figure 2 is a block diagram of QMF analysis and complex synthesis filter banks. The analysis bank 30 divides the signal into N complex subbands, which are then subsampled internally at the factor N. The stylized frequency response is shown in Figure 3. The QMF Synthesis Filter Bank 31 takes N complex subband signals as output signals and generates the actual PCM output signal. According to the developers' observations, when the QMF filter bank with complex transformation is used, it is possible to obtain a de-correlated signal which is perceived as very close to the "ideal" state. For the QMF complex transformation bank, there are implementations that are more efficient than the convolution used in MPEG-4 PDAM 2, Section 5.4.6; convolution is relatively expensive in terms of computational load and memory consumption. Furthermore, advantageously, the use of the complex transform QMF filter bank also allows the use of an effective combination of parametric stereo signal generation techniques and frequency band replication ("SBR") techniques. The idea behind SBR is that higher frequencies can be reproduced from lower frequencies using very little auxiliary information. In practice, this reconstruction is carried out using a bank of quadrature filters with complex transformation (QMF). To effectively obtain a de-correlated signal in the subband frequency domain, embodiments of the invention use frequency-dependent delay (or subband index) in the subband frequency domain. Since the complex QMF filter bank is not critically sampled, no additional measures are required to compensate for the distortion caused by spectral overlaps. Furthermore, because the delay is small, the total RAM demand of such an embodiment is small. It should be noted that in an SBR decoder such as disclosed by Ekstrand, the QMF analysis bank covers only 32 bands, while the QMF synthesis bank covers 64 bands because the specific data decoder operates at half the sampling frequency compared to the whole audio decoder. However, the corresponding encoder uses the QMF analysis bank covering 64 bands covering the full frequency range
Using the total number of delayed subband sample signals as a de-correlated signal causes "blurring" in the time domain, i.e. positioning of the signals in
The time domain is not preserved. This effect contributes to the formation of artifacts around transients, that is, in those cases where the change in signal strength is above a certain threshold. Signal strength can be measured as amplitude, energy, etc. In a preferred embodiment of the invention, the artifacts around transients are reduced by determining a de-correlation signal in a transient environment using partial delays instead of total delays. Partial delays is a delay less than the time between two consecutive subband samples and can be easily implemented using phase shift. The transition between partial delay and total delay and vice versa may result in a de-correlated discontinuity. To prevent such discontinuities, the preferred embodiment of the invention provides a smooth transition between using a partially delayed decorrelated signal and a completely delayed decorrelated signal.
The subject of the invention has been presented in a preferred embodiment in the drawing in which:
Fig. 1 is a block diagram of a parametric stereo decoder;
Fig. 2 is a block diagram of an N band QMF filter bank with complex transformation (left) and synthesis (right);
Fig. 3 shows the stylized frequency response of N band QMF complex transform banks of Fig. 2;
Fig. 4 shows the spectrogram of the impulse response used in the MPEG-4 PDAM 2 scheme, Chapter 5.4.6 for generating a de-correlated signal, with the time (samples) on the x axis and normalized frequencies on the y axis;
Fig. 5 is a block diagram illustrating an apparatus according to an embodiment of the invention;
Fig. 6 shows the delay expressed as subband samples as a function of the subband index according to an embodiment of the invention;
Fig. 7 shows a preferred audio decoder according to an embodiment of the invention that combines parametric decoding and frequency spectrum replication techniques,
Fig. 8 illustrates the occurrence of transient echo caused by mixing of a delayed de-correlated signal;
Fig. 9 shows an example of mixing coefficients, a value of 1 means that the signal is de-correlated with an overall delay, and a value of 0 means that the signal is de-correlated with a partial delay;
Fig. 10 shows the resulting audio output signal generated using the coefficient of Fig. 9; and
Fig. 11 shows the audio decoder of Fig. 7, with another partial delay delay unit.
The drawings show only those elements that are necessary to understand the invention.
In the following, a preferred embodiment of the invention is described for generating a stereo audio output signal based on a mono audio input signal using parametric stereo coding techniques. The input audio signal contains a set of input subband signals. The set of input subband signals is delayed in the set of delay units, providing greater delay for lower frequency subbands than for higher frequency subbands. Delayed subband signals represent in the subband domain the de-correlated signal necessary to generate the stereo output signal.
In the MPEG-4 PDAM 2 schemes, Chapter 5.4.6, the decorrelated signal is obtained by calculating the phase characteristic charakterystyki, which for sampling frequency f<sub>s</sub> of 44.1 kHZ equals:
n (k -1) K (1) where φ<sub>0</sub> has the value π / 2, K is equal to 256, ak = 0 ... 256. Based on the phase response function, the filter impulse response is calculated using the inverse fast Fourier transform FFT. The transformation maintains a linear delay. This delay can be approximated by
EP 1 621 047 B1 (2) where d is the delay in the samples, af is the frequency in radians.
Preferably, the input subband signals are obtained from a complex transform QMF analysis bank that may be in a remote encoder. Because the output of the complex transformation bank QMF is subsampled by the factor N, it is not possible to accurately map the requested time domain delay to the delay for each of the subbands. Approximation acceptable to the recipient can be obtained using rounded versions of the delay function (2) as described above. For example, the delay within each of the subbands for N = 64 subbands is shown in Figure 6. For this particular implementation, only 136 complex values are stored, which allows the decorrelated signal to be formed. It should be noted that for higher frequencies, a delay of a single subband sample is still used, although the delay function, as described above, is set to 0 for half the sampling frequency. The delay of a single subband sample means that the signal is maximally de-correlated.
Fig. 5 shows a block diagram of an apparatus 50 according to an embodiment of the invention for generating a set of delayed subband signals. The device 50 is placed between the QMF analysis filter bank 30 and the QMF synthesis filter bank 31 and includes a set of delay units 501, 502, 503 and 504. The delay unit 501 introduces a single delay unit for all subbands. A group of higher frequency subbands, e.g., 40-64 bands, is sent without further delay to the QMF 31 synthesis filter bank. A group of relatively low frequencies, e.g., 0-40 bands is further delayed in delay unit 502. Part of this group, e.g. bands 0-24 are further delayed in delay unit 503 and delay unit 504 (only bands 0-8 are subject to this delay). Thus, 4 sample groups with different delays are effectively created, with delays of 1,2, 3 or 4 delay units, respectively. The delay expressed in the subband samples as a function of the subband index is shown in Fig. 6. The QMF analysis bank 30 is usually located in the remote encoder, however, in order to use SBR techniques, the smaller M band QMF analysis filter bank can also be used in the decoder.
Fig. 7 shows a preferred audio decoder 700 according to an embodiment of the invention that combines parametric stereo coding techniques and SBR. The demultiplexer 70 receives the bit stream of the encoded audio signal and determines on its basis SBR parameters, parametric stereo coding coefficients and the proper encoded audio signal. The actual encoded audio signal is decoded using the proper decoder 71, which may be a standard MPEG-1 Layer III (mp3) decoder or AAC decoder. Typically, such a decoder operates at half the output sampling frequency (fs / 2). The resulting properly decoded audio signal is fed to the M complex transformation bank QMF 72. Filter ban 72 generates M complex samples per M real input samples and is thus effectively re-sampled by a factor of 2, as described previously. In the high frequency generator 73 (HF), the higher NM frequency subbands that are not covered by the proper audio decoder are generated by replication (specific parts) of the M subbands. The output of the high frequency generator 73 connects to the lower M subbands to form N complex signals subband. Then the envelope regulator 74 adapts the reconstructed high frequency subband signals to the desired envelope, and the additional component 75 adding the additional sinusoidal component and the noise component determined by the SBR parameters. All N subband signals are sent to delay units 76, which may be the same as devices 509 shown in Fig. 5, to generate delayed subband signals. N delayed subband signals and N input subband signals are processed at the combining unit 77 depending on the stereo parameters such as the ICC parameter, obtaining N output subband signals for the first output channel and N output subband signals for the second output channel. N output subband signals for the first output channel are fed to the N band filter bank QMF 78 with complex transformation, generating the first PCM output signals for the left channel L. N output subband signals for the second channel are fed through the N band QMF synthesis 79 filter bank
EP 1 621 047 B1, which generates the first PCM output signals for the right R channel. In practical embodiments, N = 64 and M = 32.
The approach presented above finds its application in stationary signals. However, in the case of non-stationary signals, i.e. transients, problems arise when using this approach. This situation is shown in Fig. 8, the figure shows a single channel signal in the form of a castanet signal obtained using the decelerated signal with the total delay in Figs. 5 and 6 as the basis for generating the audio output signal. Typically, in a signal with strong transients, for example a castanet sound, the correlation between the left and right channels immediately after the transient is relatively low, because the signal is mainly composed of reverberation. The de-correlation signal is therefore largely mixed. This leads to the creation of a pure echo appearing immediately after the real transient castanet sound. Although the result of masking in the time domain, this effect is not seen as another transient state, it still contributes to the undesirable sound discoloration effect. In a preferred embodiment of the invention, these types of artifacts are mitigated by the creation of a de-correlated signal in a transient environment using a partial delay. Partial delay can be implemented efficiently using phase shift. In another embodiment, to prevent discontinuities in the entire de-correlated signal, the partially-delayed signal or phase-shifted (slowly) signal smoothly over time passes into the total-de-correlated signal.
Thus, it is proposed to use a partially delayed or phase-shifted version of the original signal instead of a completely delayed signal in a frequency-dependent manner, starting from where the transient occurs. Due to the masking properties of the human hearing aid, it is not particularly important how the de-correlated signal is determined. Thus, the de-correlated signal can be obtained by applying a 90 degree phase shift in each of the original signal subbands.
To prevent discontinuities in the de-correlated signal starting from the location of the transient, smooth transitions between the use of the completely delayed signal and the phase-shifted signal are preferably used. A smooth transition can be made as d:, \<sup>n</sup>\ = <sup>m</sup>\ .n \ ddelay \ n \ + <sup>(</sup>1- m \ n<sup>])</sup>drotation \ n \ where n is the sample index (subband), m [n] is the mixing coefficient or smooth transition, d<sub>d</sub>El<sub>and</sub>y [n] is a de-correlated signal (subband) created using a frequency-dependent total delay, d<sub>r</sub>ot<sub>and</sub>thion [n] is a de-correlated signal (subband) created using partial delay or phase shift and dhybrid [n] is the resulting hybrid de-correlated signal. The mixing coefficient m [n] is zero when the transient is started. Then its value remains at zero for a period of time typically corresponding to about 20 ms (about 12 ms corresponds to the length of the delay and 8 ms is the transient). The transition between zero and one typically occurs between 10 and 20ms. The function of the mixing factor m [n] may or may not be limited to a linear function or a function periodically with pieces of linear. It should be noted that the mixing factor m [n] may also depend on the frequency. Since the delay is typically shorter for higher frequencies, it is preferable, due to the perception of the output signal, to keep the mixing times shorter for higher frequencies than for lower frequencies.
Fig. 11 shows the audio decoder of Fig. 7, it uses a partial delay unit 110 with partial delays to determine partial delayed subband signals. Delay unit 76 generates subband signals with a frequency-dependent delay. In practice, the partial delay delay unit 110 may work in parallel with the delay units 76, however it is also possible to turn off the next delay unit 110 when the delay units 76 are working and vice versa. Preferably, switching is performed between the partially delayed subband signals and the frequency dependent delay subband signals in the switching unit 111. The switching unit preferably performs the smooth switching operation as described above, but it is possible to switch hard. Smooth switching depends on
EP 1 621 047 B1 detects the presence of a transient condition. The transient detection is preferably performed in the transient detector 113. Alternatively, it is possible for the encoder to enter switching markers into the encoded bit stream. Then, the bit stream demultiplexer 70 extracts the markers from the bit stream and supplies the switching markers to the switching unit 111, the switching then being dependent on the presence of the switching marker.
It should be noted that the embodiments described above are illustrative and not limiting the invention itself, moreover, one of ordinary skill in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, reference marks in brackets do not constitute a limitation of the claims. The word 'contains' does not exclude the presence of other elements or stages than those mentioned in the claim. The invention may be implemented by hardware means comprising a number of different elements, or by means of an appropriately programmed computer. In the claim regarding the device, listing a number of elements, some of these elements may be made as a single hardware element. The mere fact that specific means are mentioned in mutually different dependent claims does not mean that a combination of these features cannot be advantageously used. The scope of patent protection is determined only by the attached patent claims.
38 members in 12 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 03076134 | European Patent Office (EPO) | A | |
| 03076280 | European Patent Office (EPO) | A | |
| 04727354 | European Patent Office (EPO) | A | |
| 2004050432 | International Bureau of the World Intellectual Property Organization (WIPO) | W | |
| EP20030076134 | – | – | – |
| EP20030076280 | – | – | – |
| EP20040727354 | – | – | – |
| WO2004IB50432 | – | – | – |
Members38
| Document | Office | Kind | |
|---|---|---|---|
| WO2004093494A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2004093495A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20050121733A | Republic of Korea | A | |
| KR20050122267A | Republic of Korea | A | |
| EP1618763A1 | European Patent Office (EPO) | A1 | |
| EP1621047A1 | European Patent Office (EPO) | A1 | |
| AT355590T | Austria | T | |
| ATE355590T1 | Austria | T1 | |
| RU2005135648A | Russian Federation | A | |
| RU2005135650A | Russian Federation | A | |
| BRPI0409327A | Brazil | A | |
| BRPI0409337A | Brazil | A | |
| CN1774956A | China | A | |
| CN1774957A | China | A | |
| JP2006523859A | Japan | A | |
| JP2006524002A | Japan | A | |
| US2007038439A1 | United States of America | A1 | |
| EP1618763B1 | European Patent Office (EPO) | B1 | |
| EP1621047B1 | European Patent Office (EPO) | B1 | |
| DE602004005020D1 | Germany | D1 | |
| AT359687T | Austria | T | |
| ATE359687T1 | Austria | T1 | |
| US2007112559A1 | United States of America | A1 | |
| DE602004005846D1 | Germany | D1 | |
| PL1618763T3 | Poland | T3 | |
| PL1621047T3This record | Poland | T3 | |
| ES2281795T3 | Spain | T3 | |
| ES2282860T3 | Spain | T3 | |
| DE602004005020T2 | Germany | T2 | |
| DE602004005846T2 | Germany | T2 | |
| JP4597967B2 | Japan | B2 | |
| KR20110044281A | Republic of Korea | A | |
| CN1774956B | China | B | |
| JP4834539B2 | Japan | B2 | |
| KR101169596B1 | Republic of Korea | B1 | |
| KR101200776B1 | Republic of Korea | B1 | |
| US8311809B2 | United States of America | B2 | |
| BRPI0409327B1 | Brazil | B1 |
Numbers
- Publication, DOCDB
- 1621047
- Publication, EPODOC
- PL1621047T
- Application
- 727354
- Application, DOCDB
- 04727354
- Application, EPODOC
- PL20040727354T
Titles2
- English
- AUDIO SIGNAL GENERATION
- Polish
- Generowanie sygnału audio
Classification
- CPC, 5
- G10L19/00
- H04S5/00
- H04S1/007
- H04S2420/03
- H04S1/00
- IPC, 3
- H04S5 00
- G10L19 00
- H04S1 00