Audio encoding/decoding scheme having a switchable bypass
Abstract
This record has no abstract on file.
Term
2.8 yearsto projected expiry
Projected expiry 6 July 2029, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
19 claims: 12 independent, 7 dependent
- 1Patent claims Zastrzeżenia patentowe 1. An apparatus for encoding an audio signal to obtain an encoded audio signal, which audio signal is in the first domain, comprising:1. Urządzenie do kodowania sygnału audio do uzyskania zakodowanego sygnału audio, który to sygnał audio jest w domenie pierwszej, zawierające: a first domain converter (510) for converting the audio signal from the first domain to the second domain;pierwszy konwerter domeny (510) do konwersji sygnału audio z domeny pierwszej do domeny drugiej;przełączalne obejście (50) do omijania pierwszego konwertera domeny (510), jeśli przełączalne obejście jest ustawione w stanie aktywnym, lub do spowodowania konwersji sygnału audio przez pierwszy konwerter domeny (510), jeśli przełączalne obejście ustawione jest w stanie nieaktywnym w odpowiedzi na sygnał sterujący (51) obejściem;switchable bypass (50) to bypass the first domain converter (510) if the switchable bypass is set in the active state, or to cause the audio signal to be converted by the first domain converter (510) if the switchable bypass is set in the inactive state in response to the control signal (51) bypass;a second domain converter (410) for converting the audio signal received from the switchable bypass (50) or first domain converter (510) to a third domain, which third domain is different from the second domain;drugi konwerter domeny (410) do konwersji sygnału audio, odebranego z przełączalnego obejścia (50) lub pierwszego konwertera domeny (510), do domeny trzeciej, która to domena trzecia jest różna od domeny drugiej;a first processor (420) for encoding the third-domain audio signal according to the first encoding algorithm;and a second processor (520) for encoding the audio signal received from the first domain converter (510) if the switchable bypass is set in an inactive state, according to a second encoding algorithm, different from the first encoding algorithm, to obtain a second processed signal, encoded the signal in the audio signal portion includes either the first processed signal or the second processed signal. pierwszy procesor (420) do kodowania sygnału audio w domenie trzeciej zgodnie z pierwszym algorytmem kodowania;i drugi procesor (520) do kodowania sygnału audio odebranego z pierwszego konwertera domeny (510), jeśli przełączalne obejście ustawione jest w stanie nieaktywnym, zgodnie z drugim algorytmem kodowania, różnym od pierwszego algorytmu kodowania, w celu uzyskania drugiego przetworzonego sygnału, przy czym zakodowany sygnał w części sygnału audio zawiera albo pierwszy przetworzony sygnał, albo drugi przetworzony sygnał.
- 4An apparatus according to one of the preceding claims, in which the second processor (520) operates to generate an encoded output signal such that the encoded output signal is in the same domain as the input signal to the second processor (520). 4. Urządzenie według jednego z poprzednich zastrzeżeń, w którym drugi procesor (520) działa w celu generowania zakodowanego sygnału wyjściowego, tak że zakodowany sygnał wyjściowy jest w tej samej domenie co sygnał wejściowy do drugiego procesora (520).
- 5An apparatus according to one of the preceding claims, in which the first processor (420) comprises a quantizer and an entropy encoder, and wherein the second processor (520) comprises a source coder based coder. 5. Urządzenie według jednego z poprzednich zastrzeżeń, w którym pierwszy procesor (420) zawiera kwantyzator i koder entropijny, oraz w którym drugi procesor (520) zawiera koder źródła oparty na książce kodowej.
- 6An apparatus according to one of the preceding claims, in which the first processor (420) is based on the information outlet model and the second processor (520) is based on the information source model. 6. Urządzenie według jednego z poprzednich zastrzeżeń, w którym pierwszy procesor (420) oparty jest na modelu ujścia informacji, zaś drugi procesor (520) oparty jest na modelu źródła informacji.
- 7The device according to one of the preceding claims, further comprising a switching member (200) switched between the output of the first domain converter (510) and the input of the second domain converter (410) and the input of the second processor (520), the switching member (200) being adapted for switching between the input of the second domain converter (410) and the input of the second processor (520) in response to the control signal of the switching member. 7. Urządzenie według jednego z poprzednich zastrzeżeń, zawierające ponadto człon przełączający (200) włączony między wyjściem pierwszego konwertera domeny (510) i wejściem drugiego konwertera domeny (410) i wejściem drugiego procesora (520), przy czym człon przełączający (200) przystosowany jest do przełączania między wejściem drugiego konwertera domeny (410) i wejściem drugiego procesora (520) w odpowiedzi na sygnał sterujący członu przełączającego.
- 8The device according to one of the preceding claims, in which the switchable bypass output (50) is connected to the output of the first domain converter (510) and the switchable bypass input (50) is connected to the input of the first domain converter (510). 8. Urządzenie według jednego z poprzednich zastrzeżeń, w którym wyjście przełączalnego obejścia (50) połączone jest z wyjściem pierwszego konwertera domeny (510), zaś wejście przełączalnego obejścia (50) połączone jest z wejściem pierwszego konwertera domeny (510).
- 9The device according to one of the preceding claims, further comprising a signal classifier for controlling the switchable bypass (50) in the audio signal portion depending on the result of the analysis for that audio signal portion. 9. Urządzenie według jednego z poprzednich zastrzeżeń, zawierające ponadto klasyfikator sygnału do sterowania przełączalnym obejściem (50) w części sygnału audio w zależności od wyniku analizy dla tej części sygnału audio.
- 10An apparatus according to one of the preceding claims, in which the second domain converter (410) operates to convert the input signal in blocks, and in which the second domain converter operates to switch blocks in response to audio signal analysis, so that the second domain converter (410) is so controlled that blocks of different lengths are converted depending on the content of the audio signal. 10. Urządzenie według jednego z poprzednich zastrzeżeń, w którym drugi konwerter domeny (410) działa w celu konwersji sygnału wejściowego blokami, oraz w którym drugi konwerter domeny działa w celu przełączania bloków w odpowiedzi na analizę sygnału audio, tak że drugi konwerter domeny (410) jest tak sterowany, że bloki o różnych długościach poddawane są konwersji w zależności od zawartości sygnału audio.
- 11A method of encoding an audio signal to obtain an encoded audio signal, which audio signal is in the first domain, comprising:11. Sposób kodowania sygnału audio w celu uzyskania zakodowanego sygnału audio, który to sygnał audio jest w domenie pierwszej, obejmujący: converting (510) the audio signal from the first domain to the second domain;bypassing (50) the conversion step (510) of the audio signal from the first domain to the second domain or causing conversion of the audio signal from the first domain to the second domain in response to the control signal (51) of the bypass switch;konwersję (510) sygnału audio z domeny pierwszej do domeny drugiej;ominięcie (50) etapu konwersji (510) sygnału audio z domeny pierwszej do domeny drugiej lub spowodowanie konwersji sygnału audio z domeny pierwszej do domeny drugiej w odpowiedzi na sygnał sterujący (51) przełącznika obejścia;converting (410) the audio signal from the bypass (50) in the second domain to a third domain, wherein the third domain is different from the second domain;encoding (420) the third domain audio signal generated in the conversion step (410) of the bypass audio signal (50) or the second domain audio signal according to the first encoding algorithm;and encoding (520) the second domain's audio signal if the bypass (50) is not activated according to a second encoding algorithm, different from the first encoding algorithm, to obtain a second processed signal in which the encoded signal includes or first processed part of the audio signal signal or second processed signal. konwersję (410) sygnału audio z obejścia (50) w domenie drugiej do domeny trzeciej, przy czym domena trzecia jest różna od domeny drugiej;kodowanie (420) sygnału audio domeny trzeciej generowanego w etapie konwersji (410) sygnału audio z obejścia (50) lub sygnału audio w domenie drugiej, zgodnie z pierwszym algorytmem kodowania;i kodowanie (520) sygnału audio domeny drugiej, jeśli obejście (50) nie zostało aktywowane zgodnie z drugim algorytmem kodowania, różnym od pierwszego algorytmu kodowania, w celu uzyskania drugiego przetworzonego sygnału, w którym zakodowany sygnał, w części sygnału audio zawiera albo pierwszy przetworzony sygnał, albo drugi przetworzony sygnał.
- 12An apparatus for decoding an encoded audio signal, which encoded signal comprises a first processed signal located in a third domain and a second processed signal located in a second domain, wherein the second domain and the third domain are different, comprising:12. Urządzenie do dekodowania zakodowanego sygnału audio, który to zakodowany sygnał zawiera pierwszy przetworzony sygnał znajdujący się w domenie trzeciej i drugi przetworzony sygnał, znajdujący się w domenie drugiej, przy czym domena druga i domena trzecia są od siebie różne, zawierające: a first inverse processor (430) for inverse processing the first processed signal;pierwszy procesor odwrotny (430) do odwrotnego przetwarzania pierwszego przetworzonego sygnału;a second reverse processor (530) for inversely processing the second processed signal;drugi procesor odwrotny (530) do odwrotnego przetwarzania drugiego przetworzonego sygnału;a second converter (440) for converting the domain of the first inversely processed signal from the third domain to another domain;drugi konwerter (440) do konwersji domeny pierwszego odwrotnie przetworzonego sygnału z domeny trzeciej do domeny innej;a first converter (540) for converting the second inverse processed signal to the first domain or for converting the first inverse processed signal that has been converted to another domain to the first domain when the other domain is not the first domain;and a bypass (52) to bypass the first converter (540) when the other domain is the first domain. pierwszy konwerter (540) do konwersji drugiego odwrotnie przetworzonego sygnału do domeny pierwszej lub do konwersji pierwszego odwrotnie przetworzonego sygnału, który został przekonwertowany do innej domeny, do domeny pierwszej, kiedy domena inna nie jest domeną pierwszą;i obejście (52) do omijania pierwszego konwertera (540), kiedy domena inna jest domeną pierwszą.
- 14The decoding device according to any of claims 12 or 13, further comprising an input interface (900) for extracting from the encoded audio signal, the first processed signal, the second processed signal and the control signal indicating whether for some first inverse processed signal, the first converter (540) is to be bypassed or not. 14. Urządzenie do dekodowania według któregokolwiek z zastrzeżeń 12 albo 13, zawierające ponadto interfejs wejściowy (900) do ekstrakcji z zakodowanego sygnału audio, pierwszego przetworzonego sygnału, drugiego przetworzonego sygnału i sygnału sterującego wskazującego, czy dla pewnego pierwszego odwrotnie przetworzonego sygnału, pierwszy konwerter (540) ma zostać ominięty przez obejście, czy nie.
- 18A method of decoding an encoded audio signal, which encoded audio signal comprises a first processed signal in a third domain and a second processed signal in a second domain, wherein the second domain and the third domain are different, including:18. Sposób dekodowania zakodowanego sygnału audio, który to zakodowany sygnał audio zawiera pierwszy przetworzony sygnał znajdujący się w domenie trzeciej i drugi przetworzony sygnał znajdujący się w domenie drugiej, przy czym domena druga i domena trzecia są od siebie różne, obejmujący: inverse processing (430) the first processed signal;odwrotne przetwarzanie (430) pierwszego przetworzonego sygnału;inverse processing (530) the second processed signal;odwrotne przetwarzanie (530) drugiego przetworzonego sygnału;a second domain conversion (440) of the first inversely processed signal from the third domain to another domain;drugą konwersję domeny (440) pierwszego odwrotnie przetworzonego sygnału z domeny trzeciej do domeny innej;pierwszą konwersję domeny (540) drugiego odwrotnie przetworzonego sygnału do domeny pierwszej lub konwersję pierwszego odwrotnie przetworzonego sygnału, który był przekonwertowany do domeny innej, do domeny pierwszej, kiedy domena inna nie jest domeną pierwszą;oraz ominięcie (52) etapu pierwszej konwersji domeny (540), kiedy domena inna jest domeną pierwszą. the first conversion of the domain (540) of the second inverse processed signal to the first domain or the conversion of the first inverse processed signal that was converted to another domain into the first domain when the other domain is not the first domain;and bypassing (52) the first domain conversion step (540) when the other domain is the first domain.
Independent claims12
200 paragraphs in 1 section, as filed
[0001] The present invention relates to encoding of an audio signal and, in particular, methods of encoding a low bit rate audio signal.
[0002] Frequency domain coding methods such as MP3 or AAC are known in the art. Such frequency domain coders are based on conversion from a time domain to a frequency domain, followed by a quantization stage in which the quantization error is controlled using information from the psychoacoustic module, and a coding stage in which the quantized spectral coefficients along with the corresponding additional information coded are entropy using code tables.
[0003] On the other hand, there are encoders very well suited for speech signal processing, such as the AMR-WB + encoder described in the 3GPP TS 26.290 standard. Such speech coding algorithms perform Linear Prediction (LP) of the time domain signal. Such LP filtering is obtained from linear predictive analysis of the input time domain signal. The obtained LP filter coefficients are then coded and sent as additional information. This algorithm is known as Linear Prediction Coding (LPC). At the filter output, the prediction signal or prediction error signal, also known as the excitation signal, is encoded using the analysis steps by synthesis of the ACELP encoder, or alternatively, it is encoded using a transformer encoder using Fourier transform with pleats. The choice between ACELP coding and excitation transform coding, also called TCX coding, is done using either a closed loop algorithm or an open loop algorithm.
[0004] Frequency domain audio coding algorithms, such as the high performance AAC coding algorithm, which combines the AAC coding algorithm with the spectral band replication technique, can also be combined in a combined stereo coding or multi-channel coding tool, which is known as " MPEG surround. "
[0005] On the other hand, speech signal coders such as AMR-WB + also include a high frequency enrichment step, and a stereo function.
[0006] Frequency domain coding algorithms are advantageous in that they offer high quality for music signals at a low bit rate. However, the quality of speech signals at low bit rates is problematic.
[0007] Speech signal coding algorithms offer high quality even at low bit rates, but poor quality of music signals at such low bit rates.
[0008] WO 2008/071353 A2 discloses a hybrid speech and music codec combining time domain and frequency domain coding / decoding. Domain switching takes place depending on whether the audio signal is more like speech or music. Switchable bypassing of overlap-add synthesis (used in frequency domain coding) occurs in the decoder if the signal to be decoded has been encoded in the time domain.
[0009] The object of the present invention is to provide an improved coding / decoding concept.
[0010] This object is achieved by an audio coding apparatus according to claim 1, a method of coding an audio signal according to claim 11, a device for decoding an encoded audio signal according to claim 12, a method of decoding an encoded audio signal according to claim 18, or by computer program in accordance with claim 19.
[0011] In the encoder according to the present invention, two domain converters are used, wherein the first domain converter processes the audio signal from the first domain, such as time domain, to the second domain, such as LPC domain. The second domain converter works to process from the input domain to the output domain and the second domain converter receives, as input, the output signal from the first domain converter or the output signal of the switchable bypass which is combined to bypass the first domain converter. In other words, this means that the second domain converter receives, as its input, the audio signal of a first domain such as a time domain, or alternatively, the output signal of a first domain converter, i.e. an audio signal that has already been processed from one domain to another domain. The output of the second domain converter is processed by the first processor to generate the first processed signal and the output of the first domain converter is processed by the second processor to generate the second processed signal. Preferably, the switchable bypass may also be connected to a second processor such that the input to the second processor is a time domain audio signal rather than the output of the first domain converter.
[0012] This extremely flexible coding concept is particularly useful in high quality and bit rate coding of an audio signal because it enables coding of an audio signal in at least three different domains, and when the switchable bypass is additionally connected also to a second processor, even in four domains. This can be achieved by controlling switchable bypass to bypass or bridge the first domain converter over a portion of the time domain audio signal, or no bypass or bridging. Even if the first domain converter is omitted, there are still two different options for encoding the time domain audio signal, e.g., through the first processor connected to the second domain converter, or through a second processor.
[0013] Preferably, the first processor and the second domain converter form together an information output model encoder, such as a psychoacoustic controlled audio encoder known from PMEG 1 Layer 3 or MPEG 4 (AAC).
[0014] Preferably, the second encoder, i.e. the second processor, is a time-domain encoder, which is, for example, a residual encoder known from the ACELP encoder, where the LPC residual signal is encoded using a residual encoder, such as a vector quantization coder for the residual signal LPC or time domain signal. In a variant of the invention, this time domain encoder receives the LPC domain signal as input when the bypass is open. Such an encoder is an encoder for an information source model because, unlike the encoder for an information outlet model, the information source model encoder is specifically designed to utilize the specificity of a speech generation model. However, when the bypass is closed, the input signal to the second processor will be a time domain signal, rather than an LPC domain signal.
[0015] However, if the switchable bypass is deactivated, which means that the audio signal from the first domain is converted to the second domain before further processing, there are still two different possibilities, i.e. either encoding the output of the first domain converter in the second domain, which may for example, being an LPC domain, or alternatively, processing a signal in a second domain to a third domain, which may for example be a spectral domain.
[0016] Preferably, the spectral domain converter, i.e. the second domain converter, is adapted to implement the same algorithm, regardless of whether the input signal of the second domain converter is in a first domain, such as a time domain, or a second domain, such as as LPC domain.
[0017] On the decoder side, there are two different decoding paths, where one decoding path includes a domain converter, i.e. a second domain converter, while the other decoding path only includes an inverse processor, but does not include a domain converter. Depending on the current bypass setting on the encoder side, i.e. whether the bypass was active or not, the first converter in the decoder is bypassed or not. In particular, the first converter in the decoder is bypassed when the output of the second converter is already in the target domain, such as the first domain or the time domain. However, if the output of the second converter in the decoder is in a domain other than the first domain, the decoder bypass is deactivated and the signal is converted from another domain to the destination domain, i.e. to the first domain in the preferred embodiment. The second signal is processed, in one variant of the invention, in the same domain, i.e. in the second domain, but in other variants of the invention in which the switchable encoder side bypass is also connectable to the second processor, the output of the second reverse processor on the decoder side can also already be in the first domain. In this case, the first converter is bypassed using a switchable decoder side bypass, so that the decoder output combiner receives input signals that represent different parts of the audio signal and that are in the same domain. These signals can be time multiplexed by the combinator or they can be mutually penetrated by the decoder output combinator.
[0018] In a preferred embodiment of the invention, the coding apparatus comprises a typical preprocessing member for compressing the input signal. This typical pre-processing member may include a multi-channel processor and / or a spectrum band replication processor, so that the output of a typical pre-processing member for all different encoding modes is a compressed version relative to the input to the pre-processing member. Thus, the output signal on the decoder combinator side can be post-processed in a typical post-processing member that, for example, works to perform the synthesis of spectral band replication and / or multi-channel expansion operations, such as multi-channel upmix operation, which is preferably controlled using parametric multi-channel information sent from the encoder page to the decoder page.
In a preferred embodiment of the invention, the first domain in which the input signal input into the encoder and the output signal output from the decoder are located is the time domain. In a preferred embodiment, the second domain in which the output of the first domain converter is located is an LPC domain, such that the first domain converter is a member of the LPC analysis. In another embodiment of the invention, the third domain, i.e. the domain in which the output of the second domain converter is located is the spectral domain or is the spectral domain of the signal in the LPC domain generated by the first domain converter. The first processor attached to the second domain converter is preferably implemented as an information output coder, such as a quantizer / converter with an entropy reduction encoder, such as a psychoacoustic controlled quantizer connected to a Huffman encoder or arithmetic encoder, which perform the same functions, regardless of whether the input signal is in the spectral domain or in the LPC spectral domain.
[0020] In another embodiment of the invention, the second processor for processing the output of the first domain converter or for processing the switchable bypass output in a device with full functionality is a time domain encoder, such as a residual signal encoder, used in the ACELP encoder or in any other CELP encoder .
[0021] Preferred variants of the present invention are described below with reference to the accompanying drawings, in which:
Fig. 1a is a block diagram of an encoding method in accordance with the first aspect of the present invention;
Fig. 1b is a block diagram of a decoding method in accordance with the first aspect of the present invention;
Fig. 1c is a block diagram of a coding method in accordance with another aspect of the present invention;
Fig. 1d is a block diagram of a decoding method in accordance with another aspect of the present invention;
Fig. 2a is a block diagram of a coding method in accordance with the second aspect of the present invention;
Fig. 2b is a block diagram of a decoding method in accordance with the second aspect of the present invention;
Fig. 2c is a block diagram of the preferred co-processing of Fig. 2a; and
Fig. 2d is a block diagram of the preferred co-processing of Fig. 2b;
Fig. 3a is a block diagram of a coding method in accordance with another aspect of the present invention;
Fig. 3b is a block diagram of a decoding method in accordance with another aspect of the present invention;
Fig. 3c is a schematic representation of a coding device / method with cascade switches;
Fig. 3d shows a block diagram of an apparatus or method for decoding in which cascade combinators are used;
Fig. 3e is an example of a time domain signal and a corresponding representation of the encoded signal illustrating the short penetration areas that are included in both encoded signals;
Fig. 4a is a block diagram with a switch positioned in front of the coding paths;
Fig. 4b is a block diagram of an encoding method with a switch positioned immediately after the encoding paths;
Fig. 4c is a block diagram of a preferred combinator variant;
Fig. 5a shows the wave form of a time-domain speech signal segment as a quasi-periodic signal segment or pulse type signal segment;
Fig. 5b shows the spectrum of the segment of Fig. 5a;
Fig. 5c illustrates a time domain speech signal segment for unvoiced speech as an example of a noise type segment or a fixed segment;
Fig. 5d shows the time domain wave spectrum of Fig. 5c;
Fig. 6 is a block diagram of analysis by CELP encoder synthesis;
Figures 7a to 7d show voiced and voiceless excitation signals as examples of pulse type and solid type signals;
Fig. 7e illustrates the coder side LPC member providing short-term prediction information and the prediction error signal;
Fig. 7f shows a further variant of the LPC device for generating a weighted signal;
Fig. 7g illustrates the implementation for converting a weighted signal into an excitation signal by applying an inverse weighing operation followed by an analysis of the excitation in accordance with the requirements of the converter 537 in Fig. 2b;
Fig. 8 is a block diagram of a combined multi-channel algorithm in accordance with an embodiment of the present invention;
Fig. 9 shows a preferred variant of the bandwidth extension algorithm;
Fig. 10a shows a detailed description of the switch when making an open loop decision; and
Fig. 10b shows an example of a switch operating in closed-loop decision mode [0022] Fig. 1a shows a variant of the invention in which there are two domain converters 510, 410 and a switchable bypass 50. The switchable bypass 50 is adapted to be in an active or inactive state in response on control signal 51 which is input to the switching control input of the switchable bypass 50. If switchable bypass is active, the audio signal at input 99, 195 of the audio signal is not supplied to the first domain converter 510, but is provided to switchable bypass 50, so that the second domain converter 410 receives the audio signal directly at input 99, 195. W one embodiment of the invention which will be discussed with reference to Fig. 1c and 1d, the switchable bypass 50 is alternatively connectable to the second processor 520 without connecting to the second domain converter 410, so that the output signal of the switchable bypass 50 is processed only by the second processor 520.
[0023] However, if the switchable bypass 50 is set to inactive by means of control signal 51, the audio signal at the output 99 or 195 of the audio signal is fed into the first domain converter 510 and is, at the output of the first domain converter 510, or fed into the second converter domain 410, or to the second processor 520. The decision whether the output of the first domain converter is input to the second domain converter 410 or to the second processor 520 is preferably also based on the switching control signal, but may alternatively be taken by other means such as metadata or based on signal analysis. Alternatively, the signal of the first domain converter 510 can even be input to both devices 410, 520, and the choice of which processing signal is input to the output interface to present the signal over a certain portion of time is made using a switch connected between the processors and the output interface, discussed with reference to Fig. 4b. On the other hand, the decision as to which signal is to be input into the output data stream can also be made in the output interface 800 itself.
[0024] As shown in Fig. 1a, a device according to the invention for encoding an audio signal to obtain an encoded audio signal, wherein the audio signal at the 99/195 input is in the first domain, comprises a second domain converter for converting the audio signal from the first domain into second domain. In addition, a switchable bypass 54 is provided, bypassing the first domain converter 510, or operating to cause the conversion of the audio signal by the first domain converter in response to the bypass switching control signal 51. Thus, in the active state, the switchable bypass bypasses the first domain converter, and in inactive state, the audio signal is fed into the first domain converter.
[0025] Furthermore, a second domain converter 410 is provided for processing the audio signal received from the switchable bypass 50 or from the first domain converter to the third domain. The third domain is different from the second domain. In addition, a first processor 420 is provided for encoding the third domain audio signal according to the first encoding algorithm to obtain the first processed signal. In addition, a second processor 520 is provided for encoding the audio signal received from the first domain converter according to the second coding algorithm, the second coding algorithm being different from the first coding algorithm. The second processor provides the second processed signal. In particular, the device is adapted to have at its output an encoded audio signal for a portion of the audio signal in which the encoded signal either comprises a first processed signal or a second processed signal. There may be demarcation areas, of course, but with the aim of improving coding performance, the goal is to keep the demarcation areas as small as possible and eliminate them whenever possible to achieve maximum bit rate compression.
[0026] Fig. 1b shows a decoder corresponding to the encoder of Fig. 1a in a preferred embodiment. The apparatus for decoding an encoded audio signal in Fig. 1b receives, as an input, an encoded audio signal comprising a first processed signal in a third domain and a second processed signal in a second domain, wherein the second domain and the third domain are different from each other. In particular, the signal input into the input interface 900 is similar to the output from the interface 800 of Fig. 1a. The decoding apparatus includes a first inverse processor 430 for inverse processing of the first processed signal and a second inverse processor 530 for inverse processing of the second processed signal. In addition, a second converter 440 is provided for converting the domain of the first inversely processed signal from the third domain to another domain. In addition, a first converter 540 is provided for converting the second inverse processed signal into the first domain or for converting the first inverse processed signal to the first domain when the other domain is not the first domain. This means that the first inversely processed signal is processed by the first converter when the first processed signal is no longer in the first domain, i.e. in the target domain in which the audio signal or intermediate audio signal is to be decoded for the pre-processing / post-processing circuit. In addition, the decoder includes bypass 52 to bypass the first converter 540 when the other domain is the first domain. The circuit in Fig. 1b further includes a combiner 600 for connecting the output from the first converter 540 and the bypass output, i.e. a signal output by bypass 52 to obtain a total decoded audio signal 699 that can be used as it is, or which can even be decompressed using a typical post processing member, as will be discussed later.
[0027] Fig. 1c shows a preferred variant of the audio encoder according to the invention, in which the signal classifier in the psychoacoustic model 300 is provided to classify the audio signal introduced into the typical pre-processing member formed by the MPEG Surround encoder 101 and the processor 102 of improved spectral band replication. In addition, the first domain converter 510 is an LPC analysis element, and the switchable bypass is connected between the input and output of the LPC analysis element 510 which is the first domain converter.
[0028] The LPC device generally outputs an LPC domain signal, which can be any LPC domain signal, such as the excitation signal in Fig. 7e or the weighted signal in Fig. 7f, or any other signal that has been generated by applying LPC filter coefficients for audio signal. In addition, the LPC device can also determine these coefficients and can also quantize / code these coefficients.
[0029] In addition, a switch 200 is provided at the output of the first domain converter, so that a signal at the common bypass output 50 and LPC link 510 is sent to either the first coding path 400 or the second coding path 500. The first coding path 400 includes a second domain converter 410 and the first processor 420 of Fig. 1a, and the second encoding path 500 includes the second processor 520 of Fig. 1a. In the encoder variant of Fig. 1c, the input of the second domain converter 510 is connected to the input of the switchable bypass 50, and the output of the switchable bypass 50 is connected to the output of the first domain converter 510 to create a common output, and this common output is the input of switch 200, wherein the switch has two outputs, but it may even contain additional outputs for additional coding processors.
[0030] Preferably, the second domain converter 410 in the first coding path 400 includes an MDCT transformation which is additionally connected to a switchable time warping function (TW). The MDCT spectrum is coded using a scaler / quantizer that quantizes the input values based on information provided from the psychoacoustic model in block 300 of the signal classifier. On the other hand, the second processor includes a time domain encoder for encoding the time domain input signal. In one embodiment of the invention, the switch 200 is controlled such that in the case of an active / closed bypass 50, the switch 200 is automatically set to the upper coding path 400. However, in another embodiment of the invention, the switch 200 can also be controlled independently of the switchable bypass 50, even when the bypass is active / closed, such that the time domain encoder 520 can directly receive the time domain audio input.
[0031] Fig. 1d shows an analogous decoder in which the LPC synthesis block 540 corresponds to the first converter of Fig. 1b and can be bypassed by bypass 52, which is preferably a switchable bypass controlled by a bypass signal generated by the bit stream demultiplexer 900. The bit stream demultiplexer 900 may generate this signal as well as all other control signals for coding paths 430, 530 or SBR synthesis block 701 or MPEG Surround decoder block 702 from input bit stream 899, or may receive data for these control lines from signal analysis or any another separate source of information.
[0032] A more detailed explanation of the variant of Fig. 1c for the encoder and of Fig. 1d for the decoder will be given below.
[0032] A preferred variant consists of a hybrid audio encoder that combines the advantages of successful MPEG technology such as AAC, SBR and MPEG Surround with effective speech coding technology. The resulting codec contains common pre-processing for all signal categories consisting of MPEG Surround and improved SBR (eSBR). The information output encoder architecture controlled by the psychoacoustic model or derived from the source, depending on the signal category, is selected for each frame.
[0034] The proposed codec preferably uses encoding tools such as MPEG Surround, SBR and the AAC base encoder. These tools have made changes and improvements to improve the performance of speech signals and at very low bit rates. At higher bit rates, at least the AAC performance is matched as the new codec can switch to very close AAC mode. An improved noiseless coding mode has been introduced that provides on average slightly better noiseless coding quality. At a bit rate of about 32 kbps and lower, additional tools are activated to improve the quality of the base encoder for speech and other signals. The main components of these tools are LPC-based frequency shaping, more alternative window length options for the MDCT based encoder, and a time domain encoder. The new bandwidth extension technique is used to develop the SBR tool, which is better suited to low crossover frequencies and for speech signals. The MPEG Surround tool provides a parametric representation of a stereo or multi-channel signal by providing a downmix and a parameterized stereo scene. In given test cases, it is only used to encode stereo signals, but is also suitable for multi-channel input signals, thanks to the use of existing MPEG Surround functionality with MPEG-D.
[0035] All tools in the codec chain except the MDCT encoder are preferably used only for low bit rates.
[0036] The MPEG Surround technique is used to transfer N audio input channels via M audio transmission channels. So the system is naturally adapted to multi-channel systems. MPEG Surround has been enriched to improve quality at low bit rates and for speech type signals.
[0037] The basic mode of operation is to create a high quality mono downmix from the stereo input signal. In addition, the set of spatial parameters is extracted. On the decoder side, a stereo output signal is generated using a decoded mono downmix in combination with the extracted and transmitted spatial parameters. A low bit rate 2-1-2 mode has been added to existing 5-x-5 or 7-x-7 operating points in MPEG Surround, using a simple tree structure consisting of a single OTT tray (one-to-two, one to two) in MPEG Surround upmix. Some components have been modified to better suit speech reproduction. For higher data flows, such as 64 kbps and higher, the core encoder uses discrete stereo coding (Mid / Side or L / R), and MPEG Surround is not used at this operating point.
[0038] The bandwidth extension proposed in this technology application is based on the MPEG SBR technique. The filter bank used is identical to the QMF filter bank in MPEG Surround and SBR, offering the option of sharing samples in the domain
QMF between MPEG Surround and SBR without additional synthesis / analysis. Compared to the standardized SBR tool, eSBR introduces an enhanced processing algorithm that is optimal for both speech and audio content. An SBR extension is also included which is better suited to very low bit rates and low crossover frequencies.
[0039] As is known from the combination of SBR and AAC, this property can be deactivated globally, leaving the entire frequency range coding to the core coder.
[0040] Part of the core coder of the proposed system can be seen as a combination of an optional LPC filter and switchable between the frequency domain and the time domain of the core coder.
[0041] As is known from speech coder architectures, the LPC filter provides the basis for the human speech source model. LPC processing can be enabled or disabled (bypassed) globally or for each frame.
[0042] After the LPC filter, the signal in the LPC domain is encoded using the encoder architecture of either a time domain or a transformation-dependent frequency domain. Switching between these two paths is controlled by an extended psychoacoustic model.
[0043] The time domain coder architecture is based on the ACELP technique, providing optimal coding performance, especially for speech signals at low bit rates.
[0044] The frequency codec base path is based on an MDCT architecture with a scalar quantizer and entropy coding.
[0045] Optionally, a time warping tool is available to improve coding performance for speech signals at higher bit rates (such as 64 kbps and higher) through a more compact signal representation.
[0046] MDCT-based architecture provides good quality at lower bit rates and scales closer to transparency, as is known from existing MPEG techniques. At higher bit rates it can change to AAC mode.
The buffering requirements are identical to those in AAC, i.e. the maximum number of bits in the input buffer is 6144 per core encoder channel: 6144 bits per mono channel element, 12288 bits per stereo channel pair element.
[0048] The bit container is controlled in an encoder that allows the encoding process to be adapted to the current bit requirements. The properties of the bit container are identical to those in AAC.
[0049] The encoder and decoder are controlled to operate at different bit rates from 12 kbps mono to 64 kbps stereo.
[0050] The complexity of the decoder is defined in the PCU range. A complexity of approximately 11.7 PCU is required for the base decoder. In the case where a time-matching tool is used, as in the 64 kbps test mode, the decoder complexity increases to 22.2 PCU.
[0051] The RAM and ROM demand for the preferred stereo decoder is:
RAM: -24 kWords (thousand words)
ROM: -150 kWords (thousand words) [0052] Notification of the entropy encoder allows to obtain the total size
RAM only -98 kWords.
[0053] In the case where a time warping tool is used, the RAM demand increases by -3 kWords, the ROM demand increases by -40 kWords.
[0054] The theoretical algorithmic delay depends on the tools used in the codec chain (e.g. MPEG Surround etc.):
The algorithmic delay of the proposed technique is presented on the operating point for the codec sampling frequency. The following values do not include the frame delay, i.e. the delay required to fill the encoder input buffer with the number of samples necessary to process the first frame. The framing delay is 2048 samples for all specified operating modes. The following tables contain both the minimum algorithmic delay and the delay for the implementation used. The additional 48 kHz sampling delay of input PCM files for the codec sampling frequency is specified in parentheses (.).
<td>No. of test</td><td>Theoretical Minimal algorithmic delay (Sample)</td><td>Algorithmic delay implemented (samples)</td>
<td>Test 1, stereo, 64 kbps</td><td> 8278</td><td> 8278 (+44)</td>
<td>Test 2, stereo, 32 kbps</td><td> 9153</td><td> 11201 (+44)</td>
<td>Test 3, stereo, 24 kbps</td><td> 9153</td><td> 11200 (+45)</td>
<td>Test 4, stereo, 20 kbps</td><td> 9153</td><td> 9153 (+44)</td>
<td>Test 5, stereo, 16 kbps</td><td> 11201</td><td> 11201 (+44)</td>
<td>Test 6, mono, 24 kbps</td><td> 4794</td><td> 5021 (+45)</td>
<td>Test 7, mono, 20 kbps</td><td> 4794</td><td> 4854 (+44)</td>
<td>Test 8, mono, 16 kbps</td><td> 6842</td><td> 6842 (+44)</td>
<td>Test 9, mono, 12 kbps</td><td> 6842</td><td> 6842 (+44)</td>
[0055] The basic attributes of this codec can be summarized as follows:
The proposed technique advantageously uses the latest coding technology for speech and audio signals, without sacrificing the quality of both speech and music content.
As a result, a codec is created that is able to provide the best available quality of speech, music and mixed content for bit rates ranging from very low speeds (12 kbps) and increasing to high data rates such as 128 kbps and higher, with whose codec achieves transparent quality.
[0056] A mono, stereo or multi-channel signal is input into the typical pre-processing stage 100 in Fig. 2a. A typical pre-processing algorithm may have a combined stereo function, a surround function, and / or a bandwidth extension function. At the output of block 100 there is a mono channel, stereo channel, or multiple channels that are connected to bypass 50 and 510 converter of many units of this type.
[0057] A bypass assembly 50 and converter 510 may be placed for each output from the member 100 when the member 100 has more than one output, i.e. when a stereo or multi-channel signal is output from the member 100. For example, the first stereo signal channel may be a speech signal channel and the second stereo signal channel may be a music channel. In this situation, the decision at the same time in the decision section may be different for both channels.
[0058] Bypass 50 is controlled by the decision member 300. The decision member receives the input to block 100 or the output from block 100 at its input. Alternatively, the decision member 300 may also receive additional information that is contained in the mono signal, stereo signal, or multi-channel signal, or where such information - which, for example, was generated during the original production of the mono, stereo or multi-channel signal - exists, is at least associated with such signals.
[0059] In one variant, the decision member does not control the processing member 100, and the arrow between block 300 and block 100 does not exist. In a further variation, the processing block 100 is controlled to some extent by the decision member 300 to set at least one of the parameters in block 100 based on the decision. However, this does not affect the overall algorithm in block 100, so that the basic functions in block 100 are active regardless of the decision in block 300.
[0060] Decision member 300 activates bypass 50 to pass the results of a typical preprocessing member, either to frequency coding portion 400 shown in the upper path in Fig. 1a, or to LPC domain converter 510, which may be part of the second coding portion 500 shown in bottom path in Fig. 2a and containing elements 510, 520.
[0061] In one variant, the bypass bypasses a single domain converter. In another variant, there may be additional domain converters for different encoding paths, such as a third encoding path or even a fourth encoding path, or there may be even more encoding paths. In a variant with three coding paths, the third coding path may be similar to the second coding path, but may include a excitation coder different from the excitation coder 520 in the second path 500. In this variant, the second path includes the LPC term 510 and a codebook based excitation coder, such as the ACELP encoder, and the third path includes the LPC term and the excitation coder operating on the spectral representation of the LPC term output signal.
[0062] The basic element of the frequency domain coding path is spectral conversion block 410 that converts a typical output signal from the pre-processing element into a spectral domain signal. The spectral conversion block may include an MDCT, QMF, or FFT algorithm, a wavelet analysis or a filter bank, such as a critically sampled filter bank having a number of filter bank channels, where the subband signals in such filter bank can be real value signals or complex value signals. The output from spectral conversion block 410 is encoded using an audio signal spectral encoder 420, which may include processing blocks known from the AAC encoding algorithm.
[0063] The basic element in the lower coding path 500 is a source model analyzer such as LPC 510, which in this variant is a domain converter 510, and which produces two types of signals. One signal is the LPC information signal, which is used to control the characteristics of the LPC synthesis filter. Such LPC information is sent to the decoder. Another output of the LPC 510 is the excitation signal or LPC domain signal that is input into the excitation encoder 520. The excitation encoder 520 can be any source filter model encoder, such as a CELP encoder, ACELP encoder, or any other LPC domain signal encoder.
[0064] Another preferred implementation of the excitation encoder is the transform coding of the excitation signal or LPC domain signal. In this variant, the excitation signal is not coded using the ACELP codebook mechanism, but is converted to a spectral representation, and spectral representation values, such as subband signals in the case of a filter bank or frequency factors in the case of a transformation such as FFT, are coded to obtaining data compression. The implementation of such an excitation encoder is a TCX coding mode known from AMR-WB +. This mode is obtained by connecting the output of 510 LPC to the spectral converter 410. TCX mode, known from 3GPP TS 26.290, abolishes the perceptual processing of the weighted signal in the transformation domain. The weighted Fourier transform signal is quantized using split-rate multi-speed quantization (algebraic VQ) with noise factor quantization. The transformation is calculated in windows containing 1024, 512 or 256 samples. The excitation signal is recovered by inverse filtering the quantized weighted signal by the inverse weight filter.
[0065] In Figs. 1a or 1c, followed by the LPC block 510 is a time domain coder which may be an ACELP block or a transformation domain coder which may be a TCX block 527. ACELP is described in 3GPP TS 26.190, while TCX is described in 3GPP TS 26.290. Generally, the ACELP block receives the LPC excitation signal calculated in the procedure shown in Fig. 7e. The TCX block 527 receives the weighted signal generated in Fig. 7f.
[0066] In TCX, the transformation is applied to a weighted signal calculated by filtering the input signal in an LPC-based weighting filter. The weighting filter used in preferred embodiments of the invention may be represented as (1-Α (ζ / Υ) / (1-μζ<sup>ί</sup>'). Thus, the signal is weighted by a signal in the LPC domain and its transform is the LPC spectral domain. The signal processed by block 526 ACELP is an excitation signal and is different from the signal processed by block 527, but both signals are in the LPC domain.
[0067] On the decoder side, the reverse of the weighting filter, i.e. (1 ^ z ^ / Az /), is used after the inverse spectral transformation. Then, the signal is filtered by (1A (z)) to introduce it into the LPC excitation domain. Thus, conversion to the LPC domain and TCX operation<sup>-1</sup> includes reverse transformation followed by filtering to convert from the weighted signal domain to the excitation domain.
[0068] Although item 510 illustrates a single block, block 510 can output different signals as long as the signals are in the LPC domain. The current mode of block 510, such as excitation signal mode or weighted signal mode may depend on the current state of the switch. Alternatively, block 510 may have two parallel processing devices, where one device is implemented similarly to Fig. 7e and the other device is implemented as in Fig. 7f. Hence, the LPC domain at the output of block 510 may represent either the LPC excitation signal or the LPC weighted signal, or any other signal in the LPC domain.
[0069] In LPC mode, when the bypass is inactive, i.e. when ACELP / TCX coding takes place, preferably the signal is pre-emphasized by a 1-0.68z filter<sup>-1</sup> before coding. In the ACELP / TCX decoder, the synthesis signal is withdrawn by the filter 1 / (1-0.68z<sup>-1</sup>). Pre-emphasizing may be part of an LPC block 510 in which the signal is pre-highlighted before LPC analysis and quantization. Similarly, withdrawal of embossment may be part of LPC synthesis block 540<sup>-1</sup>.
[0070] There are several LPC domains. The first LPC domain represents LPC excitation, and the second LPC domain represents the LPC weighted signal. That is, the signal of the first LPC domain is obtained by filtering (1-A (z)) for conversion to the residual / excitation LPC domain, while the signal of the second LPC domain is obtained by filtering with the filter (1-A (zY) / (1 -pz "<sup>1</sup>) to convert to a weighted LPC domain.
[0071] The decision in the decision term may be adaptable to the signal, such that the decision term distinguishes music from speech and controls the bypass 50 and, if present, the switch 200 of Fig. 1c in such a way that the music signals are fed into the upper paths 400, and speech signals are input into bottom path 500. In one variant, the decision member forwards the decision information to the output bit stream, so that the decoder can use this information to perform the correct decoding operation.
[0072] Such a decoder is shown in Fig. 2b. The output signal from the spectral audio encoder 420, after transmission, is input into the spectral audio decoder 430. The output signal from the spectral audio decoder 430 is input into the time domain converter 440. Similarly, the output signal of the excitation encoder 520 of Fig. 2a is input to the excitation decoder 530, which produces the output signal in the LPC domain. The LPC domain signal is introduced into the LPC synthesis stage 540, which receives, on the next input, LPC information generated by the appropriate LPC analysis stage 510. The output signals from the time domain converter 440 and / or from the LPC synthesis term 540 are input into switchable bypass 52. The bypass 52 is controlled by a switch control signal that has been generated, for example, in the decision member 300, or which has been provided externally, for example, by an entity producing the original mono, stereo, or multi-channel signal.
[0073] The output signal from bypass 540 or link 540 is a complete mono signal which is then input into a typical post-processing member 700 in which combined stereo processing or bandwidth extension processing, etc. can take place. Depending on the specific function of the typical post processing member, it produces a mono signal, stereo signal or multi-channel signal, which, when the post processing member 700 performs the bandwidth extension operation, has a larger bandwidth than the input signal to block 700.
[0074] In one variant, the bypass 52 is adapted to bypass a single converter 540. In another variant, there may be additional converters forming additional decoding paths, such as a third decoding path and even a fourth decoding path, or there may be even more paths decoding. In the variant with three decoding paths, the third decoding path may be similar to the second decoding path, but may include a different excitation decoder than the excitation decoder 530 in the second path 530, 540. In this variant, the second path includes LPC 540 and the excitation decoder based on a codebook, such as ACELP, and the third track includes the LPC term and the excitation decoder operating on the spectral representation of the output signal of the LPC term 540.
[0075] As previously stated, Fig. 2c shows a preferred coding algorithm in accordance with the second aspect of the present invention. The typical processing algorithm 100 in Fig. 1a now includes a stereo surround / combined block 101 that generates the total stereo parameters at its output and the mono output signal that is generated by the downmix operation of the input signal, which is a signal having at least two channels. Generally, the signal at the output of block 101 may also be a signal with more channels, but due to the downmix function of block 101, the number of channels at the output of block 101 will be less than the number of channels input at block 101.
[0076] The output of block 101 is provided to the input of the band extension block 102, which in the encoder in Fig. 2c, produces a limited band signal at its output, such as a signal in the low frequency band, or a signal after lowpass filtering. In addition, for the high frequency band of the signal input into block 102, bit stream parameters, bandwidth extension parameters such as spectral envelope parameters, reverse filtering parameters, background noise parameters, etc., known from the HE-AAC profile of the MPEG standard are generated and transmitted to the multiplexer 800 bit stream. 4.
[0077] Preferably, the decision member 300 receives the signal input into block 101 or input into block 102 to make a choice, for example, as to the music mode or the speech mode. In the music mode, the upper coding path 400 is selected, and in the speech mode, the lower coding path 500. Preferably, the decision term further controls the stereo combined block 101 and / or the band extension block 102 to match the function of these blocks to a particular signal. Thus, when the decision term determines that a certain time segment of the input signal is of the first type, like the music type, then the decision term 300 may control the specific features of block 101 and / or block 102. Alternatively, when the decision term 300 determines that the signal is in in speech mode, or generally in coding mode in the LPC domain, then specific features of blocks 101 and 102 may be controlled according to the result of the decision term.
[0078] Depending on the switch decision that can be obtained from the input signal of the switch 200 or from any external source, such as the manufacturer of the original audio signal being the base of the input signal to the member 200, the switch switches between frequency coding path 400 and coding path 500 LPC. The frequency coding path 400 includes a spectral conversion term, and the next quantization / coding term. The quantization / coding term may contain any of the functions known in modern frequency domain encoders, such as the AAC encoder. In addition, the quantization operation in the quantization / coding term can be controlled by means of a psychoacoustic model that generates psychoacoustic information, such as a psychoacoustic frequency masking threshold, which information is entered into the term.
[0079] Preferably, the spectral conversion is performed using an MDCT operation, which, more preferably, is a time warped MDCT transformation operation, where the power, or generally, the matching power can be adjusted from zero to high matching power. At zero match power, the MDCT operation in block 410 of Fig. 1c is the usual MDCT operation known in the art. The time warping power, together with the additional time warping information, can be sent / input to the multiplexer 800 bit stream as additional information. Therefore, if the TW-MDCT operation is used, the additional time warping information should be sent to the bit stream, which is designated as 424 in Fig. 1c, and - on the decoder side - the additional time warping information should be received from the bit stream , which is designated 434 in Fig. 1d.
[0080] In the LPC coding path, the encoder in the LPC domain may comprise an ACELP core calculating tone gain, tone delay and / or code book information, such as code book index and code gain.
[0081] In the first coding path 400, the spectral converter preferably includes a specially tailored MDCT operation having certain window functions followed by a quantization / entropy coding member which may be a vector quantization member but is preferably a quantizer / encoder similar quantizer / encoder in the frequency domain coding path.
[0082] Fig. 2d shows a decoding scheme corresponding to the encoding scheme of Fig. 2c. The bit stream generated by the bit stream multiplexer is introduced into the bit stream demultiplexer. Depending on the information obtained, for example, from the bit stream by the mode detection block, the decoder side switch is set to send to the block 701 bandwidth extension, signals from the upper path or from the lower path. The band extension block 701 receives additional information from the bit stream demultiplexer and based on this additional information, and based on the mode decision result, reproduces the high frequency band based on the low frequency band transmitted for example by the combiner 600 of Fig. 1d.
[0083] The full-band signal generated by block 701 is input to the stereo / surround combined processing member 702, which reconstructs two stereo channels, or several channels of multi-channel signal. Generally, block 702 will leave more channels than has been input into it. Depending on the application, there may be as many as two channels at the input of block 702, as in stereo mode, and even more channels, provided that the output of this block has more channels than at its input.
[0084] The switch 200 of Fig. 1c has been shown to switch between both paths so that only one of the paths receives the signal for processing and the other path does not receive the signal for processing as shown generally in Fig. 4a. However, in the alternative variant, shown in Fig. 4b, the switch can also be located immediately after, for example, the audio signal encoder 420 and the excitation encoder 520, which means that both paths 400, 500 in parallel process the same signal. However, in order to avoid doubling the bit rate, only one output signal 400 or 500 is selected from these two paths to enter into the output bit stream. The decision member will operate in such a way that the signal introduced into the bit stream minimizes a certain cost function, where the cost function can be generated bit rate, or perceptual distortion generated, or a combined bit rate / distortion cost function. Therefore, in this mode, or in the mode shown in the Figures, the decision member may also operate in a closed loop mode to ensure that only the path that ultimately offers the smallest bit rate at the given distortion is finally introduced into the bit stream. given bit rate, provides the smallest perceptual distortions.
[0085] In general, the processing in path 400 is processing in the perception-based model or in the information exit model. So this path represents the human hearing receiving the sound. In contrast, processing in path 500 is intended to generate a signal in the excitation domain, residual domain, or LPC domain. Generally, processing in path 500 is processing in the speech model, or in the information generation model. For speech signals, this model is a model of sound generation by the human speech / sound generation system. However, if you want to encode a sound from another source that requires a different sound generation model, the processing path 500 may look different.
[0086] Although Figures 1a to 4c show block diagrams of the device, these figures simultaneously illustrate a method where the block functions correspond to the method steps.
[0087] Fig. 3c shows an audio encoder for encoding an input audio signal 195. The input audio signal 195 is present in a first domain, which may for example be a time domain, but which may also be any other domain such as frequency domain, LPC domain , LPC spectral domain, or any other domain. Generally, conversion from one domain to another domain is performed by a type of conversion algorithm, such as any of the well-known time-frequency conversion algorithms or frequency-time conversion algorithms.
[0088] An alternative transform from a time domain, for example in the LPC domain, is the result of LPC filtering of a time domain signal, which results in a residual LPC signal or excitation signal, or another signal in the LPC domain. Any other filtering operation that produces a filtered signal has an impact on a significant number of signal samples before - as may be the case - transformation can be used as a transformation algorithm. Therefore, weighing an audio signal using an LPC-based weighting filter is another transformation that generates an LPC domain signal. In time-frequency transformation, modification of a single spectral value will affect all values in the time domain before transformation. Similarly, modification of any time domain sample will affect any frequency domain sample. Similarly, modification of the excitation signal sample for the LPC domain will, due to the length of the LPC filter, affect a significant number of samples before LPC filtering. Similarly, sample modification before LPC transformation will affect many samples obtained by this LPC transformation due to the natural memory effect of the LPC filter.
[0089] The audio encoder of Fig. 3c includes a first encoding path 522 that generates a first encoded signal. This first signal may be encoded in a fourth domain, which in a preferred embodiment is a time-frequency domain, i.e. a domain that is obtained when the time-domain signal is processed in time-frequency conversion.
[0090] Therefore, the first encoding path 522 for encoding the audio signal uses a first encoding algorithm to obtain a first encoded signal, wherein the first encoding algorithm may or may not include a time-frequency conversion algorithm.
[0091] The audio signal encoder further includes a second encoding path 523 for encoding the audio signal. Second coding path 523 uses the second coding algorithm to obtain a second encoded signal that is different from the first coding algorithm.
[0092] The audio encoder further includes a first switch 521 for switching between the first encoding path 522 and the second encoding path 523, 524, so that for a portion of the audio signal, the encoder output signal includes either the first encoded signal at the output of block 522 or the second encoded signal at the output of the second coding path. Thus, when for some portion of the input audio signal 195, the first encoded signal in the fourth domain is included in the encoder output signal, the second encoded signal, which is either the first processed signal in the second domain or the second processed signal in the third domain, is not included. in the encoder output signal. This is to ensure that this decoder is bit-rate efficient. In embodiments of the invention, any time portions of the audio signal that are contained in two different encoded signals are small compared to the frame length, as will be discussed with reference to Fig. 3e. These small parts are useful for penetrating from one encoded signal to another encoded signal in the event of a switching event to reduce the artifacts that can occur in the absence of penetration. Therefore, apart from the interference area, each block of the time domain is represented by a signal encoded in only one domain.
[0093] As shown in Fig. 3c, the second coding path 523 is next downstream of the converter 521 for converting the first domain audio signal, i.e. signal 195 to the second domain, and after bypass 50. In addition, the first processing path 522 obtains a first processed signal, which is preferably also in the second domain, such that the first processing path 522 does not change the domain, or which is in the first domain.
[0094] The second coding path 523, 524 performs the conversion of the audio signal into a third domain or a fourth domain that is different from the first domain and which is also different from the second domain to obtain a second processed signal at the output of the second processing path 523, 524.
[0095] Furthermore, the encoder includes a switch 521 for switching between the first switching path 522 and the second switching path 523, 524, the switch corresponding to the switch 200 of Fig. 1c.
[0096] Fig. 3d shows a corresponding decoder for decoding the encoded audio signal generated by the encoder of Fig. 3c. Generally, each block of a first domain audio signal is represented by either a second or first domain signal or an encoded third or fourth domain signal, outside an optional diffusion area, which is preferably short compared to the length of one frame, to obtain a system that is closest to the critical sampling limit. The encoded audio signal comprises a first encoded signal, a second encoded signal, wherein the first encoded signal and the second encoded signal relate to different time portions of the decoded audio signal and wherein the second domain, third domain and first domain for the decoded audio signal are different from each other.
[0097] The decoder includes a first decoding path for decoding based on the first encoding algorithm. The first decoding path is designated 531 in Fig. 3d.
[0098] The decoder of Fig. 3d further includes a second decoding path 533, 534, which includes several elements.
[0099] The decoder further includes a first combiner 532 for combining the first inverse processed signal and the second inverse processed signal to obtain a signal in the first or second domain, said combined signal being in the first time moment influenced only by the first inverse processed signal, and in later time, influenced only by the second inversely processed signal.
[0100] The decoder further includes a converter 540 for converting the combined signal to the first domain, and a switchable bypass 52.
[0101] Finally, the decoder shown in Fig. 3d includes a second combiner 600 for combining the decoded first signal from the bypass 52 and converter output 540 to obtain the decoded output signal in the first domain. Again, the decoded output signal in the first domain at the first time moment is only influenced by the signal output from the 540 converter, and at a later time moment is only influenced by the bypass signal.
[0102] This situation is illustrated, from the encoder side, in Fig. 3e. The upper part of Fig. 3e is a schematic representation of a first domain audio signal, such as a time domain audio signal, where the time index increases from left to right and position 3 can be considered as the stream of audio samples representing the signal 195 of Fig. 3c. FIG. 3e shows frames 3a, 3b, 3c, 3d, which can be generated by switching between the first encoded signal and the second encoded signal, as indicated by reference number 4 in Fig. 3e. The first encoded signal and the second encoded signal are all in different domains. To ensure that switching between different domains does not cause artifacts on the decoder side, frames 3a, 3b, 3c, ... signal in the time domain contain an overlap area, which is designated as a diffusion area. However, there is no such interfering area between 3d and 3c frames, which means that the 3d frame can also be represented by a signal in the same domain as the previous 3c signal and there is no domain change between frame 3c and 3d.
Therefore, in general, it is preferable not to provide the penetration area where there is no change of domain and to provide the penetration area, i.e. the portion of the audio signal that is encoded by two consecutive encoded / processed signals when the domain change occurs, i.e. switching action of any of the two switches.
[0104] In an embodiment in which the first encoded signal or the second processed signal has been generated by MDCT processing with e.g. a 50% overlap, each time domain sample is contained in two successive frames. However, due to the properties of MDCT, this does not result in overhead, because MDCT is a critically sampled system. In this context, critical sampling means that the number of spectral values is the same as the number of time domain values. MDCT is advantageous in that the penetration effect is provided without a specific penetration area, so that the transition from the MDCT block to the next MDCT block is provided without any overhead that would interfere with the critical sampling requirement.
[0105] Preferably, the first coding algorithm in the first coding path is based on the information output model, and the second coding algorithm in the second coding path is based on the information source model or SNR model. The SNR model is a model that is not specifically associated with a particular sound generation mechanism, but is one of the coding modes that can be selected from many coding modes based e.g. on a closed loop decision. Thus, the SNR model is any available coding model that does not necessarily have to be associated with the physical form of the sound generator, but which is any parameterized coding model other than the information output model that can be selected in a closed loop decision and, in particular, by comparing the results of different SNRs from different models.
[0106] As shown in Fig. 3c, a controller 300, 525 is provided. This controller may include the functionality of the decision member 300 of Fig. 1c. Generally, the controller is used to control the bypass and switch 200 of Fig. 1c in a signal-matched manner. The controller operates to analyze the signal input to the bypass or the signal output from the first or second coding path or the signals obtained by coding and decoding the first and second coding path with respect to the target function. Alternatively or additionally, the controller operates to analyze the signal input to the switch or derived from the first processing path or second processing path or obtained by processing and inverse processing from the first processing path and from the second processing path, again, with respect to the target function.
[0107] In one embodiment, the first coding path or second coding path includes an introducing aliasing time-frequency conversion algorithm, such as an MDCT or MDST algorithm, which differs from a simple FFT transformation that does not introduce an aliasing effect. In addition, one or both paths contain a quantizer / entropy coder block. In particular, only the second processing path of the second coding path includes the time-frequency converter introducing the aliasing operation, and the first processing path of the second coding path includes the quantizer and / or entropy encoder and does not introduce any aliasing effects. The aliasing time-frequency converter preferably includes a window module for inputting the analysis window and the MDCT transformation algorithm. In particular, the windowed module works to introduce the window function into successive frames in a folded manner, so that a windowed signal sample appears in at least two successive windowed frames.
[0108] In one embodiment, the first processing path comprises an ACELP encoder, and the second processing path comprises an MDCT spectral converter and a quantizer for quantizing the spectral components to obtain quantized spectral components, each quantized spectral component being zero or defined by one quantization index from many different possible quantization indexes.
[0109] As previously stated, both coding paths operate to encode the audio signal in a block fashion in which a bypass or switch operates on the blocks, such that the switching or bypass operation occurs at least after a block of a predetermined number of signal samples, wherein the set number creates the frame length for the corresponding switch. Thus, the resolution for passing through the bypass can be, for example, a block of 2048 or 1028 samples, and the frame length on which the bypass is switched can be variable, but it is preferably fixed permanently for such a fairly long period.
[0110] In contrast, the block length for switch 200, i.e. when the switch 200 switches from one mode to another, is clearly smaller than the block length for the first switch. Preferably, both block lengths for the switches are selected such that the length of the longer block is an integer multiple of the shorter block length. In a preferred embodiment of the invention, the block length of the first switch is 2048 and the block length of the second switch is 1024, or more preferably 512, and even more preferably 256, even more preferably 256 or even 128 samples, so that the maximum switch can switch 16 times when the bypass changes only one time.
[0111] In another embodiment of the invention, the controller 300 operates to distinguish speech from music for the first switch in such a way that the music decision is favored over the speech decision. In this variant, the speech decision is made even if the part less than 50% of the frame for the first switch is speech, and the part greater than 50% of the frame is music.
[0112] In addition, the controller operates to switch to speech mode when a fairly small portion of the first frame is speech, in particular when a portion of the first frame is speech that represents 50% of the length of the smaller second frame. Thus, a favorable switch decision favoring speech results in switching to speech mode even when, for example, only 6% or 12% of the block corresponding to the frame length of the first switch is speech.
[0113] This procedure preferably aims to take full advantage of the possibilities of saving the bit rate of the first processing path, which has a voiced speech core in one variant, and avoiding any loss of quality, even in the remainder of the large first frame that does not contain speech. due to the fact that the second processing path includes a converter and is therefore useful for audio signals that also contain non-speech signals. Preferably, this second processing path includes a bookmarking MDCT transformation that is critically sampled, and which, even at small window sizes, provides very efficient and aliasing-free operation due to time domain processing that eliminates aliasing, such as overlap-add () after decoder side. In addition, a large block length in the first coding path is useful, which is preferably an AAC type MDCT coding path, because non-speech signals are normally quite stationary and the long transformation window provides high frequency resolution and thus high quality and, in addition, provides bit rate performance thanks to the psychoacoustic controlled quantization module, which can also be used in a transformation based coding mode in the second processing path of the second coding path.
[0114] Referring to the illustration of the decoder in Fig. 3d, it is preferred that the transmitted signal includes a clear indicator as additional information 4a, as shown in Fig. 3e. Additional information 4a is extracted by a bit stream parser, not shown in Fig. 3d to send the corresponding first processed signal or second processed signal to a suitable processor, such as the first inverse processing path or the second inverse processing path in Fig. 3d. Therefore, the encoded signal not only contains the encoded / processed signals, but also contains additional information regarding these signals. In other variants, however, there may be implicit assay allowing the decoder bit stream parser to distinguish different signals. Referring to Fig. 3e, it is noted here that the first processed signal or second processed signal is the output of the second coding path, and thus of the second encoded signal.
[0115] Preferably, the first decoding path and / or the second inverse processing path comprises an MDCT transformation for conversion from the spectral domain to the time domain. To this end, an overlap-add module is provided to perform the function of eliminating time domain aliasing, which at the same time provides a penetration effect to avoid artifacts. Generally, the first decoding path converts the signal encoded in the fourth domain to the first domain, while the second inverse processing path converts from the third domain to the second domain and the converter then connected to the first combiner provides the conversion from the second domain to the first domain, such that Combiner 600 only has first domain signals that represent the decoded output signal in the variant of Fig. 3d.
[0116] Fig. 4c illustrates another aspect of a preferred decoder variant. In order to avoid audible artifacts, especially in a situation where the first decoder is a time aliasing generating decoder, or generally speaking, a frequency domain decoder, and the second decoder is a time domain device, the boundaries between the frame blocks produced by the first 450 and the second decoder 550 decoder should not be completely continuous, especially in a switching situation. Therefore, when the first block from the first decoder 450 is derived and when the block from the second decoder is derived for the next time segment, it is preferable to perform the fade operation, shown as fade member 607. To this end, the fade block 607 can be implemented like this as illustrated in Fig. 4c with members 607a, 607b and 607c. Each track may contain a weighing module with a weighing factor m1 between 0 and 1 on a standardized scale, whereby the weighing factor may change as shown in diagram 609, and the permeation rule thus determined ensures that continuous and gentle penetration takes place, which additionally provides, that the user will not experience any volume fluctuations. Instead of a linear crossfade principle, a nonlinear crossfade principle can be used, such as the sin crossfade principle<sup>2</sup>.
[0117] In some cases, the last block of the first decoder was generated using a window that actually muted that block. In this case, the weighting factor m1 in block 607 is 1 and in fact no weighing is required for this path. When switching from the second decoder to the first decoder occurs, and when the second decoder contains a window that actually mutes the output signal at the end of the block, the weighing module labeled as m2 is not required, or the weighting factor can be set to 1 for the entire penetration area.
[0118] When the first block after switching was generated using window operations and this window actually performed the level rise operation, the corresponding weighting factor may also be set to 1, so that the weighing module is not actually needed. Therefore, when the last block is windowed to be muted by the decoder, and when the first block after switching is windowed using the decoder to enter the level rise, weighing modules 607a, 607b in general are not needed and only the operation of adding by combiner 607c is sufficient .
[0119] In this case, the muted portion of the last frame, and the portion with the increasing level of the next frame, define the cross-fade area, designated as block 609. Furthermore, in such a situation, it is preferable that the last block of one decoder overlap the first block of the other to some extent decoder.
[0120] If a crossfade operation is not required, is not possible, or not desired, and if there is only a simple switch from one decoder to another decoder, it is preferable that the switchover takes place in the noiseless portions of the audio signal, or at least in the portions of the signal. acoustic energy, i.e. fragments that are perceived as noiseless or almost noiseless. Preferably, the decision member 300 ensures in that variant that the switch 200 is activated only when the respective time segment after the switching event has energy that is, for example, less than the average energy of the audio signal, and preferably is less than 50% of the average acoustic signal energy referring to, for example, two or even more time periods / frames of the acoustic signal.
[0121] Preferably, the second encoding / decoding rule is an algorithm based on the LPC technique. In LPC-based speech coding, there is a distinction between segments or parts of a quasi-periodic pulse type excitation signal, and segments or parts of a noise type excitation signal. This is done for LPC vocoders with very low bit rates (2.4 kbps) as in Fig. 7b. However, in CELP encoders for average bit rates, excitation is obtained for adding scaled vectors from the adaptive code book and constant code book.
[0122] Quasi-periodic pulse type excitation signal segments, i.e. signal segments of a particular tone, are coded on a different basis than noise type excitation signals. While quasi-periodic impulse-type excitation signal segments are associated with voiced speech, noise-type signals are associated with voiceless speech.
[0123] For example, reference will be made to Figs. 5a to 5d. Examples of quasi-periodic segments or portions of an impulse type excitation signal, and segments or portions of an excitation signal in a noise type are discussed herein. Specifically, Fig. 5a shows a time domain voiced speech signal, Fig. 5b the same frequency domain signal as an example of a quasi-periodic segment of an impulse type excitation signal, and Fig. 5c and 5d show the segment of the voiceless speech signal as an example of a part of the excitation signal in the noise type. Speech in general can be classified as voiced, voiceless or mixed. Graphs of the sampled voiced and voiceless segments in the time domain and frequency domain are shown in Figs. 5a to 5d. Voiced speech is quasi-periodic in the time domain, and has a harmonic structure in the frequency domain, while unvoiced speech is a random and broadband signal. The momentary spectrum of voiced speech is characterized by a clean and control structure. Pure harmonic structure is the result of quasi-periodic speech and can be attributed to the properties of vibrating vocal cords. The control structure (spectral envelope) is the result of the interaction of the source and vocal tract. The vocal tract consists of the throat and mouth. The shape of the spectral envelope that "matches" the instantaneous spectrum of voiced speech is associated with the transmission properties of the vocal tract, and the spectral tilt (6 dB / octave) resulting from glottis vibration. The spectral envelope is described by a set of peaks, called controls. Controls are resonant modes of the vocal tract. In a typical voice apparatus, there are three to five controls below 5kHz. The amplitudes and locations of the first three controls, usually appearing below 3 kHz, are very important in both speech synthesis and reception. Higher controls are also important in broadband and unvoiced speech representations. Speech properties are associated with the physical speech production system as follows. Voiced speech is produced by stimulating the larynx with quasi-periodic air impulses in the glottis generated by vibrating vocal cords. The frequency of periodic pulses is called the base frequency or tone. Voiceless speech is produced by forcing airflow through constrictions in the larynx. Nasal sounds are the result of a sonic connection of the nasal passage to the larynx, while explosive sounds are produced by the sudden release of air pressure generated behind the larynx.
[0124] Thus, the noise-type audio signal portion does not exhibit either a time domain impulse structure or a frequency domain harmonic structure as shown in Figs. 5c and 5d, which is different from the quasi-periodic impulse portion shown for example in Fig. 5a and 5b. However, as will be outlined later, the differences between the parts in the type of noise, and the quasi-periodic parts may also be observed in the excitation signal after LPC processing. The LPC technique is a method of mapping the vocal tract and extracts larynx stimulation from the signal.
[0125] Furthermore, the quasi-periodic pulse-type parts, and the noise-type signal parts may appear chronologically, which means that at some point in time the part of the audio signal is noisy, while at another section the audio signal is quasi-periodic, i.e. tonal. Alternatively or additionally, the signal characteristics may be different in different frequency bands. Therefore, determining whether an audio signal is noisy or tonal can also be frequency selective, so that a certain frequency band or set of certain frequency bands are considered noisy, while other frequency bands are considered tonal. In this case, some of the audio signal may contain tonal components and noise components.
[0126] Fig. 7a shows a linear model of a speech production system. In this system, two-stage stimulation is assumed, i.e. a pulse train for voiced speech as shown in Fig. 7c, and random noise for unvoiced speech as shown in Fig. 7d. The vocal tract is mapped as an all-field filter 70 that processes the pulses or noise of Figs. 7c or 7d generated by the 72 glottis model. Hence, the system of Fig. 7a can be reduced to the all-field filter model of Fig. 7b including the gain member 77, the transfer path 78, the feedback path 79, and the sum member 80. In the feedback path 79 there is a predictive filter 81, and the entire source model synthesis system shown in Fig. 7b can be represented using domain functions (z) as follows:
<img file="PL2301024T3_D0001.tif" />
where g is gain, A (z) is a predictive filter determined by LP analysis, X (z) is an excitation signal, and S (z) is a speech synthesis output.
[0127] Figs. 7c and 7d are graphical descriptions of the domain of voiceless synthesis using a linear source system model. This system and excitation parameters in the above equation are unknown and must be determined based on a finite set of speech signal samples. A (z) coefficients are obtained using linear input signal prediction and filter coefficient quantization. In linear forward prediction of step p, the current sample of the speech signal sequence is predicted based on the linear combination of p of the previous samples. Prediction coefficients can be determined using generally known algorithms, such as the Levinson-Dubrin algorithm, or generally by autocorrelation or reflection.
[0128] Fig. 7e shows a more detailed implementation of the LPC 510 analysis block. An audio signal is input into the filter determination block which determines the filter information A (z). This information is output as the short-term prediction information required by the decoder. This information is quantized in a quantizer 81, known for example from the AMR-WB + specification. Short-term prediction information is required by the actual 85 predictor filter. In the subtraction system 86, a current sample of the audio signal is introduced, from which the predicted value of the current sample is subtracted, so that for this sample the prediction error signal is generated on line 84. The sequence of such prediction error signal samples is very schematically shown in Fig. 7c or 7d. Thus, it can be seen that Figs. 7c, 7d show a purified pulse type signal.
[0129] When Fig. 7e shows a preferred method of calculating the excitation signal, Fig. 7f shows a preferred method of calculating the weighted signal. In contrast to Fig. 7e, the filter 85 is different if γ is different from 1. For γ a value less than 1 is preferred. In addition, block 87 is present and μ is preferably a number less than 1. Generally, the elements in Fig. 7e and 7f can be implemented, as in 3GPP TS 26.190 or 3GPP TS 26.290.
[0130] Fig. 7g shows inverse processing that can be applied on the decoder side, as in element 537 of Fig. 2b. In particular, block 88 generates an unweighted signal from the weighted signal, and block 89 calculates the excitation from the unweighted signal. Generally, all signals except the unweighted signal of Fig. 7g are in the LPC domain, but the excitation signal and the weighted signal are different signals in the same domain. Block 89 outputs an excitation signal which can then be used together with the output of block 536. Then a common LPC inverse transformation can be performed in block 540 of Fig. 2b.
[0131] Next, a CELP encoder based on the "analysis by synthesis" principle will be discussed with reference to Fig. 6 to illustrate the modifications introduced to this algorithm. Such a CELP encoder is discussed in detail in the publication "Speech Coding: A Tutorial Review", Andreas Spanias, Proceedings of the IEEE, Vol. 82, No. 10, October 1994, pages 1541-1582. The CELP encoder shown in Fig. 6 includes element 60 of long-term prediction, and element 62 of short-term prediction. In addition, a code book is used, indicated by reference number 64. Reference number 66 indicates the weighting filter W (z) used, and reference number 68 indicates the error minimization driver used. s (n) is a time domain input signal. After passing through perceptual weighing, the weighted signal is introduced into the subtraction system 69, which calculates the deviation between the weighted synthesis signal at the output of block 66 and the original weighted signal sw (n). Generally, the short-term prediction filter coefficients A (z) are calculated by the LP analysis term, and its coefficients are quantized in A (z) as shown in Fig. 7e. The long-term prediction information AL (z), including the g-gain of the long-term prediction, and the quantization vector index, i.e. reference to the codebook, is calculated in the prediction error signal at the output of the LPC analysis term, designated 10a in Fig. 7e. LTP parameters are tone delay and gain. In CELP, this is usually implemented as an adaptive code book containing the past excitation signal (not residual). Adaptive CB delay and gain are determined by minimizing the mean square weighted error (tone search in closed loop).
[0132] The CELP algorithm then encodes the residual signal obtained after short- and long-term prediction using a code book, for example Gaussian sequences.
The ACELP algorithm, where "A" means "algebraic" has a specific code book with algebraic structure.
[0133] The code book may contain more or less vectors, where each of the vectors has a length of a number of samples. The gain factor g scales the code vector, and the strengthened code is filtered by the long-term prediction synthesis filter and by the short-term prediction synthesis filter. The "optimal" code vector is selected in such a way that the perceptually weighted mean square error at the output of the subtraction circuit is minimized 69. The search process in the CELP algorithm is carried out by means of optimization analysis by synthesis, as shown in Fig. 6.
[0134] In special cases, when the frame is a mixture of voiced and unvoiced speech, or when speech is background music, TCX coding may be more suitable for encoding excitation in the LPC domain. TCX coding processes a frequency domain weighted signal without making any assumptions about the excitation. The TCX coding is therefore more general than CELP coding and is not limited to the model of the source of voiced or unvoiced excitation. TCX is still coding based on the source filter model, using a linear prediction filter to model the speech signal control.
[0135] In AMR-WB + coding, the choice between different TCX and ACELP modes is as described in the AMR-WB + codec. TCX modes differ in the length of the block discrete Fourier transformation, and the best mode can be selected in the analysis by synthesis, or in the mode of direct feedforward.
[0136] As discussed with reference to Figs. 2c and 2d, a typical pre-processing member 100 preferably includes a combined multi-channel (surround / stereo combined) member 101 and additionally a bandwidth extension member 102. Accordingly, the decoder includes a bandwidth extension member 701 and a combined multi-channel member 702 attached thereto. Preferably, the combined multi-channel member 101, with respect to the encoder, is located in front of the bandwidth extension member 102, and on the decoder side, the band extension member 701 is located in front of the combined multi-channel member 702 with respect to the signal processing direction. Alternatively, however, a typical pre-processing member may comprise a combined multi-channel member without a bandwidth extension member attached thereto, or a bandwidth extension member without a combined multi-channel member attached.
[0137] A preferred example of the combined multi-channel element 101a, 101b on the encoder side, and 702a, 702b on the decoder side is shown in Fig. 8. The number E of the original input channels is introduced into the downmix element 101a, such that the downmix member generates K transmitted channels, where the number K is greater than or equal to unity, and less than E or equal to E.
[0138] Preferably, E input channels are input into the combined multi-channel analyzer 101b, which generates parametric information. This parametric information is preferably entropy coded, for example by other coding followed by Huffman coding, or alternatively arithmetic coding. The encoded parametric information output from block 101d is sent to parametric decoder 702b, which may be part of position 702 in Fig. 2b. The parametric decoder 702b decodes the transmitted parametric information and forwards this decoded parametric information to the upmix member 702a. Upmix member 702a receives K transmitted channels and generates L output channels, where the number L is greater than or equal to K, and less than E, or equal to E.
[0139] Parametric information may include inter-channel level differences, inter-channel time differences, inter-channel phase differences, and / or inter-channel coherence measures, known from the BCC technique, or known and described in detail in the MPEG surround standard. The number of channels transmitted may be one mono channel for extremely low bit rates applications, or may include compatible stereo applications, or may include a compatible stereo signal, i.e. two channels. Typically, the number of E input channels can be 5 or even greater than 5. Alternatively, the number E of input channels can also mean E audio objects known in the context of spatial audio object coding (SAOC).
[0140] In one application, the downmix member performs weighted or unweighted addition of original E input channels, or addition of E input acoustic objects. In the case of acoustic objects as input channels, the combined multi-channel parameter analyzer 101b will calculate the parameters of the acoustic objects, such as the correlation matrix between the acoustic objects, preferably for each time segment, and more preferably for each frequency band. To this end, the entire frequency range may be divided into at least 10, and preferably 32 or 64 frequency bands.
[0141] Fig. 9 shows a preferred embodiment of the bandwidth extension member 102 of Fig. 2a, and the corresponding bandwidth extension member 701 of Fig. 2b. On the coder side, the band extension block 102 preferably includes the low pass filter block 102b, the sampling frequency reduction block that is downstream of the low pass filter, or which is part of the inverse QMF that affects only half of the QMF bands, and the high frequency band analyzer 102a. The original audio signal introduced into the band extension block 102 is low-pass filtered to generate a signal in the low frequency band, which is then input into the coding paths and / or the switch. The low-pass filter has a cut-off frequency that can be in the range of 3 kHz to 10 kHz. In addition, band extension block 102 further includes a high frequency band analyzer for calculating bandwidth parameters, such as information about spectral envelope parameters, background noise parameter information, information about inverse filtering parameters, additional parametric information regarding certain harmonic lines in the high frequency band, and additional parameters, which are discussed in detail in the MPEG-4 standard, in the chapter regarding the replication of spectral bands.
[0142] At the decoder side, block 701 includes a transposition module 701a, a matching module 701b, and a connecting module 701c. The connecting module 701c combines the decoded low frequency signal, and the reconstructed and matched high frequency signal output from the matching module 701b. The input signal to the matching module 701b is provided by a transposition module that operates to obtain the high frequency band signal from the low frequency band signal using the spectral band replication technique, or generally by means of bandwidth extension. Transposition carried out in the transposition module 701a can be a transposition carried out harmonically or non-harmonically. The signal generated by the transposition module 701a is then matched in the matching module 701b using transmitted parametric bandwidth extension information.
[0143] As shown in Fig. 8 and Fig. 9, the described blocks may in a preferred embodiment have a mode control input. The mode control input signal is obtained from the output signal of decision member 300. In such a preferred variant, the characteristics of the respective block can be matched to the output signal from the decision member, i.e. whether the 'speech' decision is made for a given time segment of the acoustic signal, or "music" decision. Preferably, mode control only applies to at least one function of these blocks, but not to all of their functions. For example, the decision may only apply to the transposition module 701a, but may not apply to other blocks in Fig. 9, or may, for example, only apply to the combined multi-channel parametric analyzer 101b in Fig. 8, but not to other blocks in Fig. 8. This implementation is preferably implemented in such a way that greater flexibility and higher quality, as well as a lower output signal rate are obtained by providing flexibility in a typical pre-processing stage. On the other hand, however, the use of algorithms in the typical pre-processing segment for both types of signal allows the implementation of an efficient coding / decoding algorithm.
[0144] Figs. 10a and 10b show two different implementations of the decision member 300. In Fig. 10a, an open loop decision is shown. The signal analyzer 300a in the decision term has certain rules for deciding whether a given time segment or frequency segment of the input signal has characteristics that require coding in the first coding path 400 or in the second coding path 500. To this end, the signal analyzer 300a may analyze the audio signal at the input of a typical preprocessing member, or may analyze the audio signal at the output of a typical preprocessing member, i.e., an intermediate audio signal, or may analyze the intermediate signal within a typical preprocessing member, such as a signal the output after the downmix operation, which may be a mono signal, or it may be a signal having k channels shown in Fig. 8. On the output side, the signal analyzer 300a generates a switching decision to control the switch 200 on the encoder side, and the corresponding switch 600, or a connecting module 600 on the decoder side.
[0145] Alternatively, the decision member 300 may implement a closed loop decision, which means that both coding paths perform their tasks on the same audio signal portion, and both encoded signals are decoded by the corresponding decoding paths 300c, 300d. The output signal from the 300c and 300d devices is introduced into the comparison module 300b, which compares the output signals of the decoding devices with the corresponding fragment of, for example, an intermediate audio signal. Then, depending on the cost function, such as the signal-to-noise ratio for the path, the switching decision is made. This closed-loop decision is more complex compared to the open-loop decision, but this complexity exists only on the encoder side, and on the decoder side there are no difficulties due to this process, because the decoder can effectively use the result of this encoding decision. Therefore, in applications where the complexity of the decoder is not a problem, such as in media transmission applications where there is only a small number of encoders and a large number of decoders, which should also be fast and cheap, closed loop mode is preferred because of the complexity and quality.
[0146] The cost function used by the comparison module 300d may be a cost function controlled by qualitative factors or it may be a cost function controlled by noise related aspects, or it may be a cost function controlled by flow aspects, or it may be a combined cost function controlled by a combination of bit rates, quality, noise (introduced by coding artifacts, especially in quantization), etc.
[0147] Preferably, the first coding path or second coding path includes a time warping function on the encoder side and the decoder side, respectively. In one variant, the first coding path includes a time warping module for calculating the variable time warping characteristics depending on the audio signal fragment, a resampling module for performing the resampling according to a fixed time warping characteristic, a converter from a time domain to a frequency domain, and an entropy encoder for conversion the result of the conversion from a time domain to a frequency domain into an encoded representation. The variable time warping characteristics are contained in an encoded audio signal. This information is read in the decoding path enriched with time warping functions and processed to finally get the output signal in the wrinkled time scale. For example, the coding path performs entropy decoding, inverse quantization, and conversion from the frequency domain back to the time domain. In the time domain, an inverse time warping may be applied, followed by an appropriate resampling operation to finally obtain a discrete acoustic signal on an unwrinkled time scale.
[0148] Depending on the requirements of individual implementations of the methods of the present invention, these methods can be implemented in hardware or in software. The implementation can be carried out using digital storage media, in particular a DVD or CD containing electronically readable control signals that interact with a programmable computer system in such a way that the methods of the present invention are implemented. In general, the present invention is therefore a computer program product with a program code stored on a machine readable medium, which program code may operate to implement the inventive methods when the computer program product is running on a computer. In other words, the methods of the invention are thus a computer program with program code for implementing at least one of the methods of the invention when the computer program is running on the computer.
[0149] The encoded audio signal of the invention may be stored on a digital storage medium or may be transmitted using transmission means such as wireless transmission means or wired transmission means such as the Internet.
[0150] The above described embodiments are merely illustrative for the principles of the present invention. It should be understood that modifications and variants of the arrangements and details described herein are obvious to those skilled in the art. It is therefore intended that the restrictions arise only from the scope of the following claims and not from the specific details provided for the description and explanation of the present variants of the invention.
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV
VoiceAge Corporation
Proxy::
EP 2 301 024 B1 Z-9777
34 members in 17 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 8158608 | United States of America | P | |
| 09002270 | European Patent Office (EPO) | A | |
| 09797423 | European Patent Office (EPO) | A | |
| 2009004875 | European Patent Office (EPO) | W | |
| EP20090002270 | – | – | – |
| EP20090797423 | – | – | – |
| US20080081586P | – | – | – |
| WO2009EP04875 | – | – | – |
Members34
| Document | Office | Kind | |
|---|---|---|---|
| EP2146344A1 | European Patent Office (EPO) | A1 | |
| AU2009270524A1 | Australia | A1 | |
| WO2010006717A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201009814A | Taiwan Province of China | A | |
| CA2727883A1 | Canada | A1 | |
| HK1138673A1 | Hong Kong, China | A1 | |
| AR072551A1 | Argentina | A1 | |
| EP2301024A1 | European Patent Office (EPO) | A1 | |
| MX2011000534A | Mexico | A | |
| KR20110055515A | Republic of Korea | A | |
| CN102099856A | China | A | |
| US2011202355A1 | United States of America | A1 | |
| JP2011528129A | Japan | A | |
| AU2009270524B2 | Australia | B2 | |
| HK1156143A1 | Hong Kong, China | A1 | |
| RU2010154749A | Russian Federation | A | |
| EP2301024B1 | European Patent Office (EPO) | B1 | |
| CN102099856B | China | B | |
| US8321210B2 | United States of America | B2 | |
| ES2391715T3 | Spain | T3 | |
| PL2301024T3This record | Poland | T3 | |
| KR101224884B1 | Republic of Korea | B1 | |
| US2013066640A1 | United States of America | A1 | |
| RU2483364C2 | Russian Federation | C2 | |
| TWI441167B | Taiwan Province of China | B | |
| CA2727883C | Canada | C | |
| JP5613157B2 | Japan | B2 | |
| US8959017B2 | United States of America | B2 | |
| EP2146344B1 | European Patent Office (EPO) | B1 | |
| PT2146344T | Portugal | T | |
| ES2592416T3 | Spain | T3 | |
| PL2146344T3 | Poland | T3 | |
| BRPI0910999A2 | Brazil | A2 | |
| BRPI0910999B1 | Brazil | B1 |
Numbers
- Publication, DOCDB
- 2301024
- Publication, EPODOC
- PL2301024T
- Application
- 797423
- Application, DOCDB
- 09797423
- Application, EPODOC
- PL20090797423T
Titles2
- English
- AUDIO ENCODING/DECODING SCHEME HAVING A SWITCHABLE BYPASS
- Polish
- Sposób kodowania/dekodowania sygnału audio obejmujący przełączalne obejście
Classification
- CPC, 6
- G10L19/18
- G10L19/173
- G10L19/0017
- G10L19/008
- G10L19/0212
- G10L2019/0008
- IPC, 1
- G10L19 14