Audio coding method and apparatus
Abstract
Encoder and built-in method for encoding an audio signal, wherein a frequency spectrum of the audio signal is divided into a first and second region, wherein at least the second region comprises a number of bands. In addition, the spectral peaks in the first region are encoded by a first coding method. The method comprises: for a segment of the audio signal: determining a relation between an energy of a band in the second region and an estimate of energy of the first region. The method further comprises determining a relationship between the energy of the band in the second region and an energy of neighboring bands in the second region. The method further comprises determining whether an available number of bits is sufficient to encode at least one non-peak segment of the first region and the band in the second region. Further, when the relations meet a respective predetermined criterion and the number of bits is sufficient, the band in the second region and at least one segment of the first region are encoded using a second coding method. Otherwise, the band in the second region is subjected to BWE or noise filling.

Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
13 claims: 9 independent, 4 dependent
- 1REIVINDICACIONES 1. Un método para codificar una señal de audio, donde un espectro de frecuencia espectro de la señal de audio se divide en al menos una primera y una segunda región, donde al menos la segunda región comprende un número de bandas, y donde los picos espectrales en la primera región se codifican medíante un primer método de codificación, y donde cada segmento de pico espectral comprende un pico y un número determinado de compartimientos MDCT vecinos, caracterizado porque comprende:para un segmento de la señal de audio: -determinar (301) una relación entre una energía de una banda en la segunda región y una estimación de energía de la primera región;-determinar (302) una relación entre la energía de la banda en la segunda región y una energía de bandas vecinas en la segunda región;-determinar (303, 305) si un número disponible de bits es suficiente para codificar al menos un segmento de no pico de la primera región y la banda en la segunda región;y cuando las relaciones cumplan un criterio predeterminado respectivo (304) y el número de bits es suficiente (305): -codificar (306) la banda en la segunda región y al menos un segmento no pico de la primera región utilizando un segundo método de codificación que es diferente del primer método de codificación, y sino: -someter (307) la banda en la segunda región a la extensión de ancho de banda BWE o relleno de ruido.
- 2El método de acuerdo con la reivindicación 1, caracterizado porque el primer método de codificación es un método de codificación basado en pico que comprende codificar una posición del pico, una amplitud, y signo de un pico y un vector de forma que representa compartimientos de MDCT vecinos.
- 3El método de acuerdo con la reivindicación 1 o 2, caracterizado porque la estimación de energía de la primera región se basa en las energías de los picos espectrales en la primera región.
- 4El método de acuerdo con cualquiera de las reivindicaciones precedentes, caracterizado porque si la determinación del número de bits es suficiente (305) aplica una prioridad para codificar la región perceptualmente más relevante.
- 5El método de acuerdo con cualquiera de las reivindicaciones precedentes, caracterizado porque si la determinación del número de bits es suficiente para codificar la banda en la segunda región considera el mínimo número de bits requeridos para codificar al menos un coeficiente de la banda en la segunda región.
- 6El método de acuerdo con cualquiera de las reivindicaciones precedentes, caracterizado porque el segundo método de codificación comprende cuantificación vectorial o cuantificación vectorial de pirámide.
- 7Un codificador para codificar una señal de audio, donde un espectro de frecuencia de la señal de audio se divide en al menos una primera y una segunda región, donde al menos la segunda región comprende un número de bandas, el codificador configurándose para codificar segmentos de picos espectrales en la primera región utilizando un primer método de codificación, donde cada segmento de pico espectral comprende un pico y un número determinado de compartimientos MDCT vecinos, caracterizado porque el codificador además está configurado para:para un segmento de la señal de audio: -determinar una relación entre una energía de una banda en la segunda región y una estimación de energía de la primera región;-determinar una relación entre la energía de la banda en la segunda región y una energía de bandas vecinas en la segunda región;-determinar si un número disponible de bits es suficiente para codificar al menos un segmento no pico de la primera región y la banda en la segunda región;y cuando las relaciones cumplen con un criterio predeterminado respectivo y el número de bits es suficiente: -codificar la banda en la segunda región y al menos un segmento no pico de la primera región utilizando un segundo método de codificación que es diferente del primer método de codificación;y sino: -someter la banda en la segunda región a la extensión ancho de banda BWE o relleno de ruido.
- 8El codificador de acuerdo con la reivindicación 7, caracterizado porque el primer método de codificación es un método de codificación basado en pico que comprende codificar una posición del pico, una amplitud y signo de un pico y un vector de forma que representa los compartimientos MDCT vecinos.
- 9El codificador de acuerdo con la reivindicación 7 o 8, caracterizado porque la estimación de la primera región se basa en las energías de los picos espectrales en la primera región.
- 10El codificador de acuerdo con cualquiera de las reivindicaciones 7-9, caracterizado porque si la determinación del número de bits es suficiente (305) aplica una prioridad para codificar la región perceptivamente más relevante.
- 11El codificador de acuerdo con cualquiera de las reivindicaciones 7-10, caracterizado porque si la determinación del número de bits es suficiente para codificar la banda en la segunda región considera el mínimo número de bits requeridos para codificar al menos un coeficiente de la banda en la segunda región.
- 12El codificador de acuerdo con cualquiera de las reivindicaciones 7-11, caracterizado porque el segundo método de codificación comprende cuantificación vectorial o cuantificación vectorial de pirámide.
- 13Dispositivo de comunicación que comprende un codificador de acuerdo con cualquiera de las reivindicaciones 7-12.
Independent claims13
163 paragraphs in 5 sections, as filed
METHOD AND ENCODER FOR CODING AN AUDIO SIGNAL, AND DEVICE THAT INCLUDES SUCH ENCODER
TECHNICAL FIELD
The proposed technology generally refers to encoders and methods for audio coding.
The embodiments herein refer generally to audio coding where parts of the spectrum cannot be encoded due to bit rate limitations. In particular, it refers to bandwidth extension technologies where a significantly less important band is reconstructed using, for example, a parametric representation and approximations of a significantly more significantly encoded band.
BASIS
Most existing telecommunications systems operate in a limited audio bandwidth. Based on the limitations of fixed-line telephone systems, most voice services are limited to only transmitting the lower end of the spectrum. Although limited audio bandwidth is sufficient for most conversations, there is a desire to increase audio bandwidth to improve intelligibility and sense of presence. Although the capacity in telecommunications networks is continually increasing, it is still of great interest to limit the bandwidth required per communication channel. In the smallest transmission bandwidths of mobile networks for each call that produces lower power consumption on both the mobile device and the base station. This translates into energy and cost savings for the mobile phone operator, while the end user will experience prolonged battery life and longer talk time. In addition, with less bandwidth consumed per user, the mobile network can serve a greater number of users in parallel.
A property of the human auditory system is that perception is frequency dependent. In particular, our hearing is less accurate for higher frequencies. This has inspired the techniques of so-called bandwidth extension (BWE), where a high frequency band is reconstructed from a low frequency band using a low number of transmitted parameters.
Conventional BWE uses a parametric representation of the band signal
<img file="AR110293A2_D0001.tif" />
high, such as spectral envelope and temporal envelope, and reproduces the fine structure of the signal spectrum through the use of generated noise or a modified version of the low band signal. If the high band envelope is represented by a filter, the fine structure signal is often called the excitation signal. An exact representation of the high band envelope is significantly more important than the fine structure. Consequently, it is common that the resources available in terms of bits are spent on the envelope representation while the fine structure is reconstructed from the encoded low band signal without additional lateral information.
BWE technology has been applied in a variety of audio coding systems. For example, the 3GPP AMR-WB + uses a time domain BWE based on a low-band encoder that switches between the voice coding of the Excited Linear Code Indicator (CELP) and residual encoder transformer coding (TCX). Another example is the 3GPP eAAC based audio codec transformer that performs the transformation of the BWE domain variant called Spectral Band Replication (SBR).
Although division into a low band and a high band is often perceptually motivated, it may be less suitable for certain types of signals. As an example, if the high band of a particular signal is significantly more important than the lower band, the majority of the bits spent in the lower band will be in vain while the upper band will be represented with little precision. In general, if a portion of the spectrum is set to be encoded while other parts are not encoded, there may always be signals that do not fit the hypothesis a priori. The worst case scenario would be that all the signal energy is contained in the uncoded part which would produce very poor performance.
SHORT DESCRIPTION
It is an object to provide more flexible audio coding schemes. This and other objects are fulfilled through realizations of the proposed technology.
The proposed technology refers to adding decision logic to include a band or bands, a-priori assumed not to be important, in the fine structure coding. The decision logic is designed to maintain "conventional" behavior for signals where the a priori hypothesis for the coded boundaries and BWE boundaries is valid, while including parts of the non-important BWE region assumed a priori BWE in the coded region for the remaining signals
<img file="AR110293A2_D0002.tif" />
out of this group
An advantage of the proposed technology is to maintain the beneficial structure of a partially coded band based on a priori knowledge while extending it to handle specific cases of signals.
Other advantages will be appreciated when the detailed description is read.
According to a first aspect, a method is provided for encoding an audio signal, where a frequency spectrum of the audio signal is divided into at least a first and a second region, where at least the second region comprises a number of bands. In addition, the spectral peaks in the first region are encoded by a first coding method. The method provided here comprises: for a segment of the audio signal: determining a relationship between an energy of a band in the second region and an estimate of the energy of the first region. The method further comprises determining a relationship between the energy of the band in the second region and an energy of neighboring bands in the second region. The method further comprises determining if an available number of bits is sufficient for the coding of at least one non-peak segment of the first region and the band in the second region. In addition, when the relationships meet a respective predetermined criterion and the number of bits is sufficient, the band in the second region and the segment of at least one of the first region are encoded using a second coding method. Otherwise, the band in the second region is subjected instead to a BWE or noise fill.
According to a second aspect, an encoder is provided to encode an audio signal, where a frequency spectrum of the audio signal is divided into at least a first and second region, where at least the second region comprises a number of bands . The encoder is configured to encode spectral peaks in the first region using a first coding method. The encoder is further configured to: for a segment of the audio signal: determine a relationship between an energy of a band in the second region and an estimate of the energy of the first region; determine a relationship between the energy of the band in the second region and an energy of neighboring bands in the second region; determine if an available number of bits is sufficient to encode at least one non-peak segment of the first region and the band in the second region. The decoder is also configured: when the relationships meet a respective predetermined criterion and the number of bits is sufficient: encode the second band in the second region and the segment of at least one of the first region using a second coding method, and otherwise, to fasten the band in the second region to the BWE Extension or noise fill.
According to a third aspect, a communication device is provided, which comprises an encoder according to the second aspect.
According to a fourth aspect, a computer program is provided, which comprises instructions that, when executed on at least one processor, cause the processor to carry out at least one method according to the first and / or second aspect .
According to a fifth aspect, a carrier is provided, which contains the fourth aspect computer program.
BRIEF DESCRIPTION OF THE DRAWINGS
The foregoing and other objects, features, and advantages of the technology described herein will be apparent from the following more particular description of embodiments as illustrated in the accompanying drawings. The drawings are not necessarily to scale, the emphasis is instead placed on illustrating the principles of the technology described in this document.
Figure 1 is an example of a harmonic spectrum directed by the concept of coding presented. By comparison, the figure in the figure illustrates audio spectrum with slow spectral envelope variation;
Figure 2a is a structural view of the four different types of coding regions of the MDCT spectrum;
Figure 2b is an example of LF encoded region that models the space between spectral peaks;
Figure 3 is a flow chart illustrating a method according to an exemplary embodiment;
Figure 4 illustrates an introduction of a band encoded in the BWE region;
Figures 5a-c illustrate implementations of an encoder according to exemplary embodiments.
Figure 6 illustrates an embodiment of an encoder;
Figure 7 illustrates an embodiment of a computer impiementation of an encoder;
Figure 8 is a schematic block diagram illustrating an embodiment of an encoder comprising a group of function modules; Y
Figure 9 illustrates an embodiment of a coding method
DETAILED DESCRIPTION
The proposed technology is intended to be implemented in a codec, that is, an encoder and a corresponding decoder (often abbreviated as a codec). An audio signal is received and encoded by the encoder. The resulting encoded signal is supplied, and typically transmitted to a receiver, where it is decoded by a corresponding decoder. In some cases, the encoded signal is instead stored in a memory for later recovery.
The proposed technology can be applied to an encoder and / or decoder, for example, of a user terminal or user equipment, which can be a cable or wireless device. All alternative devices and nodes described in this document are summarized in the term "communication device", in which the solution described herein could be applied.
As used herein, the non-limiting terms "User Equipment" and "wireless device" may refer to a mobile phone, a cell phone, a personal digital assistant, PDA, equipped with radio communication capabilities, a Smart phone, a laptop or personal computer, PC, equipped with an internal or external mobile broadband modem, a Tablet PC with radio communication capabilities, a target device, an EU device to device, a type of EU or EU machine capable of machine-to-machine communication, PAD, local customer equipment, CPE, portable integrated equipment, LEE, portable mounted equipment, LME, USB device, a device Portable electronic radio communication, a sensor device equipped with radio communication capabilities or the like. In particular, the term "UE" and the term "wireless device should be construed as non-limiting terms comprising any type of wireless device that communicates with a radio network node in a cellular or mobile communication system or any device equipped with Radio circuits for wireless communication according to any relevant standard for communication within a cellular or mobile communication system.
As used herein, the term "wired device" may refer to any device configured or prepared for wired connection to a network. In particular, the wired device may be at least some of the previous devices, with or without radio communication capability, when configured for wired connection.
The proposed technology can also be applied to an encoder and / or decoder of a radio network node. As used herein, the non-limiting term "radio network node" may refer to base stations, network control nodes such as network controllers, radio network controllers, base station controllers, and the like In particular, the term "base station may cover different types of radio base stations including standardized base stations such as Bs Node, or evolved Bs node, eNBs, and also macro / micro / peak radio base stations, home base stations , also known as femto base stations, relay nodes, repeaters, radio access points, base transceiver stations, BTSs, and even radio control nodes controlling one or more remote radio units, RRUs, or the like.
As for the relative terminology of the frequency spectrum of the audio signal to be encoded, here we will try to explain some of the terms used. As described above, audio frequencies are often divided into a so-called “low band,” (LB), or “low frequency band,” (LF); and a so called “high band” (HB), or “high frequency band” (HF). Typically, the high band is not encoded in the same way as the low band, but instead subjected to BWE. The BWE may comprise the coding of a spectral and a temporal envelope, as described above. However, an extended high bandwidth bandwidth can still be referred to as not encoded herein. In other words, a "high band without coding" can still be associated with some coding of, for example, envelopes, but this coding can be assumed to be associated with much less bits than the coding of the encoded regions.
In this document, the terminology "a first region" and "a second region" will be used, referring to the parts of the audio spectrum. In a preferred embodiment, the first region can be assumed to be low band and the second region can be assumed to be high band, as in conventional audio coding using BWE. However, there may be more than two regions, and the regions can be configured differently.
The proposed technology is embedded in the context of an audio codec that is focusing signals with a strong harmonic content. An illustration of audio signals is presented in Figure 1. The upper audio spectrum in Figure 1 is an example of a harmonic spectrum, that is, an example of a spectrum of an audio signal with a strong harmonic content. . For comparison, the background spectrum in Figure 1 illustrates an audio spectrum with a spectral envelope that varies slowly.
In an exemplary embodiment, encoding and decoding is performed in the frequency domain transformed using the Modified Discrete Cosine Transform (MDCT). The harmonic structure is modeled using a specific peak coding method in the so-called low band, which is complemented by a vector quantifier (VQ) focusing on the important low frequency coefficients (LF) of the MDCT spectrum and a region of BWE where the highest frequencies are generated from low band synthesis. An overview of this system is represented in Figure 2a and 2b.
Figure 2a shows a structural view of the four different types of coding regions of an MDCT spectrum. In the low band, spectral peaks are encoded using a peak-based coding method. In the high band, the BWE (dotted line) is applied, which may involve coding, that is, some parametric representation, of information related to the spectral envelope and the temporal envelope. The region marked LF encoded in Figure 2a (double line) is encoded using a gain-form coding method, that is, not the same coding method as the one used for the peaks. The region of encoded LF is dynamic in the sense that it depends on the remaining amount of bits, on a budget of bits, which are available to be encoded when the peaks have been encoded. In Figure 2b, the same regions as in Figure 2a can be seen, but here it can also be seen that the region of encoded LF extends between encoded peaks. In other words, also parts of the low band spectrum located between the peaks can be modeled by the gain-form coding method, depending on the peak positions of the target spectrum and the available number of bits. Parts of the spectrum that comprise the encoded peaks are excluded from the gain-form coding of the low frequency region, that is, of the encoded LF region. The parts of the low band that remain decoded when the available bits are spent on the peak coding and the LF coding are subjected to the noise fill (dashed line in Figure 1).
Assuming that the previous structure, that is, a first region where the peaks and important coefficients / parts without peaks are encoded, and a second region, possibly also denoted BWE region, of which there is an a priori assumption of not understanding as perceptually important information as the first region,
<img file="AR110293A2_D0003.tif" />
A novel technique is proposed to add spectral component coding in the BWE region. The idea is to introduce an encoded band in the BWE region (see figure 3) if certain requirements are completed. More than one encoded band could also be introduced when appropriate.
Since it is an objective to maintain a structure of a coded region, such as a low frequency part of the spectrum, and a second region, such as a high frequency part of the spectrum that extends bandwidth for most signals , a band encoded in the second region should, in one embodiment, only be introduced if certain conditions with respect to the band being met. The conditions, or criteria, for a candidate band, in a second region, evaluated for coding can be formulated as follows:
1. The energy in the candidate band, for example, a frequency band in a high frequency part of the spectrum, must be relatively high compared to an estimate of the energy of a coded peak region, for example, in the lower part of the spectrum of frequency. This energy ratio indicates an audible, and therefore perceptually relevant, band in the second region.
two. The candidate band must have a relatively high energy compared to the neighboring bands in the second region. This indicates a peak structure in the second region that cannot be modeled well with the BWE technique.
3. The resources, that is, bits, to encode the candidate band should not compete with the most important components (cf. LF Coded in Figure 2a and 2b) in the coding of parts of the encoded region.
Exemplary Embodiments
Next, exemplary embodiments related to a method for encoding an audio signal with reference to Figure 3 will be described. The frequency spectrum of the audio signal is divided into at least a first and second region, where at least the second region It comprises a number of bands, and where the spectral peaks in the first region are encoded by a first coding method. The method to be performed by an encoder with a corresponding method in the decoder. The encoder and decoder can be configured to be compatible with one or more standards for audio coding and decoding. The method comprises, a segment of the audio signal:
-determine 301 a relationship between an energy of a band in the second
<img file="AR110293A2_D0004.tif" />
region and an estimate of the energy of the first region;
-determine 302 a relationship between the energy of the band in the second region and an energy of neighboring bands in the second region; - determine 303, 305 if an available number of bits is sufficient to encode at least one non-peak segment of the first region and the band in the second region; and, when the relationships complete 304 a respective predetermined criterion and the number of bits is sufficient 305:
-coding 306 the band in the second region and at least one segment of the first region using a second coding method; if not:
- subject 307 the band in the second region to the BWE or noise fill.
The first region would typically be a lower part of the frequency spectrum than the second region. The first region can, as previously mentioned, be called low band, and the second region can be called high band. The regions do not overlap, and may be adjacent. Other regions are also possible, which can, for example, separate the first and second regions.
At least one segment of the first region, cf. part of "LF Coded" in Figure 2a and 2b, and the candidate band selected to be encoded in the second region is encoded using the same second coding method. This second coding method may comprise vector quantification or vector quantification pyramid. Since the energy envelope or the gains of the bands in the second region are already encoded to aid BWE technology, it is beneficial to complement this coding with a quantifier of form applied to the fine structure of the selected candidate band. In this way, a gain coding of the selected candidate band is achieved. In some preferred embodiments, the spectral peaks in the first region are encoded by a first coding method, as mentioned above. The first coding method is preferably a peak-based coding method, as described in, for example, 3GPP TS 26.445, section 5.3.4.2.5. A second coding method is exemplified in the same document in section 5.3.4.2.7.
Figure 4 illustrates a possible result of applying an embodiment of the method described above. In figure 4, a band, B<sub>H</sub>b, it is coded in the second region instead of being attached to the BWE (as in Figures 2a and 2b), which would have been the case if an embodiment of the method was not applied. Band B<sub>H</sub>b is encoded using the same coding method as used for the parts marked "Coded LF" in Figure 4. The peaks in the first region, marked "encoded peak", however, are
<img file="AR110293A2_D0005.tif" />
encode using another coding method, which is preferably based on peak. Note that since the content of the second region is not strictly populated using the BWE or other spectral filling techniques, the a priori assumption of an encoded band and an uncoded band is no longer true. For this reason it may be more appropriate to call the filling strategy a noise filling. The term noise filling is most often used for spectral filling in regions that can appear anywhere in the spectrum and / or between coded parts of the spectrum.
The determination of the relationships between energies and the sufficiency of a number of bits available for encoding corresponds to the three conditions, listed 1-3, described above. Examples of how the determination can be made will be described below. The evaluation is described for a candidate band in the second region.
Condition Assessment 1
The first condition refers to the energy in the candidate band should have a certain relationship with an estimate of the energy of a coded peak region. This relationship is described herein as that the energy of the candidate band should be relatively high compared to the energy estimate of the first region. Assuming, as in the example, that coding is carried out in the frequency transformed domain using the Cosine Transformer
Discrete Modified, where the MDCT coefficients are calculated as:
<td>ΓΓ</td><td>π</td><td></td><td>Ϊ <sup>1</sup></td><td></td><td>π</td>
<td></td><td> —</td><td>eos</td><td>n + - + -</td><td>k H—</td><td> —</td>
<td>Ll 2;</td><td>L</td><td></td><td> .1 2 2)</td><td> < 2/</td><td>L</td>
<img file="AR110293A2_D0006.tif" />
where x (rí) denotes a frame of input audio samples with frame index i. Here, 77, is an index of time domain samples, and k the index of frequency domain coefficients. For simplicity of notation, the frame rate i will be omitted when all calculations are made within the same frame. In general, it should be understood that all calculations derived from the input audio frame x (n) will be executed on a frame basis and all of the following variables could be denoted with an index i.
The logarithm energies of band E (jj of the second region, eg, high band region, can be defined as:
<img file="AR110293A2_D0007.tif" />
E (j) = 21og
<img file="AR110293A2_D0008.tif" />
(2) where is the first coefficient in the band j and N j refers to the number of MDCT coefficients in the band. A typical number for a high frequency region is 24-64 coefficients per band. It should be noted that el21og<sub>2</sub>(-) is merely an example that was found suitable in the directed audio coding system and that other logarithm bases and scale factors can be used. Using other logarithm bases and scale factors would give different absolute logarithm energy values and require different threshold values, but the method in other aspects remains the same.
As previously described, the spectral peaks in the first region are preferably encoded using a peak based on the coding method. The coded peaks of the first region, e.g., lower frequency region, are modeled in this example using the peak position p (m), an amplitude, including the sign,
G (m) that is set to match the MDCT container at the given position r (pi mj), and a shape vector V (m) representing the neighboring peaks, e.g., the four neighboring MDCT containers, where m = I ..N<sub>peaks</sub>and N<sub>peaks</sub> It is the number of peaks used in the representation of the first region.
To assess compliance with condition 1 above we would like to make an estimate of the energy in the first region, to compare to the candidate band energy. Assuming that most of the energy in the first region is contained within the modeled peaks, an estimate of the energy in the first region, E<sub>peak</sub>(i), from the framework i can be derived as;
peak (z) = 21og<sub>2</sub> (3)
Now, condition 1 can be assessed by setting a threshold for the envelope energy E (j) of the candidate band j, such as;
Axis<sub>peak</sub>^> T<sub>x</sub> (5) where T \ is a logarithmic threshold energy to pass, that is, complying with condition 1. Due to the computational complexity of the logarithm function, the following mathematically equivalent alternative can be used:
<img file="AR110293A2_D0009.tif" />
(6)
2^0)/2,2<sup>_7</sup>ΐ<sup>/2</sup> > 2<sup>α</sup>'^<sup>(</sup>'<sup>)/2</sup> ο
2 ^ (/) / 2 qTpeakd! 2 ^ / 2
The threshold value should be set such that it corresponds to the perceptual importance of the band. The current value may depend on the band structure. In an exemplary embodiment, a suitable value for 2<sup>1</sup> it was found to be what<sup>-5</sup>. Condition Evaluation 2
The second condition refers to the fact that the energy in the candidate band should have a certain relationship with an energy of neighboring bands in the second region. This relationship is expressed in this document, as the candidate band should have a relatively high energy compared to neighboring bands in the second region.
An example of how to assess compliance with condition 2 is to compare the logarithm energy of the candidate band with the average logarithm energy of the second entire region, eg high band. First, an average log energy of the second region can be defined as:
<img file="AR110293A2_D0010.tif" />
(8)
Then, an expression for condition 2 can be formulated as:
~ E> T¿ (gj where T<sub>2</sub> denotes the logarithm energy threshold to pass condition 2. Equivalently, as for condition 1, this can be formulated in the energy domain rather than the logarithm domain, cf. Equation (6), if this is seen as beneficial for an aspect of computational complexity. In an exemplary embodiment, a suitable value for T<sub>2</sub> It was found to be 3. As an alternative to using the average log energy of the second entire region, only parts of the second region can be used, eg, a number of bands surrounding the candidate band.
Condition Evaluation 3
The third condition refers to whether an available number of bits is sufficient to encode at least one non-peak segment of the first region and the band in the second region. Otherwise the band in the second region should not be encoded.
Condition 3 refers to the intended coding method, denoted "second coding method" above, which is a gain form coding.
<img file="AR110293A2_D0011.tif" />
The general VQ directed for the LF region encoded, that is, non-peak portions of the first region, is in accordance with an embodiment configured to also cover selected bands in the second region, eg, high frequency region. However, since the first region, typically a low frequency region, is sensitive to MDCT domain coding, it should be ensured that some resources, bits, are allocated to encode at least part of this frequency range. Since the preferred general pyramid vector quantifier (PVQ) that is targeted for non-peak part coding of the first region (cf. "LF encoded" in Figures 2a and 2b) is operating in a target spectrum divided into bands, this requirement is met by ensuring that at least one band is allocated for the first region, that is:
N<sub>hanJ</sub>> l (10) where N<sub>hand</sub> denotes the number of bands in the destination signal for the coded LF part. These bands are not the same type of band as the bands in the second region. The band N<sub>hand</sub> here is a band with a width given by the encoder, and the band comprises a part of the first region that is not encoded by a peak coding method.
In case there are enough bits available to encode at least one non-peak part of the first region, and a selected band, which meets the conditions 1-2 above, the selected band can be encoded together with at least one non-part peak of the first region using the second coding method (gain form). Another useful condition to avoid wasting resources is to ensure that the bit range for the selected band is high enough to represent the band with acceptable quality. If not, the bits spent in encoding the selected band would be wasted, and would be better spent in encoding more than the low frequency part of the first region (cf. plus LF encoded in Figure 2a) In an exemplary embodiment, the coding of a noptic part of the first region is handled using PVQ, which has an explicit relationship between the number of pulses, vector length and the required bit rate defined by the pulses2bits function (W, where W<sub>¡</sub> denotes the bandwidth of the selected band and P<sub>rain</sub> It is the minimum number of pulses that should be represented. Assume that B<sub>lasl</sub> denotes the number of bits assigned to the last band in the target vector for the PVQ encoder, then the condition to avoid wasting resources can be written as:
<img file="AR110293A2_D0012.tif" />
(11) <sup>B</sup>i<sub>ust</sub> > pulses2bits (W<sub>/</sub>, )
The minimum number of pulses P<sub>min</sub> It is an adjustment parameter, but it must be at least /<sub>nin</sub> > 1. Equations (10) and (11) together fulfill condition 3 in an exemplary embodiment.
A novel part of embodiments described herein is a decision logic for the evaluation of whether to encode a band in a BWE region or not. By BWE region is meant here a region, defined, for example in frequency, that an encoder without the functionality of this suggested document would have been attached to the BWE. For example, the BWE region could be frequencies above 5.6 kHz, or above 8 kHz.
As an example of embodiments described above, it suggests a structure where the so-called low band is coded and the so-called "high band" extends from the low band. The terms "low band" and "high band" refer to the parts of a frequency spectrum that is divided at a certain frequency. That is, a frequency spectrum divided into a lower part, a "low band" and a higher part, a "high band" at a certain frequency, eg 5.6 or 8 kHz. The solution described herein is not limited, however, to such frequency partitioning but can also be applied to other encoded and uncoded distributions, that is, estimates, regions, where the encoded and estimated regions or parts are decided. , based on a priori knowledge about the source and perceptual importance of the signal in the hand.
An exemplary embodiment of a method for encoding an audio signal comprises receiving an audio signal and further analyzing at least a portion of the audio signal. The method further comprises determining, based on the analysis, whether to encode a high band region of a frequency spectrum of the audio signal together with a low band region of the frequency spectrum. The exemplary method further comprises encoding the audio signal for transmission over a junction in a communication network based on the determination of whether to encode the high band region.
The analysis described above can also be performed on the quantized and reconstructed parameters in the encoder. Logarithmic energies £ (j) in that case would be replaced with their quantified counterpart ¿(/ jen equation (8) and peak gains G (m) would be replaced with quantified peak gains G (m) er \ equation (3 ). Using the quantified parameters allows the method described above to be implemented in the same way in the corresponding encoder and decoder, since the quantized parameters are available for both. That is, the method described above is also performed in the decoder, in order to determine how to decode and reconstruct the audio signal. The advantage of this configuration is that no additional information needs to be transported from the encoder to the decoder, indicating whether a band in the second region has been encoded or not. A solution where information is transported, which indicates whether a band in the second region is encoded or not is also possible.
A method for decoding an audio signal, corresponding to the method for encoding an audio signal described above, will be described below. As before, a frequency spectrum of the audio signal is divided into at least a first and a second region, where at least the second region comprises a number of bands, and where the spectral peaks in the first region are decoded using a first coding method The method, which will be performed by a decoder, comprises, for a segment of the audio signal:
-determine a relationship between an energy of a band in the second region and an estimate of the energy of the first region;
-determine a relationship between the energy of the band in the second region and an energy of the neighboring bands in the second region;
-determine if an available number of bits is sufficient to encode at least one non-peak segment of the first region and the band in the second region. The method also includes:
when the relationships meet a respective predetermined criterion (304) and the number of bits is sufficient:
-decoding a band in the second region and at least one segment of the first region using a second coding method; if not
-rebuild the band in the second region based on the BWE or noise fill.
Implementations
The method and techniques described above can be implemented in encoders and / or decoders, which can be part of eg communication devices.
Encoder, fiquras 5a-5c
An exemplary embodiment of an encoder is illustrated in a general manner in
<img file="AR110293A2_D0013.tif" />
Figure 5a. By encoder refers to an encoder configured to encode audio signal. The encoder could possibly also be configured to encode other types of signals. The encoder 500 is configured to perform at least one of the embodiments of the method described above with reference, eg, figure 3. The encoder 500 is associated with the same technical characteristics, objects and advantages as the embodiments of the method described above. The encoder can be configured to be compatible with one or more standards for audio coding. The encoder will be described shortly in order to avoid unnecessary repetition.
The encoder can be implemented and / or described as follows:
The encoder 500 is configured to encode an audio signal, where a frequency spectrum of the audio signal is divided into at least a first and a second region, where at least the second region comprises a number of bands, and where the peaks Spectral in the first region are encoded by a first coding method. The encoder 500 comprises processing circuits, or processing means 501 and a communication interface 502. The processing circuit 501 is configured to cause the encoder 500, for a segment of the audio signal: to determine a relationship between an energy of a band in the second region and an estimate of the energy of the first region. The processing circuit 501 is further configured to cause the encoder to determine a relationship between the energy of the band in the second region and an energy of neighboring bands in the second region. The processing circuit 501 is further configured to cause the encoder to determine if an available number of bits is sufficient to encode at least one non-peak segment of the first region and the band in the second region. The processing circuit 501 is further configured to cause the encoder, when the ratios meet a respective predetermined criterion and the number of bits is sufficient, encode the band in the second region and at least one segment of the first region using a second method. of coding. Otherwise, when at least one of the relationships does not meet the predetermined criteria and / or when the number of bits is not sufficient, the band in the second region is subject to the BWE or noise fill. Communication interface 502, which can also be denoted, for example, input / output interface (I / O), includes an interface for sending data and receiving data from other entities or modules.
The processing circuit 501 could, as illustrated in Figure 5b, comprise processing means, such as a processor 503, for example, a
<img file="AR110293A2_D0014.tif" />
CPU, and a 504 memory to store or hold instructions. The memory would then comprise instructions, for example, in the form of a computer program 505, which when executed through the processing means 503 causes the encoder 500 to perform the actions described above.
An alternative implementation of the processing circuit 501 is shown in Figure 5c. The processing circuit here comprises a first determination unit 506, configured to cause the encoder 500: to determine a relationship between an energy of a band in the second region and an energy estimate of the first region. The processing circuit further comprises a second determining unit 507 configured to cause the encoder to determine a relationship between the energy of the band in the second region and an energy of neighboring bands in the second region; The processing circuit further comprises a third determining unit 508, configured to cause the encoder to determine if an available number of bits is sufficient to encode at least one non-peak segment of the first region and the band in the second region. The processing circuit further comprises an encoding unit, configured to cause that encoder, when the relationships meet a respective predetermined criterion and the number of bits is sufficient, encode the band in the second region and at least one segment of the first region using a first coding method. The processing circuit 501 could comprise more units, such as a decision unit configured to cause the encoder to decide whether the determined relationships meet the criteria or not. This task could alternatively be performed by one or more of the other units.
The encoders, or codees, described above could be configured for the different embodiments of the method described herein, such as using different methods of encoding profit forms such as the second coding method; different peak coding methods to encode the peaks in the first region, operating in different domains to transform, etc.
The encoder 500 can be assumed to further comprise functionality, to perform regular encoder functions.
Figure 6 illustrates an embodiment of an encoder. An audio signal is received and the bands of a first region, typically the low frequency region, are encoded. In addition, at least one band of a second region, Typically, the high frequency region, exclusive to the first region, is encoded. Depending on the conditions discussed above, you can decide whether or not to include the
<img file="AR110293A2_D0015.tif" />
coding of the band in the second region in the final encoded signal. The final encoded signal is typically provided to a receiving party, where the encoded signal is decoded into an audio signal. The UE or network node may also include radio circuitry for communication with one or more other nodes, including transmission and / or reception of information.
In the following, an example of an equipment implementation will be described with reference to Figure 7. The encoder comprises processing circuits such as one or more processors and a memory. In this particular example, at least some of the steps, functions, procedures, modules and / or blocks described in this document are implemented in a computer program, which is loaded into memory for execution by the processing circuit. The processing and memory circuit are interconnected to allow normal software execution. The optional input / output device can also be interconnected to the processing circuit and / or memory to allow the input and / or output of relevant data as input parameter (s) and / or resulting output parameter (s). An encoder could alternatively be implemented using function modules, as illustrated in Figure 8.
An exemplary embodiment of an encoder for encoding an audio signal could be described as follows:
The encoder comprises a processor; and a memory to store Instructions that, when executed by the processor, cause the encoder to: receive an audio signal; analyze at least a portion of the audio signal; and: based on the analysis, determine whether to encode a high band region of a frequency spectrum of the audio signal together with a low band region of the frequency spectrum; and also: based on the determination of whether to encode the high band region, encode the audio signal for transmission over a link in a communication network.
The encoder could be understood in a user equipment for operation in a wireless communication network.
The term 'computer' should be interpreted in a general sense as any system or device capable of executing program code or computer program instructions to perform a particular process, determination or calculation task.
In a particular embodiment, the computer program comprises instructions, which when executed by at least one processor, causes the (the)
<img file="AR110293A2_D0016.tif" />
processor (s) encode the bands of a first frequency region, to encode at least one band of a second region, and depending on the specified conditions decide whether or not to encode the band in the second region in the encoded signal final.
It will be appreciated that the methods and devices described herein can be combined and re-arranged in a variety of ways. For example, the embodiments can be implemented in hardware, or in software for execution by suitable processing circuits, or a combination thereof.
The steps, functions, procedures, modules and / or blocks described herein can be implemented in the hardware using any conventional technology, such as discrete circuit or integrated circuit technology, including both general-purpose electronic circuits and circuitry. specific application
Particular examples include one or more configured digital signal processors and other known electronic circuits, for example, interconnected discrete logic gates to perform a specialized function, or Application Specific Integrated Circuits (ASICs).
Alternatively, at least some of the steps, functions, procedures, modules and / or blocks described herein may be implemented in the software such as a computer program for execution by suitable processing circuits such as one or more. processors or processing units.
The flow chart or diagrams presented in this document can therefore be considered as a computer flow chart or diagrams, when performed by one or more processors. A corresponding device can be defined as a group of function modules, where each step performed by the processor corresponds to a function module. In this case, the function modules are implemented as a computer program that runs on the processor.
It should also be noted that in some alternative implementations, the functions / acts indicated in the blocks may occur outside the order indicated in the flowcharts. For example, two blocks shown in succession can in fact be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order, depending on the functionality / acts involved. On the other hand, the functionality of a given block of the diagrams of
<img file="AR110293A2_D0017.tif" />
Flow and / or block diagrams can be separated into multiple blocks and / or the functionality of two or more blocks of the flow diagrams and / or block diagrams can at least be partially integrated. Finally, other blocks may be added / inserted between the blocks illustrated, and / or blocks / operations may be omitted without departing from the scope of the inventive concepts.
It is to be understood that the choice of the interaction units, as well as the nomenclature of the units within this disclosure are for an exemplary purpose only, and the nodes suitable for executing any of the methods described above may be configured in a plurality of ways. alternatives in order to be able to execute the suggested procedural actions.
It should also be noted that the units described in this disclosure should be considered as logical entities and not with the need as separate physical entities.
Examples of processing circuits include, but are not limited to, one or more microprocessors, one or more digital signal processors, DSPs for their acronym in English, one or more central processing units, CPUs for their acronym in English, acceleration hardware video, and / or any suitable programmable logic circuit such as one or more programmable door arrays, FPGAs for its acronym in English, or one or more programmable logic controllers, PLC for its acronym in English.
It should also be understood that it may be possible to reuse the general processing capabilities of any conventional device or unit in which the proposed technology is implemented. It may also be possible to reuse existing software, e.g. by reprogramming existing software or adding new software components.
The proposed technology provides an encoder usable in a UE or network node configured to encode audio signals, where the encoder is configured to perform the necessary functions.
In a particular example, the encoder comprises a processor and a memory, the memory comprising instructions executable by the processor, whereby the apparatus / processor is operative to perform the coding and decision steps.
The proposed technology also provides a carrier comprising the computer program, wherein the carrier is one of an electronic signal, a
<img file="AR110293A2_D0018.tif" />
optical signal, an electromagnetic signal, a magnetic signal, an electrical signal, a radio signal, a microwave signal, or a computer-readable storage medium.
The software or computer program can therefore be realized as a computer program product, which is normally performed or stored in a computer-readable medium. The computer-readable medium may include one or more removable or non-removable memory devices including, but not limited to a read-only memory, ROM, random access memory, RAM, a compact disk, CD, a digital versatile disk, DVD, a Blueray disc, a universal serial device, USB, memory, a hard disk drive, HDD storage device, flash memory, a magnetic tape, or any other conventional memory device. The computer program can therefore be loaded into the operating memory of a computer or equivalent processing device for execution by the processing circuit thereof. That is, the software could be transported by a vehicle, such as an electronic signal, an optical signal, a radio signal, or a computer-readable storage medium before and / or during the use of the computer program on the network nodes. .
For example, the computer program stored in memory includes program instructions executable by the processing circuit, whereby the processing circuit is capable or operative to execute the steps, functions, procedures and / or blocks described above. The encoder is thus configured to perform, when the computer program is executed, well-defined processing tasks such as those described herein. The computer or processing circuit does not have to devote itself to just executing the steps, functions, procedures and / or blocks described above, but it can also perform other tasks.
As indicated herein, the encoder can alternatively be defined as a group of function modules, where the function modules are implemented as a computer program by running at least one processor. Figure 8 is a schematic block diagram illustrating an example of an encoder comprising a processor and associated memory. The computer program that resides in memory can therefore be organized as appropriate function modules configured to perform, when executed by the processor, at least part of the steps and / or tasks described herein. A
<img file="AR110293A2_D0019.tif" />
An example of such function modules is illustrated in Figure 6.
Figure 8 is a schematic block diagram illustrating an example of an encoder comprising a group of function modules.
The embodiments described above are merely given as examples, and it should be understood that the proposed technology is not limited to this. It should be understood by those skilled in the art that various modifications, combinations and changes can be made to the embodiments without departing from the present scope as defined by the appended claims. In particular, different part solutions in the different embodiments can be combined in other configurations, where technically possible.
Contents5
29 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29
47 members in 12 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 201461953331 | United States of America | P | |
| 201461953331 | United States of America | P | |
| 61953331 | United States of America | – | |
| 61953331 | – | – | – |
| US201461953331P | – | – | – |
Members47
| Document | Office | Kind | |
|---|---|---|---|
| WO2015136078A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AR099761A1 | Argentina | A1 | |
| US2016254004A1 | United States of America | A1 | |
| IL247337A0 | Israel | A0 | |
| IL247337D0 | Israel | D0 | |
| MX2016011328A | Mexico | A | |
| CN106104685A | China | A | |
| EP3117432A1 | European Patent Office (EPO) | A1 | |
| BR112016020988A2 | Brazil | A2 | |
| US9741349B2 | United States of America | B2 | |
| US2017316788A1 | United States of America | A1 | |
| MX353200B | Mexico | B | |
| US10147435B2 | United States of America | B2 | |
| US2019057707A1 | United States of America | A1 | |
| AR110293A2This record | Argentina | A2 | |
| EP3117432B1 | European Patent Office (EPO) | B1 | |
| IL265424A | Israel | A | |
| TR2019007596T4 | Türkiye | T4 | |
| TR201907596T4 | Türkiye | T4 | |
| EP3518237A1 | European Patent Office (EPO) | A1 | |
| IL265424B | Israel | B | |
| IL268543A | Israel | A | |
| PL3117432T3 | Poland | T3 | |
| MX369614B | Mexico | B | |
| CN106104685B | China | B | |
| CN110619884A | China | A | |
| MX2019012777A | Mexico | A | |
| US10553227B2 | United States of America | B2 | |
| ES2741506T3 | Spain | T3 | |
| CN110808056A | China | A | |
| IL268543B | Israel | B | |
| US2020126573A1 | United States of America | A1 | |
| BR112016020988A8 | Brazil | A8 | |
| BR112016020988B1 | Brazil | B1 | |
| EP3518237B1 | European Patent Office (EPO) | B1 | |
| DK3518237T3 | Denmark | T3 | |
| ES2930366T3 | Spain | T3 | |
| EP4109445A1 | European Patent Office (EPO) | A1 | |
| CN110619884B | China | B | |
| CN110808056B | China | B | |
| EP4109445B1 | European Patent Office (EPO) | B1 | |
| EP4109445C0 | European Patent Office (EPO) | C0 | |
| EP4465296A2 | European Patent Office (EPO) | A2 | |
| EP4465296A3 | European Patent Office (EPO) | A3 | |
| ES2995635T3 | Spain | T3 | |
| US12236967B2 | United States of America | B2 | |
| US2025201254A1 | United States of America | A1 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant, registrationFG | FG |
Numbers
- Publication
- 110293
- Publication, DOCDB
- 110293
- Publication, EPODOC
- AR110293
- Application
- 103355
- Application, DOCDB
- P170103355
- Application, EPODOC
- AR2017P103355
Titles2
- Spanish
- MÉTODO Y CODIFICADOR PARA CODIFICAR UNA SEÑAL DE AUDIO, Y DISPOSITIVO QUE COMPRENDE DICHO CODIFICADOR
- English
- METHOD AND ENCODER FOR CODING AN AUDIO SIGNAL, AND DEVICE COMPRISING SUCH ENCODER
Classification
- CPC, 11
- G10L19/0204
- G10L19/028
- G10L25/51
- G10L19/002
- G10L19/18
- G10L25/21
- G10L19/02
- G10L19/038
- G10L25/06
- G10L25/18
- H04L2012/5632
- IPC, 2
- G10L19 00
- G10L19 02