Method and apparatus for obtaining an attenuation factor
Abstract
A method for treating a voice signal synthesized in packet loss concealment, which method comprises: obtaining a tendency to change the voice signal, which comprises obtaining a relationship between the energy of the last periodic tonal voice signal and the energy of a previous periodic tonal voice signal, of the voice signal; obtain an attenuation factor according to the tendency to change the signal; and obtain a lost frame, reconstructed after attenuation according to the attenuation factor.

Term
2.1 yearsto projected expiry
Projected expiry 5 November 2028, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
15 claims: 4 independent, 11 dependent
- 1ES 2 340 975 T3 IS 2 340 975 T3 CLAIMS REIVINDICACIONES 1. A method for handling a synthesized voice signal in packet loss concealment, which method comprises:1. Un método para tratar una señal de voz sintetizada en ocultación de pérdida de paquetes, cuyo método comprende: obtaining a trend of change of the speech signal, comprising obtaining a relationship between the energy of the last periodic tonal speech signal and the energy of a previous periodic tonal speech signal, of the speech signal;obtener una tendencia al cambio de la señal de voz, que comprende obtener una relación entre la energía de la última señal de voz tonal periódica y la energía de una señal de voz tonal periódica previa, de la señal de voz;obtain an attenuation factor according to the tendency to change the signal;and obtaining a lost frame, reconstructed after fading according to the fading factor. obtener un factor de atenuación de acuerdo con la tendencia al cambio de la señal;y obtener una trama perdida, reconstruida después de atenuación de acuerdo con el factor de atenuación.
- 11Un aparato para tratar una señal de voz sintetizada en ocultación de pérdida de paquetes, cuyo aparato comprende:eleven. An apparatus for processing a synthesized voice signal in packet loss concealment, which apparatus comprises: a unit for obtaining a tendency to change, comprising an energy obtaining sub-unit, intended to obtain energy from a last periodic tonal voice signal and energy from a previous periodic tonal voice signal, from the voice signal;una unidad de obtención de una tendencia al cambio, que comprende una sub-unidad de obtención de energía, destinada a obtener energía de una última señal de voz tonal periódica y energía de una señal de voz tonal periódica previa, de la señal de voz;a sub-unit for obtaining an energy ratio, intended to obtain a ratio between the energy of the last periodic tonal voice signal and the energy of the previous periodic tonal voice signal, of the voice signal;una sub-unidad de obtención de una relación energética, destinada a obtener una relación entre la energía de la última señal de voz tonal periódica y la energía de la señal de voz tonal periódica previa, de la señal de voz;ES 2 340 975 T3 a unit for obtaining an attenuation factor, intended to obtain the attenuation factor according to the relation obtained by the sub-unit for obtaining an energy relation;and a lost frame reconstruction unit, intended to obtain a lost frame, reconstructed after attenuation according to the attenuation factor. ES 2 340 975 T3 una unidad de obtención de un factor de atenuación, destinada a obtener el factor de atenuación de acuerdo con la relación obtenida por la sub-unidad de obtención de una relación energética;y una unidad de reconstrucción de tramas perdidas, destinada a obtener una trama perdida, reconstruida después de atenuación de acuerdo con el factor de atenuación.
- 14A speech decoder, comprising:a low-band decoder unit, a high-band decoder unit, and a quadrature specular filtering unit, wherein: 14. Un descodificador de voz, que comprende: una unidad descodificadora de banda baja, una unidad descodificadora de banda alta y una unidad de filtrado especular en cuadratura, en el que: la unidad descodificadora de banda baja está destinada a descodificar una señal de voz de descodificación de banda baja recibida, y a compensar una señal de voz de banda baja perdida;the low-band decoding unit is adapted to decode a received low-band decoding speech signal, and to compensate for a lost low-band speech signal;la unidad descodificadora de banda alta está destinada a descodificar una señal de voz de descodificación de banda alta recibida, y a compensar una señal de voz de banda alta perdida;the high band decoding unit is adapted to decode a received high band decoding speech signal, and to compensate for a lost high band speech signal;la unidad de filtrado especular en cuadratura está destinada a obtener una señal de voz de salida final sintetizando la señal de voz de descodificación de banda baja y la señal de voz de descodificación en banda alta;the quadrature specular filtering unit is adapted to obtain a final output speech signal by synthesizing the low-band decoding speech signal and the high-band decoding speech signal;la unidad descodificadora de banda baja comprende una sub-unidad de descodificación de banda baja, una subunidad de codificación predictiva lineal basada en la repetición tonal y una sub-unidad de desvanecimiento cruzado;the low-band decoding unit comprises a low-band decoding sub-unit, a linear predictive coding sub-unit based on tonal repetition and a cross-fade sub-unit;en el que la sub-unidad de descodificación de banda baja está destinada a descodificar una señal de voz de flujo de banda baja recibida;wherein the low-band decoding sub-unit is adapted to decode a received low-band stream speech signal;la sub-unidad de codificación predictiva lineal (LPC) basada en la repetición tonal, está destinada a generar una señal de voz sintetizada correspondiente a una trama perdida;the linear predictive coding (LPC) sub-unit based on tonal repetition, is intended to generate a synthesized speech signal corresponding to a lost frame;The cross-fade sub-unit is intended to apply the cross-fading to the speech signal processed by the low-band decoding sub-unit and the synthesized speech signal corresponding to the lost frame generated by the LPC-based sub-unit. in tonal repetition;la sub-unidad de desvanecimiento cruzado está destinada a aplicar el desvanecimiento cruzado a la señal de voz tratada por la sub-unidad de descodificación de banda baja y la señal de voz sintetizada correspondiente a la trama perdida generada por la sub-unidad de LPC basada en la repetición tonal;la sub-unidad de LPC basada en la repetición tonal comprende un módulo analizador y un aparato de acuerdo con las reivindicaciones 11 a 13, en el que el módulo analizador está destinado a analizar una señal de voz histórica, y a generar una señal de voz con trama perdida reconstruida. the LPC sub-unit based on tonal repetition comprises an analyzer module and an apparatus according to claims 11 to 13, in which the analyzer module is intended to analyze a historical speech signal, and to generate a speech signal with Lost plot rebuilt.
- 15Un producto programa de ordenador que comprende códigos de programa de ordenador que permiten que un ordenador ejecute las operaciones de una cualquiera de las reivindicaciones 1 a 10, cuando los códigos de programa de ordenador son ejecutados por el ordenador. fifteen. A computer program product comprising computer program codes that enable a computer to execute the operations of any one of claims 1 to 10, when the computer program codes are executed by the computer.
Independent claims4
129 paragraphs in 10 sections, as filed
IS 2 340 975 T3
DESCRIPTION
Method and apparatus for obtaining an attenuation factor.
This application claims the priority of Chinese Patent Application No. 200710169618.0, entitled "Method and apparatus to obtain an attenuation factor", presented on November 5, 2007, at the State Intellectual Property Office of the PRC.
Invention field
The present invention relates to the field of signal processing and, in particular, to a method and an apparatus for obtaining an attenuation factor.
Background of the invention
A voice data transmission is required to be real-time and reliable in a real-time voice communication system, for example a VoIP (voice over IP) system. Due to the unreliable characteristics of a network system, data packets can be lost or not reach their destination in time in a transmission procedure, from a sending end to a receiving end. These two kinds of situations are considered, by the receiving end, as network packet losses. Network packet loss is inevitable. Also, network packet loss is one of the most important factors influencing voice quality. Therefore, a robust method of hiding packet loss is needed in order to recover lost data packets in the real-time communication system so that good speech quality is still obtained in the packet loss situation. of the network.
In existing real-time voice communication technology, at the sending end an encoder splits broadband voice into a high subband and a low subband, and makes use of ADPCM (Differential Pulse Code Modulation). , adaptive) to encode the two sub-bands, respectively, and send them together over the network to the receiving end. At the receiving end, the two subbands are decoded, respectively, by the ADPCM decoder, and then the final signal is synthesized using a QMF synthesis filter (Quadrature Specular Filter).
For two different sub-bands, different Packet Loss Concealment (PLC) methods are adopted. For a low subband, in the situation where there is no packet loss, a reconstruction signal does not undergo any change during the crossfade. In the situation where there is packet loss, for the first lost frame the historical signal is analyzed (the historical signal is a voice signal prior to the lost frame in the document of the present application) using a short-term predictor and a long-term predictor, and information about the voice classification is extracted. The lost frame signal is reconstructed using LPC (Linear Predictive Coding) based on the tonal repetition method, the predictor and the classification information. The ADPCM status will be updated, too, synchronously until a good frame is found. In addition, not only does the signal corresponding to the lost frame have to be generated, but also a signal section has to be generated that accommodates the Cross Fading. Thus, once a good frame is received, Cross Fading is performed to handle the good frame signal and the signal section. It is to be noted that this kind of crossfade only occurs after the receiving end loses a frame and receives the first good frame.
During the process of putting the present invention into practice, the inventor encountered at least the following problems in the prior art. In the prior art, the energy of the synthesized signal is controlled using a static, self-adapting attenuation factor. Although the defined attenuation factor changes gradually, its attenuation speed, that is, the attenuation factor value is the same relative to the same voice rating. However, human voices are different. If the attenuation factor does not match the characteristic of human voices, an uncomfortable noise will appear in the reconstruction signal, particularly at the end of stable vowels. The self-adapting static attenuation factor cannot adapt to the characteristic of various human voices.
The situation shown in figure 1 is taken as an example, in which T<sub>0</sub> is the tonal period of the historical signal. The upper signal corresponds to an original signal, that is, a schematic waveform diagram in the situation where there is no packet loss. The lower signal represented by the dashed line is a signal synthesized according to the prior art. As can be seen in the figure, the synthesized signal does not maintain the same attenuation speed as the original signal. If the same tonal repetition exists too many times, the synthesized signal will produce obvious musical noise so that the difference between the location of the synthesized signal and the desirable location is large.
EP 1 291 851 A2 describes a method and system for attenuation of error-corrupted rate frame waveforms.
Summary
In order to achieve the aforementioned object, an embodiment of the present invention provides a method of processing a synthesized speech signal in packet loss concealment as defined in claim 1.
IS 2 340 975 T3
An embodiment of the present invention also provides an apparatus for processing a synthesized speech signal in packet loss concealment according to claim 11.
An embodiment of the present invention also provides a speech decoder according to claim 14.
An embodiment of the present invention further provides a computer program product as defined in claim 15.
Compared to the prior art, the embodiments of the present invention have the following advantages:
A self-adapting attenuation factor is dynamically adjusted using the changing trend of a historical signal. The smooth transition from the historical data to the data received last is done so that the rate of attenuation between the compensated signal and the original signal is kept as consistent as possible to accommodate the characteristic of various human voices.
Brief description of the drawings
Figure 1 is a schematic diagram illustrating the original signal and the synthesized signal according to the prior art;
Fig. 2 is a flow chart illustrating a method for obtaining an attenuation factor in accordance with Embodiment 1 of the present invention;
Figure 3 is a schematic diagram illustrating the principles of the encoder;
Fig. 4 is a schematic diagram illustrating the modulus of an LPC based on the pitch repetition sub-unit of the low-band decoder unit;
Fig. 5 is a schematic diagram illustrating an output signal after adopting the dynamic attenuation method according to Embodiment 1 of the present invention;
Figures 6A and 6B are schematic diagrams illustrating the structure of the apparatus for obtaining an attenuation factor according to Embodiment 2 of the present invention;
Fig. 7 is a schematic diagram illustrating the application scene of the apparatus for obtaining an attenuation factor according to Embodiment 2 of the present invention;
Figures 8A and 8B are schematic diagrams illustrating the structure of the signal processing apparatus according to Embodiment 3 of the present invention;
Figure 9 is a schematic diagram illustrating the speech decoder module according to Embodiment 4 of the present invention;
Fig. 10 is a schematic diagram illustrating the low-band decoder unit module of the speech decoder according to Embodiment 4 of the present invention;
Fig. 11 is a schematic diagram illustrating the modulus of the LPC based on a tonal repeating sub-unit, in accordance with Embodiment 4 of the present invention.
Detailed description
The present invention will be described in more detail with reference to the drawings and embodiments.
A method for obtaining an attenuation factor is provided in Embodiment 1 of the present invention, intended to treat the synthesized signal in packet loss concealment, as shown in Figure 2, and includes the following operations.
Operation s101, you get a trend to the change of a signal:
Specifically, the tendency to change can be expressed by the following parameters: (1) the relationship between the energy of the last periodic tone signal and the energy of the previous periodic tone signal of the signal; (2) the ratio of the difference between the maximum value of the amplitude and the minimum value of the amplitude of the last periodic tone signal and the difference between the maximum value of the amplitude and the minimum value of the amplitude of the periodic tone signal pre-signal.
Step s102, an attenuation factor is obtained according to the trend to change.
IS 2 340 975 T3
The specific treatment method of Embodiment 1 of the present invention will be described together with a specific application scene.
A method of obtaining the attenuation factor that is intended to treat the synthesized signal in packet loss concealment is provided in Embodiment 1 of the present invention.
As shown in Figure 3, different PLC methods are adopted for two different sub-bands. The PLC method for the low band part is shown as the ® part in a frame represented in dashed line in figure 3. On the other hand, a ® frame represented in dashed line in figure 3 corresponds to the PLC algorithm for the high band. For a high-band signal, zh (n) is a high-band signal finally output. After obtaining the low-band signal zl (n) and the high-band signal zh (n), the QMF is run for the low-band signal and the high-band signal and a y (n) band signal is synthesized finally broadcasted as output.
Only the low band signal is described in detail as follows.
In the situation where there is no loss of frames, the signal xl (n) is obtained, where n = 0, ..., L-1 after decoding the current frame received by the low-band ADPCM decoder and the output is zl (n), where n = 0, ..., L-1 corresponding to the current frame. In this situation, the reconstruction signal does not change during the crossover-fade, that is, zl (n) = xl (n), where n = 0, ..., L-1, where L is the length of the frame ;
In the situation where there is loss of frames, as regards the first lost frame, the historical signal zl (n) is analyzed, where n <0 using a short-term predictor and a long-term predictor, and it is extracted voice classification information. Adopting the above-mentioned predictors and classification information, the yl (n) signal is generated using a tonal repetition-based LPC method. And the lost frame signal zl (n) is reconstructed as zl (n) = yl (n), where n = L, ..., L-1. Additionally, the ADpCm status will also update synchronously until a good frame is found. It will be observed that not only the signal corresponding to the lost frame has to be generated, but also that a signal of 10 ms and l (n) has to be generated, where n = L, ..., L + M-1 that adapts to Crossfade, where M is the number of signal sample points included in the process when energy is calculated. Thus, once a good frame is received, Cross Fading is performed for xl (n), where n = L, ..., L + M-1 and yl (n), where n = L , ..., L + M-1. It is to be noted that this kind of crossfade only occurs after a loss of frames and when the receiving end receives the data from the first good frame.
An LPC based on the tonal repetition method of Figure 3 is as shown in Figure 4.
When the data frame is a good frame, zl (n) is buffered for future use.
When the first missing frame is found, the final signal yl (n) has to be synthesized in two stages. In a first, the historical signal zl (n) is analyzed, where n = -297, ..., - 1. Then, the signal yl (n) is synthesized, where n = 0, ..., L-1, according to the result of the analysis, where L is the frame length of the data frame, that is, the number of sample points corresponding to a signal frame, Q is the length of the signal that is needed to analyze the historical signal.
The LPC node based on tonal repetition specifically includes the following parts.
(1) An LP analysis (linear prediction)
Short-term analysis filter A (z) and synthesis filter 1 / A (z) are linear prediction filters (LP) based on an order P. The LP analysis filter is defined as:
A (z) = l + aiz '<sup>1</sup>+ a2z '<sup>2</sup>+. . . + apz
By means of the LP analysis of the historical signal zl (n), where n = -Q, ..., - 1 with the filter A (z), a residual signal e (n) is obtained, where n = -Q, ..., - 1, corresponding to the historical signal zl (n), where n = -Q, ..., - 1:
pe (n) = zl (η) + ^ a, zZ (nz), where n = -Q,. . . , -1 i = \
ES 2 340 975 T3 (2) An analysis of the historical signal
The lost signal is compensated for by a tonal repetition method. Therefore, first a tonal period T has to be estimated<sub>0</sub> corresponding to the historical signal zl (n), where n = -Q, ..., - 1. The operations are as follows: the zl (n) is pre-treated to eliminate an unnecessary low frequency ingredient in a LTP (long-term prediction) analysis and the tonal period T<sub>0</sub> of zl (n) can be obtained by LTP analysis. Voice classification is obtained by combining a signal classification module after obtaining the tonal period T<sub>0</sub>.
The voice ratings are as shown in the following Table 1:
TABLE 1
Voice ratings
<td>Classification name</td><td>Explanation</td>
<td>TRANSIENT</td><td>for voices with strong energetic variation (for example, explosive consonants)</td>
<td>UNVOICED</td><td>for voiceless signals</td>
<td>VUV_TRANSITION</td><td>for a transition between voice and voiceless signals</td>
<td>WEAKLY_VOICED</td><td>for weak voice signals (for example, initial vowels and finals)</td>
<td>VOICED</td><td>voice signals (for example, stable vowels)</td>
(3) A tonal repeat
A tonal repetition module is intended to estimate a residual LP signal e (N), where n = 0, ..., L-1 of a lost frame. Before tonal repetition is performed, if the voice classification is not VOICED, the following formula is adopted to limit the amplitude of a sample:
e (n) = min (max (Ie (nT<sub>0</sub> + i) I), Ie (η) I) xsign (e (n)), /=-2,...,+2 where n = -T0, ..., - 1 where
<img file="ES2340975T3_D0001.tif" />
If the voice classification is VOICED, the residual signal e (n) is obtained, where n = 0, ..., L-1 corresponding to the lost signal, adopting the step of repeating the residual signal corresponding to the signal of the last tonal period of the signal of a newly received good frame, i.e .:
<img file="ES2340975T3_D0002.tif" />
IS 2 340 975 T3
In relation to other classifications of voices, to avoid that the periodicity of the generated signal is too strong (with respect to the signal without speech, if the periodicity is too strong, some uncomfortable musical noise may be heard), the residual signal e (n ), where n = 0, ..., L-1 corresponding to the lost signal, it is generated using the following formula:
<img file="ES2340975T3_D0003.tif" />
In addition to generating the residual signal corresponding to the lost frame, the residual signals e (n) continue to be generated, where n = L, ..., L + N-1 of N additional samples in order to generate a signal destined for fading crossed, in order to ensure the smooth division between the lost frame and the first good frame after the lost frame.
(4) An LP synthesis
After generating the residual signal e (n) corresponding to the lost frame and the crossfade, a reconstruction of the lost frame signal yl is obtained.<sub>pre</sub>(n), where n = 0, ..., L-1 using the following formula:
ylpre (n) = e (n) - aiyl (ni) / = 1 where the residual signal e (n), where n = 0, ..., L-1 is the residual signal obtained from the previous stages of tonal repetition.
In addition, they are generated and l<sub>pre</sub>(n), where n = L, ..., L + N-1 with N samples destined for Cross Fading using the above formula.
(5) Adaptive silencer
To perform a smooth energy transition, before running QMF with the high-band signal, the low-band signal has to be also cross-faded, the rules are shown in the table below:
<td colspan="2" rowspan="2"></td><td colspan="2">Running plot</td>
<td>Bad plot</td><td>Good plot</td>
<td rowspan="2">Previous plot</td><td>Plot Bad</td><td>Zl (n) = yl (n), Where n = 0, ..., Τ- Ι</td><td>η n zl (n) - xl (n) + (l-) yl 7V-1 7V-1 Where n = 0, ..., N-1 and Zl (n) = xl (n), where n = N, ... , L-1</td>
<td>Good plot</td><td>Zl (n) = yl (n), Where n = 0, ..., Σ- Ι</td><td>Zl (n) = xl (n), Where n = 0, ..., Ll</td>
In the above table, zl (n) is a signal finally emitted as output, corresponding to the current frame, xl (n) is the signal of the good frame corresponding to the current frame; and l (n) is a synthesized signal corresponding to the same instant of the current frame, where L is the length of the frame, N is the number of samples that execute the Cross Fading.
Aiming at different voice classifications, the signal energy in yl<sub>pre</sub>(n) is controlled before by executing Crossfade according to the coefficient corresponding to each sample. The value of the coefficient changes according to the different voice classifications and the packet loss situation.
In detail, in the case that the last two-tone periodic signal of the received historical signal is the original signal as shown in figure 5, the self-adaptive dynamic attenuation factor is dynamically adjusted according to the tendency to change of the last two-tone period of the historical signal. The detailed setting method includes the following operations:
Operation s201, the trend of the signal change is obtained.
IS 2 340 975 T3
The tendency to change the signal can be expressed by the relationship between the energy of the last periodic tonal signal and the energy of the previous periodic tonal signal of the signal, that is, the energy E, and the energy E<sub>2</sub> of the last two-tone periodic signal of the historical signal, and the relationship between the two energies is calculated.
<img file="ES2340975T3_D0004.tif" />
Where E is the energy of the last periodic tonal signal, E<sub>2</sub> the energy of the previous periodic tone signal, and T<sub>0</sub> the tonal period corresponding to the historical signal.
Optionally, the trend to change of the signal can be expressed by the relationship between the picovalle differences of the last two tonal periods of the historical signal.
Pi = max (xl (i)) - min (xl (j))
P<sub>2</sub>= max (xl (i)) - min (xl (j)) (i, j) = -T<sub>0</sub>,. . . , -1 (i, j) = - 2T<sub>0</sub>,. . . , - (To + 1) where P, is the difference between the maximum value of the amplitude and the minimum value of the amplitude of the last periodic tone signal, P<sub>2</sub> is the difference between the maximum value of the amplitude and the minimum value of the amplitude of the previous periodic tone signal, and the ratio is calculated as:
<img file="ES2340975T3_D0005.tif" />
Step s202, the synthesized signal is dynamically attenuated in accordance with the obtained trend of change of the signal.
The calculation formula is shown as follows:
yl (n) = yl<sub>pre</sub> (n) * (1-C * (n + 1)), where n = 0, ..., Nl where yl<sub>pre</sub> (n) is the reconstruction of the lost frame signal, N is the length of the synthesized signal, and C is the self-adaptive attenuation coefficient, whose value is:
<img file="ES2340975T3_D0006.tif" />
In the situation where the attenuation factor is 1-C * (n + 1) <0, it is necessary to make 1-C * (n + 1) = 0, in order to avoid the appearance of a situation in which the attenuation factor corresponding to the samples is negative.
In particular, to avoid the situation in which the value of the amplitude corresponding to a sample may overflow in the situation of R> 1, the synthesized signal is dynamically attenuated by using the formula of operation s202 of the present embodiment, which it can only take into account the situation of R <1.
IS 2 340 975 T3
In particular, in order to avoid the situation where the attenuation speed of the signal with less energy is too high, only in the situation where E<sub>1</sub> exceeds a certain limit value, in the present embodiment the synthesized signal is dynamically attenuated by using the formula of step s202.
In particular, to avoid that the attenuation rate of the synthesized signal is too high, especially in the situation of continuous loss of frames, an upper limit value is set for the attenuation coefficient C. When C * (n + 1) exceeds a limit value, the attenuation coefficient is set as the upper limit value.
In particular, in the situation where the network environment is bad and the loss of frames is continuous, a certain condition can be set to avoid too high fading rate. For example, it can be taken into account that, when the number of lost frames exceeds a designated number, for example two frames; or when the signal corresponding to the lost frame exceeds a designated length, for example 20 ms; or in at least one of the above conditions where the current attenuation coefficient 1-C * (n + 1) reaches a designated threshold value, the attenuation coefficient C has to be adjusted in order to avoid the attenuation speed too high. high which may result in the situation where the output signal becomes silent voice.
For example, in the sampling situation at the frequency of 8 kHz and a frame length of 40 samples, the number of lost frames can be set to 4 and, after the attenuation factor 1-C * (n + 1) becomes less than 0.9, the attenuation coefficient C is adjusted to have a smaller value. The rule for setting the lowest value is as follows.
Hypothetically, the current attenuation coefficient is predicted to be C and the attenuation factor value to be V, and the attenuation factor V can attenuate to 0 after V / C samples. However, the most desirable situation is that the attenuation factor V attenuates to 0 after M (M ^ V / C) samples. Thus, the attenuation factor C is adjusted to
C = V / M
As shown in Figure 5, the upper signal is the original signal; the average signal is the synthesized signal. As can be seen from the figure, although the signal has a certain degree of attenuation, it continues to have an intensive sonic characteristic. If the duration is too long, the signal may show up as musical noise, especially at the end of the sound. The lower signal is the signal after the use of dynamic attenuation in the embodiment of the present invention, which can look very similar to the original signal.
According to the method provided by the above-mentioned embodiment, the self-adaptive attenuation factor is dynamically adjusted using the changing trend of the historical signal, so that the smooth transition from the historical data to the last received data can be performed. The attenuation speed is kept consistent, as far as possible, between the compensated signal and the original signal in order to tailor the characteristic of various human voices as much as possible.
In Embodiment 2 of the present invention, an apparatus is provided for obtaining an attenuation factor, for treating the synthesized signal in packet loss concealment, including:
a changing trend obtaining unit 10, designed to obtain a changing trend of a signal;
an attenuation factor obtaining unit 20, intended to obtain an attenuation factor in accordance with the tendency to change obtained by means of the unit 10 for obtaining the tendency to change.
The unit 20 for obtaining the attenuation factor also includes: a sub-unit 21 for obtaining the attenuation coefficient, intended to generate the attenuation coefficient according to the trend to change obtained by the unit 10 for obtaining the tendency to change; a sub-unit 22 for obtaining the attenuation factor, destined to obtain an attenuation factor according to the c attenuation coefficient generated by the subunit 21 for obtaining the attenuation factor. The unit 20 for obtaining the attenuation factor further includes: a subunit 23 for adjusting the attenuation coefficient, intended to adjust the value of the attenuation coefficient obtained by the sub-unit 21 for obtaining the attenuation coefficient to a value given in given conditions, the conditions of which include at least one of the following: the value of the attenuation coefficient exceeds an upper limit value; the situation of continuous loss of frames occurs; and the dimming speed is too high.
The method of obtaining an attenuation factor in the above embodiment is the same as the method of obtaining the attenuation factor of the method embodiments.
In detail, the change trend obtained by the change trend obtaining unit 10 can be expressed by the following parameters: (1) the relationship between the energy of the last periodic tone signal and the energy of the previous periodic tone signal, Of the signal; (2) the relationship between the difference between the maximum amplitude value and the minimum amplitude value of the last periodic tone signal, with the difference between the maximum amplitude value and the minimum amplitude value of the periodic tone signal pre-signal.
IS 2 340 975 T3
When the change tendency is expressed as the energy ratio of (1), the structure of the apparatus to obtain an attenuation factor is as shown in Fig. 6A. Unit 10 for obtaining the trend to change also includes:
a sub-unit 11 for obtaining energy, intended to obtain the energy of the last periodic tone signal and the energy of the previous periodic tone signal;
a sub-unit 12 for obtaining the energy ratio, intended to obtain the relationship between the energy of the last periodic tonal signal and the energy of the previous periodic tonal signal obtained by the sub-unit 11 for obtaining energy and use the ratio to show the trend of the signal change.
When the change tendency is expressed as the amplitude difference ratio of (2), the structure of the apparatus to obtain an attenuation factor is as shown in Fig. 6B. The unit 10 for obtaining trend to change also includes:
a subunit 13 for obtaining the difference in amplitude, designed to obtain the difference between the maximum value of amplitude and the minimum value of amplitude of the last periodic tonal signal, and the difference between the maximum value of amplitude and the minimum value amplitude of the previous periodic tonal signal;
a sub-unit 14 for obtaining the amplitude difference relationship, designed to obtain the relationship that maintains the difference between the maximum amplitude value and the minimum amplitude value of the last periodic tonal signal, with the difference between the maximum value of amplitude and the minimum value of amplitude of the previous periodic tone signal, and use such relation to show the tendency to change of the signal.
A schematic diagram illustrating the application scene of the apparatus for obtaining an attenuation factor according to Embodiment 2 of the present invention, is as illustrated in Fig. 7. The self-adaptive attenuation factor is dynamically adjusted using the changing trend of the historical landmark.
By using the apparatus provided by the aforementioned embodiment, the self-adaptive attenuation factor is dynamically adjusted using the changing trend of the historical signal so as to carry out the smooth transition from the historical data to the last received data. The attenuation speed is kept constant as much as possible between the compensated signal and the original signal to adapt, as much as possible, to the characteristic of various human voices.
In Embodiment 3 of the present invention there is provided a signal processing apparatus for processing the synthesized signal in packet loss concealment, as shown in FIG. 8A and FIG. 8B. Based on Embodiment 2, a lost frame reconstruction unit 30 is added correlative with the attenuation factor obtaining unit. The lost frame reconstruction unit 30 obtains a lost frame after fading according to the attenuation factor obtained by the attenuation factor obtaining unit 20.
Using the apparatus provided by the aforementioned embodiment, the self-adaptive attenuation factor is dynamically adjusted by using the changing trend of the historical signal, and a lost frame reconstructed after attenuation is obtained according to the attenuation factor. , so that the smooth transition is made from the historical data to the last received data. The attenuation speed is kept as consistent as possible between the compensated signal and the original signal to adapt, as far as possible, to the characteristic of various human voices.
Embodiment 4 of the present invention provides a speech decoder, as shown in Figure 9. The speech decoder includes: a high-band decoder unit 40 for decoding a received high-band decoding signal and compensating for a lost high-band signal; a low-band decoding unit 50 for decoding a received low-band decoding signal and compensating for a lost low-band signal; and a quadrature specular filtering unit 60 for obtaining a final output signal by synthesizing the low-band decoding signal and the high-band decoding signal. The high-band decoder unit 40 decodes the high-band stream signal received by the receiving end and synthesizes the lost high-band signal. The low-band decoder unit 50 decodes the low-band stream signal received by the receiving end and synthesizes the lost low-band signal. The quadrature specular filtering unit 60 obtains the final decoding signal by synthesizing the low-band decoding signal output by the low-band decoding unit 50 and the high-band decoding signal output by the low-band decoding unit 40. high band.
As for the low-band decoder unit 50, as shown in Fig. 10, it includes the following units. An LPC sub-unit 51 based on tonal repetition, which is intended to generate a synthesized signal corresponding to the lost frame, a low-band decoding sub-unit 52, which is intended to decode a low-band stream signal received, and a 53 cross-fade sub-unit, which is intended to achieve cross-fading of the signal decoded by the low-band decoding sub-unit and the synthesized signal corresponding to the lost frame generated by the LPC unit based on tonal repetition.
IS 2 340 975 T3
The low-band decoding sub-unit 52 decodes the received low-band stream signal. Tonal repetition-based LPC subunit 51 generates the synthesized signal by executing an LPC over the lost low-band signal. And finally, the cross-fade sub-unit 53 applies the cross-fading to the signal processed by the low-band decoding sub-unit 52 and the synthesized signal in order to obtain a final decoding signal after offsetting. the lost plot.
The LPC sub-unit 51 based on tonal repetition, as shown in FIG. 10, further includes an analysis module 511 and a signal processing module 512. The analysis module 511 analyzes a historical signal and generates a reconstructed lost frame signal; the signal processing module 512 obtains a change trend of a signal and obtains an attenuation factor according to the change trend of the signal, and attenuates the reconstructed lost frame signal and obtains a lost frame, reconstructed after the attenuation.
The signal processing module 512 further includes an attenuation factor obtaining unit 5121 and a lost frame reconstruction unit 5122. The attenuation factor obtaining unit 5121 obtains a change trend of a signal and obtains an attenuation factor in accordance with the change tendency; the lost frame reconstruction unit 5122 attenuates the reconstructed lost frame signal according to the attenuation factor and obtains a reconstructed lost frame after attenuation. The signal processing module 512 includes two structures, corresponding to schematic diagrams illustrating the structure of the signal processing apparatus of Figures 8A and 8B, respectively.
The attenuation factor obtaining unit 5121 includes two structures, corresponding to schematic diagrams illustrating the structure of the apparatus for obtaining an attenuation factor of Figures 6A and 6B, respectively. The specific functions and the means for their implementation in practice of the aforementioned modules and units may refer to the content set forth in the method embodiments. Unnecessary details will not be repeated in this document.
By describing the above-mentioned embodiments, those skilled in the art can clearly understand that the present invention can be practiced depending on the necessary and general software and hardware platform and can certainly also be practiced by software. However, in most situations, the former is a preferable embodiment. Based on such an understanding, the essence or the prior art contributing part of the technical scheme of the present invention may be incorporated in the form of a software product stored on a storage medium, and the software product includes certain instructions for making a device implements the embodiments of the present invention.
While the illustration and description of the present disclosure have been made with reference to their embodiments, it should be appreciated by those of ordinary skill in the art that various changes can be made, in form and detail, without deviating from the scope defined by the claims. attached.
Contents10
13 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13
61 members in 14 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 200710169618 | China | A | |
| 200710169618 | China | A | |
| 20071016961808168328 | – | – | – |
| CN20071169618 | – | – | – |
Members61
| Document | Office | Kind | |
|---|---|---|---|
| CN101207459A | China | A | |
| CN101207665A | China | A | |
| EP2056291A1 | European Patent Office (EPO) | A1 | |
| EP2056292A2 | European Patent Office (EPO) | A2 | |
| US2009116486A1 | United States of America | A1 | |
| US2009119098A1 | United States of America | A1 | |
| KR20090046713A | Republic of Korea | A | |
| KR20090046714A | Republic of Korea | A | |
| WO2009059497A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009059498A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2056292A3 | European Patent Office (EPO) | A3 | |
| JP2009116332A | Japan | A | |
| JP2009175693A | Japan | A | |
| CN100550712C | China | C | |
| CN101578657A | China | A | |
| US2009292542A1 | United States of America | A1 | |
| CN101601217A | China | A | |
| US2009316598A1 | United States of America | A1 | |
| EP2056291B1 | European Patent Office (EPO) | B1 | |
| AT456126T | Austria | T | |
| EP2056292B1 | European Patent Office (EPO) | B1 | |
| EP2157572A1 | European Patent Office (EPO) | A1 | |
| EP2161719A2 | European Patent Office (EPO) | A2 | |
| DE602008000579D1 | Germany | D1 | |
| AT458241T | Austria | T | |
| PT2056291E | Portugal | E | |
| EP2161719A3 | European Patent Office (EPO) | A3 | |
| DE602008000668D1 | Germany | D1 | |
| DK2056292T3 | Denmark | T3 | |
| ES2340975T3This record | Spain | T3 | |
| PL2056292T3 | Poland | T3 | |
| JP2010176142A | Japan | A | |
| DE202008017752U1 | Germany | U1 | |
| EP2161719B1 | European Patent Office (EPO) | B1 | |
| AT484052T | Austria | T | |
| US7835912B2 | United States of America | B2 | |
| DE602008002938D1 | Germany | D1 | |
| JP4586090B2 | Japan | B2 | |
| CN101207665B | China | B | |
| HK1142713A1 | Hong Kong, China | A1 | |
| KR101023460B1 | Republic of Korea | B1 | |
| US7957961B2 | United States of America | B2 | |
| CN102122511A | China | A | |
| CN102169692A | China | A | |
| EP2157572B1 | European Patent Office (EPO) | B1 | |
| AT529854T | Austria | T | |
| JP4824734B2 | Japan | B2 | |
| ES2374043T3 | Spain | T3 | |
| HK1154696A1 | Hong Kong, China | A1 | |
| HK1155844A1 | Hong Kong, China | A1 | |
| KR101168648B1 | Republic of Korea | B1 | |
| CN102682777A | China | A | |
| CN101578657B | China | B | |
| US8320265B2 | United States of America | B2 | |
| CN101601217B | China | B | |
| JP5255585B2 | Japan | B2 | |
| CN102682777B | China | B | |
| CN102122511B | China | B | |
| CN102169692B | China | B | |
| BRPI0808765A2 | Brazil | A2 | |
| BRPI0808765B1 | Brazil | B1 |
Numbers
- Publication, DOCDB
- 2340975
- Publication, EPODOC
- ES2340975T
- Application
- 8168328
- Application, DOCDB
- 08168328
- Application, EPODOC
- ES20080168328T
Titles2
- English
- METHOD AND APPLIANCE TO OBTAIN A DAMAGE FACTOR.
- Spanish
- METODO Y APARATO PARA OBTENER UN FACTOR DE ATENUACION.
Classification
- CPC, 4
- G10L19/005
- G10L19/0204
- G10L19/097
- G10L19/04
- IPC, 2
- G10L19 005
- G10L25 12