Signal processing method, processing appartus and voice decoder
Abstract
A method of signal processing to treat a signal synthesized in packet loss concealment, comprising: receiving (101) a good frame following a lost frame, characterized in that the method further comprises: obtaining (102) a ratio of energy between the energy of the good frame and the energy of the corresponding synthesized signal at the same time of the good frame; and adjusting (103), by changing the energy scale, said corresponding synthesized signal at the same time of the good frame according to the energy ratio.

Term
2.1 yearsto projected expiry
Projected expiry 4 November 2028, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
13 claims: 4 independent, 9 dependent
- 1ES 2 374 043 T3 ES 2 374 043 T3 CLAIMS REIVINDICACIONES 1. A signal processing method for treating a synthesized signal in packet loss concealment, comprising:1. Un método de tratamiento de señales para tratar una señal sintetizada en ocultación de pérdida de paquetes, que comprende: receiving (101) a good frame following a lost frame, characterized in that the method further comprises: recibir (101) una trama buena a continuación de una trama perdida, caracterizado porque el método comprende, además: obtaining (102) an energy ratio between the energy of the good frame and the energy of the synthesized signal corresponding to the same instant of the good frame;and adjusting (103), by changing the energy scale, said synthesized signal corresponding to the same instant of the good frame according to the energy ratio. obtener (102) una relación de energía entre la energía de la trama buena y la energía de la señal sintetizada correspondiente al mismo instante de la trama buena;y ajustar (103), mediante cambio de escala de energía, dicha señal sintetizada correspondiente al mismo instante de la trama buena de acuerdo con la relación de energía.
- 8A signal processing apparatus for processing a synthesized signal in packet loss concealment, configured to:8. Un aparato de tratamiento de señales destinado a tratar una señal sintetizada en ocultación de pérdida de paquetes, configurado para: receiving (101) a good frame following the lost frame;recibir (101) una trama buena a continuación de la trama perdida;caracterizado porque el aparato de tratamiento de señales está configurado además para: characterized in that the signal processing apparatus is further configured to: obtaining (102) an energy ratio between the energy of the good frame and the energy of the synthesized signal corresponding to the same instant of the good frame;Y obtener (102) una relación de energía entre la energía de la trama buena y la energía de la señal sintetizada correspondiente al mismo instante de la trama buena;y ES 2 374 043 T3 adjusting (103), by changing the energy scale, the synthesized signal corresponding to the same instant of the good frame according to the energy ratio. ES 2 374 043 T3 ajustar (103), mediante cambio de escala de energía, la señal sintetizada correspondiente al mismo instante de la trama buena de acuerdo con la relación de energía.
- 12A speech decoder, comprising a low-band decoder unit, a high-band decoder unit, and a quadrature mirror filter unit;12. Un descodificador de voz, que comprende una unidad descodificadora de banda baja, una unidad descodificadora de banda alta y una unidad de filtro de espejo en cuadratura;en el que la unidad descodificadora de banda baja está configurada para descodificar una señal de descodificación de banda baja recibida y compensar una trama de señal de banda baja perdida;wherein the low-band decoder unit is configured to decode a received low-band decode signal and compensate for a lost low-band signal frame;la unidad descodificadora de banda alta está configurada para descodificar una señal de descodificación de banda alta recibida y compensar una trama de señal de banda alta perdida;the high band decoding unit is configured to decode a received high band decoding signal and compensate for a lost high band signal frame;la unidad de filtro de espejo en cuadratura está configurada para sintetizar una señal descodificada de banda baja y una señal descodificada de banda alta para obtener una señal de salida final;the quadrature mirror filter unit is configured to synthesize a decoded low-band signal and a decoded high-band signal to obtain a final output signal;la unidad descodificadora de banda baja incluye una sub-unidad descodificadora de banda baja, una sub-unidad codificadora predictiva lineal basada en la repetición tonal, una sub-unidad de tratamiento de señales y una subunidad de desvanecimiento cruzado;the low-band decoder unit includes a low-band decoder sub-unit, a pitch repetition-based linear predictive encoder sub-unit, a signal processing sub-unit and a cross-fade sub-unit;en el que la sub-unidad descodificadora de banda baja está configurada para descodificar una señal de flujo de código de banda baja recibida;wherein the low-band decoder sub-unit is configured to decode a received low-band code stream signal;la sub-unidad codificadora predictiva lineal basada en la repetición tonal está configurada para generar una señal sintetizada correspondiente a una trama perdida;the linear predictive coding sub-unit based on tonal repetition is configured to generate a synthesized signal corresponding to a lost frame;la sub-unidad de tratamiento de señales es de acuerdo con una cualquiera de las reivindicaciones 9-11;y la sub-unidad de desvanecimiento cruzado está configurada para someter a desvanecimiento cruzado a la señal descodificada de banda baja descodificada por la sub-unidad descodificadora de banda baja y la señal sintetizada ajustada después de ajustar la energía mediante la sub-unidad de tratamiento de señales. the signal processing sub-unit is according to any one of claims 9-11;and the cross-fade sub-unit is configured to cross-fade the low-band decoded signal decoded by the low-band decoder sub-unit and the synthesized signal adjusted after adjusting the power by the power processing sub-unit. signs.
- 13A computer program product comprising computer program code, wherein the computer program code causes a computer to execute the steps of any one of claims 1-7, when the program code is executed by the computer. 13. Un producto programa de ordenador que comprende código de programa de ordenador, en el que el código de programa de ordenador hace que un ordenador ejecute los pasos de una cualquiera de las reivindicaciones 1-7, cuando el código de programa es ejecutado por el ordenador.
Independent claims4
158 paragraphs in 9 sections, as filed
ES 2 374 043 T3
DESCRIPTION
Signal processing method, processing apparatus and speech decoder
FIELD OF THE INVENTION
The present invention relates to the field of signal processing and, more particularly, to a signal processing method, a processing apparatus and a speech decoder.
BACKGROUND
In a real-time voice communication system, it is necessary to transmit voice data in time and reliably, such as a VoIP (voice over IP) system. However, given the unreliability of the network system itself, during the transmission process from a transmitter to a receiver, the data packet may be lost or may not reach its destination in time. Both situations are considered by the receiver as network packet loss. Network packet loss is inevitable and is one of the main factors influencing the quality of voice communications. Therefore, in the real-time voice communication system, a method is necessary to forcibly hide the packet loss in order to restore a lost data packet and preserve a good quality of the voice communication in the situation in network packet loss occurs.
In previous real-time voice communications technologies, an encoder in the transmitter divides a broadband voice into two sub-bands, a high band and a low band, encoding the two sub-bands respectively using pulse code modulation. Adaptive Differential (ADPCM) and sends the two encoded subbands to the receiver over the network. At the receiver, the two sub-bands are decoded respectively by an ADPCM decoder and are synthesized to obtain a final signal by a quadrature mirror filter (QMF).
For two different sub-bands, different Packet Loss Concealment (PLC) methods are used. For the low-band signal, when there is no packet loss, a reconstructed signal does not change during a crossfade. When there is packet loss, a short-term predictor and a long-term predictor are used to analyze a past signal (the past signal, in the present application, means the voice signal before a lost frame), and information is extracted about voice class. And the signal of the lost frame is reconstructed following the linear predictive coding (LPC) method based on the tonal repetition, and using the predictors and the information about speech class. The ADPCM status should update synchronously until a good frame appears. Furthermore, not only the corresponding lost frame signal must be generated, but also a signal for the crossfade. And, after a good frame is received, crossfade can be performed for the signal of the good frame and the aforementioned signal. It is to be noted that cross fading only occurs when a good frame is received after a loss of frame by the receiver.
During the process of putting the present invention into practice, the inventor finds that the following problems are encountered in the prior art: the reconstructed signal of the lost frame is synthesized using the past signal. The waveform and energy are more similar to the historical buffer signal, that is, the signal before the lost frame, even at the end of the synthesized signal, but they do not resemble the newly decoded signal. This can cause an abrupt waveform change or an abrupt energy change of the synthesized signal to occur at the junction between the lost frame and the first frame following the lost frame. The abrupt change is represented in figure 1. In figure 1, three signal frames are included, separated by two vertical lines. Frame N is a lost frame and the other two frames are good frames. The upper signal corresponds to an original signal. None of the other three data frames is lost in transmission. And a broken midline corresponds to a signal synthesized using frames N-1, N-2, etc., prior to frame N. The signal in the lowest row corresponds to the signal synthesized using the above techniques. From figure 1 it can be seen that there is a sudden energy change in the transition of frame N and frame N + 1 of the final output signal, especially at the end of speech and with longer frames. And excessive repetition of the same tonal repetition signal can result in musical noises.
WO 03/102921 A describes a method of controlling the energy in packet loss concealment, in which the energy of the LP filter of the first non-erased frame is set to a ratio between the energy of the impulse response of the LP filter of the last good frame and that of the first good frame.
ITU-T G.722 Appendix IV, A Low Complexity Algorithm for Packet Loss Concealment with G.722, dated November 1, 2006, describes a modified G.722 decoder that includes a mechanism to hide erasure screen, in which crossfade is used.
Document US 2006 / 206318A1 describes an apparatus for phase adaptation of frames in speech signal coders. The apparatus includes a decoder comprising a synthesizer having at least one input
ES 2 374 043 T3 operating connected with the output of a speech signal encoder. In it, the decoder comprises a memory and the decoder is intended to execute instructions stored in the memory comprising the phase adaptation and contraction in time of a speech frame.
In the paper dealing with A Linear Prediction Based on a Packet Loss Concealment Algorithm for PCM Encoded Speech, from IEEe Transactions on Speech and Audio Processing, vol. 9, no. 8, of November 1, 2001, an energy value from the end of the previous packet is used to perform a scale change.
SUMMARY
Embodiments of the present invention provide a signal processing method for treating a synthesized signal in packet loss concealment to render the waveform of a junction between a lost frame and a first frame of the synthesized signal smooth transmission. .
Embodiments of the present invention provide a signal processing method according to claim 1.
Embodiments of the present invention also provide a signal processing apparatus according to claim 8.
Embodiments of the present invention also provide a speech decoder for decoding a speech signal, including a low-band decoding unit, a high-band decoding unit, and a quadrature mirror filter unit.
The low band decoding unit is configured to decode a received low band decoding signal and to compensate for a lost low band signal frame.
The high band decoding unit is configured to decode a received high band decoding signal and to compensate for a lost high band signal frame.
The quadrature mirror filter unit is configured to synthesize the decoded low-band decoding signal and the decoded high-band decoding signal to obtain a final output signal.
The low-band decoding unit includes a low-band decoding sub-unit, a pitch repetition-based linear predictive coding sub-unit, a signal processing sub-unit and a cross-fade sub-unit.
The low-band decoding sub-unit is configured to decode a received low-band code stream signal.
The pitch repetition based linear predictive coding subunit is configured to generate a synthesized signal corresponding to a lost frame.
The signal processing sub-unit is configured according to any one of claims 9-11.
The cross-fade sub-unit is configured to achieve cross-fading of the signal decoded by the low-band decoding sub-unit and the signal after adjusting the power by the signal processing sub-unit.
Embodiments of the present invention also provide a computer program product that includes computer program code. The computer program code can cause a computer to execute any step of the signal processing method in packet loss concealment when the program code is executed by the computer.
Compared to the prior art, embodiments of the present invention have the following advantages:
The synthesized signal is adjusted according to the energy ratio between the energy of the first good frame that follows the lost frame and the energy of the synthesized signal to ensure that there is no abrupt change in waveform or abrupt change in waveform. the energy at the point where the lost frame and the first good frame that follows the lost frame merge into the synthesized signal, to smoothly transition the waveform and to avoid musical noise.
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1 is a schematic diagram illustrating a sharp change in the waveform or a sharp change in
ES 2 374 043 T3 energy at the place where, in the prior art, a lost frame and a first good frame that follows the lost frame are joined;
Figure 2 is a flow chart of a signal processing method according to a first embodiment of the present invention;
Figure 3 is a main schematic diagram of a signal processing method according to a first embodiment of the present invention;
Figure 4 is a schematic diagram of a linear predictive coding module based on tonal repetition;
Figure 5 is a schematic diagram of different signals according to a first embodiment of the present invention;
Figure 6 is a schematic diagram illustrating a phase discontinuity situation that occurs when a method based on tonal repetition is used to synthesize a signal in accordance with a second embodiment of the present invention;
Figure 7 is a main schematic diagram of a signal processing method according to a second embodiment of the present invention;
Figure 8 is a schematic structural diagram of a first signal processing apparatus according to a third embodiment of the present invention;
Figure 9 is a schematic structural diagram of a second signal processing apparatus according to a third embodiment of the present invention;
Figure 10 is a schematic structural diagram of a third signal processing apparatus according to a third embodiment of the present invention;
Fig. 11 is a schematic diagram illustrating an application case of a treatment apparatus according to a third embodiment of the present invention;
Figure 12 is a schematic module diagram of a speech decoder according to a fourth embodiment of the present invention; and Fig. 13 is a schematic module diagram of a low-band decoding unit of a speech decoder according to a fourth embodiment of the present invention.
DETAILED DESCRIPTION
Embodiments of the present invention are described in greater detail in conjunction with the accompanying drawings.
A first embodiment of the present invention provides a signal processing method for processing a synthesized signal in packet loss concealment. As shown in Figure 2, the method comprises the following steps:
Step s101, a frame following a lost frame is detected as a good frame.
Step s102, the energy ratio between the energy of a good frame signal and the energy of the synchronized synthesized signal is obtained.
Step s103, the synthesized signal is adjusted according to the energy ratio.
In step s102, the synchronized synthesized signal is the synthesized signal corresponding to the same time of the good frame. The synchronized synthesized signal appearing elsewhere in the present application can be understood in the same way.
The signal processing method of the first embodiment of the present invention is described as follows in combination with specific application cases.
In the first embodiment of the present invention, a signal processing method is provided for processing the synthesized signal in packet loss concealment. The main schematic diagram is shown in figure 3.
In the case that a current frame is not lost, a low-band ADPCM decoder decodes the received current frame to obtain a signal xl (n), n = 0, ..., L-1 and an output corresponding to the frame current is zl (n),
ES 2 374 043 T3 n = 0, ..., L-1. In this condition, the reconstructed signal does not change when cross-faded. That is: zl [n] = xl [n], n = 0, ..., L-1, where L is the length of the frame.
In the event that a current frame is lost, a synthesized signal yl '(n), n = 0, ..., L-1 corresponding to the current frame is generated using the linear predictive coding method based on tonal repetition. . Depending on whether or not a following frame is lost after the current frame, a different treatment is performed:
When the following frame is lost after the current frame:
In this condition, no energy scaling treatment is performed for the synthesized signal. The output signal corresponding to the first lost frame zl (n), n = 0, ..., L-1 is the synthesized signal yl '(n), n = 0, ..., L-1, that is , zl [n] = yl [n] = yl '[n], n = 0, ..., L-1.
When the following frame is not lost after the current frame:
Suppose that, when the power scaling is performed, the good frame (which is the next frame after the first missing frame) that is used is the good frame xl (n), n = L, ..., L + M-1, which is obtained after the one that is decoded by the ADPCM decoder, where M is the number of signal samples when energy is calculated. The synthesized signal used that corresponds to the same instant of the signal of the good frame is the signal yl '(n), n = L, ..., L + M-1 which is generated by the repetition-based linear predictive coding tonal. The energy of the signal yl '(n), n = 0, ..., L + N-1 is scaled to obtain the signal yl (n), n = 0, ..., L + N-1 , whose energy can coincide with that of the signal xl (n), n = L, ..., L + N-1, where N is the length of the crossfade signal. The output signal zl (n), n = 0, ..., L-1 corresponding to the current frame is zl (n) = yl (n), n = 0, ..., L-1.
The xl (n), n = L, ..., L + N-1 is updated as signal zl (n) obtained by cross fading of the xl (n), n = L, ..., L + N- 1 and yl (n), n = L, ..., L + N-1.
The linear predictive coding method based on tonal repetition involved in figure 3 is shown in figure 4:
Before finding a lost frame, zl (n) is stored in a buffer for future use, when a received frame is a good frame.
When a first lost frame appears, two steps are required to synthesize the final signal yl '(n). First, the signal passed zl (n), n = -Q, ...- 1 is analyzed, and then the signal yl '(n) is synthesized in combination with the result of the analysis, where Q is the necessary length of the signal when the past signal is analyzed.
The module for linear predictive coding based on tonal repetition specifically comprises the following parts:
(1) Linear Prediction Analysis (LP)
Short-term analysis A (z) and 1 / A (z) synthesis filters are based on P-order LP filters. The LP analysis filter is defined as:
A (z) = 1 + a-iz '<sup>1</sup>+ a2z '<sup>2</sup>+ ... + apz<sup>-P</sup>
After the LP analysis of the filter A (z), the residual signal e (n), n = -Q, ..., - 1 corresponding to the passed signal zl (n), n = -Q, ... is obtained. , -1, using the following formula:
P e (n) = zl (n) + aizl (ni), n = -Q, ..., - 1 i = 1 (2) Analysis of the passed signal
The tonal repeat method is used to compensate for the lost signal. Therefore, a tonal period T0 corresponding to the past signal zl (n), n = -Q, ..., - 1 has to be estimated. The steps in detail are as follows: First, zl (n) is pretreated to remove an unnecessary low-frequency part in the long-term prediction (LTP) analysis; then the tonal period T0 of zl (n) could be obtained by LTP analysis; and the voice class could be obtained by combining with a signal class module, after the tonal period T0 is obtained.
The voice classes are shown in Table 1:
ES 2 374 043 T3
Table 1: voice classes
<td>Class name</td><td>Description</td>
<td>TRANSIENT</td><td>for voice that is transient with large energy variation (e.g., stops)</td>
<td>UNVOICED (NO VOICE)</td><td>for non-voice signals</td>
<td>VUV-TRANSITION</td><td>corresponding to a transition between voice signals and non-voice signals</td>
<td>WEAKLY VOICED</td><td>the beginning or end of voice signals</td>
<td>VOICED (VOICE)</td><td>voice cues (eg, regular)</td>
(3) Tonal repeat
A tonal repetition modulus is used to estimate the residual signal LP e (n), n = 0, ..., L-1 corresponding to the lost frame. Before tonal repetition, if the voice class is not VOICED, the magnitude of each sample will be limited by the following formula:
e (n) = minI max (| e (n - T + i) |), e (n) | I xsign (e (n)), n = -Tü, ..., - 1 \ i = —2, ···, +2 J where sign (x) = <
if x> 0 if x <0
If the voice class is VOICED, the residual e (n), n = 0, ..., L-1 corresponding to the lost signal will be obtained by repeating the residual signal corresponding to the last tonal period in a signal just received from a frame good, that is: e (n) = e (n-Tü).
For other kinds of voice, in order to avoid the periodicity that the generated data is too strong (for the UNVOICED signal, if the periodicity is too strong, it will sound like musical noises or other uncomfortable noises), the following formula is used to generate the residual signal e (n), n = 0, ..., L-1 corresponding to the lost signal:
e (n) = e (n-Tc + (-1)<sup>n</sup>).
In addition to generating the residual signal corresponding to the lost frame, in order to guarantee a smooth connection between the lost frame and the first good frame that follows the lost frame, the residual signal e (n), n = L will be continuously generated. , ..., L + N-1, from the additional N sample, in order to generate a signal for crossfade.
(4) LP synthesis
After generating the residual signal e (n) corresponding to the lost frame and the signal for crossfade, the reconstructed signal of the lost frame is given by:
ylpre (n) = e (n) -a¡yl (ni) i = 1 where e (n), n = 0, ..., L-1 is the residual signal obtained in the tonal repetition. In addition, N samples of ylpre (n), n = L, ..., L + N-1 are generated using the above formula; these samples are used for cross fade.
(5) Adaptive squelch
The energy of the ylpre (n) is controlled according to the different classes of voices provided in Table 1. That is:
yl (n) = gsilence (n) xylpre (n), n = 0, ..., L + M-1, gsilence (n) e [0 1] where gsilence (n) corresponds to a silencing factor corresponding to each sample. The value of gsilence (n) changes according to different voice classes and the situation of packet loss. In what follows is
ES 2 374 043 T3 offers an example:
For voices with large energy variation, for example plosives, corresponding to the TRANSIENT class and VUV_TRANSITION class voice in Table 1, the fade rate may be a bit high. For voices with small energy variation, the fading rate may be a bit slow. For convenient description, a 1 ms signal is assumed to include R samples.
Specifically, for the TRANSIENT class voice, within 10 ms (in total S = 10 * R samples), making gsilence (1) = 1, gsilence (n) fades from 1 to 0. gsilence (n) corresponding to samples after 10 ms it is 0, which can be displayed using a formula like:
<sup>g</sup>silence<sup>(n)</sup>= <sup><</sup><sup>g</sup> silence n = 0, ..., S -1 n> S
For the VUV_TRANSITION class voice, the fade rate within the initial 10 ms may be a bit slow and the voice fades to 0 quickly within the next 10 ms, which can be illustrated using a formula like:
<sup>g</sup>s<sup>il</sup>and<sup>n</sup>c<sup>i</sup>or<sup>(n)</sup>= 'gsilence <sup>(n -1) -</sup> , <sub>1λ</sub> 0.024 (n +1) gsilence <sup>(n-1)</sup>--<sub>χ |</sub>----- silence <sup>(S -</sup> 1) ' <sup>(n + 1 - S)</sup> n = 0, ..., S -1 n = S, ..., 2S -1 n> 2S
For voice of other classes, the fading rate within the initial 10 ms may be a little slow, the fading rate within the next 10 ms may be a little higher, and the voice fading to 0 rapidly within next 20 ms, which can be illustrated using a formula like:
<sup>S</sup> silence <sup>(n -</sup>1) <sup>-</sup>
0.024 (n +1)
S + 1 <sup>g</sup>silence<sup>(n)</sup>= <sup><</sup><sup>S</sup> silence <sup>(n 1)</sup>
0.048 · (n +1 - S) S +1 <sub>σ</sub> or _<sub>n</sub>_H.H<sub>I</sub>ienc<sub>¡</sub>or (2 S - 1) (n +1 - 2 S) <sup>S</sup> silence <sup>(n 1)</sup>
2S +1 n = 0, ..., S -1 n = S, ..., 2S -1 n = 2S, ..., 4S -1 n> 4S
The energy scale change in figure 3 consists of:
The detailed method to perform the energy scaling to yl '(n), n = 0, ..., L + N-1 according to xl (n), n = L, ..., L + M-1 and yl '(n), n = L, ..., L + M-1 includes the following steps, with reference to Figure 3.
Step s201, an energy E1 corresponding to the synthesized signal yl '(n), n = L, ..., L + M-1 and an energy E2 corresponding to the signal xl (n), n = L, ..., L + M-1.
L + M-1 L + M-1
Specifically, E1 = yl '<sup>2</sup> (i) and E2 = xl<sup>2</sup>(i) i = L i = L where M is the number of signal samples when the energy is calculated. The value of M could be flexibly set according to specific cases. For example, in the circumstances where the length of the frame is a bit short, such as the length L of the frame is less than 5 ms, M = L is recommended; In the circumstances where the length of the frame is a bit long and the tonal period is less than the length of one frame, M could be set as the corresponding length of a tone period signal.
Step s202, the energy ratio R between E1 and E2 is calculated.
Specifically,
ES 2 374 043 T3
R = sign (Ei-E2) <sup>AND</sup>1 - <sup>AND</sup>2
Ei where the sign () function is a symbolic function and is defined as follows:
sign (x) = <
if x> 0 if x <0
Step s203, the signal magnitude yl '(n), n = 0, ..., L + N-1 is adjusted according to the energy ratio R.
Specifically,
R yl (n) = yl '(n) * (1 --------- * n) n = 0, ..., L + N-1,
L + N where N is a length used for crossfade by the current frame. The value of N could be set flexibly according to specific cases. In this circumstance where the length of the frame is a bit short, N could be set as the length of one frame, that is, N = L.
In order to avoid the circumstance that the magnitude of the energy has an excessive flow (the magnitude of the energy exceeds the maximum allowable value of the corresponding magnitudes of the samples) when Ei <E2 using the above method, only The previous formula is used for the fading of the signal yl '(n), n = 0, ..., L + N-1 when Ei <E2.
When the previous frame is a lost frame and the current frame is also a lost frame, it is not necessary to scale the energy for the previous frame, that is, the corresponding yl (n) for the previous frame is :
yl (n) = yl '(n) n = 0, ..., L-1.
Specifically, the crossfade in Figure 3 is:
In order to make a smooth energy transition, after generated yl (n), n = 0, ..., L + N-1 executing a change of energy scale by means of the synthesized signal yl '(n), n = 0, ..., L + N-1, it is necessary to treat low band signals by cross fade. The standard is shown in Table 2.
Table 2: The Cross Fade Norm
<td colspan="2" rowspan="2"></td><td colspan="2">running plot</td>
<td>lost plot</td><td>good plot</td>
<td rowspan="2">previous plot</td><td>lost plot</td><td>zl (n) = yl (n), n = 0, ..., L-1</td><td>nn zl (n) = ------- xl (n) + (1 --------) yl (n), N -1 N -1 n = 0, ..., N-1 <sup>Y</sup>zl (n) = xl (n), n = N, ..., L-1</td>
<td>good plot</td><td>zl (n) = yl (n), n = 0, ..., L-1</td><td>zl (n) = xl (n), n = 0, ..., L-1</td>
In Table 2, zl (n) is the signal that corresponds to the signal corresponding to the current frame finally emitted as output. xl (n) is the signal of the good frame corresponding to the current frame. and l (n) is a signal synthesized at the same time corresponding to the current frame.
The schematic diagram of the above processes is shown in figure 5.
The first row is an original sign. The second row is the synthesized signal displayed as a broken line. The bottom row is an output signal represented as a dashed line, which is the signal after power adjustment. Frame N is a lost frame and frame N-1 and frame N + 1 are both good frames. First, the energy ratio between the energy of the received signal of the N + 1 frame and the energy of the synthesized signal corresponding to the N + 1 frame is calculated, and then the synthesized signal is faded according to the
ES 2 374 043 T3 energy ratio, to obtain the output signal in the lowest row. The fading method can refer to step s203 above. Lastly, the crossfade treatment is executed. For frame N, an output signal after the fading of frame N is taken as the output of frame N (it is assumed in this case that the output of the signal is allowed to have at least a delay of one frame, that is that is, frame N could be output after frame N + 1 is received as input). For the N + 1 frame, according to the principle of cross fading, the output signal of the N + 1 frame after fading multiplied by a descending interval, is superimposed on the original signal received from the N + 1 frame multiplied by an ascending interval. The signal obtained by superposition is taken as the output of the N + 1 frame.
In a second embodiment of the present invention, a signal processing method is provided which is adapted to process the synthesized signal in packet loss concealment. The difference between the treatment methods of the first embodiment and the second embodiment is that in the first embodiment above, when the method based on the tonal period is used to synthesize the signal yl '(n), the state of phase discontinuity, as shown in figure 6.
As shown in Figure 6, the signal between two solid vertical lines corresponds to one signal frame. Given the diversity and variation of the human voice, the tonal period corresponding to the voice cannot remain unchanged but is constantly changing. Therefore, when the last tonal period of the past signal is repeatedly used to synthesize the signal of the lost frame, the situation will arise where the waveform between the end of the synthesized signal and the beginning of the current frame is discontinuous. . The waveform has a sharp change, namely the phase mismatch situation. From figure 6 it can be seen that the distance from the starting point of the current frame to the minimum distance adaptation points on the left of the synthesized signal is, and that the distance from the starting point of the frame current to the minimum distance matching points to the right of the synthesized signal is dc. In the prior art, a method is provided to achieve phase adaptation by performing interpolation for the synthesized signal. For example, the corresponding phase separation d is -de when the frame length is L (if the optimum adaptation point is to the left of the starting point of the current frame, and the distance between the optimum point and the point starting point of the current frame is, then d = -de; if the optimum adaptation point is to the right of the starting point of the current frame, and the distance between the optimum point and the starting point of the current frame is dc, so d = dc). And then the signal from L + d samples is interpolated to generate the signal from N samples by the interpolation method.
The signal is synthesized based on the tonal repetition in figure 6, so the situation of phase mismatch inevitably occurs as well. In order to avoid the situation, a method is provided and the principle schematic diagram is shown in Figure 7. The difference between this embodiment and the first embodiment is that the energy scaling treatment can be executed after executing the phase adaptation for the linear predictive coding signal based on the tonal repetition. The phase adaptation is performed for the signal yl '(n), n = 0, ..., L + N-1 before the energy scale change. For example, an interpolated signal yl (n), n = 0, ..., L + N-1 can be obtained by executing the interpolation on yl '(n), n = 0, ..., L + N-1 , using the above interpolation method, and the signal yl (n) can be obtained by performing the energy scaling for yl (n) in combination with the signal xl (n) and the signal yl (n). Finally, the cross fade step is the same step as the first embodiment.
Through the use of the signal processing method provided by the embodiments of the present invention, adjusts the synthesized signal according to the energy ratio between the energy of the first good frame following the lost frame and the energy of the synthesized signal to ensure that there is no abrupt change in the waveform not a sudden change in energy at the place where the lost frame and the first frame following the lost frame meet for the synthesized signal, This allows you to achieve a smooth transition of the waveform and avoid musical noise.
A third embodiment of the present invention also provides a signal processing apparatus which is adapted to process the synthesized signal in packet loss concealment. The schematic diagram of the structure is shown in figure 8. The apparatus includes:
a detection module 10, configured to notify a power obtaining module 30, when it is detected that a frame following a lost frame is a good frame;
the energy obtaining module 30, configured to obtain an energy ratio between the energy of the good frame signal and the energy of the synchronized synthesized signal when the notification sent by the detection module 10 is received;
a synthesized signal adjustment module 40, configured to adjust the synthesized signal according to the energy ratio obtained by the energy obtaining module 30.
Specifically, the module 30 for obtaining energy also includes:
ES 2 374 043 T3 a sub-module 21 for obtaining energy from the good frame signal, configured to obtain the energy from the good frame signal;
a sub-module 22 for obtaining energy from the synthesized signal, configured to obtain the energy from the synthesized signal; and a sub-module 23 for obtaining an energy ratio, configured to obtain the energy ratio between the energy of the good frame signal and the energy of the synchronized synthesized signal.
Furthermore, the apparatus for signal processing also comprises:
a phase adaptation module 20, configured to perform phase adaptation for the synthesized signal admitted as input and send the synthesized signal, after phase adaptation, to the energy obtaining module 30, represented in figure 9, as a second signal processing apparatus provided by the third embodiment of the invention.
Furthermore, as shown in Figure 10, the phase adaptation module 20 can also be arranged between the energy obtaining module 30 and the synthesized signal adjustment module 40, configured to obtain the energy ratio between the energy of the good frame signal and the energy of the synthesized signal corresponding to the same instant of the good frame and execute phase adaptation for a signal admitted as input to the phase adaptation module 20 and send the signal, after phase adaptation, to the synthesized signal adjustment module 40.
A specific application case of the processing apparatus of the third embodiment of the present invention is shown in Fig. 11. In the case that a current frame is not lost, a low-band ADPCM decoder decodes the received current frame to obtain a signal. xl (n), n = 0, ..., L-1 and an output corresponding to the current frame is zl (n), n = 0, ..., L-1. In this condition, the reconstruction signal is not changed in the crossfade. Namely:
zl [n] = xl [n], n = 0, ..., L-1 where L is the frame length.
In the event that the current frame is lost, a synthesized signal yl '(n), n = 0, ..., L-1 corresponding to the current frame is generated using the linear predictive coding method based on tonal repetition. . Depending on whether a following frame is lost or not lost after the lost frame, a different treatment is executed:
When the following frame is lost after the current frame:
In this condition, the signal processing apparatus of the embodiments of the invention does not process the synthesized signal yl '(n), n = 0, ..., L-1. The output signal zl (n), n = 0, ..., L-1 corresponding to a first lost frame is the synthesized signal yl '(n), n = 0, ..., L-1, that is , zl [n] = yl [n] = yl '[n], n = 0, ..., L-1
When the following frame is not lost after the current frame:
When the synthesized signal yl '(n), n = 0, ..., L + N-1 is processed using the signal processing apparatus of embodiments of the invention, the good frame (i.e., the next frame below of the first lost frame) that is used is the good frame xl (n), n = L, ..., L + M-1 obtained after decoding the ADPCM decoder, where M is the number of signal samples when calculating energy. The synthesized signal that is used corresponding to the same instant of the good signal is the signal yl '(n), n = L, ..., L + M-1 generated by linear predictive coding based on tonal repetition. The yl '(n), n = 0, ..., L + N-1 is treated to obtain the signal yl (n), n = 0, ..., L + N-1 that can coincide with the signal xl (n), n = L, ..., L + N-1 in its energy, where N is the length of the signal to perform the crossfade. The output signal zl (n), n = 0, ..., L-1 corresponding to the current frame is zl (n) = yl (n), n = 0, ..., L-1.
Update xl (n), n = L, ..., L + N-1 to the signal zl (n), which is obtained by cross fading of xl (n), n = L, ..., L + N-1 and yl (n), n = L, ..., L + N-1.
Using the signal processing apparatus provided by embodiments of the present invention, the synthesized signal is adjusted according to the energy ratio between the energy of the first good frame following the lost frame and the energy of the signal. synthesized, to ensure that there is no sudden waveform change or sudden energy change where the lost frame and the first frame following the lost frame meet for the synthesized signal, allowing for a smooth transition of the waveform and avoid musical noises.
A fourth embodiment of the present invention provides a speech decoder, as shown in FIG. 12, that includes a high-band decoder unit 50 configured to decode a decode signal.
ES 2 374 043 T3 received high band and compensate for a lost high band signal frame; a low-band decoding unit 60 configured to decode a received low-band decoding signal and compensate for a lost low-band signal frame; a quadrature mirror filter unit 70 configured to synthesize a low-band decoded signal and a high-band decoded signal to obtain a final output signal. The high-band decoding unit 50 decodes the received high-band code stream signal and synthesizes the lost high-band signal frame. The low-band decoding unit 60 decodes the received low-band code stream signal and synthesizes the lost low-band signal frame. The quadrature mirror filter unit 70 synthesizes the low-band decoded signal output from the low-band decoder unit 60 and the high-band decoded signal output from the high-band decoder unit 50, to obtain a signal decoded final.
The low-band decoding unit 60, as shown in Figure 13, specifically includes the following modules: a linear predictive encoder sub-unit 61 based on pitch repetition configured to generate a synthesized signal corresponding to a lost frame; a low-band decoder sub-unit 62 configured to decode a received low-band code stream signal; a signal processing sub-unit 63 configured to adjust the synthesized signal; a cross-fade sub-unit 64 configured to cross-fade the signal decoded by the low-band decoder sub-unit and the signal adjusted by the signal processing sub-unit 63.
The low-band decoder sub-unit 62 decodes a received low-band signal. The linear predictive coding sub-unit 61 based on tonal repetition obtains a synthesized signal by linear predictive coding for the low-band signal frame. The signal processing sub-unit 63 adjusts the synthesized signal to make the magnitude of the energy of the synthesized signal consistent with the magnitude of the energy of the decoded signal processed by the low-band decoding sub-unit 62, and avoid the appearance of musical noises. The cross-fade sub-unit 64 crossfades the decoded signal processed by the low-band decoder sub-unit 62 and the synthesized signal adjusted by the signal processing sub-unit 63 to obtain the final decoded signal after recording. compensation for lost plot.
The structure of the signal processing sub-unit 63 has three different shapes corresponding to schematic structural diagrams of the signal processing apparatus shown in Figures 8 to 10, and their detailed description is omitted.
Thanks to the description of the above embodiments, the person skilled in the art could clearly understand that the present invention could be achieved using software and the required general hardware platform, or by hardware, but the former is, in many cases, a better embodiment. Based on such an understanding, the substantial issue of the technical solution of the present invention or the part that contributes to the prior art could be implemented in the form of software products. Computer software products are stored on a storage medium and comprise various instructions for causing the apparatus to implement the method described in each embodiment of the present invention.
While an illustration and description of the present disclosure have been provided in conjunction with its preferred embodiments, it should be appreciated by those of ordinary skill in the art that various changes can be made to its shape and details, without departing from the scope of this disclosure, which are defined in the appended claims.
Contents9
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
61 members in 14 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 200710169616 | China | A | |
| 200710169616 | China | A | |
| 200710169616 | China | – | |
| 200710169616 | – | – | – |
| CN20071169616 | – | – | – |
Members61
| Document | Office | Kind | |
|---|---|---|---|
| CN101207459A | China | A | |
| CN101207665A | China | A | |
| EP2056291A1 | European Patent Office (EPO) | A1 | |
| EP2056292A2 | European Patent Office (EPO) | A2 | |
| US2009116486A1 | United States of America | A1 | |
| US2009119098A1 | United States of America | A1 | |
| KR20090046713A | Republic of Korea | A | |
| KR20090046714A | Republic of Korea | A | |
| WO2009059497A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009059498A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2056292A3 | European Patent Office (EPO) | A3 | |
| JP2009116332A | Japan | A | |
| JP2009175693A | Japan | A | |
| CN100550712C | China | C | |
| CN101578657A | China | A | |
| US2009292542A1 | United States of America | A1 | |
| CN101601217A | China | A | |
| US2009316598A1 | United States of America | A1 | |
| EP2056291B1 | European Patent Office (EPO) | B1 | |
| AT456126T | Austria | T | |
| EP2056292B1 | European Patent Office (EPO) | B1 | |
| EP2157572A1 | European Patent Office (EPO) | A1 | |
| EP2161719A2 | European Patent Office (EPO) | A2 | |
| DE602008000579D1 | Germany | D1 | |
| AT458241T | Austria | T | |
| PT2056291E | Portugal | E | |
| EP2161719A3 | European Patent Office (EPO) | A3 | |
| DE602008000668D1 | Germany | D1 | |
| DK2056292T3 | Denmark | T3 | |
| ES2340975T3 | Spain | T3 | |
| PL2056292T3 | Poland | T3 | |
| JP2010176142A | Japan | A | |
| DE202008017752U1 | Germany | U1 | |
| EP2161719B1 | European Patent Office (EPO) | B1 | |
| AT484052T | Austria | T | |
| US7835912B2 | United States of America | B2 | |
| DE602008002938D1 | Germany | D1 | |
| JP4586090B2 | Japan | B2 | |
| CN101207665B | China | B | |
| HK1142713A1 | Hong Kong, China | A1 | |
| KR101023460B1 | Republic of Korea | B1 | |
| US7957961B2 | United States of America | B2 | |
| CN102122511A | China | A | |
| CN102169692A | China | A | |
| EP2157572B1 | European Patent Office (EPO) | B1 | |
| AT529854T | Austria | T | |
| JP4824734B2 | Japan | B2 | |
| ES2374043T3This record | Spain | T3 | |
| HK1154696A1 | Hong Kong, China | A1 | |
| HK1155844A1 | Hong Kong, China | A1 | |
| KR101168648B1 | Republic of Korea | B1 | |
| CN102682777A | China | A | |
| CN101578657B | China | B | |
| US8320265B2 | United States of America | B2 | |
| CN101601217B | China | B | |
| JP5255585B2 | Japan | B2 | |
| CN102682777B | China | B | |
| CN102122511B | China | B | |
| CN102169692B | China | B | |
| BRPI0808765A2 | Brazil | A2 | |
| BRPI0808765B1 | Brazil | B1 |
Numbers
- Publication
- 2374043
- Publication, DOCDB
- 2374043
- Publication, EPODOC
- ES2374043T
- Application
- 9176498
- Application, DOCDB
- 09176498
- Application, EPODOC
- ES20090176498T
Titles2
- Spanish
- METODO DE TRATAMIENTO DE SENALES, APARATO DE TRATAMIENTO Y DESCODIFICADOR DE VOZ.
- English
- SIGNAL TREATMENT METHOD, TREATMENT DEVICE AND VOICE DECODER.
Classification
- CPC, 2
- G10L19/005
- G10L19/04
- IPC, 2
- G10L19 00
- G10L19 005