Method and apparatus for obtaining an attenuation factor
Abstract
method and apparatus for obtaining an attenuation factor. the present invention relates to a method for obtaining an attenuation factor. the method is adapted to process the synthesized signal in packet loss cancellation, and includes: obtaining a signal changing tendency; obtain an attenuation factor according to the signal's changing tendency. the present invention also describes an apparatus for obtaining an attenuation factor. a self-adaptive attenuation factor is dynamically adjusted using the latest trend of changing a history signal using the present invention. the smooth transition of the historical data to the last received data is performed so that the attenuation speed is kept as constant as possible between the compensated signal and the original signal to adapt to the characteristic of varied human voices.
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
15 claims: 3 independent, 12 dependent
- 1CLAIMS REIVINDICAÇÕES 1. Method for processing a synthesized voice signal in hiding packet loss, comprising:1. Método para processamento de um sinal de voz sintetizado em ocultação de perda de pacotes, compreendendo: obtain a tendency to change the voice signal;obter uma tendência de mudança do sinal de voz;caracterizado pelo fato de que obter uma tendência de mudança do sinal de voz compreende: characterized by the fact that obtaining a tendency to change the voice signal comprises: obtain a power ratio of a last pitch periodic voice signal to energy of a previous pitch periodic voice signal in the voice signal, or obtain a ratio of a difference between a maximum amplitude value and a minimum amplitude value of the last pitch periodic voice signal for a difference between a maximum amplitude value and a minimum span value of the previous tone periodic voice signal in the voice signal;obter uma razão de energia de um último sinal de voz periódico de tom para energia de um sinal de voz periódico de tom anterior no sinal de voz, ou obter uma razão de uma diferença entre um valor de amplitude máximo e um valor de amplitude mínimo do último sinal de voz periódico de tom para uma diferença entre um valor de amplitude máximo e um valor de amplitude mínimo do sinal de voz periódico de tom anterior no sinal de voz;obtain an attenuation factor according to the tendency to change the voice signal;and obtain a lost picture reconstructed after attenuating according to the attenuation factor. obter um fator de atenuação de acordo com a tendência de mudança do sinal de voz;e obter um quadro perdido reconstruído após atenuar de acordo com o fator de atenuação.
- 12Apparatus for processing a synthesized voice signal in hiding packet loss, in which the apparatus comprises:12. Aparelho para processar um sinal de voz sintetizado em ocultação de perda de pacotes, em que o aparelho compreende: a change trend obtaining unit adapted to obtain a change in the voice signal;and an attenuation factor acquisition unit adapted to obtain an attenuation factor according to the change trend obtained by the change trend obtaining unit;uma unidade de obtenção de tendência de mudança adaptada para obter uma tendência de mudança do sinal de voz;e uma unidade de obtenção de fator de atenuação adaptada para obter um fator de atenuação de acordo com a tendência de mudança obtida pela unidade de obtenção de tendência de mudança;a lost frame reconstruction unit adapted to obtain a lost frame after attenuating according to the attenuation factor;uma unidade de reconstrução de quadro perdido adaptada para obter um quadro perdido após atenuar de acordo com o fator de atenuação;caracterizado pelo fato de que a unidade de obtenção de tendência de mudança compreende: characterized by the fact that the unit for obtaining a trend of change comprises: a sub-unit for obtaining energy adapted for uma subunidade de obtenção de energia adaptada para Petition 870200044915, of 4/8/2020, p. 28/37 Petição 870200044915, de 08/04/2020, pág. 28/37 4/6 obtaining energy from a last periodic tone voice signal and energy from a previous tone periodic voice signal in the voice signal;and an energy ratio subunit adapted to obtain an energy ratio of the last tone periodic voice signal to the energy of the previous tone periodic voice signal in the voice signal obtained by the energy source subunit, where the reason is used to express the tendency to change the voice signal;or a subunit of obtaining amplitude difference adapted to obtain the difference between a maximum amplitude value and a minimum amplitude value of a last periodic tone voice signal, and the difference between a maximum amplitude value and an amplitude value minimum of a periodic voice signal of previous tone in the voice signal;and a subunit of obtaining amplitude difference ratio adapted to obtain a ratio of the difference of the last periodic tone signal to the difference of the previous periodic tone signal in the voice signal, where the difference of the last tone signal periodic pitch voice and the difference from the previous pitch periodic voice signal are obtained by the amplitude difference obtaining subunit, and the ratio is used to express the tendency for the voice signal to change. 4/6 obter energia de um último sinal de voz periódico de tom e energia de um sinal de voz periódico de tom anterior no sinal de voz;e uma subunidade de obtenção de razão de energia adaptada para obter uma razão da energia do último sinal de voz periódico de tom para a energia do sinal de voz periódico de tom anterior no sinal de voz obtido pela subunidade de obtenção de energia, em que a razão é usada para expressar a tendência de mudança do sinal de voz;ou uma subunidade de obtenção de diferença de amplitude adaptada para obter a diferença entre um valor de amplitude máximo e um valor de amplitude mínimo de um último sinal de voz periódico de tom, e a diferença entre um valor de amplitude máximo e um valor de amplitude mínimo de um sinal de voz periódico de tom anterior no sinal de voz;e uma subunidade de obtenção de razão de diferença de amplitude adaptada para obter uma razão da diferença do último sinal de voz periódico de tom para a diferença do sinal de voz periódico de tom anterior no sinal de voz, em que a diferença do último sinal de voz periódico de tom e a diferença do sinal de voz periódico de tom anterior são obtidas pela subunidade de obtenção de diferença de amplitude, e a razão é usada para expressar a tendência de mudança do sinal de voz.
- 15Voice decoder, comprising:a low band decoding unit, a high band decoding unit and a quadrature mirrored filter unit, where: 15. Decodificador de voz, compreendendo: uma unidade de decodificação de banda baixa, uma unidade de decodificação de banda alta e uma unidade de filtro espelhado em quadratura, em que: a unidade de decodificação de banda baixa é adaptada para decodificar um sinal de voz de decodificação de banda baixa recebido, e compensar um sinal de voz de banda baixa perdido;the low band decoding unit is adapted to decode a received low band decode speech signal, and to compensate for a lost low band speech signal;a unidade de decodificação de banda alta é adaptada para decodificar um sinal de voz de decodificação de banda alta recebido, e compensar um sinal de voz de banda alta perdido;the high band decoding unit is adapted to decode a received high band decoding voice signal, and to compensate for a lost high band voice signal;a unidade de filtro espelhado em quadratura é adaptada para obter um sinal de voz de saída final ao sintetizar o sinal de voz de decodificação de banda baixa e o sinal de voz de decodificação de banda alta;the quadrature mirrored filter unit is adapted to obtain a final output speech signal by synthesizing the low band decoding speech signal and the high band decoding speech signal;a unidade de decodificação de banda baixa compreende uma subunidade de decodificação de banda baixa, uma codificação de predição linear baseada em uma subunidade de repetição de tom e uma the low band decoding unit comprises a low band decoding subunit, a linear prediction encoding based on a tone repeat subunit and a Petition 870200044915, of 4/8/2020, p. 30/37 Petição 870200044915, de 08/04/2020, pág. 30/37 6/6 cross fading subunit;6/6 subunidade de fading cruzado;em que a subunidade de decodificação de banda baixa é adaptada para decodificar um sinal de voz de fluxo de banda baixa recebido;wherein the low band decoding subunit is adapted to decode a received low band flow voice signal;a codificação de predição linear (LPC) baseada na subunidade de repetição de tom é adaptada para gerar um sinal de voz sintetizado correspondendo a um quadro perdido;linear prediction coding (LPC) based on the tone repetition subunit is adapted to generate a synthesized speech signal corresponding to a lost frame;a subunidade de fading cruzado é adaptada para realizar o fading cruzado do sinal de voz processado pela subunidade de decodificação de banda baixa e o sinal de voz sintetizado correspondendo ao quadro perdido gerado pela LPC baseada na subunidade de repetição de tom;the cross-fading subunit is adapted to cross-fade the voice signal processed by the low-band decoding subunit and the synthesized voice signal corresponding to the lost frame generated by the LPC based on the tone repetition subunit;caracterizado pelo fato de que a LPC baseada na subunidade de repetição de tom compreende um módulo de análise e um aparelho como definido em qualquer uma das reivindicações 12 a 14, em que o módulo de análise é adaptado para analisar um histórico de sinal de voz, e gerar um sinal de voz reconstruído do quadro perdido. characterized by the fact that the LPC based on the tone repetition subunit comprises an analysis module and an apparatus as defined in any of claims 12 to 14, in which the analysis module is adapted to analyze a voice signal history, and generate a reconstructed voice signal from the lost frame.
Independent claims3
193 paragraphs in 7 sections, as filed
Descriptive Report of the Invention Patent for METHOD AND APPARATUS FOR PROCESSING A SYNTHESIZED VOICE SIGNAL IN HIDDEN PACKAGE LOSS AND VOICE DECODER.
CROSS REFERENCE TO RELATED DEPOSIT REQUESTS
[001] This order claims the priority of International Order
N_PCT / CN2008 / 070807, filed on April 25, 2008, which claims priority benefit from Chinese Patent Application No. 2007101696180, filed on November 5, 2007. The contents of the applications identified above are hereby incorporated by reference in its entirety.
TECHNOLOGY FIELD
[002] The present invention relates to the field of signal processing, and particularly to a method and apparatus for obtaining an attenuation factor.
BACKGROUND
[003] It is required that a voice data transmission be in real time and reliable in a real time voice communication system, for example, a VoIP (Voice over IP) system. Due to the unreliable characteristics of a network system, a data packet may be lost or not reach its destination in time in a procedure for transmitting from a sending end to a receiving end. These two types of situations are considered to be loss of network packet by the receiving end. Network packet loss is inevitable. However, the loss of network packet is one of the most important factors that influence the speech quality of the voice. Therefore, a resistant packet loss hiding method is necessary to recover the lost data packet in the real-time communication system.
Petition 870200044915, of 4/8/2020, p. 4/37
2/22 so that a good speech quality is still obtained under the network packet loss situation.
[004] In the existing real-time voice communication technology, at the sending end, an encoder divides a broadband voice into a high subband and a low subband, and uses ADPCM (Adaptive Pulse Code Modulation) Differentials) to encode the two subbands respectively and send them together to the receiving end via the network. At the receiving end, the two subbands are decoded respectively by the ADPCM decoder, and then the final signal is synthesized using a QMF synthesis filter (Mirrored Quadrature Filter).
[005] Different Packet Loss Concealment (PLC) methods are adopted for two different sub-bands. For a low band signal, under the situation without packet loss, a reconstruction signal is not altered during FADING-CROSS. Under the packet loss situation, for the first lost frame, the history signal (the history signal is a voice signal before the lost frame in the present order document) is analyzed using a short-term indicator and an a long-term, and voice classification information is extracted. The lost frame signal is reconstructed using a LPC (linear predictive encoding) based on the tone repetition method, the indicator and the classification information. The ADPCM status will also be updated synchronously until a good frame is found. In addition, not only the signal corresponding to the lost frame needs to be generated, but also a signal section adapted for FADING-CRUZADO needs to be generated. In this way, once a satisfactory frame is received, FADING-CROSSED is performed to process the satisfactory frame signal and the signal section. It is observed that this type of FADING-CROSS occurs only after the receiving end loses a frame and receives the first
Petition 870200044915, of 4/8/2020, p. 5/37
3/22 satisfactory picture.
[006] During the process of carrying out the present invention, the inventor realized at least the following problems in the prior art: the energy of the synthesized signal is controlled using a static auto-adaptive attenuation factor in the prior art. Although the defined attenuation factor changes gradually, its attenuation speed, that is, the attenuation factor value, is the same with respect to the same voice classification. However, human voices are varied. If the attenuation factor does not match the characteristic of human voices, an uncomfortable noise will occur in the reconstruction signal, particularly at the end of the fixed vowels. The static self-adaptive attenuation factor may not be adapted to the characteristic of different human voices.
[007] The situation shown in figure 1 is taken as an example, where To is the tone period of the historical signal. The upper signal corresponds to an original signal, that is, a schematic diagram in wave form under the situation without packet loss. The lower signal with a dashed line is a signal synthesized according to the prior art. As can be seen from the figure, the synthesized signal does not maintain the same attenuation speed as the original signal. If there is often the same tone repetition, the synthesized signal will produce obvious musical noise so that the difference between the situation of the synthesized signal and the desired situation is large.
SUMMARY
[008] One embodiment of the present invention provides a method and apparatus for obtaining an attenuation factor adapted to obtain a self-adaptive and dynamically adjustable attenuation factor used in synthetic signal processing.
[009] One embodiment of the present invention provides a method for obtaining the attenuation factor adapted to process the signal
Petition 870200044915, of 4/8/2020, p. 6/37
4/22 synthesized in hiding packet loss, including:
[0010] obtain a tendency to change a signal; and
[0011] obtain an attenuation factor according to the signal's changing tendency.
[0012] One embodiment of the present invention also provides an apparatus for obtaining an attenuation factor for processing a synthesized signal in hiding packet loss. The device to obtain an attenuation factor is configured to:
[0013] obtain a tendency to change a signal; and
[0014] obtain an attenuation factor according to the obtained trend of change.
[0015] An embodiment of the present invention also provides a method and apparatus for obtaining an attenuation factor adapted to smoothly transition the historical data to the last received data.
[0016] To achieve the above objective, an embodiment of the invention provides a method for signal processing, adapted to process a synthesized signal in hiding packet loss, including:
[0017] obtain a tendency to change a signal;
[0018] obtain an attenuation factor according to the trend of changing the signal; and
[0019] to obtain a lost picture reconstructed after the attenuation according to the attenuation factor.
[0020] One embodiment of the present invention also provides a signal processing apparatus for processing a synthesized signal in hiding packet loss, including:
[0021] the device to obtain an attenuation factor to process a signal synthesized in hiding packet loss; and
[0022] a lost frame reconstruction unit adapted
Petition 870200044915, of 4/8/2020, p. 7/37
5/22 to obtain a lost picture reconstructed after attenuation according to the attenuation factor.
[0023] One embodiment of the present invention also provides a speech decoder adapted to decode the speech signal, including a low band decoding unit, a high band decoding unit and a quadrature mirrored filter unit.
[0024] The low band decoding unit is adapted to decode a received low band decode signal, and to compensate for a lost low signal.
[0025] The high band decoding unit is adapted to decode a high band decoding signal, and to compensate for a lost high band signal.
[0026] The quadrature mirrored filter unit is adapted to obtain a final output signal synthesizing the low band decoding signal and the high band decoding signal.
[0027] The low band decode signal unit includes a low band decode signal subunit, an LPC based on the tone repeat subunit and a cross fading subunit.
[0028] The low band decoding subunit is adapted to decode a received low band flow signal.
[0029] The LPC based on the tone repetition subunit is adapted to generate a synthesized signal corresponding to the lost frame.
[0030] The cross-fading subunit is adapted to cross-fade the signal processed by the low-band decoding subunit and synthesized signal corresponding to the lost frame generated by the LPC based on the tone repetition subunit.
[0031] The LPC based on the tone repeat subunit includes a
Petition 870200044915, of 4/8/2020, p. 8/37
6/22 analysis module and a signal processing module.
[0032] The analysis module is adapted to analyze a historical signal and generate a reconstructed lost frame signal.
[0033] One embodiment of the present invention further provides a computer program product, including computer program codes that allow a computer to perform any step in the method to obtain the attenuation factor adapted to process the signal synthesized in hiding packet loss or any step in the signal processing method to process a signal synthesized in hiding packet loss when computer program codes are executed by the computer.
[0034] Compared with the prior art, the modalities of the present invention have the following advantages:
[0035] a self-adaptive attenuation factor is dynamically adjusted using the trend of changing a historical signal. The smooth transition from historical data to the last received data is performed so that the attenuation speed between the compensated signal and the original signal is kept as constant as possible to adapt the characteristic of different human voices.
BRIEF DESCRIPTION OF THE DRAWING (S)
[0036] Figure 1 is a schematic diagram showing the original signal and the signal synthesized according to the prior art;
[0037] figure 2 is a flow chart illustrating a method for obtaining an attenuation factor in accordance with Mode 1 of the present invention;
[0038] figure 3 is a schematic diagram that illustrates the principles of the encoder;
[0039] figure 4 is a schematic diagram illustrating the module of an LPC based on the unit's tone repetition subunit
Petition 870200044915, of 4/8/2020, p. 9/37
7/22 low band decoding;
[0040] figure 5 is a schematic diagram that illustrates an output signal after adopting the dynamic attenuation method according to Modality 1 of the present invention;
[0041] figures 6A and 6B are schematic diagrams that illustrate the structure of the apparatus to obtain an attenuation factor according to Modality 2 of the present invention;
[0042] figure 7 is a schematic diagram that illustrates the application scenario of the device to obtain an attenuation factor according to Modality 2 of the present invention;
[0043] figures 8A and 8B are schematic diagrams illustrating the structure of the apparatus for signal processing according to Modality 3 of the present invention;
[0044] figure 9 is a schematic diagram illustrating the speech decoder module according to Modality 4 of the present invention;
[0045] figure 10 is a schematic diagram illustrating the module of the low band decoding unit in the voice decoder according to Mode 4 of the present invention;
[0046] Figure 11 is a schematic diagram illustrating the LPC module based on the tone repetition subunit according to Modality 4 of the present invention.
DETAILED DESCRIPTION
[0047] The present invention will be described in more detail with reference to the drawings and modalities.
[0048] A method is provided to obtain an attenuation factor in Mode 1 of the present invention, adapted to process the signal synthesized in hiding packet loss, as shown in Figure 2, including the following steps.
[0049] Step s101, a change trend is obtained from one
Petition 870200044915, of 4/8/2020, p. 10/37
8/22 signal;
[0050] Specifically, the trend of change can be expressed in the following parameters: (1) the energy ratio of the last periodic tone signal to the energy of the previous tone periodic signal in the signal; (2) the ratio of the difference between the maximum amplitude value and the minimum amplitude value of the last periodic tone signal to the difference between the maximum amplitude value and the minimum amplitude value of the previous periodic tone signal in the signal.
[0051] Step s102, an attenuation factor is obtained according to the trend of change.
[0052] The specific Modality 1 processing method of the present invention will be described together with the specific application scenario.
[0053] A method for obtaining an attenuation factor that is adapted to process the signal synthesized in hiding packet loss is provided in Mode 1 of the present invention.
[0054] As shown in figure 3, different PLC methods are adopted for two different sub-bands. The PLC method for the low band part is shown as part 1 in a dashed frame in figure 3. While a dashed frame 2 in figure 3 corresponds to the high band PLC algorithm. For a high band signal,<sup>zh (n</sup>) is a high band signal finally produced. After obtaining the low band signal<sup>zl (n</sup>) and the high band signal <sup>zh (n</sup>), QMF runs for the low-band signal and the high-band signal and a finally produced broadband signal <sup>y (n</sup>) is synthesized.
[0055] Only the low band signal is described in detail below.
[0056] Under the situation without loss of frame, the signal <sup>xl (n</sup>)’<sup>n</sup> = °’···’<sup>L</sup> ~<sup>1</sup> is obtained after decoding the current frame received by the low band ADPCM decoder, and the output is <sup>zl</sup>(<sup>n</sup>)’<sup>n</sup> = °’···’<sup>L</sup>~<sup>1</sup>
Petition 870200044915, of 4/8/2020, p. 11/37
9/22 corresponding to the current table. In this situation, the reconstruction signal is not changed during the FADING-CRUZADO, which is<sup>zl [</sup>n \ = xlη], n = o, ..., L-1, where L is the length of the frame;
[0057] Under the situation with loss of frames, with respect to the first lost frame, the history signal <sup>zl</sup>(<sup>n</sup>), <sup>n</sup> < <sup>0</sup> it is analyzed using a short-term indicator and a long-term indicator, and voice classification information is extracted. Adopting the above indicators and classification information, the signal<sup>yl</sup>(<sup>n</sup>) is generated using an LPC method based on tone repetition. And the lost frame signal<sup>zl</sup>(<sup>n</sup>) is reconstructed as <sup>zl</sup>(<sup>n</sup>) = <sup>yl</sup>(<sup>n</sup>, <sup>n</sup>=<sup>0</sup>, *. In addition, the status of ADPCM will also be updated synchronously until a satisfactory picture is obtained. It is observed that not only the signal corresponding to the lost frame needs to be generated, but also a 10ms signal<sup>yl</sup>(<sup>n</sup>),<sup>n</sup> = <sup>L</sup>,·” <sup>-1</sup> adapted for
FADING-CRUZADO needs to be generated, the M is the number of signal sampling points that are included in the process when calculating the energy. In this way, once a satisfactory frame is received, FADING-CRUZADO is performed to<sup>xl</sup> (<sup>n</sup>), <sup>n</sup> = <sup>Ly</sup><sup>-</sup> 1 and <sup>yl</sup> (<sup>n</sup>), <sup>n</sup> = <sup>L</sup> ,’’ <sup>-</sup> 1. It is observed that this type of
FADING-CROSSING occurs only after a loss of frame and when the receiving end receives the first satisfactory frame data.
[0058] An LPC based on the tone repetition method in the figure is as shown in figure 4.
[0059] When the data frame is a satisfactory frame, <sup>zl</sup> (<sup>n</sup>) is stored in a buffer for future use.
[0060] When the first missing frame is found, the final signal <sup>yl (n)</sup> needs to be synthesized in two steps. At first, the signal
Petition 870200044915, of 4/8/2020, p. 12/37
10/22 history <sup>zl</sup>(<sup>n</sup>), <sup>n</sup>= <sup>29Π</sup>·>'·'·> <sup>1</sup> is analyzed. So, the signal<sup>yl (n</sup>), <sup>n</sup> = <sup>0</sup>, ···, L-1 is synthesized according to the result of the analysis, where L is the frame length of the data frame, that is, the number of sampling points corresponding to a signal frame, Q is the length of the signal that is needed to analyze the history signal. [0061] The LPC module based on tone repetition specifically includes the following parts.
(1) An analysis of LP (Linear Prediction)
[0062] The short-term analysis filter <sup>A (z</sup>) and the synthesis filter
1/<sup>A (z</sup>) are Linear Prediction (LP) filters based on the P order. The LP analysis filter is defined as:
A (z) = 1 + az<sup>-</sup> + a2 z <sup>-2</sup> + ----- + ap z<sup>- P</sup>
[0063] Through the LP analysis of the historical signal <sup>zl (n</sup>), <sup>n</sup> Q ··· with the filter <sup>A (z</sup>), a residual signal <sup>e (n</sup>), <sup>n Q,</sup>'corresponding to the history signal <sup>zl (n</sup>), <sup>n Q</sup>,' is obtained:
P e (n) = zl (n) + ^ azl (n - i), n = - Q, ..., - 1 i = 1 (2) A historical signal analysis
[0064] The lost signal is compensated by a repetition method. ^ of tone. So first, a period of tone<sup>0</sup> corresponding to the history signal <sup>zl (n</sup>), <sup>n Q</sup>, 'needs to be estimated. The steps are as follows:<sup>zl (n</sup>) is pre-processed to remove a useless low-frequency ingredient in an LTP (long-term prediction) analysis, and the tone period <sup>T0</sup> of <sup>zl (n</sup>) can be obtained by analyzing LTP. Voice classification is obtained despite combining a ....... ....... T ... T signal classification module after obtaining the tone period<sup>0</sup>.
[0065] The voice classifications are as shown in table 1 below:
Petition 870200044915, of 4/8/2020, p. 13/37
11/22
Table 1 Voice classifications
<td>Classification Name</td><td>Explanation</td>
<td>TRANSIENT</td><td>for voices with a wide range of energy (for example, stops)</td>
<td>UNVOICED</td><td>for speechless signals</td>
<td>VUV TRANSITION</td><td>for a transition between voice and speechless signals</td>
<td>WEAKLY_VOICED</td><td>for weak voice signals (for example, initial or final vowels)</td>
<td>VOICED</td><td>voice signals (for example, fixed vowels)</td>
(3) One tone repetition
[0066] A tone repetition module is adapted to estimate a residual LP signal <sup>e (n</sup>), <sup>n</sup>= °, 'of a lost frame. Before tone repetition is performed, if the voice classification is not VOICED, the following formula is adopted to limit the amplitude of a sample:
e (n) = mini max (e (n - T<sub>O</sub> + i)) | e (n)]
Ç i = -2, -, + 2 ....... J x sign (e (n)), n = -T °, ---, - 1
[0067] where, sign (x) = if x> ° if x <°
[0068] If the voice classification is VOICED, the <sup>e (n</sup>),<sup>n</sup> = °,''',<sup>L 1</sup> residual signal corresponding to the lost signal is obtained by adopting a repeat step of the residual signal corresponding to the signal of the last tone period in the signal of a recently received satisfactory frame, which is:
[0069] With respect to other classifications of voices, to avoid that the periodicity of the generated signal is too intense (in relation to the signal without voice, if the periodicity is very intense, you can hear some uncomfortable noise like a noise), the residual signal <sup>and (n)</sup> ,
Petition 870200044915, of 4/8/2020, p. 14/37
12/22 <sup>n</sup> - ο, -, L 1 corresponding to the lost signal is generated using the following formula:
e (n) - e (n - Tο + (-1) <sup>n</sup> )
[0070] In addition to generating the residual signal corresponding to the lost frame, the residual signals <sup>and (n)</sup>, <sup>n -</sup> L -'-<sup>L + N-1</sup> of additional samples N continue to be generated to generate a signal adapted to FADING-CROSSED, in order to guarantee the smooth combination between the lost frame and the first satisfactory frame after the lost frame. (4) An LP analysis
[0071] After generating the residual signal <sup>and (n)</sup> corresponding to the lost frame and the FADING-CRUZADO, a reconstruction lost frame signal <sup>yl</sup>and<sup>(n)</sup>, <sup>n-0</sup>,’ <sup>L-1</sup> is obtained using the following formula: 8 yljree<sup>(n</sup>) <sup>-</sup> and<sup>(n</sup>) <sup>-</sup> Σ ^^<sup>(n</sup> ~ i) i-1
[0072] where the residual signal <sup>and (n)</sup>, <sup>n-0</sup>,’ <sup>L-1</sup> is the residual signal obtained from the tone repeat steps above.
[0073] Also, <sup>yl</sup>P<sup>re (n)</sup>, <sup>n -</sup> L -'-<sup>L + N-1</sup> with samples <sup>N</sup> adapted for FADING-CRUZADO are generated using the formula above.
(5) An adaptive silencing
[0074] To carry out a smooth energy transition, before executing QMF with the high band signal, the low band signal also needs to perform FADING-CRUZADO, the rules are shown as the following table:
Petition 870200044915, of 4/8/2020, p. 15/37
13/22
<td colspan="2" rowspan="2"></td><td colspan="2">Current framework</td>
<td>Painting unsatisfactory</td><td>Satisfactory picture</td>
<td>Painting previous</td><td>Painting unsatisfactory</td><td>zl (n) = yl (n) , n = 0, -, L - 1</td><td>_ n. n.<sup>zl (n</sup>) = ττ<sup>-</sup>;·<sup>xl (n</sup>)<sup>+(1—</sup>τ<sup>-</sup>/) yl<sup>(n</sup>) N-1 N-1 , n = 0, -, N-1 and zl (n) = xl (n) n = N, -, L - 1,</td>
<td></td><td>Satisfactory picture</td><td>zl (n) = yl (n) , n = 0, -, L - 1</td><td>zl (n) = xl (n) n = 0, -, L - 1,</td>
[0075] In the table above, <sup>zl (n)</sup> it is a signal finally produced corresponding to the current picture; <sup>xl (n)</sup> is the sign of the satisfactory picture corresponding to the current picture; <sup>yl (n)</sup> is a synthesized signal corresponding to the current frame at the same time, where L is the frame length, N is the number of samples that perform FADING-CROSSED.
[0076] With respect to different voice classifications, the signal energy in <sup>yl</sup>P<sup>re</sup>(<sup>n</sup>) is controlled before executing FADING-CRUZADO according to the coefficient corresponding to each sample. The value of the coefficient changes according to different voice classifications and the packet loss situation.
[0077] In detail, in the case where the last two periodic tone signals in the received history signal is the original signal as shown in figure 5, the self-adaptive dynamic attenuation factor is dynamically adjusted according to the changing trend of the last two tone periods in the history signal. The detailed adjustment method includes the following steps:
[0078] Step s201, the trend of changing the signal is obtained.
[0079] The trend of signal change can be expressed by the ratio of the energy of the last tone period signal to the energy of the
Petition 870200044915, of 4/8/2020, p. 16/37
14/22 previous tone period signal in the signal, that is, the energy Ei and E2 of the last two tone period signals in the history signal, and the ratio of the two energies is calculated.
T
E 1 = £ xl <sup>2</sup>(-ii) i = 1
To
E2 = £ xl<sup>2</sup>(-i - To) i = 1
[0080] <sup>E1</sup> is the energy of the last tone period signal, <sup>E2</sup> and the . t energy of the previous tone period signal, and<sup>O</sup> is the pitch period corresponding to the history signal.
[0081] Optionally, the trend of signal change can be expressed by the reason of the peak-valley differences of the last two tone periods in the historical signal.
Pi = max (xl (i)) - min (xl (j)) (i, j) = -To, ..., - 1
P2 = <sup>m</sup>The<sup>x (xl (i</sup>)) <sup>-</sup> min (xl<sup>(</sup>j)) <sup>(</sup>i, j) = <sup>-</sup>2<sup>T</sup>O,...,<sup>- (T</sup>the +1)
[0082] where, <sup>P1</sup> is the difference between the maximum amplitude value and the minimum amplitude value of the last periodic tone signal, <sup>P</sup>two is the difference between the maximum amplitude value and the minimum amplitude value of the previous tone periodic signal, and the ratio is calculated as:
[0083] Step s202, the synthesized signal is dynamically attenuated according to the change trend obtained from the signal.
[0084] The calculation formula is shown as follows:
yl (n) = yl<sub>pre</sub> (n) * (1 - C * (n + 1)) n = o, .., N -1
[0085] where, <sup>yl</sup>P<sup>re (n</sup>) is the sign of lost picture of reconstruction,
N is the length of the synthesized signal, and C is the attenuation coefficient
Petition 870200044915, of 4/8/2020, p. 17/37
Self-adaptive 15/22 whose value is:
<sub>Ç</sub>_1 - R
T <sup>2</sup>0
[0086] Under the attenuation factor situation <sup>1 Ç</sup>*(<sup>n + 1</sup>)<<sup>0</sup>, you need to adjust <sup>1-C</sup>*(<sup>n + 1</sup>)<sup>=0</sup>, to avoid showing a situation where the attenuation factor corresponding to the samples is negative.
[0087] In particular, to avoid the situation where the amplitude value corresponding to a sample is exceeded under the situation of R> 1, the synthesized signal is dynamically attenuated using the formula of step s202 in the present modality that takes into account only the situation of R <1.
[0088] In particular, to avoid the situation where the signal attenuation speed with less energy is very fast, only under the situation where <sup>E1</sup> exceeds a certain limit value, the synthesized signal is dynamically attenuated using the formula of step s202 in the present modality.
[0089] In particular, to prevent the attenuation speed of the synthesized signal from being too fast, especially under the situation of continuous loss of frame, an upper limit value is adjusted for the attenuation coefficient <sup>Ç</sup>. When<sup>Ç</sup>*(<sup>n + 1</sup>) exceeds a limit value, the attenuation coefficient is set as the upper limit value.
[0090] In particular, under the situation of a bad network environment and continuous loss of frame, a certain condition can be established to avoid a very fast attenuation speed. For example, one can take into account that, when the number of frames lost exceeds a fixed number, for example, two frames; or when the signal corresponding to the lost frame exceeds a fixed length, for example, 20ms; or in at least one of the conditions above the current attenuation coefficient<sup>1 - Ç</sup> *(<sup>n +</sup> 1) achieve
Petition 870200044915, of 4/8/2020, p. 18/37
16/22 a fixed threshold value, the attenuation coefficient <sup>Ç</sup> needs to be adjusted to avoid the very fast attenuation speed which can result in the situation where the output signal is silenced.
[0091] For example, under situation sampling at a frequency of 8k and the frame length of 40 samples, the number of frames lost can be adjusted to 4, and after the attenuation factor <sup>1 _ Ç</sup> *(<sup>n +</sup> 1) becomes less than 0.9, the attenuation coefficient <sup>Ç</sup> is adjusted to be a smaller value. The rule of adjusting the smallest value is as follows.
[0092] Hypothetically, it is predicted that the current attenuation coefficient is C and the attenuation factor value is V, and the attenuation factor V can be attenuated to 0 after the samples. <sup>V</sup>/<sup>Ç</sup>. While a more desirable situation is one where the attenuation factor V must be attenuated to 0 after samples M (M *<sup>V</sup> /<sup>Ç</sup>). So, the attenuation coefficient<sup>Ç</sup> is set to:
C = V / M
[0093] As shown in figure 5, the upper signal is the original signal;
the intermediate signal is the synthesized signal. As seen in the figure, although the signal has some degree of attenuation, the signal still remains sound intensive. If the duration is too long, the signal can be shown as a musical noise, especially at the end of the sound. The lower signal is the signal after using dynamic attenuation in the mode of the present invention, which can be very similar to the original signal.
[0094] According to the method provided by the aforementioned modality, the self-adaptive attenuation factor is dynamically adjusted using the trend of changing the historical signal, so that the smooth transition from historical data to the last received data can be performed . The attenuation speed is kept as constant as possible between the signal
Petition 870200044915, of 4/8/2020, p. 19/37
17/22 compensated and the original signal to adapt the characteristic of varied human voices.
[0095] An apparatus for obtaining an attenuation factor is provided in Modality 2 of the present invention, adapted to process the synthesized signal in hiding packet loss, including:
[0096] a changing trend obtaining unit 10, adapted to obtain a changing trend of a signal;
[0097] an attenuation factor 20 unit, adapted to obtain an attenuation factor according to the change trend obtained by the change trend unit 10.
[0098] The unit for obtaining attenuation factor 20 additionally includes: a unit for obtaining attenuation coefficient 21, adapted to generate the attenuation coefficient according to the change trend obtained by the unit for obtaining change tendency 10; an attenuation factor 22 subunit, adapted to obtain an attenuation factor according to the attenuation coefficient generated by the attenuation factor 21 subunit. The attenuation factor 20 unit also includes: an attenuation coefficient adjustment subunit 23, adapted to adjust the attenuation coefficient value obtained by the attenuation coefficient subunit 21 to a certain value under certain conditions that include at least one of the following characteristics: if the value of the attenuation coefficient exceeds an upper limit value; if there is a situation of continuous loss of staff; and if the attenuation speed is too fast.
[0099] The method for obtaining an attenuation factor in the above modality is the same method for obtaining an attenuation factor in
Petition 870200044915, of 4/8/2020, p. 20/37
18/22 methods of the method.
[00100] In detail, the trend of change obtained by the unit of obtaining trend of change 10 can be expressed in the following parameters: (1) the ratio of the energy of the last periodic tone signal to the energy of the previous periodic tone signal in the signal; (2) the ratio of the difference between the maximum amplitude value and the minimum amplitude value of the last periodic tone signal to the difference between the maximum amplitude value and the minimum amplitude value of the previous periodic tone signal in the signal.
[00101] When the trend of change is expressed in the energy ratio in (1), the structure of the device to obtain an attenuation factor is as shown in figure 6A. The unit for obtaining a trend of change 10 also includes:
[00102] an energy obtaining subunit 11 adapted to obtain the energy of the last periodic tone signal and the energy of the previous tone periodic signal;
[00103] an energy ratio obtaining subunit 12 adapted to obtain the energy ratio of the last periodic tone signal to the energy of the previous tone periodic signal obtained by the energy obtaining subunit 11 and uses the reason to show the trend of changing the signal.
[00104] When the trend of change is expressed in the amplitude difference ratio in (2), the structure of the device to obtain an attenuation factor is as shown in figure 6B. The unit for obtaining a trend of change 10 also includes:
[00105] a subunit of obtaining amplitude difference 13, adapted to obtain the difference between the maximum amplitude value and the minimum amplitude value of the last periodic tone signal, and the difference between the maximum amplitude value and the value of minimum amplitude of the previous tone periodic signal;
Petition 870200044915, of 4/8/2020, p. 21/37
19/22
[00106] a subunit of obtaining amplitude difference ratio 14, adapted to obtain the ratio of the difference between the maximum amplitude value and the minimum amplitude value of the last periodic tone signal for the difference between the maximum amplitude value and the minimum amplitude value of the previous tone periodic signal, and use the reason to show the changing trend of the signal.
[00107] A schematic diagram illustrating the application scenario of the device to obtain an attenuation factor according to Modality 2 of the present invention is as shown in figure 7. The self-adaptive attenuation factor is dynamically adjusted using the trend of change of the history sign.
[00108] Using the device provided by the aforementioned modality, the self-adaptive attenuation factor is dynamically adjusted using the trend of changing the historical signal so that the smooth transition of the historical data to the last received data is carried out. The attenuation speed is kept as constant as possible between the compensated signal and the original signal to adapt the characteristic of varied human voices.
[00109] An apparatus for signal processing in Modality 3 of the present invention is provided, adapted to process the synthesized signal in concealment of packet loss, as shown in figure 8A and figure 8B. Based on Modality 2, a lost frame reconstruction unit 30 related to an attenuation factor obtaining unit is added. The lost frame reconstruction unit 30 obtains a lost reconstructed frame after attenuation according to the attenuation factor obtained by the attenuation factor 20 unit.
[00110] Using the device provided by the aforementioned modality, the self-adaptive attenuation factor is dynamically adjusted using the trend of changing the signal.
Petition 870200044915, of 4/8/2020, p. 22/37
20/22 history, and a lost frame reconstructed after the attenuation is obtained according to the attenuation factor, so that a smooth transition is made from the historical data to the last received data. The attenuation speed is kept as constant as possible between the compensated signal and the original signal to adapt to the characteristic of varied human voices.
[00111] A voice decoder by Modality 4 of the present invention is provided, as shown in figure 9. The speech decoder includes: a high-band decoding unit 40 is adapted to decode a received high-band decode signal and to compensate for a lost high-band signal; a low band decoding unit 50 is adapted to decode a received low band decode signal and compensate for a lost low band signal; and a quadrature mirrored filter unit 60 is adapted to obtain a final output signal synthesizing the low band decode signal and the high band decode signal. The high band decoding unit 40 decodes the high band flow signal received by the receiving end, and synthesizes the lost high band signal. The low band decoding unit 50 decodes the low band flow signal received by the receiving end and synthesizes the lost low band signal. The square mirrored filter unit 60 obtains the final decoding signal by synthesizing the low band decoding signal sent by the low band decoding unit 50 and the high band decoding signal sent by the high band decoding unit 40. [00112] For the low band decoding unit 50, as shown in figure 10, the following units are included. An LPC based on the tone repeat subunit 51 that is adapted to generate a synthesized signal corresponding to the lost frame, a low band decoding subunit 52 that is adapted to
Petition 870200044915, of 4/8/2020, p. 23/37
21/22 decode a received low band flow signal, and a cross fading subunit 53 that is adapted to cross fade the signal decoded by the low band decoding subunit and the synthesized signal corresponding to the lost frame generated by the LPC based in the tone repeat subunit.
[00113] The low band decoding subunit 52 decodes the received low band flow signal. The LPC based on the tone repeat subunit 51 generates the synthesized signal by executing an LPC on the lost low band signal. And finally, the cross fading subunit 53 crossfades the signal processed by the low band decoding subunit 52 and the synthesized signal to obtain a final decoding signal after lost frame compensation.
[00114] The LPC based on the tone repetition subunit 51, as shown in figure 10, further includes an analysis module 511 and a signal processing module 512. The analysis module 511 analyzes a history signal, and generates a reconstructed lost frame signal; the signal processing module 512 obtains a signal changing tendency, and obtains an attenuation factor according to the signal changing tendency, and attenuates the reconstructed lost frame signal, and obtains a reconstructed lost frame after attenuation .
[00115] The signal processing module 512 also includes an attenuation factor obtaining unit 5121 and a lost frame reconstruction unit 5122. The unit of obtaining attenuation factor 5121 obtains a tendency to change a signal, and obtains an attenuation factor according to the tendency to change; the lost frame reconstruction unit 5122 attenuates the lost frame signal reconstructed according to the attenuation factor, and obtains a lost frame reconstructed after the attenuation. The signal processing module 512 includes two structures corresponding to the
Petition 870200044915, of 4/8/2020, p. 24/37
22/22 schematic diagrams illustrating the structure of the apparatus for signal processing in figures 8A and 8B, respectively.
[00116] The unit for obtaining attenuation factor 5121 includes two structures corresponding to the schematic diagrams that illustrate the structure of the apparatus to obtain an attenuation factor in figure 6A and 6B, respectively. The specific functions and means of implementing the above modules and units may refer to the content revealed in the method modalities. Unnecessary details will not be repeated here.
[00117] Through the description of the modalities mentioned above, those skilled in the art can clearly understand that the present invention can be realized depending on the software plus the general and necessary hardware platform, and certainly can also be accomplished by hardware. However, in most situations, the trainer is a preferable modality. Based on such an understanding, the essence or part that contributes to the prior art in the technical scheme of the present invention can be expressed in the form of a software product that is stored in a storage medium, and the software product includes some instructions for instruct a device to perform the modalities of the present invention.
[00118] Although the illustration and description of this description are determined with reference to the modalities of this, it should be evaluated by experts in the technique that various changes in shapes and details can be made without abandoning the scope of the description.
Contents7
7 priority claims, no other members on record
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 200710169618 | China | A | |
| 2007101696180 | China | – | |
| 2008070807 | China | W | |
| 2007101696180 | – | – | – |
| CN20071169618 | – | – | – |
| PCTCN2008070807 | – | – | – |
| WO2008CN70807 | – | – | – |
Numbers
- Publication
- PI0808765
- Publication, DOCDB
- PI0808765
- Publication, EPODOC
- BRPI0808765
- Application
- 8765
- Application, DOCDB
- PI0808765
- Application, EPODOC
- BR2008PI08765
Titles2
- Portuguese
- método e aparelho para processamento de um sinal de voz sintetizado em ocultação de perda de pacotes e decodificador de voz.
- English
- METHOD AND APPARATUS FOR PROCESSING A SYNTHESIZED VOICE SIGNAL IN HIDING PACKAGE LOSS AND VOICE DECODER.
Classification
- CPC, 3
- G10L19/005
- G10L19/0204
- G10L19/097
- IPC, 4
- G10L19 005
- G10L19 02
- G10L19 097
- G10L25 12