Transmission error concealment in an audio signal
Abstract
Process of concealment of transmission error in an audio-numerical signal in which in the detection (3) of missing or erroneous samples in a signal, synthesis samples (5) are generated with the help of at least one prediction operator short term and at least for sound sounds an estimated long term prediction operator based on decoded samples of a past decoded signal, said decoded samples being memorized (6) above when the data transmitted from said past signal is valid, characterized in that the energy of the synthesis signal generated in this way is controlled with the help of a calculated and adapted gain sample by sample according to a law of adaptation that depends on at least one parameter of said stored decoded samples.

Term
Term ended
Projected expiry passed 5 September 2021, 5.1 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
18 claims: 10 independent, 8 dependent
- 1ES 2 298 261 T3 REIVINDICACIONES 1. Proceso de disimulación de error de transmisión en una señal audio-numérica en la cual en la detección (3) de muestras faltantes o erróneas en una señal, se generan muestras de síntesis (5) con la ayuda de al menos un operador de predicción a corto plazo y al menos para los sonidos sonoros un operador de predicción a largo plazo estimado en función de muestras descodificadas de una señal descodificada pasada, dichas muestras descodificadas siendo memorizadas (6) anteriormente cuando los datos transmitidos de dicha señal pasada son válidos, caracterizado porque se controla la energía de la señal de síntesis generada de esta manera con la ayuda de una ganancia calculada y adaptada muestra por muestra según una ley de adaptación que depende de al menos un parámetro de dichas muestras descodificadas memorizadas.
- 2Proceso según la reivindicación 1, caracterizado porque la ganancia para el control de la señal de síntesis es calculada en función de al menos uno de los parámetros siguientes:valores de energía previamente memorizados para las muestras que corresponden a los datos válidos, período fundamental para los sonidos sonoros, o cualquier parámetro que caracteriza el espectro de frecuencias.
- 3Proceso según una de las reivindicaciones anteriores, caracterizado porque la ganancia aplicada a la señal de síntesis decrece progresivamente en función de la duración durante la cual las muestras de síntesis son generadas.
- 4Proceso según una de las reivindicaciones anteriores, caracterizado porque se discrimina en los datos válidos los sonidos estacionarios y los sonidos no estacionarios y se ponen en práctica las leyes de adaptación de la ganancia que permiten controlar la señal de síntesis diferentes por una parte para las muestras generadas a continuación de datos válidos que corresponden a sonidos estacionarios y por otra parte para las muestras generadas a continuación de datos válidos que corresponden a sonidos no estacionarios.
- 5Proceso según una de las reivindicaciones anteriores, caracterizado porque se actualiza en función de las muestras de síntesis generadas el contenido de memorias utilizadas para el tratamiento de descodificación.
- 6Proceso según la reivindicación 5, caracterizado porque se pone en práctica al menos parcialmente sobre las muestras sintetizadas una codificación análoga a aquella puesta en práctica en el emisor seguida eventualmente de una operación de descodificación al menos parcial, los datos obtenidos sirviendo para regenerar las memorias del descodificador.
- 7Proceso según la reivindicación 6, caracterizado porque se regenera la primera trama borrada por medio de esta operación de codificación-descodificación, explotando el contenido de las memorias del descodificador antes del corte, cuando dichas memorias contienen informaciones explotables en esta operación.
- 8Proceso según una de las reivindicaciones anteriores caracterizado porque se genera a la entrada del operador de predicción a corto plazo una señal de excitación que, en zona sonora, es la suma de una componente armónica y de una componente débilmente armónica o no armónica, y en zona no sonora, limitada por una componente no armónica.
- 9Proceso según la reivindicación 8, caracterizado porque la componente armónica es obtenida poniendo en práctica una filtración por medio del operador de predicción a largo plazo aplicado sobre una señal residual calculada poniendo en práctica una filtración a corto plazo inversa sobre las muestras memorizadas.
- 10Proceso según la reivindicación 9, caracterizado porque la otra componente es determinada con la ayuda de un operador de predicción a largo plazo en el cual se aplican perturbaciones seudo-aleatorias.
- 11Proceso según una de las reivindicaciones 8 a 10, caracterizado porque para la generación de una señal de excitación sonora, la componente armónica está limitada a bajas frecuencias del espectro, mientras que, la otra componente está limitada a altas frecuencias.
- 12Proceso según una de las reivindicaciones anteriores, caracterizado porque el operador de predicción a largo plazo es determinado a partir de muestras de tramas válidas memorizadas, con un número de muestras utilizadas para esta estimación que varía entre un valor mínimo y un valor igual a al menos dos veces el período fundamental estimado para el sonido sonoro.
- 13Proceso según una de las reivindicaciones anteriores, caracterizado porque la señal residual es tratada de manera no lineal para eliminar los picos de amplitud.
- 14Proceso según una de las reivindicaciones anteriores, caracterizado porque detecta la actividad vocal estimando los parámetros de ruido y porque se hacen tender los parámetros de la señal sintetizada hacia los del ruido estimado.
- 15Proceso según la reivindicación 14, caracterizado porque se estima la envoltura espectral del ruido de las muestras descodificadas válidas y se genera una señal sintetizada que evoluciona hacia una señal que posee la misma envoltura espectral. ES 2 298 261 T3
- 16Proceso de tratamiento de señales de sonidos, caracterizado porque se pone en práctica una discriminación entre los sonidos sonoros y los sonidos musicales y cuando se detectan los sonidos musicales, se pone en práctica un proceso según una de las reivindicaciones anteriores sin estimación de un operador de predicción a largo plazo.
- 17Dispositivo de disimulación de error de transmisión en una señal audio-numérica que recibe a la entrada una señal descodificada que le transmite un descodificador y que genera muestras faltantes o erróneas en esta señal descodificada, caracterizado porque comprende medios de tratamiento aptos para poner en práctica el proceso según una de las reivindicaciones anteriores.
- 18Sistema de transmisión que comprende al menos un codificador, al menos un canal de transmisión, un módulo apto para detectar qué datos transmitidos se han perdido o son fuertemente erróneos, al menos un descodificador y un dispositivo de disimulación de errores que recibe la señal descodificada, caracterizado porque este dispositivo de disimulación de errores es un dispositivo según la reivindicación 17.
Independent claims18
231 paragraphs in 12 sections, as filed
ES 2 298 261 T3
DESCRIPTION
Concealment of transmission errors in an audio signal.
1. Technical domain
The present invention concerns techniques for concealing consecutive transmission errors in transmission systems that use any type of numerical encoding of the speech and / or sound signal.
Two main categories of encoders are classically distinguished:
- the so-called temporal encoders, which compress the numbered signal samples sample by sample (this is the case of the MIC or MICDA [DAUMER] [MAITRE] encoders for example)
- and the parametric encoders that analyze the successive frames of samples of the signal to be encoded in order to extract, in each of these frames, a certain number of parameters that are then encoded and transmitted (in the case of [TREMAIN] vocoders, of which IMBE encoders [HARDWICK], or transform encoders [BRANDENBURG]).
There are intermediate categories that complete the encoding of the representative parameters of the parametric encoders by encoding a residual temporal waveform. For simplicity, these encoders can be arranged in the category of parametric encoders.
In this category are predictive encoders and particularly the family of analysis by synthesis encoders such as RPE-LTP ([HELLWING]) or CELP ([ATAL]).
For all these encoders, the encoded values are then transformed into a binary stream that will be transmitted on a transmission channel. Depending on the quality of this channel and the type of transport, disturbances can affect the transmitted signal and produce errors on the bit stream received by the decoder. These errors can intervene in an isolated way in the binary stream but are very frequently produced by bursts. This is then a packet of bits that corresponds to a complete portion of the signal that is erroneous or not received. These types of problems are found, for example, in transmissions over mobile networks. They are also found in transmissions on packet networks and in particular on internet-type networks.
When the transmission system or the loaded reception modules allow detecting that the received data is strongly erroneous (for example in mobile networks), or that a data block has not been received (case of packet transmission systems for example ), error concealment procedures are then put into practice. These procedures make it possible to extrapolate to the decoder the samples of the missing signal from the available signals and data coming from the preceding frames and eventually following the erased areas.
Such techniques have been put into practice mainly in the case of parametric encoders (techniques for recovering erased frames). They make it possible to strongly limit the subjective degradation of the signal perceived in the decoder in the presence of erased frames. Most of the algorithms developed rely on the technique used by the encoder and decoder, and are in fact an extension of the decoder.
A general objective of the invention is to improve, for any speech and sound compression system, the subjective quality of the speech signal restored in the decoder when, due to poor quality of the transmission channel or following loss or not receiving a packet in a packet transmission system, a set of consecutive encoded data has been lost.
For this purpose, it proposes a technique that makes it possible to hide successive transmission errors (error packets) regardless of the coding technique used, the proposed technique being able to be used, for example, in the case of temporary coders whose structure is less suitable. good a priori for the concealment of error packages.
2. Prior state of the art
Most of the predictive coding algorithms propose erased frame recovery techniques ([GSM-FR], [REC G.723.1A], [SALAMI], [HONKANEN], [COX-2], [CHEN- 2], [CHEN-3], [CHEN-4], [CHEN-5], [CHEN-6], [CHEN-7], [KROON-2], [WATKINS]). The decoder is informed of the occurrence of an erased frame in one way or another, for example in the case of radio-mobile systems by transmitting the frame erasure information that comes from the channel decoder. The purpose of recovery devices for erased frames is to extrapolate the parameters of the erased frame from the last previous frames considered valid. Certain parameters manipulated or encoded by predictive coders show a strong inter-frame correlation (case of short-term prediction parameters, also called “LPC” of “Linear Predictive Coding” (see [RABINER]) that represent the spectral envelope, and long-term prediction parameters for voiced sounds, for example). Due to the fact of this correlation
ES 2 298 261 T3 it is much more advantageous to reuse the parameters of the last valid frame to synthesize the erased frame than to use erroneous or random parameters.
For the CELP coding algorithm (from “Code Excited Linear Prediction”, consult [RABINER]), the parameters of the erased frame are classically obtained as follows:
- the LPC filter is obtained from the LPC parameters of the last valid frame either by re-copying the parameters or by introducing a certain damping (cf. encoder G723.1 [REC G.723.1A]).
- the voicing is detected to determine the degree of harmonicity of the signal at the level of the erased frame ([SALAMI], this detection occurs as follows:
in the case of a non-audible signal:
an excitation signal is randomly generated (pull of a code word and gain of the past excitation slightly damped [SALAMI], random selection in the past excitation [CHEN], use of the codes transmitted eventually totally wrong [HONKANEN ] ...) in the case of an audible signal:
the LTP term is generally the term calculated in the previous frame, eventually with a slight fluctuation ([SALAMI]), the LTP gain being taken very close to 1 or equal to 1. The excitation signal is limited to long-term prediction effected from past excitement.
In all the examples cited above, the procedures for concealing the erased frames are strongly linked to the decoder and use modules of this decoder, such as the signal synthesis module. They also use intermediate signals available within this decoder as the excitation signal passed and memorized during the processing of the valid frames that precede the erased frames.
Most of the methods used to conceal the errors produced by lost packets during the transport of data encoded by time-type encoders refer to waveform substitution techniques such as those presented in [GOODMAN], [ERDOL] , [AT&T]. Methods of this type reconstitute the signal by selecting portions of the decoded signal before the lost period and do not cite synthesis models. Lysing techniques are also used to avoid artifacts produced by the concatenation of different signals.
For transform encoders, the reconstruction techniques for erased frames also rely on the coding structure used: algorithms, such as [PICTEL, MAHIEUX-2], aim to regenerate the lost transformed coefficients from the values taken. by these coefficients before deletion.
The method described in [PARIKH] can be applied to any type of signals; it is based on the construction of a sinusoidal model from the decoded valid signal that precedes the erasure, to regenerate the part of the lost signal.
Finally, there is a family of erased frame cloaking techniques developed in conjunction with channel coding. These methods, such as those described in [FINGSCHEIDT], make use of information provided by the channel decoder, for example information concerning the degree of reliability of the received parameters. They are fundamentally different from the present invention that does not presuppose the existence of a channel encoder.
A prior art that can be considered as the closest to the present invention is that described in [COMBESCURE], which proposed a method of concealment of erased frames equivalent to that used in CELP encoders for a transform encoder. The drawbacks of the proposed method were the introduction of audible spectral distortions ("synthetic" voice, parasitic resonances, ...), mainly due to the use of poorly controlled long-term synthesis filters (single harmonic component in sonic sounds, generation of the excitation signal limited to the use of portions of the residual signal passed). Furthermore, the energy control was carried out in [COMBESCURE] at the level of the excitation signal, the energy target of this signal was kept constant throughout the duration of the erasure, which also generated annoying artifacts. The same considerations apply to US5884010.
3. Presentation of the invention
The invention as defined in claims 1, 17 and 18 allows for the concealment of erased frames without marked distortion at the highest error rates and / or by longer erased intervals.
It mainly proposes a transmission error concealment process in an audio-numerical signal according to which a decoded signal is received after transmission, the decoded samples are memorized when the transmitted data is valid, and at least one operator of short-term forecast and at least one
ES 2 298 261 T3 is a long-term prediction operator based on the valid samples stored and any missing or erroneous samples are generated in the decoded signal with the help of the operators estimated in this way.
According to a first particularly advantageous aspect of the invention, the energy of the synthesis signal thus generated is controlled with the aid of a calculated and adapted sample-by-sample gain.
This contributes in particular to improving the performances of the technique in the erasing areas of a longer duration.
Mainly, the gain for the control of the synthesis signal is advantageously calculated as a function of at least one of the following parameters: energy values previously memorized by the samples that correspond to the valid data, fundamental period for the audible sounds, or any parameter that characterizes the frequency spectrum.
Also advantageously, the gain applied to the synthesis signal progressively decreases as a function of the duration during which the synthesis samples are generated.
Likewise, in a preferred manner, the stationary sounds and the non-stationary sounds are discriminated in the valid data and laws of adaptation of this gain (decay rate, for example) are put into practice, different on the one hand for the samples generated after valid data corresponding to stationary sounds and on the other hand for the samples generated following valid data corresponding to non-stationary sounds.
According to another independent aspect of the invention, the content of the memories used for the decoding treatment is updated as a function of the synthesis samples generated.
In this way, on the one hand, the possible desynchronization of the encoder and decoder is limited (see paragraph 5.1.4 below), and sharp discontinuities between the erased area reconstructed according to the invention and the samples that follow this area are avoided.
Mainly, a coding analogous to that carried out at the transmitter is carried out at least partially on the synthesized samples, eventually followed by a decoding operation (possibly partial), the data obtained serving to regenerate the memories of the decoder.
In particular, this possibly partial encoding-decoding operation can be advantageously used to regenerate the first erased frame because it makes it possible to exploit the content of the decoder's memories before cutting, when these memories contain information not provided by the last decoded valid samples (for example, example in the case of add-overlay transform encoders, see paragraph
5.2.2.2.1 point 10).
According to a different aspect of the invention, an excitation signal is generated at the input of the short-term prediction operator, which, in the sound zone, is the sum of a harmonic component and a weakly harmonic or non-harmonic component, and in Limited sound zone in the non-harmonic component.
Mainly, the harmonic component is advantageously obtained by implementing a filtering by means of the long-term prediction operator applied on a calculated residual signal by implementing an inverse short-term filtering on the stored samples.
The other component can be determined with the help of a long-term prediction operator in which pseudo-random disturbances (eg gain or period disturbances) are applied.
In a particularly preferred manner, for the generation of a sound excitation signal, the harmonic component represents the low frequencies of the spectrum, while the other component the high frequency part.
According to yet another aspect, the long-term prediction operator is determined from the valid frame samples stored, with a number of samples used for this estimation that varies between a minimum value and a value equal to at least twice the period estimated fundamental for voiced sound.
On the other hand, the residual signal is advantageously modified by non-linear type treatments to eliminate amplitude peaks.
Likewise, according to another advantageous aspect, the speech activity is detected by estimating the noise parameters when the signal is considered as not active, and the parameters of the synthesized signal are made to tend towards those of the estimated noise.
Also preferentially, the spectral envelope of the noise of the valid decoded samples is estimated and a synthesized signal is generated that evolves towards a signal that has the same spectral development.
ES 2 298 261 T3
The invention also proposes a sound signal processing process, characterized in that a discrimination between the word and musical sounds is implemented and when musical sounds are detected, a precipitous type process is implemented without estimation of an operator of long-term prediction, the excitation signal being limited to a non-harmonic component obtained for example by generating a uniform white noise.
The invention also concerns a device for concealing transmission error in an audio-numeric signal that receives at the input a decoded signal transmitted by a decoder and that generates missing or erroneous samples in that decoded signal, characterized in that it comprises processing means suitable for implement the aforementioned process.
It also comprises a transmission system comprising at least one encoder, at least one transmission channel, a module capable of detecting which transmitted data has been lost or is strongly erroneous, at least one decoder and an error concealment device that receives the decoded signal, characterized in that this error concealment device is a device of the aforementioned type.
Four. Presentation of the figures
Other characteristics and advantages of the invention will also result from the description that follows, which is purely illustrative and not limiting and should be read in relation to the attached drawings in which:
FIG. 1 is a synoptic diagram illustrating a transmission system according to a possible embodiment of the invention;
FIG. 2 and FIG. 3 are synoptic diagrams illustrating an implementation according to a possible mode of the invention;
Figures 4 to 6 schematically illustrate the windows used with the error concealment process according to a possible way of putting the invention into practice;
Figures 7 and 8 are schematic representations illustrating a possible way of putting the invention into practice in the case of musical signals.
5. Description of one or more possible embodiments of the invention
5.1 Principle of a possible embodiment
Figure 1 presents a device for encoding and decoding the digital audio signal, comprising an encoder 1, a transmission channel 2, a module 3 that allows detecting that transmitted data has been lost or is strongly erroneous, a decoder 4, and a module 5 for concealing errors or lost packets according to a possible embodiment of the invention.
It will be noted that this module 5, in addition to indicating the erased data, receives the decoded signal in a valid period and transmits signals used for its updating to the decoder.
More precisely, the treatment put into practice by module 5 is based on:
1. storing the decoded samples when the transmitted data is valid (processing 6);
2. during a block of erased data, the synthesis of the samples corresponding to the lost data (treatment 7);
3. when transmission is re-established, the lysate between the synthesis samples produced during the erased period and the decoded samples (treatment 8);
Four. the updating of the decoder's memories (treatment 9) (updating that is carried out either during the generation of the erased samples, or at the time of reestablishment of transmission).
5.1.1 In valid period
After the valid data has been decoded, the memory of the decoded samples is updated, which contains a sufficient number of samples for the regeneration of any subsequent erased periods. Typically, the order of 20 to 40 ms of signal is memorized. The energy of the valid frames is also calculated and the energies corresponding to the last valid frames processed (typically of the order of 5s) are retained in memory.
ES 2 298 261 T3
5.1.2 During a deleted data block
The following operations are carried out, illustrated by figure 3:
1. Estimation of the current spectral envelope
This spectral development is calculated in the manner of an LPC filter [RABINER] [KLEIJN]. The analysis is carried out by classical methods ([KLEIJN]) after the windowing of the memorized samples in valid period. Mainly, an LPC analysis is carried out (step 10) to obtain the parameters of a filter A (z), the inverse of which is used for the LPC filtration (step 11). As the coefficients calculated in this way are not transmitted, a high order can be used for this analysis, which allows obtaining good performances on the musical signals.
2. Detection of audible sounds and calculation of LTP parameters
A method of detecting the audible sounds (treatment 12 of figure 3: V / NV detection, by “audible / non-audible”) is used on the last memorized data. For example, the normalized correlation ([KLEIJN]) can be used for this, or the criterion presented in the exemplary embodiment that follows.
When the signal is declared voiced, the parameters that allow the generation of a long-term synthesis filter are calculated, also called the LTP filter ([KLEIJN]) (figure 3: LTP analysis, the inverse filter is defined by B (z) Calculated LTP). Such a filter is generally represented by a period corresponding to the fundamental period and a profit. The precision of this filter can be improved by the use of fractional pitch or a multicoefficient structure [KROON].
When the signal is declared not audible, a particular value is assigned to the LTP synthesis filter (see paragraph 4).
It is particularly interesting in this estimation of the LTP synthesis filter to restrict the analyzed area to the end of the period prior to erasure. The length of the analysis window varies between a minimum value and a value linked to the fundamental period of the signal.
3. Calculation of the residual signal
A residual signal is calculated by inverse LPC filtration (treatment 10) of the last stored samples. This signal is then used to generate a drive signal from the LPC synthesis filter 11 (see below).
Four. Synthesis of missing samples
The synthesis of the replacement samples is carried out by introducing an excitation signal (calculated at 13 from the output signal of the inverse LPC filter) into the LPC synthesis filter 11 (1 / A (z)) calculated at 1. This excitation signal is generated in two different ways depending on whether the signal is audible or not:
4.1 In sound zone
The excitation signal is the sum of two signals, a strongly harmonic component and the other less or not at all harmonic.
The strongly harmonic component is obtained by LTP filtration (treatment module 14) with the help of the parameters calculated in 2, of the residual signal mentioned in 3.
The second component can also be obtained by LTP filtering but made non-periodic by random modifications of the parameters, by generating a pseudo-random signal.
It is particularly interesting to limit the passband of the first component in the low frequencies of the spectrum. In the same way, it will be interesting to limit the second component at the highest frequencies.
4.2 In the silent zone
When the signal is non-voiced, a non-harmonic drive signal is generated. It is interesting to use a generation method similar to that used for voiced sounds, with parameter variations (period, gain, signs) that allow it to be made non-harmonic.
4.3 Amplitude control of the residual signal
When the signal is non-voiced, or weakly voiced, the residual signal used for the generation of the excitation is treated to eliminate the amplitude peaks significantly above the average.
ES 2 298 261 T3
5. Synthesis signal energy control
The energy of the synthesis signal is controlled with the help of a sample-by-sample calculated and adapted gain. In the case where the erase period is relatively long, it is necessary to progressively lower the energy of the synthesis signal. The gain adaptation law is calculated based on different parameters: energy values memorized before erasing (see in 1), fundamental period, and local seasonality of the signal at the moment of cutting.
If the system comprises a module that allows the discrimination of stationary (like music) and non-stationary (like speech) sounds, different adaptation laws can also be used.
In the case of add-overlay transform encoders, the first half of the memory of the last correctly received frame contains fairly precise information about the first half of the first lost frame (its weight in the add-overlay is more important than the current plot). This information can also be used to calculate the adaptive gain.
6. Evolution of the synthesis procedure over time
In the case of relatively long erasure periods, it is also possible to evolve the synthesis parameters. If the system is coupled to a voice activity detection device with estimation of noise parameters (such as [REC-G.723.1A], [SALAMI-2], [BENYASSINE]), it is particularly interesting to make the parameters generation of the signal to be reconstructed towards those of the estimated noise: in particular at the level of the spectral envelope (interpolation of the LPC filter with that of the estimated noise, the interpolation coefficients evolving over time until the noise filter is obtained) and of the energy (level that progressively evolves towards that of the noise, for example by windowing).
5.1.3 On restoration of transmission
In restoring transmission, it is particularly important to avoid brutal breaks between the erased period that has been reconstructed according to the techniques defined in the preceding paragraphs and the periods that follow, in the course of which all the transmitted information is available for use. decode the signal. The present invention performs time domain weighting with interpolation between the replacement samples prior to re-establishing communication and the valid decoded samples following the erased period. This operation is a priori independent of the type of encoder used.
In the case of add-overlay transform encoders, this operation is common with the update of the memories described in the following paragraph (see embodiment example).
5.1.4 Updating the decoder memories
When decoding of valid samples is resumed after an erasure period, there may be a degradation when the decoder uses the data normally produced in the previous and stored frames. It is important to properly update these memories to avoid these artifacts.
This is particularly important for coding structures that use recursive processes, which use information obtained after decoding the previous samples for a sample or a sequence of samples. These are for example the predictions ([KLEIJN]) that allow to extract the redundancy of the signal. This information is normally available both in the encoder, which for this must have carried out a form of local decoding for these previous samples, and in the remote decoder present at the reception. Since the transmission channel is disturbed and the remote decoder no longer has the same information as the local decoder present in the broadcast, there is desynchronization between the encoder and the decoder. In the case of strongly recursive coding systems, this desynchronization can cause audible impairments that can last for a long time, even amplifying over time if there are instabilities in the structure. In this case, it is then important to make an effort to re-synchronize the encoder and decoder, that is, to estimate the memories of the decoder as close as possible to those of the encoder. However, the resynchronization techniques depend on the encoding structure used. One will be presented whose principle is general in the present patent, but whose complexity is potentially important.
One possible method consists of introducing into the decoder at reception an encoding module of the same type as the one present at the broadcast, which allows the encoding-decoding of the signal samples produced by the techniques mentioned in the previous paragraph to be carried out during the erased periods. In this way the memories necessary to decode the following samples are completed with data a priori close (subject to a certain seasonality during the erasure period) of those that have been lost. In the case where this hypothesis of seasonality would not be respected, after a long erased period for example, there is not in any way sufficient information to act better.
In fact it is generally not necessary to carry out the complete coding of these samples, it is limited to the modules necessary to update the memories.
ES 2 298 261 T3
This implementation can be carried out at the time of production of the replacement samples, which distributed the complexity over the entire erasure area, but accumulates with the synthesis procedure described above.
When the encoding structure allows it, the above procedure can also be limited to an intermediate zone at the beginning of the period of valid data succeeding an erased period, the updating process then accumulating with the decoding operation.
5.2. Description of particular examples of realization
Particular examples of possible implementation are given below. The case of transform encoders of the TDAC or TCDM type ([MAHIEUX]) is in particular addressed.
5.2.1 Device description
Transform numerical encoding / decoding system of the TDAC type.
Encoder in amplified band (50-7000 Hz) at 24 kb / s or 32 kb / s.
20 ms frame (320 samples).
40 ms windows (640 samples) with add-overlays of 20 ms. A binary frame containing the encoded parameters obtained by the TDAC transformation on a window. After decoding these parameters, doing the TDAC inverse transformation, an output frame of 20 ms is obtained which is the sum of the second half of the previous window and the first half of the current window. On figure 4, the two parts of windows used for the reconstruction of frame n (in temporal) have been marked in thickness. In this way, a lost binary frame disturbs the reconstruction of two consecutive frames (the current one and the next, figure 5). On the contrary, by correctly replacing the lost parameters, it is possible to recover the parts of the information that come from the previous and next binary frames (figure 6), for the reconstruction of these two frames.
5.2.2 Implementation
All the operations described below are put into practice at reception, according to figures 1 and 2, either within the module for concealing the erased frames that communicate with the decoder, as well as in the decoder itself (memory update decoder).
5.2.2.1 In valid period
In correspondence with paragraph 5.1.2, the memory of the decoded samples is updated. This memory is used for the LPC and LTP analysis of the passed signal in the case of an erasure of a binary frame. In the example presented here, the LPC analysis is done over a signal period of 20 ms (320 samples). In general, LTP analysis needs more samples to memorize. In our example, to be able to do the LTP analysis correctly, the number of memorized samples is equal to twice the maximum value of the pitch. For example, if the maximum value of the MaxPitch pitch is set to 320 samples (50 Hz, 20 ms), the last 640 samples will be memorized (40 ms of the signal). The energy of the valid frames is also calculated and stored in a circular buffer of length 5s. When an erased frame is detected, the energy of the last valid frame is compared with the maximum and with the minimum of this circular buffer to know its relative energy.
5.2.2.2 During a deleted data block
When a binary frame is lost, two different cases are distinguished:
5.2.2.2.1 First binary frame lost after valid period
First, an analysis of the memorized signal is made to estimate the parameters of the model that serve to synthesize the regenerated signal. This model then allows us to synthesize 40 ms of signal, which corresponds to the lost 40 ms window. By doing the TDAC transformation followed by the inverse TDAC transformation on this synthesized signal (without coding-decoding the parameters), the 20 ms output signal is obtained. Thanks to these reverse TDAC-TDAC operations, the information that comes from the correctly received previous window is exploited (see figure 6). At the same time, the decoder memories are updated. In this way, the next binary frame, if well received, can be decoded normally, and the decoded frames will be automatically synchronized (figure 6).
The operations to be carried out are the following:
1. Window of the memorized signal. For example, an asymmetric Hamming window of 20 ms can be used.
2. Calculation of the autocorrelation function on the windowed signal.
ES 2 298 261 T3
3. Determination of the coefficients of the LPC filter. For this, the iterative Levinson-Durbin algorithm is classically used. The order of analysis can be high, especially when the encoder is used to encode music sequences.
Four. Loudness detection and long-term analysis of the memorized signal for modeling the eventual periodicity of the signal (voiced sounds). In the presented embodiment, the inventors limited the estimate of the fundamental period Tp to integer values, and calculated an estimate of the loudness degree in the form of the correlation coefficient MaxCorr (see below) evaluated in the selected period. Let Tm = max (T, Fs / 200), where Fs is the sampling frequency, then Fs / 200 samples correspond to a duration of 5 ms. To better model the evolution of the signal at the end of the previous frame, the correlation coefficients Corr (T) corresponding to a delay T are calculated using only 2 * Tm samples at the end of the memorized signal:
Lmem-l
Σ miiBi-T
Corr (T) = _______ i = Lmem-2Tm + T
Lmem-l Lmem-lT
Σ <sup>m</sup>i<sup>2 +</sup> Σ <sup>m</sup>i<sup>2</sup> i = Lmem-2T<sub>1</sub>r, i = Lmem-2T<sub>m</sub>+ T where m<sub>0</sub>... m<sub>Lmem-1</sub> it is the memory of the previously decoded signal. From this formula, it is seen that the length of this memory L<sub>mem</sub> It must be at least 2 times the maximum value of the fundamental period (also called “pitch”) MaxPitch.
The minimum value of the fundamental period MinPitch has also been set, which corresponds to a frequency of 600 Hz (26 samples with Fs = 16 kHz).
Corr (T) is calculated for T = 2,, MaxPitch. If T 'is the smallest delay such that Corr (T') <0 (very short-term correlations are eliminated in this way), then MaxCorr is searched, maximum of Corr (T) for T '<T <= MaxPitch . Let Tp be the period corresponding to MaxCorr (Corr (Tp) = MaxCorr). It also looks for MaxCorrMP, maximum of Corr (T) for T '<T <= 0.75 * MinPitch. If Tp <MinPitch or MaxCorrMP> 0.7 * MaxCorr and if the energy of the last valid frame is relatively weak, it is decided that the frame is not voiced, because using LTP prediction you would risk getting a very annoying resonance at high frequencies. The chosen pitch is Tp = MaxPitch / 2, and the MaxCorr correlation coefficient set at a weak value (0.25).
The plot is also considered as non-voiced when more than 80% of its energy is concentrated in the last MinPitch samples. It is then an output of the word, but the number of samples is not enough to estimate the eventual fundamental period, it is better to treat it as a non-voiced frame, even to decrease the energy of the synthesized signal more quickly (to indicate this, it is put DiminFlag = 1).
In the case where MaxCorr> 0.6, it is verified that a multiple (4, 3 or 2 times) of the fundamental period was not found. For this, the local maximum of the correlation is sought around Tp / 4, Tp / 3 and Tp / 2. The position of this maximum is noted Ti, and MaxCorrL = Corr (T<sub>1</sub>). If T<sub>1</sub> > MinPitch and MaxCorrL> 0.75 * MaxCorr, choose T<sub>1</sub> as a new fundamental period.
If Tp is less than MaxPitch / 2, it can be verified if it is really a voiced frame by looking for the local maximum of the correlation around 2 * TP (TPP) and verifying if Corr (T<sub>pp</sub>)> 0.4. If Corr (T<sub>pp</sub>) <0.4 and if the signal energy decreases, DiminFlag = 1 is set and the MaxCorr value is decreased, if the next local maximum between the current Tp and MaxPitch is not searched.
Another voicing criterion consists in verifying whether in at least 2/3 of the cases the signal delayed by the fundamental period has the same sign as the non-delayed signal.
This is verified over a length equal to the maximum between 5 ms and 2 * Tp.
It is also verified whether the signal energy has a tendency to decrease or not. If yes, DiminFlag = 1 is set and the MaxCorr value is decreased as a function of the degree of decrease.
The voicing decision also takes into account the energy of the signal: if the energy is strong, the MaxCorr value is increased, in this way it is more likely that the plot will be decided voiced. On the contrary, if the energy is very weak, the value of MaxCorr is decreased.
Finally, the voicing decision is made based on the MaxCorr value: the plot is not voiced if and only if MaxCorr <0.4. The fundamental period Tp of a non-voiced frame is defined, it must be less than or equal to MaxPitch / 2.
ES 2 298 261 T3
5. Calculation of the residual signal by inverse LPC filtration of the last stored samples. This residual signal is stored in the ResMem memory.
6. Equalization of residual signal energy. In the case of a soft or soft signal (MaxCorr <0.7), the energy of the residual signal stored in ResMem can change abruptly from one part to the other. The repetition of this excitation causes a very unpleasant periodic disturbance in the synthesized signal. To avoid this, it is ensured that no significant amplitude peaks are present in the drive of a weakly voiced frame. Since the excitation is constructed from the last Tp samples of the residual signal, this vector of Tp samples is treated. The method used in our example is the following:
The mean MeanAmpl of the absolute values of the last Tp samples of the residual signal is calculated.
If the vector of the samples to be treated contains n zero passages, it is cut into n + 1 sub-vectors, the sign of the signal in each sub-vector being then invariable.
The maximum amplitude MaxAmplSv of each sub-vector is sought. If MaxAmplSv> 1.5 * MeanAmpl, multiply the sub-vector by 1.5 * MeanAmpl / MaxAmplSv.
7. Preparation of the excitation signal of a length of 640 samples corresponding to the length of the TDAC window. Two cases are distinguished according to the sound system:
The excitation signal is the sum of two signals, a strongly band-limited harmonic component in the low frequencies of the excb spectrum and another less harmonic limited in the higher frequencies exch.
The strongly harmonic component is obtained by LTP filtering of the order 3 of the residual signal:
excb (i) = 0.15 * exc (i-Tp-1) + 0.7 * exc (i-Tp) + 0.15 * exc (i-Tp + 1)
The coefficients [0.15, 0.7, 0.15] correspond to a low-pass FIR filter with 3 dB attenuation at Fs / 4.
The second component is also obtained by a LTP filtration made non-periodic by the random modification of its fundamental period Tph. Tph is chosen as the integer part of a random real value Tpa. The initial value of Tpa is equal to Tp and then it is modified sample by sample adding a random value in [-0.5, 0.5]. Furthermore, this LTP filtration is combined with a high pass IIR filtration:
exch (i) = -0.0635 * (exc (i-Tph-1) + exc (i-Tph + 1)) +
0.1182 * exc (i-Tph) -0.9926 * exch (i-1) 0.7679 * exch (i-2)
Sound excitation is then the sum of these two components:
Exc (i) = excb (i) + exch (i)
In the case of a non-voiced frame, the excitation signal is also obtained by LTP filtering of order 3 with the coefficients [0.15, 0.7, 0.15] but it is made non-periodic by increasing the fundamental period of a value equal to 1 all the 10 samples, and inversion of the signal with a probability of 0.2.
8. Synthesis of the replacement samples by introducing the excitation signal into the LPC filter calculated at 3.
9. Control of the energy level of the synthesis signal. The energy progressively tends toward a level set in advance from the first synthesized replacement frame. This level can be defined, for example, as the energy of the weakest output frame found during the last 5 seconds prior to erasure. Two gain adaptation laws are defined that are chosen based on the DiminFlag flag calculated at 4. The rate of decrease in energy also depends on the fundamental period. There is a third more radical adaptation law that is used when it is detected that the principle of the generated signal does not correspond well to the original signal, as explained later (see point 11).
ES 2 298 261 T3
10. TDAC transformation on the synthesized signal in 8, as explained at the beginning of this chapter. The obtained TDAC coefficients replace the lost TDAC coefficients. Next, by doing the inverse TDAC transformation, the output frame is obtained. These operations have three objectives:
In the case of the first lost window, in this way the information of the correctly received previous window is exploited, which contains half of the data necessary to reconstruct the first disturbed frame (figure 6).
The decoder memory is updated to decode the next frame (encoder and decoder synchronization, see paragraph 5.1.4).
The continuous transition (without breakdown) of the output signal is automatically ensured when the first correctly received binary frame arrives after an erased period that has been reconstructed according to the techniques presented above (see paragraph 5.1.3).
eleven. The addition-coating technique allows verifying whether the synthesized sound signal corresponds well to the source signal or not because for the first half of the first lost frame the weight of the memory of the last correctly received window is more important (figure 6) . Then by taking the correlation between the first half of the first synthesized frame and the first half of the frame obtained after the reverse TDAC operations, the similarity between the lost frame and the replacement frame can be estimated. A weak correlation (<0.65) indicates that the original signal is quite different from the one obtained by the replacement method, it is better to decrease the energy of the latter quickly towards the minimum level.
5.2.2.2.2 Frames lost according to the first frame of an erased area
In the previous paragraph, points 1-6 concerning the analysis of the decoded signal that precede the first erased frame and that allow the construction of a synthesis model (LPC and eventually LTP) of this signal. For the subsequent erased frames, the analysis is not redone, the replacement of the lost signal is based on the parameters (coefficients LPC, pitch, MaxCorr, ResMem) calculated during the first erased frame. Only the operations corresponding to the synthesis of the signal and the synchronization of the decoder are then carried out, with the following modifications in relation to the first erased frame:
In the synthesis part (points 7 and 8), only 320 new samples are generated, because the TDAC transformation window covers the last 320 samples generated during the previous erased frame and these new 320 samples.
In the case where the erasure period is relatively long, it is important to evolve the synthesis parameters towards the parameters of a white noise or towards those with background noise (see point 5 in paragraph 3.2.2.2). Since the system present in this example does not include VAD / CNG, for example, you have the possibility of making one or more of the following modifications:
Progressive interpolation of the LPC filter with a flat filter to make the synthesized signal less colored.
Progressive increase in the pitch value.
In sound mode, it oscillates in non-sound mode after a certain time (for example when the minimum energy is reached).
5.3 Specific treatment for musical signals
If the system comprises a module that allows word / music discrimination, then, after selecting a music synthesis mode, a specific treatment for musical signals can be implemented. In Figure 7, the music synthesis module has been referenced by 15, the speech synthesis module by 16 and the word / music switch by 17.
Such a treatment implements, for example, for the music synthesis model the following steps, illustrated in figure 8:
1. Estimation of the current spectral envelope
This spectral envelope is calculated in the form of an LPC filter [RABINER] [KLEIJN]. The analysis is carried out by the classical methods ([KLEIJN]). After the windowing of the stored samples in valid period, an LPC analysis is carried out to calculate an LPC filter A (z) (step 19). A high order (> 100) is used for this analysis in order to obtain good performances on the musical signals.
ES 2 298 261 T3
2. Synthesis of samples / altars
The synthesis of the replacement samples is carried out by introducing an excitation signal into the LPC synthesis filter (1 / A (z)) calculated in step 19. This excitation signal - calculated in step 20 - is a white noise whose amplitude is chosen to obtain a signal that has the same energy as the last N samples memorized in valid period. In figure 8, the filtration stage is referenced by 21.
Example of residual signal amplitude control
If the excitation is presented as a uniform white noise multiplied by a gain, this gain G can be calculated as follows:
LPC filter gain estimation
The Durbin algorithm gives the energy of the residual signal. Knowing also the energy of the signal to be modeled, the gain G is estimated<sub>LPC</sub> of the LPC filter as the ratio of these two energies.
Calculation of target energy
The target energy is estimated equal to the energy of the last N memorized samples in valid period (N is typically <the length of the signal used for the LPC analysis).
The energy of the synthesized signal is the product of the energy of the white noise by G<sup>2</sup> and G<sub>LPC</sub>. G is chosen so that this energy equals the target energy.
3. Synthesis signal energy control
As for speech signals, except that the speed of reduction of the energy of the synthesis signal is much slower, and that it does not depend on the fundamental period (nonexistent):
The energy of the synthesis signal is controlled with the help of a sample-by-sample calculated and adapted gain. In the case where the erase period is relatively long, it is necessary to progressively lower the energy of the synthesis signal. The gain adaptation law can be calculated based on different parameters such as the values of the energies memorized before erasing, and local seasonality of the signal at the time of cutting.
6. Evolution of the synthesis procedure over time
As for the word signals:
In the case of relatively long erasure periods, it is also possible to evolve the synthesis parameters. If the system is coupled to a device for detecting vocal activity or musical signals with estimation of noise parameters (such as [REC-G.723.1A], [SALAMI-2], [BENYASSINE]), it will be particularly interesting make the generation parameters of the signal to be reconstructed tend towards those of the estimated noise: in particular at the level of the spectral envelope (interpolation of the LPC filter with that of the estimated noise, the interpolation coefficients evolving over time until obtaining the noise filter) and of energy (level that progressively evolves towards that of noise , for example by windowing).
6. General remark
As will have been understood, the technique that has just been described has the advantage of being usable with any type of encoder; In particular, it makes it possible to remedy the problems of bit packets lost by temporary encoders or by transformation, on speech and music signals with good performances: in fact, in the present technique, the only signals memorized during the periods where the Transmitted data are valid, they are the samples output from the decoder, information that is available whatever the coding structure used.
7. Bibliographic references
[AT&T] AT&T (DA Kapilow, RV Cox) “A high quality low-complexity algorithm for frame erasure concealment (FEC) with G.711”. Delayed Contribution D.249 (WP 3/16), ITU, May 1999.
[ATAL] BS Atal and MR Schroeder. "Predictive coding of speech signal and subjectives error criteria". IEEE Trans. on Acoustics, Speech and Signal Processing, 27: 247-254, June 1979.
[BENYASSINE] A. Benyassine, E. Shlomot and HY Su. "ITU-T recommendation G.729 Annex B: A silence compression scheme for use with G.729 optimized for V.70 digital simultaneous voice and data applications." IEEE Communication Magazine, September 97, PP. 56-63.
ES 2 298 261 T3
[BRANDENBURG] KH Brandenburg and M. Bossi. "Overview of MPEG audio: current and future standards for low-bit-rate audio coding." Journal of Audio Eng. Soc., Vol. 45-1 / 2, January / February 1997, PP.4-21.
[CHEN] JH Chen, RV Cox, YC Lin, N. Jayant, and MJ Melchner. “A low-delay CELP coder for the CCITT 16 kb / s speech coding standard”. IEEE Journal on Selected Areas on Communications, Vol. 10-5, June 1992, PP.830849.
[CHEN-2] JH Chen, CR Watkins. "Linear prediction coefficient generation during frame erasure or packet loss". Patent US5574825, EP0673018.
[CHEN-3] JH Chen, CR Watkins. "Linear prediction coefficient generation during frame erasure or packet loss". Patent 884010.
[CHEN-4] JH Chen, CR Watkins. "Frame erasure or packet loss compensation method". Patent US5550543, EP0707308.
[CHEN-5] JH Chen. "Excitation signal synthesis during frame erasure or packet loss". Patent US5615298, EP0673017.
[CHEN-6] JH Chen. "Computational complexity reduction during frame erasure of packet loss". Patent US5717822.
[CHEN-7] JH Chen. "Computational complexity reduction during frame erasure or packet loss." Patent US940212435, EP0673015.
[COX] RV Cox. "Three new speech coders from the ITU cover a range of applications". IEEE Communication Magazine, September 97, PP.40-47.
[COX-2] RV Cox. "An improved frame erasure concealment method for ITU-T Rec. G728". Delayed contribution D.107 (WP 3/16), ITU-T, January 1998.
[COMBESCURE] P. Combescure, J. Schnitzler, K. Ficher, R. Kirchherr, C. Lamblin, A. Le Guyader, D. Massaloux, C. Quinquis, J. Stegmann, P. Vary. "At 16.24.32 kbit / s Wideband Speech Codec Based on ATCELP". Proc. ofICASSP conference, 1998.
[DAUMER] WR Daumer, P. Mermelstein, X. Maítre and I. Tokizawa. "Overview of the ADPCM coding algorithm". Proc. of GLOBECOM 1984, PP.23.1.1-23.1.4.
[ERDOL] N. Erdol, C. Castelluccia, A. Zilouchian. "Recovery of Missing Speech Packets Using the Short-Time Energy and Zero-Crossing Measurements" IEEE Trans. on Speech and Audio Processing, Vol. 1-3, July 1993, PP. 295303.
[FINGSCHEIDT] T. Fingscheidt, P. Vary, "Robust speech decoding: a universal approach to bit error concealment", Proc. ofICASSP conference, 1997, pp. 1667-1670.
[GOODMAN] DJ Goodman, GB Lockhart, OJ Wasem, WC Wong. "Waveform Substitution Techniques for Recovering Missing Speech Segments in Packet Voice Communications." IEEE Trans. on Acoustics, Speech and Signal Processing, Vol. ASSP-34, December 1986, PP. 1440-1448.
[GSM-FR] Recommendation GSM 06.11. "Substitution and muting of lost frames for full rate speech traffic channels." ETSI / TC SMG, ver.:3.0.1., February 1992.
[HARDWICK] JC Hardwick and JS Lim. "The application of the IMBE speech coder to mobile communications". Proc. ofICASSP conference, 1991, PP.249-252.
[HELLWIG] K. Hellwig, P. Vary, D. Massaloux, JP Petit, C. Galand, and M. Rosso. "Speech codec for the European mobile radio system". GLOBECOM conference, 1989, PP. 1065-1069.
[HONKANEN] T. Honkanen, J. Vainio, P. Kapanen, P. Haavisto, R. Salami, C. Laflamme, and JP Adoul. "GSM enhanced full rate speech codec". Proc. ofICASSP conference, 1997, PP.771-774.
[KROON] P. Kroon, BS Atal. “On the use of pitch predictors with high temporal resolution”. IEEE Trans. on Signal Processing, Vol. 39-3, March. 1991, PP. 733-735.
[KROON-2] P. Kroon. "Linear prediction coefficient generation during frame erasure or packet loss". Patent US5450449, EP0673016.
ES 2 298 261 T3
[MAHIEUX] Y. Mahieux, JP Petit. "High quality audio transform coding at 64 kbit / s". IEEE Trans. on Com., Vol. 42-11, nov. 1994, PP.3010-3019.
[MAHIEUX-2] Y. Mahieux, "Dissimulation erreurs de transmission", Patent 92 06720 filed on June 3, 1992.
[MAITRE] X. Maitre. "7 kHz audio coding within 64 kbit / s". IEEE Journal on Selected Areas on Communications, Vol. 6-2, February 1988, PP. 283-298.
[PARIKH] VN Parikh, JH Chen, G. Aguilar. "Frame Erasure Concealment Using Sinusoidal Analysis-Synthesis and Its Application to MDCT-Based Codecs". Proc. ofICASSP conference, 2000.
[PICTEL] PictureTel Corporation, “Detailed Description of the PTC (PictureTel Transform Coder)”, Contribution ITU-T, SG15 / WP2 / Q6, 8-9 October 1996 Baltimore meeting, TD7.
[RABINER] LR Rabiner, RW Schafer. "Digital processing of speech signals". Bell Laboratoires Inc., 1978.
[REC G.723.1A] ITU-T Annex A to recommendation G.723.1 “Silence compression scheme for dual rate speech coder for multimedia communications transmitting at 5.3 & 6.3 kbit / s”.
[SALAMI] R. Salami, C. Laflamme, JP Adoul, A. Kataoka, S. Hayashi, T. Moriya, C. Lamblin, D. Massaloux, S. Proust, P. Kroon, and Y. Shoham. "Design and description of CS-ACELP: a toll quality 8kb / s speech coder". IEEE Trans. on Speech and Audio Processing, Vol. 6-2, March 1998, PP. 116-130.
[SALAMI-2] R. Salami, C. Laflamme, JP Adoul. "ITU-T G.729 Annex A: Reduced complexity 8 kb / s CSACELP codec for digital simultaneous voice and data". IEEE Communication Magazine, September 97, PP. 56-63.
[TREMAIN] TE Tremain. "The government standard linear predictive coding algorithm: LPC 10". Speech technology, April 1982, PP. 40-49.
[WATKINS] CR Watkins, JH Chen. "Improving 16 kb / s G.728 LD-CELP Speech Coder for Frame Erasure Channels". Proc. ofICASSP conference, 1995, PP. 241-244.
Contents12
3 sheets
Sheet 1 Sheet 2 Sheet 3
20 members in 11 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 0011285 | France | A | |
| 0011285 | France | A | |
| 20000011285 | France | – | |
| 001128501969857 | – | – | – |
| FR20000011285 | – | – | – |
Members20
| Document | Office | Kind | |
|---|---|---|---|
| FR2813722A1 | France | A1 | |
| WO0221515A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU8999101A | Australia | A | |
| FR2813722B1 | France | B1 | |
| EP1316087A1 | European Patent Office (EPO) | A1 | |
| IL154728D0 | Israel | D0 | |
| HK1055346A1 | Hong Kong, China | A1 | |
| US2004010407A1 | United States of America | A1 | |
| JP2004508597A | Japan | A | |
| EP1316087B1 | European Patent Office (EPO) | B1 | |
| AT382932T | Austria | T | |
| ATE382932T1 | Austria | T1 | |
| DE60132217D1 | Germany | D1 | |
| ES2298261T3This record | Spain | T3 | |
| IL154728A | Israel | A | |
| DE60132217T2 | Germany | T2 | |
| US7596489B2 | United States of America | B2 | |
| US2010070271A1 | United States of America | A1 | |
| US8239192B2 | United States of America | B2 | |
| JP5062937B2 | Japan | B2 |
Numbers
- Publication
- 2298261
- Publication, DOCDB
- 2298261
- Publication, EPODOC
- ES2298261T
- Application
- 1969857
- Application, DOCDB
- 01969857
- Application, EPODOC
- ES20010969857T
Titles2
- Spanish
- DISIMULACION DE ERRORES DE TRANSMISION EN UNA SEÑAL DE AUDIO.
- English
- REDUCTION OF TRANSMISSION ERRORS IN AN AUDIO SIGNAL.
Classification
- CPC, 1
- G10L19/005
- IPC, 3
- G10L13 00
- G10L19 005
- H04L1 00