Audio frame loss concealment
Abstract
This record has no abstract on file.
Term
7.3 yearsto projected expiry
Projected expiry 22 January 2034, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
14 claims: 7 independent, 7 dependent
- 1ZASTRZEŻENIA PATENTOWE 1. Metoda ukrywania utraconej klatki otrzymanego sygnału audio, na którą składa się:- ekstrakcja segmentu z uprzednio otrzymanego lub zrekonstruowanego sygnału audio, który wykorzystany jest jako klatka prototypowa do stworzenia klatki zastępującej utraconą klatkę audio;- transformacja ekstrahowanej klatki audio do widoku domeny częstotliwości;- wykonanie analizy sinusoidalnej (81) klatki prototypowej, na którą składa się identyfikacja częstotliwości sinusoidalnych komponentów sygnału audio;- zmienienie wszystkich spektralnych współczynników klatki prototypowej w przedziale Mk wokół sinusoidy k przez przesunięcie fazowe proporcjonalne do sinusoidalnej częstotliwości fk i różnicy w czasie pomiędzy klatką utraconą i prototypową, tym samym ewolucję w czasie sinusoidalnych komponentów klatki prototypowej aż do przedziału czasowego utraconej klatki z zachowaniem natężenia tych współczynników spektralnych;- zmiana fazy współczynnika spektralnego klatki prototypowej nie zawartego w żadnym z przedziałów powiązanych z rejonem wokół zidentyfikowanych sinusoid o wartość losową oraz zachowanie natężenia tego współczynnika spektralnego;i - wykonanie odwrotnej transformacji domeny częstotliwości przesuniętego w fazie spektrum częstotliwości klatki prototypowej i tym samym stworzenie klatki zastępczej (83) dla utraconej klatki sygnału audio.
- 2Metoda według zastrz. 1, gdzie na identyfikację częstotliwości sinusoidalnych komponentów składa się także identyfikacja częstotliwości w pobliżu pików spektrum powiązanych z zastosowaną transformatą domeny częstotliwości.
- 3Metoda według zastrz. 2, gdzie identyfikacja częstotliwości sinusoidalnych komponentów przeprowadzana jest z wyższą rozdzielczością częstotliwości niż ta używana w transformacie domeny częstotliwości.
- 4Metoda według zastrz. 3, gdzie na identyfikację częstotliwości sinusoidalnych komponentów składa się także interpolacja.
- 5Metoda według zastrz. 4, w której zastosowana interpelacja jest typu parabolicznego.
- 6Metoda według dowolnego z zastrz. 1-5 na którą składa się także ekstrakcja klatki prototypowej z dostępnego, uprzednio otrzymanego lub zrekonstruowanego sygnału z użyciem okna czasowego.
- 7Metoda według zastrz. 6, na którą składa się także przybliżenie spektrum okna czasowego takie, że spektrum klatki zastępczej składa się wyłącznie z nienakładających się na siebie części przybliżonego spektrum okna czasowego. Dekoder (1) skonfigurowany tak, by ukrywać utraconą klatkę otrzymanego sygnału audio, składający się z procesora (11) i pamięci (12), którego pamięć zawiera instrukcje wykonywane przez procesor (11), który to dekoder skonfigurowany jest do:- ekstrakcji segmentu z uprzednio otrzymanego lub zrekonstruowanego sygnału audio, który wykorzystany jest jako klatka prototypowa do stworzenia klatki zastępującej utraconą klatkę audio;- transformacji ekstrahowanej klatki audio do widoku domeny częstotliwości;- wykonanie analizy sinusoidalnej klatki prototypowej, na którą składa się identyfikacja częstotliwości sinusoidalnych komponentów sygnału audio;- zmiany wszystkich spektralnych współczynników klatki prototypowej w przedziale Mk wokół sinusoidy k przez przesunięcie fazowe proporcjonalne do sinusoidalnej częstotliwości fk i różnicy w czasie pomiędzy klatką utraconą i prototypową, tym samym ewolucję w czasie sinusoidalnych komponentów klatki prototypowej aż do przedziału czasowego utraconej klatki z zachowaniem natężenia tych współczynników spektralnych;- 12 - EP 2954517 - zmiany fazy współczynnika spektralnego klatki prototypowej nie zawartego w żadnym z przedziałów powiązanych z rejonem wokół zidentyfikowanych sinusoid o wartość losową oraz zachowanie natężenia tego współczynnika spektralnego;i - wykonania odwrotnej transformacji domeny częstotliwości przesuniętego w fazie 5 spektrum częstotliwości klatki prototypowej i tym samym stworzenie klatki zastępczej dla utraconej klatki sygnału audio.
- 89. Dekoder według zastrz. 8, w którym na identyfikację częstotliwości sinusoidalnych komponentów składają się także częstotliwości w pobliżu pików spektrum powiązanych z zastosowaną transformatą domeny częstotliwości.
- 910. Dekoder według zastrz. 8, w którym na identyfikację częstotliwości sinusoidalnych komponentów sygnału audio składa się także interpolacja paraboliczna.
- 1011. Dekoder według dowolnego z zastrz. 8 - 10 skonfigurowany tak by ekstrahować klatkę prototypową z dostępnego, uprzednio otrzymanego lub zrekonstruowanego sygnału z użyciem okna czasowego.
- 1112. Dekoder według zastrz. 11 skonfigurowany tak, by przybliżać spektrum okna czasowego tak, by spektrum klatki zastępczej składało się wyłącznie z nienakładających się na siebie części przybliżonego spektrum okna czasowego.
- 1213. Odbiornik zawierający dekoder według dowolnego z zastrz. 8-12.
- 1314. Program komputerowy (91) zawierający instrukcje, które uruchomione na procesorze sprawiają, że procesor wykonuje metodę według dowolnego z zastrz. 1-7.
- 1415. Produkt w postaci programu komputerowego (9), na który składa się odczytywalny przez komputer nośnik zawierający program komputerowy (91) zastrz. 14. KANCELARIA PRAWNO-°ATENTOWA BELLEPAT" Izabela Szych ulaka-Howranek ul Słowackiego 44, 37-700 Prz^nwśl tek (016) 7J2-37-77 fax:(016) .75-02-87 tel kom. (0608) 503-081 e-mail fc3ilepat@op.pl NIP;795-207-16-72 REGON: 1803505/6 Pełnomocnik: Pełnomocnik: KANCELARIA PRAWNO "ATENTOWA BELLEPAT" Izabela Szychulskc-Hawranek ul Słowackiego 44, 37-700 Pfwnwśl tel (016) 7Λ-37-77 fax: (016) 675-72-87 tel kom, (0608) 503-081 e-mati fcellepat@op.pl NIP: 795-207-16-72 REGON: 180350516 - 2 EP 2954517 Pełnomocnik: KANCELARIA PRAWNO PATENTOWA BELLEPAT" Izabela Szychulska-Hawranek ul Słowackiego 44, 37-700 Pfwn« śl tel (016) 7Λ-37-77 fax: (016) 675-72-87 tel kom, (0608) 503-081 e-mati fceliepat@op.pl NIP: 795-207-16-72 REGON: 180350516 - 3 EP 2954517 Pełnomocnik: KANCELARIA PRAWNO PATENTOWA BELLEPAT" Izabela Szychulska-Hawranek ul Słowackiego 44, 37-700 Pfwn« śl tel (016) 7Λ-37-77 fax: (016) 675-72-87 tel kom, (0608) 503-081 e-mati fcellepat@op.pl NIP: 795-207-16-72 REGON: 1803505(6 - 4 EP 2954517 Pełnomocnik: KANCELARIA PRAWNO PATENTOWA BELLEPAT" Izabela Szych,liska-Hawranek ul Słowackiego 44, 37-700 Przesil tel (016) 7Λ-37-77 fax: (016) .75-72-87 tel kom, (0608) 503-081 e-maii fcellepat@op.pl NIP: 795-207-16-72 REGON: 1803505(6 - 5 EP 2954517 Pełnomocnik: KANCELARIA PRAWNO PATENTOWA BELLEPAT" Izabela Szychulska-Hawranek ul Słowackiego 44 37-700 Pfwnwśl tel (016) 7Λ-37-77 fax: (016) .176-72-87 tel kom, (0608) 503-081 e-mati fcellepat@op.pl NIP: 795-207-16-72 REGON: 1803505/6 - 6 EP 2954517 Zakodowany sygnał audio Pełnomocnik: KANCELARIA PRAWNO PATENTOWA BELLEPAT" Izabela Szychulska-Hauiranek ul Słowackiego 44, 37-700 Pfwnwśl tel (016) 7j2-37-77 fax: (016) 675-72-87 tel kom, (0608) 503-081 e-mati fcellepat@op.pl NIP: 795-207-16-72 REGON: 180350516 - 7 EP 2954517 Zakodowany sygnał audio Pełnomocnik: KANCELARIA PRAWNO PATENTOWA BELLEPAT" Izabela Szychulska-Hauiranek ul Słowackiego 44, 37-700 Pfwnwśl tel (016) 7J2-37-77 fax: (016) 675-72-87 tel kom. (0608) 503-081 e-mati fcellepat@op.pl NIP: 795-207-16-72 REGON: 180350516 - 8 EP 2954517 Pełnomocnik: KANCELARIA PRAWNO PATENTOWA BELLEPAT" Izabela Szychulska-Hawranek ul Słowackiego 44, 37-700 Pfwnwśl tel (016) 7j2-37-77 fax: (016) 675-72-87 tel kom, (0608) 503-081 e-mati fcellepat@op.pl NIP: 795-207-16-72 REGON: 1803505(6
Independent claims14
143 paragraphs in 9 sections, as filed
European).
- EP 2954517
HIDE THE LOSS OF AUDIO SIGNAL CAGES
Technical area
In general, the invention relates to a method for concealing a lost frame of the received audio signal. The invention also relates to a decoder configured to hide a lost frame of the received encoded audio signal. The invention also relates to a receiver comprising a decoder and a computer program and a product in the form of a computer program.
Background
Conventional sound communication system transmits speech and audio signals in frames, this means that the sending side first divides the sound signal into short segments, i.e. frames of an audio signal, eg 20-40 ms long, which are then coded and sent as a logical unit, e.g. transmission package. The decoder on the receiver side decodes each unit and reconstructs the corresponding frames of the audio signal, which are finally reproduced as a continuous sequence of reconstructed samples (samples) of the sound signal.
Before encoding with the help of analog-to-digital (A / D conversion), you can convert an analog speech signal or sound coming from the microphone into a sequence of digital sound signal samples. On the receiving side there is a reverse final stage of the D / A conversion in which the sequence of reconstructed digital sound samples is converted into a continuous analogue signal for reproduction in the loudspeaker.
However, a typical speech and audio signal transmission system may suffer from transmission errors that may lead to a situation where one or more of the transmitted frames are not available for reconstruction at the receiver side. In this case, the decoder must create a substitute signal for each inaccessible frame. This can be done with the help of the so-called a concealment system for the loss of the frame contained in the decoder on the receiver side. The purpose of concealing the loss of a frame is to make the loss of the frame as audible as possible, thus reducing the impact of frame loss on the quality of the reproduced signal.
Conventional methods for concealing lost frames may be dependent on the structure or architecture of the codec, e.g. by repeating previously obtained codec parameters. Such repetition techniques are of course dependent on the specific parameters of the codec used and may not be easily applicable to other codecs with a different structure. Current methods of hiding frame loss can, for example, freeze or extrapolate the parameters of a previously obtained frame to create a lost replacement frame. Standardized codecs with linear predictors AMR and AMR-WB are parametric codecs that freeze previously obtained parameters or use their extrapolation type for decoding. In fact, the principle of operation is to have a coding / decoding model and its use for frozen or extrapolated parameters.
Many audio codecs use a domain frequency coding technique, which consists of a spectrum parameter coding model after frequency domain transformation. The decoder recreates the signal spectrum from the received parameters and converts the spectrum back to the time signal. Typically, the time signal is played back frame by frame, and the cages are combined with the help of application techniques and possible further processing to obtain the final reconstructed signal. The concealment of frame loss associated with this system uses the same, or at least similar, decoding model for frames lost, in which the frequency domain parameters from the previously obtained frame are frozen or respectively extrapolated and then used in frequency-time domain conversion.
However, conventional methods for concealing lost frames may suffer from quality, for example because the technique of freezing and extrapolating and re-using the same decoder model for lost frames is not always guaranteed
- 2 - EP 2954517 smooth and faithful evolution of the signal from the previously decoded cage to the lost cage. This can lead to an audible signal discontinuity and associated loss of quality. Thus, hiding frames lost with a reduced impact on quality is desirable and necessary.
Summary
The object of implementing the present invention is to respond to at least some of the problems outlined above, and the subject matter, like the others, is achieved by a method and systems according to the attached independent patent claims and the embodiments related to the dependent claims.
According to one aspect, embodiments provide methods for concealing audio frame loss, a method which consists of a sinusoidal analysis of a portion of a previously obtained or reconstructed audio signal, which consists of identifying the frequency of the sinusoidal components of the audio signal. Later, the sinusoidal model is applied to the segment of the previously obtained or reconstructed audio signal, and this segment is used as a prototype of the frame to create a cage that will replace the lost frame. Creating a replacement frame is related to the evolution during the sinusoidal components of the prototype frame up to the time frame of the lost frame in relation to the identified frequencies.
According to a second aspect, the embodiment comprises a decoder configured to hide the lost audio frame of the received signal, the decoder and processor and memory, the memory includes instructions performed by the processor, where the decoder is configured to perform a sinusoidal analysis of a portion of the previously received or reconstructed audio signal, and the sinusoidal analysis consists of the identification of the frequency of the sinusoidal components of the audio signal. The decoder is configured to apply a sinusoidal model to a previously obtained or reconstructed audio signal, and use the segment as a prototype frame to create a replacement frame for a lost audio frame and to create a replacement frame by evolving in time the sinusoidal components of the prototype frame up to the temporal aspect of the lost frame sound
According to a third aspect, the embodiment comprises a decoder configured to hide the lost audio frame of the received signal, the decoder and an input circuit configured to receive the decoded audio signal and a frame for concealing the lost frames. The system of concealing lost frames consists of means for performing a sinusoidal analysis of a part of a previously obtained or reconstructed audio signal, and the sinusoidal analysis comprises identifying the frequency of the sinusoidal components of the audio signal. The frame-hiding arrangement also includes means for applying a sinusoidal model to a segment of a previously obtained or reconstructed audio signal,
The decoder can be used in a device, such as a mobile phone.
According to a fourth aspect, a possible embodiment is a receiver comprising a decoder in accordance with the one described in the second and third aspects above.
According to a fifth aspect, feasible embodiments are also a computer program designed to conceal a lost audio frame that consists of instructions made by the processor, thereby hiding the processor of a lost audio frame in accordance with the first aspect described above.
According to the sixth aspect, the possible implementation is also a product in the form of a computer program in the form of a computer-readable medium containing a computer program described above in the fifth aspect.
An advantage of the embodiments described herein is to provide a method of concealing lost frames, which reduces the audible consequences of losing frames in the transmission of an audio signal, e.g. coded speech. The overall advantage is to ensure a smooth and faithful evolution of the reconstructed signal for the lost frame, thanks to which the audible consequences of frame loss are significantly reduced compared to conventional techniques.
Further features and advantages of the applications that are the subject of the application will become clear after reading the following description and attached drawings.
A brief description of the drawings
The embodiments will be described in more detail with reference to the attached drawings, of which:
Figure 1 shows a typical time window;
Figure 2 shows a detailed time window;
Figure 3 shows an exemplary time window intensity spectrum;
Figure 4 illustrates a linear spectrum of an example sinusoidal signal with frequency fk;
The figure shows the spectrum of the sinusoidal signal with the frequency fk; The figure shows bars corresponding to the intensity of DFT points based on frame analysis;
Figure 7 shows a parabola passing through the nodes of the DFT network;
Figure 8 is a graph of the course of the method according to the embodiments;
Figures 9 and 10 illustrate a decoder in accordance with embodiments, a
Figure 11 shows a computer program and product in the form of a computer program in accordance with embodiments.
Detailed description
The embodiments of the invention are described in more detail below. The details are disclosed for explanation and not limitation, in particular scenarios and techniques presented for in-depth understanding.
Moreover, it is evident that the exemplary method and devices described below may be implemented, at least partially, for use in software operating in connection with a programmable microprocessor or a general purpose computer and / or using dedicated integrated circuits (ASICs). Moreover, embodiments can also, at least partially, be implemented as a product in the form of a computer program in which one or more programs are stored in memory that can perform the functions described herein.
The implementation concept described below assumes hiding the lost sound frame by:
• Performing a sinusoidal analysis of at least part of the previously obtained or reconstructed audio signal, to which the analysis will consist of identifying the frequency of sinusoidal components of the audio signal;
• applying a sinusoidal model to a segment of a previously obtained or reconstructed audio signal, which will serve as a prototype frame to create a replacement frame for a lost frame, and • create a replacement frame by evolving the sinusoidal components of the prototype frame in time to equalize with the time of the lost sound frame, according to relevant frequencies identified.
- EP 2954517
Sinusoidal analysis
Hiding a lost frame in accordance with the performance assumes a sinusoidal analysis of a part of a previously obtained or reconstructed audio signal. The purpose of a sinusoidal analysis is to find the frequency of the main sinusoidal components, i.e. the sinusoid of this signal. The same underlying assumption is that the audio signal is generated according to a sinusoidal model and that it consists of a finite number of individual sinusoids, i.e. that it is a multi-sinusoidal signal of the following type:
<sup>AND</sup> fs (.nj = ^ a<sub>k</sub>· Cos (2tt · η + φ<sub>Ιι</sub>) J s (6.1) k = \
In this equation, K is the sinusoidal number assumed for a given signal from which this signal is composed. For each of the sine waves with the number k = 1 ... K, ak is the amplitude, fk is the frequency and φk is the phase. The sampling frequency is expressed by fs and the time index of the stepping intervals of signal sampling s (n) by n.
It is important to find such exact sinusoidal frequencies as possible. While a perfectly sinusoidal signal would have a spectrum coincident with fk frequencies, finding their real values would actually require infinite time to measure. Thus, finding these frequencies is in practice a difficult task, as they can only be estimated on the basis of a short measurement period corresponding to the signal segment used for the sinusoidal analysis according to the embodiment described herein; this segment of the signal will be referred to below as the analyzed one. Another difficulty is the practical possibility of signal variation over time, which means that the parameters of the above equation change with the passage of time. Thus, on the one hand, it is desirable to use a long frame analyzed to make the measurement more accurate; on the other hand, a short measurement period would be required to better cope with possible signal variation. A good compromise is the length of the analysis frame in the range of, for example, 20-40 ms.
According to the preferred embodiment of the frequencies, sine waves fk are identified by analyzing the frequency domain of the analyzed cage. To this end, the cage is transformed into a frequency domain, e.g. with the help of DFT (discrete Fourier transform) or DCT (discrete cosine transformation) or similar frequency domain transformation. When using the DFT analyzed in the cage, the spectrum is calculated from:
with,<sub>2</sub>"
X (m) = DFT (w (w) ·% («)) = e <sup>J1</sup> · In («) · x (n). (6.2) n = 0.
In this equation, w (n) is the time window, by means of which the analysis cage is extracted and weighed from the length L.
Figure 1 shows a typical time window, i.e. a rectangular window equal to 1 for n G [0 ... L-1] and 0 in other cases. It is assumed that the time indices of the previously received sound signal are such that the prototype frame finds a reference in the time indices n = 0 ... L-1. Other time windows that may be useful for spectrum analysis purposes are e.g. Hamming, Hanning, Kaiser or Blackman windows.
Figure 2 shows a more useful time window, which is a combination of a Hamming window and a rectangular window. The window shown in Figure 2 has an ascending edge, such as the left half of the Hamming window for the L1 section and the falling edge shape, as in the right part of the Hamming window for the L1 section and between the falling and rising edges the window has a value 1 on the section from L to L1 .
Peaks of the intensity of the spectrum included in the analysis box window | X (m) | they are approximately the required sinusoidal frequencies fk. The accuracy of this approximation is however limited by the DFT frequency resolution. With DFT with L block length accuracy
Λ_.
limited to 2Z This level of accuracy may be too low for the method
- in accordance with the embodiments described herein, an improvement in accuracy can be obtained based on the solution of the following factor:
The spectrum of the analyzed frame in the window is a convolution of the time window spectrum with a linear spectrum of a sinusoidal signal model 5 (Ω), then sampled at the DFT node points:
X (m) = Jć> (Ω - m · f) · (? (Ω) * 5 (Ω)) · άΩ.
(6.3)
Using the spectral titer of a sinusoidal signal model, this can be written as
<img file="PL2954517T3_D0001.tif" />
Thus, the sampled spectrum can be calculated from
<img file="PL2954517T3_D0002.tif" />
(6.5) for m = 0 ... L-1.
Based on this observed peaks of the intensity of the spectrum of the analyzed frame come from the sinusoidal window sine signal K of a sinusoid, the real sine waves being near the peaks. Thus, the identification of the sinusoidal frequency components may also consist in identifying frequencies in the vicinity of the spectral peaks associated with the frequency conversion domain used.
If m<sub>:</sub> we will assume the DFT index (node) for the observed peak k<sup>th</sup>, the frequency associated with it is
<img file="PL2954517T3_D0003.tif" />
which can be taken as approximation of the actual sinusoidal frequency fk. It should be assumed that the actual sinusoidal frequency fk
<img file="PL2954517T3_D0004.tif" />
is in the range of L For clarity, it is emphasized that the combination of the time window spectrum with the linear spectrum of a sinusoidal signal model can be understood as the overlapping of frequency-shifted version of the time window spectrum, with the frequency of the shift being sinusoidal frequencies. This overlap is then sampled at the nodes of the DFT network. The combination of the time window spectrum with the linear spectrum of the sinusoidal signal model is shown in Figures 3 to 7, of which Fig. 3 shows an example of the time window spectral spectrum, Fig. 4 the intensity spectrum (linear) of an example sinusoidal signal with one sinusoidal frequency fk. Figure 5 shows the intensity spectrum of the sinusoidal signal reproduced in the window and overlapping the frequency-shifted sinusoidal frequency spectra, and the bars in figure 6 correspond to the intensity at the DFT nodes captured in the sinusoidal time window obtained by calculating the DFT of the analyzed frame. It is worth noting that all spectra are periodic, with a normalized frequency parameter Ω where Ω = 2π which corresponds to the sampling frequency fs.
Based on the above discussion and illustration of Figure 6, we note that a better approximation of true sinusoidal frequencies can be obtained by increasing the search resolution so that it is higher than the frequency resolution of the frequency domain transform algorithm used.
Therefore, it is best to identify the frequency of sinusoidal components with a resolution higher than the frequency resolution of the frequency domain transform used, and interpolation may also be included in the identification.
- EP 2954517
An example of a preferred method for better approximation of the fk sinusoid frequency is the use of parabolic interpolation. One of the methods is to place the parabol at the node points of the DFT spectrum surrounding the peaks and calculate the parabol corresponding to the maxima, an example of choosing the amount of parabol is 2. In a more approximate way, the following procedure can be used:
1) Identification of DFT peaks included in the analysis frame window. The search for peaks will return K peaks and the corresponding DFT indices. Peak searching can be performed on the DFT spectrum or on the logarithmic DTF intensity spectrum.
2) For each peak k (at k = 1 ... K) and the corresponding DFT m index<sub>:</sub>, fits in three points {P1; P2; P3} = {(mx-1 .log (| X (mk-1) |); (mk, log (| X (mk) |); (mk + 1, log (| X (mk + 1) |) } the parabol result is the coefficients parabol bk (0), bk (1), bk (2) for the parabola described by p<sub>k</sub>=.
/ = 0
Figure 7 illustrates a parabola in DFT P1, P2 and P3 nodes.
3) For each K parabol, the frequency index interpolated mik is calculated corresponding to the value of q for which the parabola has its maximum, and fk = mik · fs / L is used as the approximation of the sinusoidal frequency fk.
Application of a sinusoidal model
The use of a sinusoidal model for stealth loss operation, in accordance with the embodiments, can be described as follows:
In the case when a given segment of the encoded signal can not be played by the decoder due to the unavailability of coded information, i.e. the frame has been lost, the available part of the signal preceding this segment can be used as a prototype frame. If y (n) for n = 0 ... N-1 is an inaccessible segment for which a replacement frame must be created from (n) iy (n) for n <0 then the previously available decoded signal is a prototype frame of L length and the initial index n-1 is cut out by the time window in (n) and converted into a frequency domain, e.g. using DFT:
, Ł- 1.
- Yv (s ·
S = 8.
The time window may be one of the windows described above for sinusoidal analysis. It is preferable, due to the simplification of calculations, that the frame transformed in the frequency domain be identical to that used in the sinusoidal analysis.
In the next step, the assumption of a sinusoidal model is imposed. According to the DFT, a prototype frame can be expressed in the following way:
<sup>K</sup> (mfmf
1> ΗΙλ '("W? + ^)) - e - * + r (2n (- ¡).<sub>e</sub>*).
* = and \ P Js <sup>L</sup> J<sub>s</sub> s
This expression was also used in the analytical part and described in detail above.
Then it turns out that the spectrum of the time frame used has significant values only in the near-zero frequency range. As shown in Figure 3, the intensity of the time window spectrum is high for near-zero frequencies and low in other areas (in the normalized frequency range from -π to π, corresponding to half of the sampling rate). Thus, it is assumed that the window spectrum W (m) is non-zero for the interval M = [-mmin, mmax], for m min and mmax being small positive numbers. In particular, the approximation of the time window of the spectrum is used so that for each k the influence of the shifted window spectrum in the above equation is always the maximum
- impact from a single adder, i.e. from one shifted window spectrum. This means that the above formula boils down to the following approximation:
<img file="PL2954517T3_D0005.tif" />
Mk = [round 'ί) for non-negative mEMk and each k. Mk is a range in integers<sup>rounding</sup>. 0 + rflmax, k]] r where m min, ki mmax, k satisfy the previously explained restrictions on the non-overlapping of compartments. A good choice is to set mmin, ki mmax, k as small integers, eg δ = 3. However, if the DFT indicators associated with two adjacent frequencies sine waves fk and fk + 1 are lower than 2δ, then δ is set to
Min.
rounded.
<img file="PL2954517T3_D0006.tif" />
rounded.
<img file="PL2954517T3_D0007.tif" />
thus guaranteeing that the compartments will not overlap. The minimum (·) function means the nearest integer lower than or equal to the function argument.
The next stage in accordance with the implementation of the use of a sinusoidal model in accordance with the above expression and evolution of its K sinusoidal time. The assumption that the time indexes of the removed segment compared to the prototype frame indicators differ by n-1 samples means that the sine wave phases should be moved forward by
Thus, the DFT spectrum of the transformed sinusoidal model can be calculated from:
WITH
<img file="PL2954517T3_D0008.tif" />
L f<sub>s</sub>
<img file="PL2954517T3_D0009.tif" />
Using the approximation again, according to which the time window spectra do not overlap, we get:
for non-negative mEMk and each k.
Comparing the DFT of the Y-1 prototype frame (m) with the DFT of the transformed sinusoidal model Y0 (m) obtained with approximation, we see that the intensity spectrum remains unchanged while the phase is shifted by fs for each mEMk. The same replacement frame can be calculated by the following formula:
<img file="PL2954517T3_D0010.tif" />
for non-negative mEMk and each k.
The particular implementation deals with the randomisation of phases of DFT indices that do not belong to any Mk interval. As described above, the intervals Mk, k = 1 ... K must be set so as not to overlap, which is obtained by using parameter δ controlling the size of the compartments. It may happen that δ will be in small relation to the frequency distance of two neighboring sine waves. Thus, there is a gap between the two compartments. As a result, the related DFT m indices do not have an assigned phase shift according to the above formula Z (m) = Y (m) · e<sup>JBX</sup>. According to this implementation, a convenient choice for such indicators is the randomization of indicators according to Z (m) = Y (m) · e<sup>f2x rand (·)</sup>where the rand (·) function returns a random number.
- EP 2954517
Based on the above figure 8, a graph illustrates an example of hiding the loss of a sound cage according to the embodiments:
In step 81, a sinusoidal analysis of a part of the previously received or decoded signal is performed, the sinusoidal analysis consists of the identification of the frequency of the sinusoidal elements, i.e. the sine of the audio signal. Then, in step 82, a sinusoidal model is applied to the segment of the previously received or decoded signal, and this segment is used as a prototype of the frame to create a replacement frame for the lost frame, and in step 83 a replacement frame is created for the frame lost by the evolution in time of sinusoidal components, i.e. the sine of the prototype frame up to the time frame of the lost audio frame in response to the associated identified frequencies.
According to a further embodiment, the audio signal consists of a limited number of single sinusoidal components, and the sinusoidal analysis is performed in the frequency domain. Furthermore, the identification of the frequency of sinusoidal components may involve the identification of frequencies adjacent to the spectrum peaks associated with the frequency domain transform used.
According to an exemplary embodiment, the method involves extracting a prototype frame from an available, previously received or reconstructed signal using a time window in which the acquired prototype frame is transformed into a frequency domain.
The further implementation assumes an approximation of the time window spectrum that the replacement frame spectrum consists only of non-overlapping parts of the approximate spectrum of the time window.
According to a further exemplary embodiment, the method relies on the evolution of sinusoidal components of the frequency spectrum of the prototype frame by shifting the phases of sinusoidal components forward relative to the frequency of each of the sinusoidal components and the time difference between the lost audio frame and the prototype frame and the spectral shift of the prototype frame in the Mk interval near the sinusoid k by the phase shift proportional to the sinusoidal frequency fk and the time difference between the lost frame and the prototype frame.
The next execution consists of a phase change of the spectral frame factor of the prototype not belonging to the recognized sine wave by a random shift or change of the phase of the spectrum of the prototype frame not belonging to any of the intervals associated with the proximity of the identified sinewaves by a random value.
The implementation also consists of an increase in the frequency transform domain for the frequency spectrum of the prototype frame.
In more detail, the method of hiding the cage loss in accordance with the performances may include the following stages:
1) Analysis of the available segment of the previously synthesized signal to obtain sinusoidal sine wave components of the sinusoidal model.
2) Extraction of the prototype y-1 cage from the previously synthesized signal and DFT calculation for this cage.
3) Calculations of the phase shift θk for each sine wave response to the sinusoidal frequency fk and progress in time between the prototype and replacement frames.
4) For each sine wave k, shift the phase of the DFT prototype frame by θk separately for DFT indicators adjacent to the sinusoidal frequency fk.
5) Calculation of the inverse DFT of the 4) spectrum thus obtained.
The embodiments described above can also be explained with the help of the following assumptions:
- EP 2954517
a) We assume that the signal can be reflected with the help of a finite number of sinusoids.
b) We assume that the replacement cage is adequately represented by these sinusoids transformed over time, compared to some previous solutions.
c) We assume an approximation of the time window spectrum such that the spectrum of the replacement frame can be constructed from non-overlapping parts of the time window of spectra shifted in frequency, and the shifted frequencies are sinusoidal.
Figure 9 is a block diagram illustrating an exemplary decoder 1 configured to hide audio frame loss according to embodiments. The presented decoder consists of one or more processors 11 and the corresponding software and corresponding storage or memory 12. The incoming decoded audio signal is received at the input (WE) to which the processor 11 and memory are connected. The decoded and reconstructed audio signal obtained from the software is provided on the output (WY). An exemplary decoder is configured to hide a lost frame of the received audio signal and consists of a processor 11 and a memory 12, in which memory instructions are provided by the processor 11 and the decoder 1 is configured to:
• performing a sinusoidal analysis of a portion of a previously obtained or reconstructed audio signal, which consists of identifying the sinusoidal frequency components of the audio signal;
• applying a sinusoidal model to a segment of a previously received or reconstructed audio signal that is used as a prototype frame to create a backup frame for a lost audio frame, and • create a replacement frame for a lost audio frame by evolving the sinusoidal components of the prototype frame in time to the time frame lost audio frames in response to the corresponding identified frequencies.
According to a further embodiment of the decoder, the superimposed sinusoidal model assumes that the audio signal consists of a limited number of individual sinusoidal components, and that the sinusoidal frequency components of the audio signal may also be identified by parabolic interpolation.
According to a further embodiment, the decoder is configured to extract a prototype frame from the available and previously received or reconstructed signal and transform it into a frequency domain through the time window.
According to yet another embodiment, the decoder is configured to evolve during the sinusoidal components of the frequency spectrum of the prototype frame by shifting the phases of the sinusoidal components in response to the frequency of each component and the time difference between the lost audio frame and the prototype frame and creating a replacement frame by performing an inverse spectrum frequency transformation frequency.
A decoder in accordance with an alternative embodiment is shown in Figure 10a and consists of an input unit configured to receive the encoded audio signal. The figure shows the concealment of frame loss by the logical hide frame loss 13 system in which the decoder 1 is configured to implement hiding the lost audio frame in accordance with the embodiments described above. The logic of concealing the loss of the frame 13 is shown in figure 10b and consists of suitable means to hide the loss of the cage, i.e. means 14 to perform a sinusoidal analysis of a portion of the previously obtained or reconstructed audio signal, which consists in identifying the frequency of the sinusoidal components of the audio signal, means to apply a sinusoidal model to the segment previously
- a sound signal obtained or reconstructed, which is used as a prototype frame to create a cage replacing the lost frame, and means 16 to create a cage replacing the cage lost by the evolution during the sinusoidal components of the prototype frame up to the time frame of the lost audio frame, regarding the frequencies identified.
The units and means connected in the decoder shown in the figures can be implemented, at least partially, in hardware, and there are a number of circuit variants that can be used and that can be created to obtain the functions of the decoder units. The subject of protected performances includes such variants. A particular example of a hardware implementation of a decoder is the implementation in a hardware digital signal processor (DSP) and integrated circuits technology, including general-purpose and specialized ones.
A computer program compliant with embodiments of the present invention consists of instructions that the processor will make the processor perform the method according to the method described in Figure 8. Figure 11 shows a computer program 9 compliant with embodiments in the form of non-volatile memory, i.e. in the EEPROM memory ( Electrically Erasable Programmable Read-Only Memory), flash or disk. The product in the form of a computer program consists of a computer program storage medium 91, which consists of computer program modules 91a, b, c, d which, when run on decoder 1, causes the decoder processor to perform the steps shown in figure 8.
A decoder in accordance with the embodiments of the present invention may be used, for example, in a receiver of a mobile device, e.g. a mobile phone or a laptop, or in a receiver of a stationary device, e.g. a personal computer.
The advantages of the embodiments described here are providing a method of concealing the loss of frames to overcome the audible consequences of losing frames in the transmission of an audio signal, e.g. coded speech. The overall advantage is to ensure a smooth and faithful evolution of the reconstructed signal of the lost frame, with a very significant reduction in the audible effects of frame loss compared to conventional techniques.
It is understandable that the selection of cooperating units or modules, as well as their naming, serves only exemplary purposes and they can be configured in a number of alternative ways to carry out the activities involved in the process. It should also be noted that the units and modules described here should be treated as logical whole, but they do not have to be separate physical entities. It is believed that the scope of the technology described herein includes all other realizations that may be thought of by those skilled in the art.
LEGAL WAREAW LAW "BELLEPAT"
Izabela Szych ulska-Hauiranek ul. Słowackiego 44, 37-700 Pfezoiwśl tel (016) 7J2-37-77 fax: (016) 675-02-87 mobile phone (0608) 503-081 e-man <a href="mailto:fcellepat@op.pl">fcellepat@op.pl</a> NIP: 795-207-16-72 REGON: 1803505 (6
Proxy:
<img file="PL2954517T3_D0011.tif" />
EP 2954517
Contents9
161 members in 22 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 201361760814 | United States of America | P | |
| 201361760814 | United States of America | P | |
| 201361760814P | – | – | – |
| US201361760814P | – | – | – |
Members161
| Document | Office | Kind | |
|---|---|---|---|
| CA2900354A1 | Canada | A1 | |
| CA2978416A1 | Canada | A1 | |
| WO2014123469A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014123470A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014123471A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2014215734A1 | Australia | A1 | |
| US2015228287A1 | United States of America | A1 | |
| SG11201505231VA | Singapore | A | |
| KR20150108419A | Republic of Korea | A | |
| PH12015501507A1 | Philippines | A1 | |
| PH12015501507B1 | Philippines | B1 | |
| KR20150108937A | Republic of Korea | A | |
| CN104969290A | China | A | |
| CN104995675A | China | A | |
| MX2015009210A | Mexico | A | |
| EP2954516A1 | European Patent Office (EPO) | A1 | |
| EP2954517A1 | European Patent Office (EPO) | A1 | |
| EP2954518A1 | European Patent Office (EPO) | A1 | |
| US2015371641A1 | United States of America | A1 | |
| US2015371642A1 | United States of America | A1 | |
| US9293144B2 | United States of America | B2 | |
| JP2016510432A | Japan | A | |
| JP2016511433A | Japan | A | |
| HK1210315A1 | Hong Kong, China | A1 | |
| KR20160045917A | Republic of Korea | A | |
| US2016155446A1 | United States of America | A1 | |
| NZ709639A | New Zealand | A | |
| KR20160075790A | Republic of Korea | A | |
| EP2954517B1 | European Patent Office (EPO) | B1 | |
| AU2014215734B2 | Australia | B2 | |
| JP5978408B2 | Japan | B2 | |
| EP2954518B1 | European Patent Office (EPO) | B1 | |
| AU2016225836A1 | Australia | A1 | |
| US9478221B2 | United States of America | B2 | |
| EP3096314A1 | European Patent Office (EPO) | A1 | |
| DK2954517T3 | Denmark | T3 | |
| PT2954518T | Portugal | T | |
| MX344550B | Mexico | B | |
| ZA201504881B | South Africa | B | |
| PL2954517T3This record | Poland | T3 | |
| ES2597829T3 | Spain | T3 | |
| EP3125239A1 | European Patent Office (EPO) | A1 | |
| JP6069526B2 | Japan | B2 | |
| ES2603827T3 | Spain | T3 | |
| RU2015137708A | Russian Federation | A | |
| SG10201700846UA | Singapore | A | |
| JP2017097365A | Japan | A | |
| BR112015017222A2 | Brazil | A2 | |
| BR112015018316A2 | Brazil | A2 | |
| US9721574B2 | United States of America | B2 | |
| RU2628144C2 | Russian Federation | C2 | |
| US2017287494A1 | United States of America | A1 | |
| CA2900354C | Canada | C | |
| US9847086B2 | United States of America | B2 | |
| EP3096314B1 | European Patent Office (EPO) | B1 | |
| NZ710308A | New Zealand | A | |
| DK3096314T3 | Denmark | T3 | |
| US2018096691A1 | United States of America | A1 | |
| ES2664968T3 | Spain | T3 | |
| KR101855021B1 | Republic of Korea | B1 | |
| KR20180049145A | Republic of Korea | A | |
| AU2018203449A1 | Australia | A1 | |
| EP3333848A1 | European Patent Office (EPO) | A1 | |
| AU2016225836B2 | Australia | B2 | |
| HUE036322T2 | Hungary | T2 | |
| CN104995675B | China | B | |
| CN104969290B | China | B | |
| CN108564958A | China | A | |
| CN108831490A | China | A | |
| CN108847247A | China | A | |
| CN108899038A | China | A | |
| JP6440674B2 | Japan | B2 | |
| RU2017124644A | Russian Federation | A | |
| JP2019061254A | Japan | A | |
| PH12018500083A1 | Philippines | A1 | |
| PH12018500083B1 | Philippines | B1 | |
| PH12018500600A1 | Philippines | A1 | |
| PH12018500600B1 | Philippines | B1 | |
| CA2978416C | Canada | C | |
| US10332528B2 | United States of America | B2 | |
| US10339939B2 | United States of America | B2 | |
| EP3125239B1 | European Patent Office (EPO) | B1 | |
| MY170368A | Malaysia | A | |
| DK3125239T3 | Denmark | T3 | |
| EP3333848B1 | European Patent Office (EPO) | B1 | |
| US2019267011A1 | United States of America | A1 | |
| US2019272832A1 | United States of America | A1 | |
| PT3125239T | Portugal | T | |
| PT3333848T | Portugal | T | |
| KR102037691B1 | Republic of Korea | B1 | |
| EP3561808A1 | European Patent Office (EPO) | A1 | |
| HK1258094A1 | Hong Kong, China | A1 | |
| EP3576087A1 | European Patent Office (EPO) | A1 | |
| PL3125239T3 | Poland | T3 | |
| AU2018203449B2 | Australia | B2 | |
| HUE045991T2 | Hungary | T2 | |
| US10559314B2 | United States of America | B2 | |
| AU2020200577A1 | Australia | A1 | |
| NZ746512A | New Zealand | A | |
| ES2750783T3 | Spain | T3 |
Numbers
- Publication
- 2954517
- Publication, DOCDB
- 2954517
- Publication, EPODOC
- PL2954517T
- Application
- 147047047
- Application, DOCDB
- 14704704
- Application, EPODOC
- PL20140704704T
Titles2
- English
- AUDIO FRAME LOSS CONCEALMENT
- Polish
- UKRYCIE UTRATY KLATKI SYGNAŁU AUDIO
Classification
- CPC, 3
- G10L19/005
- G10L19/02
- G10L25/69
- IPC, 1
- G10L19 005