Jitter buffer control, audio decoder, method and computer program
Abstract
This record has no abstract on file.
Term
7.7 yearsto projected expiry
Projected expiry 18 June 2034, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
27 claims: 4 independent, 23 dependent
- 1Zastrzeżenia patentowe 1. Sterowanie (100;350;490) buforem rozsynchronizowania w celu sterowania dostarczaniem zdekodowanej treści (312;412) audio na podstawie wejściowej treści audio, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać skalowanie czasowe oparte o ramki lub skalowanie czasowe oparte o próbki w sposób dostosowywany do sygnału, tak że decyzja o tym, czy stosowane jest skalowanie czasowe oparte o ramki, czy skalowanie czasowe oparte o próbki, dostosowana jest do właściwości sygnału audio.
- 2Sterowanie (100;350;490) buforem rozsynchronizowania według zastrz. 1, w którym ramki audio są pomijane lub wstawiane w celu sterowania głębokością bufora (320;430) rozsynchronizowania, gdy stosowane jest skalowanie czasowe oparte o ramki, i w którym wykonywane jest nakładanie i sumowanie (954;1068) części sygnału audio z przesunięciem czasowym, gdy stosowane jest skalowanie czasowe oparte o próbki.
- 3Sterowanie (100;350;490) buforem rozsynchronizowania według zastrz. 1 albo 2, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by przełączać pomiędzy skalowaniem czasowym opartym o ramki, skalowaniem czasowym opartym o próbki i wyłączaniem skalowania czasowego w sposób dostosowywany do sygnału.
- 4Sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrz. 1 do 3, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać skalowanie czasowe oparte o ramki lub skalowanie czasowe oparte o próbki w celu sterowania głębokością bufora (320;430) rozsynchronizowania.
- 5Sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrz. 1 do 4, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać wstawianie komfortowego szumu lub usuwanie (856) komfortowego szumu, jeżeli poprzednia ramka była nieaktywna.
- 6Sterowanie (100;350;490) buforem rozsynchronizowania według zastrz. 5, w którym wstawianie komfortowego szumu skutkuje wstawieniem ramki komfortowego szumu do bufora (320;430) rozsynchronizowania, zaś usuwanie komfortowego szumu skutkuje usunięciem ramki komfortowego szumu z bufora (320;430) rozsynchronizowania.
- 7Sterowanie (100;350;490) buforem rozsynchronizowania według zastrz. 5 albo 6, w którym dana ramka jest traktowana jako nieaktywna, jeżeli ramka ta zawiera informacje sygnalizujące, które wskazują generowanie komfortowego szumu.
- 8Sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrzeżeń 1 do 7, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać nakładanie i sumowanie (954;1068) części sygnału audio z przesunięciem czasowym, jeżeli poprzednia ramka była aktywna.
- 9Sterowanie (100;350;490) buforem rozsynchronizowania według zastrz. 5, w którym nakładanie i sumowanie (954;1068) części sygnału audio z przesunięciem czasowym dostosowane jest do zapewniania możliwości regulowania przesunięcia czasowego między blokami próbek audio uzyskanych na podstawie kolejnych ramek wejściowej treści audio z rozdzielczością mniejszą niż długość bloków próbek audio, albo mniejszą niż jedna czwarta długości bloków próbek audio, albo mniejszą lub równą dwóm próbkom audio.
- 10Sterowanie (100;350;490) buforem rozsynchronizowania według zastrz. 8 albo 9, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by określić (930, 936;1010, 1014), czy blok próbek audio reprezentuje aktywną, lecz niemą część sygnału audio, i w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać tryb nakładania i sumowania (962;1018), w którym przesunięcie czasowe pomiędzy blokiem próbek audio reprezentującym niemą część sygnału audio, a poprzednim lub następnym blokiem próbek audio ma z góry określoną wartość maksymalną dla bloku próbek audio reprezentującego niemą część sygnału audio.
- 11Sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrz. 8 do 10, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by określić (930, 936;1010, 1014), czy blok próbek audio reprezentuje aktywną i nie-niemą część sygnału audio, oraz by wybierać tryb nakładania i sumowania (942, 950, 954;1030, 1060, 1064, 1068), w którym przesunięcie czasowe między blokami próbek audio określone na podstawie poprzednich ramek wejściowej treści audio określane jest w sposób dostosowany do sygnału.
- 12Sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrz. 1 do 11, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać wstawianie maskowanej ramki w odpowiedzi na stwierdzenie, że wymagane jest rozciąganie czasowe i że bufor rozsynchronizowania jest pusty.
- 13Sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrz. 1 do 12, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać skalowanie czasowe oparte o ramki lub skalowanie czasowe oparte o próbki w zależności od tego, czy dla poprzedniej ramki stosowana jest lub była nieciągła transmisja w połączeniu z generowaniem komfortowego szumu.
- 14Sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrz. 1 do 13, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać skalowanie czasowe w oparciu o ramki, jeżeli obecnie stosowane jest generowanie komfortowego szumu lub było ono stosowane w odniesieniu do poprzedniej ramki, oraz by wybierać skalowanie czasowe oparte o próbki, jeżeli generowanie komfortowego szumu nie jest obecnie stosowane lub nie było stosowane w odniesieniu do poprzedniej ramki.
- 15Sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrz. 1 do 14, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać wstawianie komfortowego szumu w oparciu o ramki lub usuwanie (856) komfortowego szumu w oparciu o ramki w skalowaniu czasowym, jeżeli obecnie stosowana jest nieciągła transmisja w połączeniu z generowaniem komfortowego szumu lub była ona stosowana w odniesieniu do poprzedniej ramki;w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać operację nakładania i sumowania z wykorzystaniem z góry określonego przesunięcia czasowego (962, 1018) w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna, lecz zawiera energię sygnału mniejszą lub równą wartości granicznej energii, i jeżeli bufor rozsynchronizowania nie jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna, lecz zawierała energię sygnału mniejszą lub równą wartości granicznej energii, i jeżeli bufor rozsynchronizowania nie jest pusty;w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać operację nakładania i sumowania z wykorzystaniem dostosowywanego do sygnału przesunięcia czasowego (954;1068) w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna i zawiera energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna i zawierała energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty;i w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać wstawianie maskowanej ramki w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna i jeżeli bufor rozsynchronizowania jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna i jeżeli bufor rozsynchronizowania jest pusty.
- 16Sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrz. 1 do 15, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać operację nakładania i sumowania (942, 950, 954;1030, 1060, 1064, 1068) z wykorzystaniem dostosowywanego do sygnału przesunięcia czasowego oraz mechanizmu (950;1060, 1064, 1072, 1084) kontroli jakości w skalowaniu czasowym jeżeli bieżąca część sygnału audio jest aktywna i zawiera energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna i zawierała energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty.
- 17Dekoder (300; 400) sygnału audio przeznaczony do dostarczania zdekodowanej treści (312; 412) audio na podstawie wejściowej treści audio, gdzie dekoder sygnału audio zawiera:bufor (320;430) rozsynchronizowania skonfigurowany tak, by buforować wiele ramek audio reprezentujących bloki próbek audio;rdzeń (330;440) dekodera skonfigurowany tak, by dostarczać bloki (332;442) próbek audio uzyskane na podstawie ramek (322;432) audio odebranych z bufora rozsynchronizowania;Licznik czasu (340;450) oparty o próbki, przy czym licznik czasu oparty o próbki skonfigurowany jest tak, by dostarczać skalowane czasowo bloki próbek (342;448) audio uzyskane na podstawie bloków próbek audio dostarczanych przez rdzeń dekodera;i sterowanie (100;350;490) buforem rozsynchronizowania według jednego z zastrzeżeń 1 do 16.
- 18Dekoder (300;400) sygnału audio według zastrz. 17, w którym bufor (320;430) rozsynchronizowania skonfigurowany jest tak, by pomijać lub wstawiać ramki audio w celu przeprowadzania skalowania czasowego w oparciu o ramki.
- 19Dekoder sygnału audio według zastrz. 17 albo 18, w którym rdzeń (330;440) dekodera skonfigurowany jest tak, by wykonać generowanie komfortowego szumu w odpowiedzi na ramkę zawierającą informacje sygnalizujące, które wskazują na generowanie komfortowego szumu, i w którym rdzeń dekodera skonfigurowany jest tak, by wykonać maskowanie w odpowiedzi na pusty bufor rozsynchronizowania.
- 20Dekoder (300;400) sygnału audio według jednego z zastrzeżeń 17 do 19, w którym licznik czasu (340;450) oparty o próbki skonfigurowany jest tak, by wykonać skalowanie czasowe wejściowego sygnału audio w zależności od wyniku obliczania lub estymacji (950;1060) jakości skalowanej czasowo wersji wejściowego sygnału audio, która może zostać uzyskana przez skalowanie czasowe.
- 21Sposób (1400) sterowania dostarczaniem zdekodowanej treści na podstawie wejściowej treści audio, w którym sposób obejmuje wybieranie (1410) skalowania czasowego opartego o ramki lub skalowania czasowego opartego o próbki w sposób dostosowywany do sygnału, tak że decyzja o tym, czy stosowane jest skalowanie czasowe oparte o ramki, czy skalowanie czasowe oparte o próbki, dostosowywana jest do właściwości sygnału audio.
- 22Program komputerowy przeznaczony do wykonywania sposobu według zastrz. 21, gdy program komputerowy jest wykonywany z wykorzystaniem komputera.
- 23Sterowanie (100;350;490) buforem rozsynchronizowania według zastrz. 2, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać wstawianie komfortowego szumu w oparciu o ramki lub usuwanie (856) komfortowego szumu w oparciu o ramki w skalowaniu czasowym, jeżeli obecnie stosowana jest nieciągła transmisja w połączeniu z generowaniem komfortowego szumu lub była ona wykorzystywana w odniesieniu do poprzedniej ramki;w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać operację nakładania i sumowania z wykorzystaniem z góry określonego przesunięcia czasowego (962, 1018) w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna, lecz zawiera energię sygnału mniejszą lub równą wartości granicznej energii, i jeżeli bufor rozsynchronizowania nie jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna, lecz zawierała energię sygnału mniejszą lub równą wartości granicznej energii, i jeżeli bufor rozsynchronizowania nie jest pusty;w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać operację nakładania i sumowania z wykorzystaniem dostosowywanego do sygnału przesunięcia czasowego (954;1068) w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna i zawiera energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna i zawierała energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty;i w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać wstawianie maskowanej ramki w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna i jeżeli bufor rozsynchronizowania jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna, i jeżeli bufor rozsynchronizowania jest pusty.
- 24Sterowanie (100;350;490) buforem rozsynchronizowania według zastrz. 2, w którym sterowanie buforem rozsynchronizowania skonfigurowane jest tak, by wybierać operację nakładania i sumowania (942, 950, 954;1030, 1060, 1064, 1068) z wykorzystaniem dostosowywanego do sygnału przesunięcia czasowego oraz mechanizmu (950;1060, 1064, 1072, 1084) kontroli jakości w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna i zawiera energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna i zawierała energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty.
- 25Sposób (1400) według zastrz. 21, w którym sposób obejmuje wybieranie (1410) skalowania czasowego opartego o ramki lub skalowania czasowego opartego o próbki w sposób dostosowywany do sygnału;w którym ramki audio są pomijane lub wstawiane w celu sterowania głębokością bufora (320;430) rozsynchronizowania, gdy stosowane jest skalowanie czasowe oparte o ramki, i w którym wykonywane jest nakładanie i sumowanie (954;1068) części sygnału audio z przesunięciem czasowym, gdy stosowane jest skalowanie czasowe oparte o próbki;w którym sposób obejmuje wybieranie wstawiania komfortowego szumu lub usuwania (856) komfortowego szumu w skalowaniu czasowym, jeżeli obecnie stosowana jest nieciągła transmisja w połączeniu z generowaniem komfortowego szumu lub była ona stosowana w odniesieniu do poprzedniej ramki, wybieranie operacji nakładania i sumowania z wykorzystaniem z góry określonego przesunięcia czasowego (962, 1018) w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna, lecz zawiera energię sygnału mniejszą lub równą wartości granicznej energii, i jeżeli bufor rozsynchronizowania nie jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna, lecz zawierała energię sygnału mniejszą lub równą wartości granicznej energii, i jeżeli bufor rozsynchronizowania nie jest pusty;wybieranie operacji nakładania i sumowania z wykorzystaniem dostosowywanego do sygnału przesunięcia czasowego (954;1068) w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna i zawiera energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna i zawierała energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty;i wybieranie wstawiania maskowanej ramki w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna i jeżeli bufor rozsynchronizowania jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna i jeżeli bufor rozsynchronizowania jest pusty.
- 26Sposób (1400) według zastrz. 21, w którym sposób obejmuje wybieranie (1410) skalowania czasowego opartego o ramki lub skalowania czasowego opartego o próbki w sposób dostosowywany do sygnału;w którym ramki audio są pomijane lub wstawiane w celu sterowania głębokością bufora (320;430) rozsynchronizowania, gdy stosowane jest skalowanie czasowe oparte o ramki, i w którym wykonywane jest nakładanie i sumowanie (954;1068) części sygnału audio z przesunięciem czasowym, gdy stosowane jest skalowanie czasowe oparte o próbki;w którym sposób obejmuje wybieranie operacji nakładania i sumowania (942, 950, 954;1030, 1060, 1064, 1068) z wykorzystaniem dostosowywanego do sygnału przesunięcia czasowego oraz mechanizmu (950;1060, 1064, 1072, 1084) kontroli jakości w skalowaniu czasowym, jeżeli bieżąca część sygnału audio jest aktywna i zawiera energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty, lub jeżeli poprzednia część sygnału audio była aktywna i zawierała energię sygnału większą lub równą wartości granicznej energii i jeżeli bufor rozsynchronizowania nie jest pusty.
- 27Program komputerowy przeznaczony do wykonywania sposobu według zastrz. 25 albo 26, gdy program komputerowy jest wykonywany z wykorzystaniem komputera. Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung E.V., Niemcy Pełnomocnik:EP 3 011 692 B1 Z - 15937/17 1/14 O 'E? cś ε 1£ α L O O fZ) O 'E? cś 1£ α FIG1 EP 3 011 692 B1 Z - 15937/17 2/14 200 moduł skalowania czasowego wejściowy sygnał audio 210 X obliczanie lub estymowanie jakości skalowanej czasowo wersji wejściowego sygnału audio, która może zostać uzyskana przez skalowanie czasowe wejściowego sygnału audio skalowana czasowo wersja wejściowego sygnału audio 212 X przeprowadzanie skalowania czasowego wejściowego sygnału audio na podstawie wyniku obliczania lub estymowania jakości skalowanej czasowo wersji wejściowego sygnału audio, która może zostać uzyskana przez skalowanie czasowe FIG 2 EP 3 011 692 B1 Z - 15937/17 3/14 FIG 3 EP 3 011 692 B1 Z - 15937/17 4/14 o FIG 4 EP 3 011 692 B1 Z - 15937/17 5/14 soundCardFrameSize = sampleRate * 20/1000;while(pcmBuffer_nReadabłeSamples() soundCardFrameSize) { accessUnit = deJitterBuffer_pop();^—510 audioSamples = decoderdecodeOrConceal(accessUnit);—512 tsm_scale(audioSamples);-—514 pcmBuffer_write(aud ioSamples);^—516 l· audioSamples = pcmBuffer_read(soundCardFrameSize);—-520 soundcardpush(audioSamples);— 522 FIG 5 rtpTimeDiff = rtpTimeStamp - prevRtpTimeStamp;-—610 rcvTimeDiff = rcvTime - prevRcvTime;-—612 /* convert time base to miIIiseconds 7 rtpTimeTicks = rtpTimeStamp * sysTimeScale / rtpTimeScale;~614 rtpTimeDiff = rtpTimeDiff * sysTimeScale / rtpTimeScale;-—616 delay = rcvTimeDiff - rtpTimeDiff + prevDelay;—618 offset = rcvTime - rtpTimeTicks;—620 longTermFifo_add( delay, offset, rtpTimeTicks);^—622 /* backup values of current packet 7 prevRtpTimeStamp = rtpTimeStamp;Λ prevRcvTime = rcvTime;—624 prevDelay = delay;J FIG 6 targetMax = shortTermJitter + 60;-—710 targetMin = min( longTermJitter + 20, targetMax);—'712 targetDtx = min( longTermJitter, shortTermJitter);——714 targetStartllp = ( targetMin + targetMax) / 2;-—716 FIG 7 EP 3 011 692 B1 Z - 15937/17 6/14 Γ' ii o -=Τ LO OO OO EP 3 011 692 B1 Z - 15937/17 7/14 . ω o o cl cś cd cd CD O S en § £ o a ··» o s N O CL cś § a N .O ’5b o § brak skalowania § $ o N O \ Od o c/j O o 13 O Q_ CD -CVJ CD cn O Lł_ CD CT EP 3 011 692 B1 Z - 15937/17 1-szy blok 2-gi blok EP 3 011 692 B1 Z - 15937/17 EP 3 011 692 B1 Z - 15937/17 10/14 OQ CD CD EP 3 011 692 B1 Z - 15937/17 11/14 qMinlnitial = 1;qualityRise = 4;qualityRed = 4;p = similarityEstimation();q = c(p) * c(2*p) + c(3/2*p) * c(1/2*p);qMin = qMinlnitial - (nNotScaled * 0.1) + (nScaled * 0.2);if(q = qMin) {/* sufficient quality, apply time-scaling 7 applyOverlapAdd();if( nNotScaled 0) nNotScaled = nNotScaled -1;if( nScaled qualityRise ) nScaled = nScaled + 1;} else {/* not sufficient, postpone time-scaling 7 if( nNotScaled qualityRed ) nNotScaled = nNotScaled + 1;if( nScaled 0) nScaled = nScaled -1;} 1110 1112 1114 1116 '1118 Ί120 1122 Ί124 Ί126 FIG 11 EP 3 011 692 B1 Z - 15937/17 12/14 CM CM CM [sui] aiuaiuzodo EP 3 011 692 B1 Z - 15937/17 13/14 1300 [sui] 3MOSBZO 9TUBMOJB5JS W £ o CZ3 N U W O u w CZ3 FIG 13 EP 3 011 692 B1 Z - 15937/17 14/14 1400 Wybieranie w sposób dostosowywany do sygnału skalowania czasowego w oparciu o ramki lub skalowania czasowego w oparciu o próbki 1410 FIG 14 1510 1520 FIG 15
Independent claims27
221 paragraphs in 1 section, as filed
Technical Field [0001] Embodiments of the present invention relate to controlling a synchronization buffer to control the delivery of the decoded audio content based on the input audio content.
[0002] Further embodiments of the invention relate to an audio signal decoder for providing decoded audio content based on the input audio content.
[0003] Further embodiments of the invention relate to a method of controlling the delivery of a decoded audio content based on the input audio content.
[0004] Further embodiments of the invention relate to a computer program for carrying out said method.
2. Background of the Invention [0005] Storing and transferring audio content (containing general audio content, such as music content, speech content, and mixed general content / speech) is an important field of technology. Particular challenges are related to the fact that the listener expects continuous playback of the audio content without any interruptions, as well as without any audible artifacts associated with the storage and / or transmission of audio content. At the same time, it is desirable to maintain the smallest possible requirements regarding the means for storing data and the means for data transmission, the aim being to keep costs at an acceptable level.
[0006] Problems arise, for example, in the case of a temporary interruption or delay in reading from a data storage medium, or in the case of a temporary interruption or delay of transmission between a data source and a data mouth. For example, transmission via the Internet is not very reliable, because TCP / IP packets can be lost, and because transmission delays via the Internet may change, for example depending on the load variations of the Internet network nodes.
[0007] In order to obtain satisfactions satisfactory for the user, however, continuous playback of the audio content is required, without audible "gaps" or audible artifacts. It is furthermore desirable to avoid significant delays that could be caused by buffering significant amounts of audio information.
[0008] In light of the above discussion, a need may be felt for a concept that provides good audio quality, even in the case of discontinuous delivery of audio information. The document [AHEVS-044] is the most similar state of the art and will be discussed at the end of the description, see the Bibliography section.
3. Summary of the Invention [0009] An embodiment of the invention is the control of the synchronization buffer to control the delivery of the decoded audio content based on the input audio content. The control of the synchronization buffer is configured to obtain frame-based time scaling or time scaling based on samples carried out in a manner adapted to the signal.
[0010] This embodiment of the invention is based on the discovery that the use of temporal scaling enables delivery of continuous decoded audio content with good and at least acceptable quality, even when the input audio content has a significant out-of-sync, or in the case where parts ( for example, frames) of the input audio content have been lost. Furthermore, this embodiment of the invention is based on the discovery that frame-based time scaling is computationally efficient and provides good results in some cases, while time scaling based on samples, which is usually more computationally complex, may be in some situations recommended (and even required) to avoid the occurrence of audible artifacts of the decoded audio content. It was also found that that particularly good results can be achieved by selecting time based frame scaling and time scaling based on the samples in a manner adapted to the signal, because it is thus possible to use the most appropriate type of time scaling. This embodiment of the invention thus provides a good compromise between audio quality, delays and computational complexity.
[0011] In a preferred embodiment, when frame-based time scaling is used, the audio frames are omitted or inserted to control the depth of the synchronization buffer. When time scaling based on samples is used, the overlapping and addition of parts of the audio signal with time shift is also used. It is therefore possible to use very computationally efficient solutions, e.g. such as bypassing or inserting audio frames, if the audio signal allows such a solution without causing unacceptable distortion of the audio signal. On the other hand, more advanced and more capable of customizing the solution, namely overlapping and summation of parts of the audio signal with a time shift, is used, if a less computationally complex solution (e.g. such as bypassing or inserting audio frames) would not (or probably would not, given the result of the evaluation or analysis of audio content) obtain satisfactory audio quality. It is therefore possible to get a good compromise between computational complexity and audio quality.
[0012] In a preferred embodiment, the control of the synchronization buffer is configured to provide switching between frame-based time scaling, time scaling based on samples, and disabling time scaling in a manner adapted to the signal. By using such a solution, it is possible (at least temporarily) to avoid performing time scaling operations if it detects that the time scaling operation (even a time scaling operation based on samples) would cause unacceptable deterioration of the quality of the audio content.
[0013] In a preferred embodiment, the scratch buffer control is configured to provide a time based frame-based scaling or time-based scaling based on the samples to control the depth of the de-synchronizing buffer (which can also be called short-term "synchronization buffer"). It is thus possible to keep the "depth" (or fill) of the buffer removing the synchronization (or disjuncture buffer) in the desired range, which makes it possible to deal with the decompression without significant deterioration of the audio quality while maintaining a reasonably low delay (which usually corresponds to the depth of the buffer removing the synchronization) .
[0014] In a preferred embodiment, the control of the synchronization buffer is configured to provide for the introduction of comfortable noise or the removal of comfortable noise if the previous frame was "inactive". This concept is based on the discovery that it is sufficient to generate a comfortable noise if the frame (e.g. the previous frame) is inactive, for example when there is no speech for some time, because the speaker only listens. However, it has been found that time scaling can be performed without seriously degrading the auditory experience by introducing additional comfort noise (e.g., the entire comfort noise frame) to the audio content, because the extended duration of comfortable noise is not perceived by man as an essential artifact. Similarly, the removal of comfortable noise (for example, a comfortable noise frame) does not seriously degrade the auditory experience, because a person does not notice the "absence" of a comfortable noise frame. It is therefore possible to perform very efficient time scaling based on the frames (by introducing a comfort noise frame or removing the comfort noise frame) if the previous frame was inactive (e.g. a frame of comfortable noise). Performing a frame-based time-based scaling if the previous frame was inactive thus provides good time scaling performance (while frame-based time scaling is not usually well suited to the "active" parts of the audio content, if not present specific conditions, e.g. such as an empty buffer).
[0015] In a preferred embodiment, the introduction of comfortable noise results in the insertion of a comfort noise box into the buffer removing the synchronization (also referred to briefly as the "synchronization buffer"). Furthermore, the removal of the comfortable noise advantageously results in the removal of the comfort noise frame from the de-synchronizing buffer. It is therefore possible to perform very efficient time scaling, because the comfort noise box can be easily generated (wherein the comfort noise frame usually contains suitable signaling information intended to signal that a comfortable noise should be generated).
[0016] In a preferred embodiment, the given frame is considered inactive if the frame includes signaling information that indicates the generation of comfortable noise. It is therefore possible to carry out a small amount of means adapted to the time scaling mode selection signal (frame based time scaling or time scaling based on samples).
[0017] In a preferred embodiment, the control of the synchronization buffer is configured to ensure overlap and summation of a portion of the audio signal with a time shift if the previous frame was & quot; active & quot ;. This concept is based on the discovery that overlapping and summation of a portion of the audio signal with a time shift provides a limitation (or even elimination) of audible artifacts in the case of an active frame (e.g. a frame containing audio content, and not signaling information that indicates comfortable noise generation).
[0018] In a preferred embodiment, the overlapping and summation of a portion of the time-varying audio signal is adapted to provide the ability to adjust the time shift between blocks of audio samples obtained from a single frame or based on consecutive frames of input audio content, which is performed at a resolution less than the length an audio sample block (or less than a frame length), less than a quarter of the block length of the audio samples (or less than a quarter of the frame length), or less than two audio samples or the same. In other words, in the case of using overlapping and summation of a portion of the audio signal with a time shift, the time shift can be adjusted with a very fine resolution that can be as small as a single audio sample.
In a preferred embodiment, the scratch buffer control is configured to ensure that the block of audio samples represents an active but silent part of the audio signal (e.g. a part of the audio signal that is considered "active" because there is no generation of comfortable noise, but is treated as "silent", because the energy of the said part of the signal is less than or equal to a certain energy limit value, and even is zero), while in the case of an audio sample block representing a "quiet" (but "active") part of the audio signal , selecting an overlap and add mode in which the time shift between the block of audio samples representing the silent portion of the audio signal and the next block of audio samples has a predetermined maximum value. In the case of such parts of the audio signal, which are "active" but "silent", e.g. in accordance with the above-mentioned definitions, therefore a maximum time scaling is carried out. As a result, the maximum time scaling is obtained with respect to the portion of the audio signal where the time scaling does not cause significant audible artifacts, because "quiet" parts of the audio signal can be time scaled with little audible distortion or in the absence thereof. In addition, by using the maximum time scaling relative to the "silent" parts of the audio signal, it is sufficient to use only comparatively less time scaling with respect to the "non-silent" parts of the audio signal (e.g. "active" parts of the audio signal containing more than the limit energy). the value of energy or its equal),
[0020] In a preferred embodiment, the scratch buffer control is configured to ensure that the block of audio samples represents the active and non-silent part of the audio signal, and selecting the overlap and add mode in which the time shift between the blocks of the audio samples (which are superimposed and summed up in a way including a time shift, i.e. shifted relative to their original time position) is determined in a manner adapted to the signal. It has been found that determining the time shift between (consecutive) blocks of audio samples in a manner adapted to the signal is well suited for performing overlapping and adding in case the block of audio samples (or consecutive blocks of audio samples) represents an active and non-silent part of the audio signal,
[0021] In a preferred embodiment, the control of the shifting buffer is configured to ensure the insertion of the masked frame when it is determined that time stretching is required and the synchronization buffer is empty. During the control of the synchronization buffer, it is therefore possible to carry out specific operations in the case of an empty disjoint buffer, because other concepts of time scaling are usually not successful if the out-of-sync buffer is empty. For example, a masked frame may include audio content similar to the audio content of the previous frame (e.g. the last frame obtained before the desynchronization buffer has been emptied). In the case where the previous frame before the embolization buffer empties contains signaling information,
[0022] In a preferred embodiment, the scratch buffer control is configured to provide timing based on frames or time scaling based on samples depending on whether discontinuous transmission is currently used in combination with the generation of comfortable noise (or, equivalently, this was the case for the previous frame). This concept is based on the discovery that the use of time-based scaling on samples when discontinuous transmission is currently used in combination with the generation of comfortable noise (or was used in the previous frame) is inefficient. Therefore, information can be used in controlling the out-of-sync buffer
[0023] In a preferred embodiment, the scratch buffer control is configured to ensure that frame-based time scaling is selected, if comfort noise generation is currently used (or, equivalently, was used for the previous frame), and time based scaling selection. o samples if generation of comfortable noise is currently not used (or, equivalently, it was not used in the previous frame). The time scaling mode is therefore well suited to the signal (where in the case of the extraordinary state of the empty disjoint buffer, it is also possible to use "time scaling" based on frames using a masked frame).
selecting an overlap and add operation using a time-shifted time-shifted offset, if the current portion of the audio signal (or, equivalently, the previous portion of the audio signal) is active (e.g., comfortable noise generation is not used for it), contains energy a signal greater than or equal to the energy limit and the synchronization buffer is not empty, and selecting to insert the masked frame in time scaling, if the current portion of the audio signal (or, equivalently, the previous part of the audio signal) is active and the out-of-sync buffer is empty. It is thus possible to select, for each part (or frame) of the audio signal, the actual time scaling mode in a computationally efficient manner.
[0025] In a preferred embodiment, the control of the synchronization buffer is configured to ensure selection of the overlap and add operations using the time-shifted and quality-control mechanism in time scaling, if the current portion of the audio signal (or, equivalently, the previous portion of the audio signal). ) is active (for example, comfortable noise generation is not used for it), contains signal energy greater than or equal to the energy limit value, and the synchronization buffer is not empty. The application of overlapping and adding operations using a time-shifted and quality control mechanism that is adapted to the signal ensures the advantage of that by using a quality control mechanism it is possible to eliminate unacceptable audible artifacts. The quality control mechanism may, for example, prevent the use of time scaling, even if the current portion of the audio signal is active and contains signal energy greater than or equal to the energy limit value, and even if the time shift is desired by the control logic.
[0026] An embodiment of the invention is an audio signal decoder intended for providing decoded audio content based on the input audio content. The audio signal decoder includes a synchronization buffer configured to provide buffering of multiple audio frames representing blocks of audio samples. The audio signal decoder also includes a decoder core configured to provide blocks of audio samples obtained from the audio frames received from the out-of-sync buffer. The audio signal decoder also includes a time counter based on samples, wherein the sample based time counter is configured to provide time scaled blocks of audio samples obtained from the audio samples provided by the decoder core. The audio signal decoder further comprises controlling the synchronization buffer, such as described above. The control of the synchronization buffer is configured to select in a manner adapted to the signal time scaling based on frames carried out by the synchronization buffer or time scaling based on samples carried out by the time counter based on samples. Controlling the synchronization buffer thus provides for selecting two substantially different concepts of time scaling, where frame-based time scaling is performed before the audio frames are inserted into the decoder core (e.g. by adding or removing entire audio frames), and time scaling based on samples being carried out is for blocks of audio samples provided by the core of the decoder.
[0027] In a preferred embodiment, the out-sync buffer is configured to provide skipping or inserting audio frames to perform frame based time scaling. Thus, the time scaling can be carried out in a particularly efficient manner. For example, the skipped time frames may be time frames containing signaling information that indicate the generation of comfortable noise (and optionally also indicating "silence"). In addition, the inserted time frames can be, for example, time frames containing signaling information that indicate the need to generate a comfortable noise. Such timing frames can be easily inserted or omitted, wherein, if the omitted frame or the preceding frame preceding the inserted frame contains signaling information,
[0028] In a preferred embodiment, the core of the audio signal decoder is configured to provide comfortable noise generation in the case of a frame containing signaling information that indicates the generation of comfortable noise. Furthermore, the decoder core is preferably configured to provide masking in the case of an empty desynchronization buffer. The use of a decoder core configured to provide comfortable noise generation enables efficient audio content transfer as well as efficient time scaling. In addition, the use of a decoder core configured to mask in the case of an empty desynchronization buffer, it eliminates the problem of the possibility of interrupting the playback of audio content in the case of an empty disjoint buffer and no masking function. The use of the decoder core described thus enables an efficient and reliable delivery of the decoded audio content.
[0029] In a preferred embodiment, the sample based time counter is configured to perform time scaling of the input audio signal including the result of calculating or estimating the quality of the time scaled version of the input audio signal that can be obtained by time scaling. An additional mechanism is therefore provided to ensure that the generation of audible artifacts is limited or even eliminated at the second stage, i.e. after the decision to use time scaling based on samples has been made. In other words, the choice between time scaling based on frames and time scaling based on samples is carried out in the first stage in a manner adapted to the signal, and quality control (computational estimation of the quality of the time scaled version of the input audio signal,
[0030] An embodiment of the invention is a method of providing a decoded audio content n the basis of an input audio content. The method comprises selecting in a manner adapted to a frame-based time scaling signal or time scaling based on samples. The method is based on the same observations as the de-synchronization buffer control described above and the audio signal decoder described above.
[0031] Yet another embodiment of the invention is a computer program for carrying out said method when the computer program is executed by a computer. The computer program is based on the same observations as the method described above.
4. Brief description of the figures [0032] Embodiments of the invention will be described below with reference to the attached figures, in which:
Fig. 1 is a block diagram of a control for a synchronization buffer according to an embodiment of the present invention;
Fig. 2 shows a block diagram of a time scaling module according to an embodiment of the present invention;
Fig. 3 is a block diagram of an audio signal decoder according to an embodiment of the present invention;
Fig. 4 shows a block diagram of an audio signal decoder according to another embodiment of the present invention, wherein a review of a de-synchronization buffer management (JBM) is shown here;
Fig. 5 shows a pseudocode of a program implementing a PCM buffer level control algorithm;
Fig. 6 shows a pseudocode of a program implementing an algorithm for calculating delay values and offset values based on the time of receipt and the RTP time stamp of an RTP packet;
Fig. 7 shows a pseudocode of a program implementing an algorithm for calculating target delay values;
Fig. 8 is a flowchart of the management logic of the synchronization buffer;
Fig. 9 is a block diagram illustrating a modified WSOLA algorithm with quality control;
Figs. 10 and 10b show a flowchart of a method for controlling a time scaling module;
Fig. 11 shows a pseudocode of a program implementing a quality control algorithm in time scaling;
Fig. 12 is a graphical representation of the target delay and reproduction delay that is obtained in an embodiment of the present invention;
Fig. 13 is a graphical representation of time scaling performed in an embodiment of the present invention;
Fig. 14 is a flowchart of a method for controlling the delivery of a decoded audio content based on the input audio content; and
Fig. 15 is a flowchart of a method for providing a time-scaled version of an input audio signal according to an embodiment of the present invention.
5. Detailed description of the implementation
5.1 Control of the synchronization buffer shown in Fig. 1 [0033] Fig. 1 is a block diagram of a de-synchronization buffer control according to an embodiment of the present invention. To control the dis-synchronization buffer to control the delivery of the decoded audio content based on the input audio content, an audio signal 110 or audio information is input (the information may describe one or more characteristics of the audio signal, audio frames or other parts of the audio signal). .
[0034] Controlling 100 with the synchronization buffer further provides control information (e.g., control signal) 112 used during frame based scaling. For example, the control information 112 may include an activation signal (frame based time scaling) and / or quantitative control information (used for scaling based on samples).
[0035] Control 100 of the synchronization buffer further provides control information (e.g., control signal) 114 used during scaling based on samples. The control information 114 may, for example, include an activation signal and / or quantitative control information used for scaling based on samples.
[0036] Control 100 of the synchronization buffer is configured to provide selection in a manner adapted to a time-based scaling signal based on frames or time scaling based on samples. Controlling the synchronization buffer may thus be configured to provide evaluation of the audio signal or information regarding the audio signal 110 and to provide control information 112 and / or control information thereon 114 on the basis thereof. Decision on whether frame based time scaling or scaling is used the time based on the samples may thus be adapted to the characteristics of the audio signal, e.g. in such a way that computational simple frame-based scaling is used in the case of when based on the audio signal or on the basis of one or more audio signal properties, it is expected (or estimated) that frame-based time-based scaling will not cause significant degradation of the audio content. Conversely, in the control of the synchronization buffer, a decision is usually made to use time based scaling on samples, if it is expected or estimated (using control of the synchronization buffer) on the basis of the evaluation of the audio signal 110 that the time based scaling is required for avoiding the occurrence of audible artifacts while performing time scaling.
[0037] It should further be noted that control information may also be provided to control the synchronization buffer, for example control information, which indicates whether time scaling should be performed or not.
[0038] Some optional details of the scramble buffer control will be described in the following. For example, the scramble controller 100 may provide control information 112, 114 such that if frame-based time scaling is used, skipping or inserting audio frames is provided to control the depth of the scratch buffer, and so that when using time scaling in based on samples, overlapping and summation of parts of the audio signal with time shift is performed. In other words, the control of the de-synchronizing buffer may, for example, cooperate with the de-synchronization buffer (also called in some cases the de-synchronizing buffer) and control the out-sync buffer for performing frame-based time scaling. In this case, it is possible to control the depth of the out-of-sync buffer by skipping the frames coming from the out-of-sync buffer or by inserting frames (e.g. single frames that indicate that the frame is "inactive" and using comfortable noise to be applied) to the out-sync buffer. Also,
[0039] The out-of-synchronization buffer controller 100 may be configured to provide switching in a manner adapted to the signal between frame-based time scaling, time scaling based on samples, and disabling time scaling. In other words, the control of the desynchronization buffer usually provides not only the distinction between frame-based time scaling and time scaling based on samples, but also the selection of a state where time scaling does not occur at all. For example, the latter state may be selected in the absence of the need for time scaling, because the depth of the synchronization buffer is within an acceptable range. Speaking in other words,
[0040] In control 100, the out-of-synchronization buffer may also include information about the depth of the outage buffer, which is to decide which mode of operation (e.g. frame-based time scaling, time based scaling or no time scaling). ) should be used. For example, in controlling a synchronization buffer, a target value describing a desirable depth of a synchronization buffer (also known as a de-synchronizing buffer) can be compared to an actual value describing the actual depth of the synchronization buffer and mode selection (frame-based time scaling, time based scaling) or no temporal scaling) depending on the result of the said comparison,
[0041] The shift controller 100 may, for example, be configured to provide for the insertion of comfortable noise or for removing comfortable noise if the previous frame was inactive (which may e.g. be recognized based on the audio signal 110 itself or based on audio signal information). , for example, such as a SID tag identifying silence in the case of a discontinuous transmission mode). If time warping is desired and the previous frame (or current frame) is inactive, the control 100 of the synchronization buffer may thus ensure sending to the de-synchronization buffer (also known as a de-synchronizing buffer) a signal indicating that a comfort noise frame should be inserted. In addition, if it is desired to perform time compression, while the previous frame was inactive (or the current frame is inactive), the scramble buffer 100 may send to the disjunction buffer (or buffer to desynchronize) the delete noise frame command (e.g., the frame contains signaling information that indicates the need to generate a comfortable noise). It should be noted that a given frame may be considered inactive if the frame contains signaling information that indicates the generation of comfortable noise (and usually does not include any additional encoded audio content). In the case of a discontinuous transmission mode, such signaling information may for example be in the form of a SID tag identifying silence (SID tag). the control 100 of the synchronization buffer may send to the disjunction buffer (or buffer to remove the synchronization) a command to remove the comfort noise frame (e.g. the frame contains signaling information that indicates the need to generate a comfortable noise). It should be noted that a given frame may be considered inactive if the frame contains signaling information that indicates the generation of comfortable noise (and usually does not include any additional encoded audio content). In the case of a discontinuous transmission mode, such signaling information may for example be in the form of a SID tag identifying silence (SID tag). the control 100 of the synchronization buffer may send to the disjunction buffer (or buffer to remove the synchronization) a command to remove the comfort noise frame (e.g. the frame contains signaling information that indicates the need to generate a comfortable noise). It should be noted that a given frame may be considered inactive if the frame contains signaling information that indicates the generation of comfortable noise (and usually does not include any additional encoded audio content). In the case of a discontinuous transmission mode, such signaling information may for example be in the form of a SID tag identifying silence (SID tag).
[0042] In contrast, the control 100 of the synchronization buffer is preferably configured to ensure selecting overlapping and summation of the audio portion with a time shift if the previous frame was active (e.g., if the previous frame did not contain signaling information that indicates the need to generate comfortable noise). Such overlapping and summation of a portion of the audio signal with a time offset generally allows adjusting the time shift between blocks of audio samples obtained from successive input frames of audio information at a relatively high resolution (e.g., a resolution smaller than the block length of the audio samples or less than a quarter of the block lengths of the audio samples) . and even less than two audio samples, or the same or as small as a single audio sample). The choice of time-based scaling based on samples therefore allows very fine tuning of time scaling, which avoids the occurrence of audible artifacts for active frames.
[0043] When a buffer override control selects frame-based scaling, the scratch buffer control may also provide additional control information intended to regulate or fine tune frame based time scaling. For example, the scramble buffer control may be configured to determine whether a block of audio samples represents an active but "silent" part of the audio signal, e.g., a portion of the audio signal having relatively low energy. In this case, i.e. if the audio signal part is & quot; active & quot; (e.g. it is not part of the audio signal for which the generation of comfortable noise is used in the audio decoder, instead of more accurate decoding of the audio content), but "silent" (e.g. because the signal energy is less than or equal to a certain energy limit and even is zero), the control of the synchronization buffer may provide control information 114 to select the overlap and add mode in which the time shift between the block of audio samples representing a "silent" (but active) part of the audio signal and the next block of audio samples has a predetermined maximum value. The time counter based on samples therefore does not have to determine the appropriate degree of time scaling based on a detailed comparison of successive blocks of audio samples, but simply uses a predetermined maximum value of the time shift. You can understand this way, that the "silent" part of the audio signal does not usually produce significant artifacts during the overlap and add operation, regardless of the actual choice of the time shift. The control information 114 provided by the control of the synchronization buffer thus makes it possible to simplify the processing which is carried out by a timer based on samples.
[0044] In contrast, in the case of control in control 100, the synchronization buffer indicates that the block of audio samples represents an "active" and non-silent part of the audio signal (e.g. a part of the audio signal for which no comfortable noise generation is carried out, and which includes a signal energy greater than a predetermined limit), the control of the synchronization buffer provides control information 114 to select an overlap and summation mode in which the time shift between the blocks of the audio samples is determined in a manner adapted to the signal (e.g. o samples using the determination of similarities between consecutive blocks of audio samples).
[0045] Furthermore, information regarding the actual buffer fill may also be provided to control the dis-synchronization buffer. Controlling the synchronization buffer 100 may ensure selecting insertion of the masked frame (i.e., a frame generated using a packet loss fix mechanism, e.g. by prediction based on previously decoded frames), if it is detected that time stretching is required, and the synchronization buffer is empty. In other words, the control of the synchronization buffer may ensure the initiation of special processing in the case where it would be desirable in principle to scale time based on samples (since the previous frame or current frame is "active"), but time scaling based on samples (for example, using overlap and summation) can not be properly performed because the out-of-sync buffer (or buffer removing the out-of-sync) is empty. The control 100 of the synchronization buffer can thus be configured to provide the correct control information 112, 114 even in extraordinary cases.
[0046] To simplify the operation of the scramble buffer control 100, the scramble buffer 100 may be configured to provide time based frame scaling or time scaling based on samples depending on whether discontinuous transmission is currently being used (determined also briefly as "DTX") combined with the generation of comfortable noise (also referred to briefly as "CNG"). In other words, a scramble buffer control may, for example, provide for frame-based time based scaling when determined from an audio signal or based on audio information that the previous frame (or current frame) is an "inactive" frame, which should be used to generate comfortable noise. This can be determined, for example, by evaluating signaling information (e.g. a tag such as a so-called & quot; SID & quot; mark) that are included in the encoded representation of the audio signal. Controlling the desynchronization buffer may thus ensure that frame-based time scaling should be used if discrete transmission is currently used in conjunction with the generation of comfortable noise, since in such a case such temporal scaling can be expected to cause only slight audible distortion or it will not cause any audible distortion at all. In contrast, time scaling based on samples may be used in the opposite case (e.g.
[0047] In the case where time scaling is required, the control of the synchronization buffer may advantageously provide a choice between one of (at least) four modes. For example, the control of the synchronization buffer may be configured to provide for the insertion of a comfortable noise or the removal of comfortable noise in time scaling, if discrete transmission is currently used in combination with the generation of comfortable noise. In addition, the control of the synchronization buffer may be configured to ensure that an overlap and add operation is selected using a specific time shift in time scaling, if the current portion of the audio signal is active, but includes signal energy less than or equal to the energy limit value, and the out-sync buffer it is not empty. Also, the control of the synchronization buffer may be configured to ensure that the overlap and add operations are selected using a time-shifted time-shifting offset if the current portion of the audio signal is active and has a signal energy greater than or equal to the energy limit value, and the synchronization buffer is not it's empty. Finally, the control of the synchronization buffer may be configured to ensure that the insertion of the masked frame in time scaling is selected, if the current portion of the audio signal is active and the out-of-sync buffer is empty. It can therefore be seen that the control of the synchronization buffer can be configured as
[0048] It should further be noted that the control of the synchronization buffer can be configured to ensure selection of the overlap and add operations using the time-shifted and quality-control mechanism in time scaling, if the current portion of the audio signal is active and contains the signal energy higher. than the energy limit or its equal value, while the synchronization buffer is not empty. In other words, in the case of time scaling based on samples, an additional quality control mechanism may be used, which complements the signal-based selection between frame-based time scaling and time scaling based on the sample that is made by controlling the synchronization buffer.
[0049] In summary, the basic functionality of the scramble buffer control 100 is explained, as well as its optional improvements. It should further be noted that the control 100 of the synchronization buffer may be supplemented by any of the described features and functionalities.
5.2 The timer shown in Fig. 2. [0050] Fig. 2 is a block diagram of a time scaling module 200 according to an embodiment of the present invention. The time counter 200 is configured to receive an input audio signal 210 (e.g. in the form of a sequence of samples provided by the decoder core) and provides to a time scaled-out version 212 of the audio input signal derived therefrom. The time counter 200 is configured to calculate or estimate the quality of the time scaled version of the input audio signal that can be obtained by time scaling of the input audio signal. This functionality can be implemented, for example, by a computing unit. The time counter 200 is further configured as to perform time scaling of the input audio signal 210 depending on the result of computing or estimating the quality of the time scaled version of the input audio signal, which can be obtained by time scaling to obtain a time scaled version 212 of the input audio signal. This functionality may be implemented, for example, by a unit performing time scaling.
[0051] The time counter may thus perform quality control to ensure that no deterioration of the audio quality occurs during the time scaling. For example, the time counter may be configured to predict (or estimate) based on the input audio signal, or as a result of a predicted time scaling operation (e.g. such as an overlap and add operation performed with respect to time-shifted sample blocks (audio)) expected is to get enough good audio quality. In other words, the timer can be configured to calculate or estimate (predicted) the quality of the time scaled version of the input audio signal that can be obtained by time scaling of the input audio signal, what is done before real time scaling of the input audio signal is performed. For this purpose, the timer may, for example, compare portions of the audio signal used in the time scaling operation (said parts of the audio signal may be, for example, superimposed and summed up to thereby perform time scaling). In summary, the time counter 200 is typically configured to check if the expected time scaling can be expected to provide sufficient audio quality for the time scaled version of the input audio signal, and decide whether to perform time scaling or not. Alternatively,
[0052] Optional improvements to the time scaling module 200 will be described below.
[0053] In a preferred embodiment, the time counter is configured to perform an overlap and add operation using the first block of samples of the input audio signal and the second block of samples of the input audio signal. In this case, the timer is configured to time-shift the second block of samples relative to the first sample block, and to overlap and add the first block of samples and the time-shifted second block of samples to obtain a time-scaled version of the input audio signal. For example, if time compression is desired, the timer may enter the first number of samples of the audio input signal and provide the second number of samples of the time scaled version of the audio input derived therefrom, the second number of samples being smaller than the first number of samples. In order to achieve a reduction in the number of samples, the first number of samples may be divided into at least a first block of samples and a second block of samples (where the first block of samples and the second block of samples may overlap or overlap) and the first block of samples and the second. the sample block can be temporally offset relative to each other, whereby the temporally offset versions of the first sample block and the second sample block overlap. In the overlap region of the shifted versions of the first sample block and the second sample block, an overlap and add operation is performed. This overlap and add operation can be used without causing significant audible distortion, if the first block of samples and the second block of samples are "sufficient" similar in the overlap region (in which the overlap operation is performed), and preferably also in the vicinity of the overlap region. Due to the overlapping and summation of the signal portions that did not overlap initially, time compression is achieved, because the total number of samples is reduced by the number of samples that did not overlap each other (in the input audio signal) but which overlap in the scaled temporarily version 212 of the input audio signal.
[0054] In contrast, it is also possible to obtain a temporal extension using the overlapping and adding operations. For example, it is possible to select a first block of samples and a second block of samples overlapping that may include a first general temporal extension. Thereafter, the second block of samples may be time shifted relative to the first block of samples, whereby the overlap of the first block of samples and the second block of samples is reduced. If the time-shifted second block of samples is well suited to the first block of samples, it is possible to apply and add, wherein the overlapping region of the first sample block and the time-shifted version of the second sample block may be smaller than the original overlapping region of the first sample block and the second sample block, both for the number of samples and for time. The result of the overlap and add-on operations using the first block of samples and the time-shifted version of the second block of samples may thus include a larger time stretch (both in terms of time and number of samples) than the general stretch of the first block of samples and the second block of samples in their the original form.
[0055] It is thus seen that due to the overlapping and adding operation using the first sample block and the time-shifted version of the second sample block it is possible to obtain both time compression and time stretch where the second sample block is time shifted relative to the first sample block (or where both the first block of samples and the second block of samples are shifted in relation to each other).
The time counter 200 is preferably configured to calculate or estimate the quality of the result of the overlap and add operation performed with respect to the first sample block and the time-shifted version of the second sample block to calculate or estimate the (expected) quality of the time scaled version an input audio signal that can be obtained by time scaling. It should be noted that there are usually no audible artifacts when the overlap and summation operation is carried out with reference to sample blocks that are sufficiently similar. In other words, the quality of the result of the overlapping and adding operations greatly influences the (expected) quality of the time scaled version of the input audio signals.
The time counter 200 is preferably configured to determine a time offset of the second sample block relative to the first sample block based on the result of determining the level of similarity between the first sample block or the part (e.g. right part) of the first sample block and the time-shifted second sample block or a part (e.g. a left part) of a time-shifted second block of samples. In other words, the timer may be configured to determine which time shift between the first sample block and the second sample block is the most appropriate to obtain a sufficiently good overlap and summation result (or at least the best possible overlap and summation result). In an additional stage ("quality control"), it is possible to check
[0058] The time counter preferably defines information regarding the level of similarity between the first sample block or a part (e.g. the right part) of the first sample block and the second sample block or part (e.g., left part) with respect to a plurality of different time offsets between the first sample block and the the second block of samples and defines the (proposed) time shift to be used in the overlapping and aggregation operations, which is done on the basis of information on the level of similarity associated with many different time shifts. In other words, it is possible to conduct a search for the best match, during which information on the level of similarity associated with different time shifts can be compared to find a phase shift,
[0059] The timer is preferably configured to specify a time offset of the second block of samples relative to the first sample block, wherein the time shift is intended to be used in an overlap and add operation, this being done based on the target time shift information. In other words, information on the target time shift, which may for example be obtained from buffer fill, out of sync and optionally other additional criteria, may be considered (taken into account) when determining the time shift to be used (e.g. as a proposed offset). time) in the overlap and add operation. Application and aggregation is therefore adapted to the requirements of the system.
[0060] In some embodiments, the timer may be configured to calculate or estimate the quality of the time scaled version of the input audio signal that may be obtained by time scaling of the input audio signal, which is based on the information about the level of similarity between the first sample block. or a part (e.g. a right part) of the first block of samples and a second block of samples temporally shifted by a predetermined time shift or a part (e.g. a left part) of a second block of time shifted samples by a predetermined time shift. This information on the level of similarity provides information on the (expected) quality of the result of the overlapping and addition operations, as a result, they also provide information (at least estimated) on the quality of the time scaled version of the input audio signal that can be obtained by time scaling. In some cases, the calculated or estimated information about the quality of the time scaled version of the audio input signal that can be obtained by time scaling can be used to decide whether time scaling is actually performed or not (in the latter case, time scaling) can be postponed). In other words, the timer can be configured as based on information regarding the level of similarity between the first sample block or a part (e.g. the right part) of the first sample block and the second sample block time-shifted by a predetermined time shift or part (e.g., left part) of the second sample block temporally shifted by (proposed) time shift decided whether time scaling is actually carried out (or not). A quality control mechanism that evaluates calculated or estimated information about the quality of the time scaled version of the input audio signal that can be obtained by time scaling can actually result in the time scaling being omitted (at least in relation to the current block or frame of the audio samples),
[0061] In some embodiments, during the initial determination of the (proposed) time shift between the first sample block and the second sample block and in the final quality control mechanism, different measures of similarity may be used. In other words, the timer can be configured to provide a time shift of the second block of samples relative to the first block of samples and overlap and summation of the first block of samples and the second block of samples if computing or estimating the quality of the time scaled version of the audio input that can be obtained by time scaling, they indicate obtaining a quality that is greater than the threshold quality or the same. The time counter can be configured as to determine the (proposed) time shift of the second block of samples relative to the first block of samples based on the result of determining the level of similarity between the first sample block or the part (e.g. the right part) of the first block of samples and the second sample block or the part (e.g. left part) of the second block of samples which is evaluated using the first measure of similarity. The timer can also be configured to calculate or estimate the quality of the time scaled version of the input audio signal that can be obtained by time scaling of the input audio signal, which is made based on information regarding the level of similarity between the first sample block or a part (e.g. the right part) of the first sample block and the second sample block time-shifted by a predetermined time shift or a part (e.g., left part) of a second block of time shifted samples for a specified (proposed) time shift, which is evaluated using the second similarity measure. For example, the second similarity measure can be more computationally complex than the first measure of similarity. This concept is useful, because during scaling operations it is usually necessary to recalculate the first similarity measure (to determine the "proposed" time shift between the first sample block and the second sample block selected from the multiple possible time shift values between the first sample block and the second sample block). In contrast, the second similarity measure usually only requires a single calculation during the time shift operation, e.g. as a "final" quality check, where it is judged whether or not the proposed "time offset" determined using the first (less computationally complex) measure can be expected. quality will ensure that you get enough good audio quality. As a result, it is still possible to avoid applying and adding, if the first similarity measure indicates satisfactory good (or at least sufficient) similarity between the first block of samples (or part of it) and the time-shifted second block of samples (or part of it) for the "proposed" time shift, the second (usually more reliable or accurate) ) a measure of quality indicates that time scaling will not result in sufficiently good audio quality. The use of quality control (using the second similarity measure) thus facilitates the avoidance of audible distortions in time scaling.
[0062] For example, the first similarity measure may be cross-correlated or normalized cross-correlation, function of difference in mean values, or sum of square errors. Such similarity measures can be obtained in a computationally efficient manner and are sufficient to find a "best fit" between the first sample block (or portion thereof) and (time shifted) second sample block (or portion thereof), i.e. to determine the "proposed" offset time. In contrast, the second similarity measure may for example be a combination of multiple cross-correlations or normalized cross-correlations values for a plurality of different time shifts. This measure of similarity ensures greater accuracy and makes it easier to include additional signal components (e.g. such as harmonics) or the immutability of the audio signal when evaluating the (expected) quality of the time scaling result. The second measure of similarity, however, is more computationally demanding than the first measure of similarity, so using the second similarity measure when looking for a "proposed" time shift would be computationally inefficient.
[0063] Hereinafter, some options related to determining the second similarity measure will be described. In some embodiments, the second similarity measure may be a combination of cross correlations involving at least four different time shifts. For example, the second similarity measure may be a combination of a first cross correlation value and a second cross correlation value that have been obtained with respect to time shifts that differ from each other by the total multiple of the base period of the audio content contained in the first sample block or second sample block, and the third value of cross-correlation and the fourth value of cross-correlation, that have been obtained with respect to time shifts that differ from each other by the total multiple of the period of the fundamental frequency of the audio content. The time offset for which the first cross correlation value is obtained may differ from the time shift from which the third cross correlation value is obtained by an odd multiple of the half-period of the fundamental frequency of the audio content. If the audio content (represented by the input audio signal) is substantially immutable and a fundamental frequency dominates in it, it can be expected that the first cross correlation value and the second cross correlation value, which may be normalized, for example, will be similar to each other. Because the third cross-correlation value and the fourth cross-correlation value are obtained in relation to time shifts that differ by an odd multiple of the half-period of the fundamental frequency from the time shifts for which the first cross-correlation value is obtained and the second cross-correlation value, it can be expected that the third cross-correlation value and the fourth cross-correlation value will be opposite to the first cross-correlation value and the second cross-correlation value, if the audio content is essentially immutable and the fundamental frequency dominates in it. It is therefore possible to obtain a useful combination based on the first cross-correlation value, the second cross-correlation value,
[0064] It should be noted that particularly useful measures of similarity can be obtained by calculating a similarity measure q according to the formula q = c (p) * c (2 * p) + c (3/2 * p) * c (1/2 * p) or according to the formula q = c (p) * c (-p) + c (-1 / 2 * p) * c (1/2 * p).
[0065] In the above formulas, c (p) is the cross-correlation value between the first sample block (or portion thereof) and the second sample block (or portion thereof) that are time shifted (e.g., relative to the original time position in the input audio content) for a period p the fundamental frequency of the audio content contained in the first block of samples and / or in the second block of samples (wherein the base frequency of the audio content is usually substantially the same in the first block of samples and in the second block of samples). Speaking in other words, the cross-correlation value is calculated based on the blocks of samples from the input audio content and additionally time shifted from each other by the p-period of the basic frequency of the audio input (whereby the frequency p of the base frequency can be obtained, for example, by estimating the fundamental frequency, autocorrelation or similar way). Similarly, c (2 * p) is the cross-correlation value between the first block of samples (or part of it) and the second block of samples (or part of it) that are shifted by 2 * p. Similar definitions also apply to c (3/2) * p), c (1/2 * p), c (-p) ic (-1 / 2 * p), where argument c (.) means a time shift.
[0066] In the following, some decision-making mechanisms will be explained as to whether time scaling should be performed or not, which may optionally be implemented in the time scaling module 200. In the implementation, the time counter 200 may be configured to compare a quality value based on computing or estimating (expected) quality of the time scaled version of the input audio signal, which may be obtained by time scaling, with a variable decision-making value, or time scaling should be carried out or not. Deciding whether to perform time scaling or therefore can not be dependent on conditions, such as history representing previous time scaling.
[0067] For example, the time counter may be configured to reduce the variable limit value to thereby reduce the required quality (which must be achieved to allow for time scaling) when it is determined that the quality of the time scaling result was insufficient for one or more previous blocks of samples. Thus, it was ensured that time scaling is not prevented with respect to a long sequence of frames (or blocks of samples), which could cause buffer overflow or buffer underflow. In addition, the timer may be configured to increase the variable limit value in order to thereby increase the required quality (which must be achieved to enable time scaling) if it is found that that time scaling has been applied to one or more previous blocks or samples. It is therefore possible to prevent the use of time scaling of too many consecutive blocks or samples, if it is possible to achieve a very good quality (increased in relation to the normal quality required) of the time scaling result. It is possible to avoid the occurrence of artifacts that could occur in the case of too small quality requirements of the time scaling result. if it is possible to achieve a very good quality (increased in relation to the normal quality required) of the time scaling result. It is possible to avoid the occurrence of artifacts that could occur in the case of too small quality requirements of the time scaling result. if it is possible to achieve a very good quality (increased in relation to the normal quality required) of the time scaling result. It is possible to avoid the occurrence of artifacts that could occur in the case of too small quality requirements of the time scaling result.
[0068] In some embodiments, the timer may include a first limited range counter to count the number of sample blocks or the number of frames subjected to time scaling, as the corresponding quality requirements of the time scaled version of the input audio signal are met, which can be obtained by time scaling. . In addition, the timer may also include a second limited range counter to count the number of sample blocks or the number of frames that have not been subjected to time scaling, since the corresponding quality requirements for the time scaled version of the audio input signal that can be obtained by scaling have not been met. time. In this case, the timer can be configured as to calculate the variable limit value based on the value of the first counter and based on the value of the second counter. In order to ensure moderate computing expenditure, it is possible to include the "history" of time scaling (as well as the history of "quality").
[0069] For example, the timer may be configured to add to the initial limit value a value proportional to the value of the first counter and subtract from it a value proportional to the value of the second counter (e.g., from the result of the addition) to obtain a variable limit value. .
[0070] Some essential functionalities that may be included in some embodiments of the time scaling module 200 will be summarized below. It should be noted, however, that the functionalities described below are not basic functionalities of the time scaling module 200.
[0071] In an implementation, the timer may be configured to perform time scaling of the input audio signal based on the result of calculating or estimating the quality of the time scaled version of the input audio signal that can be obtained by time scaling. In this case, the calculation or estimation of the quality of the time scaled version of the audio input signal that can be obtained by time scaling includes computing or estimating artifacts in a time scaled version of the input audio signal, which may be a result of time scaling. It should be noted, however, that the calculation or estimation of artifacts can be carried out indirectly, for example by calculating the quality of the result of the overlap and add operations. Speaking in other words,
[0072] For example, the time counter may be configured to calculate or estimate the quality of the time scaled version of the input audio signal that can be obtained by time scaling of the input audio signal, which is based on the similarity level of successive (and optionally overlapping) blocks of audio input signal samples.
[0073] In a preferred embodiment, the timer may be configured to calculate or estimate whether there are audible artifacts in the time scaled version of the input audio signal that can be obtained by time scaling of the input audio signal. As mentioned above, the estimation of audible artifacts can be carried out indirectly.
[0074] By using quality control, time scaling can be performed at times well suited for time scaling and avoided at times that are not well suited for time scaling. For example, the time counter may be configured to postpone the time scaling to the next frame or to the next sample block if the result of calculating or estimating the quality of the time scaled version of the input audio signal that can be obtained by time scaling indicates insufficient quality (e.g. lower than the specified quality limit value). Thus, time scaling can be carried out at a time that is better suited for time scaling, resulting in fewer artifacts (in particular audible artifacts) being produced. Speaking in other words,
[0075] In summary, the timer 200 can be refined in a number of different ways, as discussed above.
[0076] It should further be noted that the time counter 200 may optionally be connected to the control 100 of the synchronization buffer, wherein the chopper buffer control 100 may ensure that a time scaling based on samples, which is usually performed by the user, should be performed. a timer of 200 times, or time scaling based on frames.
5.3. Audio signal decoder shown in Fig. 3 [0077] Fig. 3 is a block diagram of an audio signal decoder 300 according to an embodiment of the present invention.
[0078] The audio signal decoder 300 is configured to receive input audio content 310, which may be treated as an input audio representation and may for example be represented in the form of audio frames. The audio signal decoder 300 further provides decoded audio content 312, which may for example be represented in the form of decoded audio samples. The audio signal decoder 300 may, for example, comprise a desynchronization buffer 320, which is configured to receive the input audio content 310, e.g. in the form of audio frames. Out-of-sync buffer 320 is configured to buffer multiple audio frames representing blocks of audio samples (wherein a single frame may represent one or more blocks of audio samples, and audio samples represented by a single frame can be logically divided into many overlapping or non-overlapping blocks of audio samples). Out-of-sync buffer 320 also provides a "buffered" audio frame 322, wherein the audio frames 322 may include both audio frames contained in the input audio content 310, and audio frames generated or inserted by the synchronization buffer (e.g., such as "inactive" audio frames containing signaling information that indicates the generation of comfortable noise). The audio signal decoder 300 further includes a decoder core 330, which receives the buffered 322 audio frames originating from the desynchronization buffer 320 and provides sample 332 audio (e.g. blocks of audio samples associated with audio frames) obtained from the audio frames 322 coming from the out-of-sync buffer. The audio signal decoder 300 further comprises a time scaling module 340 based on samples configured to receive the audio samples 332 provided by the decoder core 330 and provides the time-based audio scales 342 forming the decoded audio content 312. The sample based scaling module 340 is configured to provide time scaled audio samples (e.g. in the form of blocks of audio samples) obtained from the audio samples 332 (i.e., based on blocks of audio samples provided by the decoder core). The audio signal decoder may further include an optional control 350. The control 350 of the desynchronization buffer used in the decoder 300 of the audio signal may be, for example, the same as the control 100 of the synchronization buffer shown in Fig. 1. In other words, the control 350 of the synchronization buffer may be configured as to ensure that frame-based time scaling is performed by a spool buffer 320 or time scaling based on a sample performed by the time scaling module 340 based on samples, which is done in a manner adapted to the signal. Therefore, input 310 audio content or information regarding the input audio content 310 may be input to control 350 of the synchronization buffer. which are an audio signal 110 or information about the audio signal 110. Controlling buffer 35 may also provide chopping buffer 320 (such as described with respect to control 100 of the synchronization buffer) to buffer 320, and control 350 may provide time scaling module 140 based on samples of control information 114 such as described in with respect to the control 100 of the synchronization buffer. The desynchronization buffer 320 may thus be configured to omit or insert audio frames to perform frame based time scaling. The decoder core 330 may further be configured to generate a comfortable noise for a frame containing signaling information, which indicate the generation of comfortable noise. Convenient noise can therefore be generated by the decoder core 330 in the case of inserting a "inactive" frame (containing signaling information that indicates the need to generate a comfortable noise) to the de-synchronization buffer 320. In other words, a simple form of frame-based time scaling may result in generating a frame containing a comfortable noise, which is initiated by inserting an "inactive" frame into the out-of-sync buffer (which may be performed in the presence of control information 112 provided by the scratch buffer control). The decoder core may further be configured to perform "masking" in the case of an empty desynchronization buffer. Such masking may include generating audio information with respect to a "missing" frame (empty out-of-sync buffer) performed based on audio information regarding one or more frames preceding the missing audio frame. For example, it is possible to use predictions with the assumption that the audio content contained in the missing audio frame is a "continuation" of the audio content contained in one or more frames preceding the missing audio frame. However, any masking concepts for frame loss known in the art can be used in the decoder core. The control 350 of the desynchronization buffer may thus provide for the desynchronization buffer (or decoder core 320) to initiate masking commands in the event of the 320 synchronization buffer being emptied.
[0079] It should further be noted that the time scaling module 340 based on the samples may be the same as the time counter 200 described with reference to Figure 2. The input audio signal 210 may thus correspond to the audio samples 332 and the time scaled version 212 of the input signal. the audio may correspond to temporally scaled 342 audio samples. The time scaling module 340 can thus be configured to perform time scaling of the input audio signal based on the result of calculating or estimating the quality of the time scaled version of the input audio signal that can be obtained by time scaling. The operation of the time scaling module 340 based on the samples may be controlled by the control 350 of the synchronization buffer, wherein the control information 114 provided by the scratch buffer control to the time scaling module 340 based on the samples may indicate whether the time scaling based on the samples should be performed or not. In addition, the control information 114 may, for example, indicate the desired amount of time scaling performed by the time scaling module 340 based on samples.
[0080] It should be noted that the time scaling module 300 may be supplemented by any of the features and functionalities described with respect to the scramble buffer control 100 and / or with respect to the time scaling module 200. In addition, the audio signal decoder 300 may also be supplemented by any of the described features and functionalities, e.g. described with reference to Fig. 4 to
15.
5.4. Audio signal decoder shown in Fig. 4 [0081] Fig. 4 is a block diagram of an audio signal decoder 400 according to an embodiment of the present invention. The audio signal decoder 400 is configured to receive packets 410 that may include a packet representation of one or more audio frames. The audio signal decoder 400 further provides the decoded audio content 412, e.g. in the form of audio samples. For example, the audio samples may be represented in the "PCM" format (i.e. in a coded form using pulse code modulation, e.g. in the form of digital sequences of values representing samples of the waveform of the audio signal).
[0082] The audio signal decoder 400 includes a depow- ning module 420 configured to receive packets 410 and provides obtaining depressed frames 422. The depalletizer module is further configured to obtain from the packets 410 a so-called & quot; SID flag & quot; indicating & quot; inactive & quot; an audio frame (i.e., an audio frame for which generation of comfortable noise should be used instead of "normal" detailed decoding of audio content). The information about the SID tag is designated 424. The depacketing module also provides a real-time transport time stamp (also marked as "RTP TS") and an arrival time stamp (also designated as "arrival TS"). The timestamp information was marked 426. The audio signal decoder 400 further comprises a disjunction clearing buffer 430 (also referred to shortly as a desynchronization buffer 430) that receives the deprived frames 422 originating from the deprancet module 420 and provides to the core 440 of the decoder buffered frames 432 (and optionally also inserted frames). The de-synchronization buffer 430 also receives control information 434 from scaling (timing) based on frames from the control logic. The desynchronization buffer 430 also provides feedback scaling information 436 used to estimate the playback delay. The audio signal decoder 400 also includes a time scaling module 450 (also designated as "TSM"), which receives decoded audio samples 442 (e.g. in the form of encoded data using pulse code modulation) derived from the decoder core 440, wherein the decoder core 440 provides the decoded 442 audio samples obtained from buffered or inserted frames 432 originating from buffer 430 removing out-of-sync . The time scaling module 450 also receives from control logic control information 444 for scaling (time) based on the samples o providing feedback scaling information 446 used to estimate the playback delay. The time scaling module 450 also provides time scaled samples 448, which may represent scaled-time audio content in a form encoded using pulse-code modulation. Audio decoder 400 also includes a PCM 460 buffer receiving time scaled samples 448 and buffering time scaled samples 448. Buffer 460 PCM also provides a buffered version of time scaled samples 448 representing a decoded 412 audio content. The PCM buffer 460 may also provide delay control information 462 to the control logic.
[0083] The audio signal decoder 400 also includes a target delay 470 estimation to which information 424 (e.g., SID tag) is inputted, as well as time signature information 426 including an RTP time stamp and an arrival time stamp. Based on this information, estimating the target delay 470 provides the target delay information 472, which describes the desired delay, e.g. a desired delay, which should be entered by the disabling buffer 430, the decoder 440, the 450 scaler module.and a 460 PCM buffer. For example, estimating the target delay 470 may provide for calculating or estimating the target delay information 472 in such a way that a delay not too large but sufficient to compensate for some disjointing of the packets 410. Furthermore, the audio signal decoder 400 includes a latitude estimation 480 of the reproduction delay configured so, that the scaling feedback information 436 is input to it from a buffer 430 that removes scattering and feedback scaling information 446 from the time scaling module 460. The scaling information 436 may, for example, describe time scaling performed by the disabling buffer. Also, feedback scaling information 446 describes time scaling performed by the time scaling module 450. With respect to feedback scaling information 446, it should be noted that the time scaling performed by the time scaling module 450 is typically adjusted to the signal, and thus the actual scaling described by the feedback scale information 446 may differ from the desired time scaling that may be described. by information 444 on scaling based on samples. In summary, feedback scaling information 436 and scaling information 446 may describe actual time scaling,
[0084] The audio signal decoder 400 further comprises a control logic 490 that provides (main) control of the audio signal decoder. Control logic 490 receives information 424 (e.g., SID tag) derived from deprancing module 420. In addition, control logic 490 receives the target delay information 472 from the target delay estimation 470, playback delay information 482 derived from estimating playback delay 480 (wherein 482 regarding the reproduction delay describes the actual delay obtained by estimating 480 the playback delay based on feedback scaling information 436 and feedback scaling information 446). Also, control logic 490 (optional) receives delay information 462 derived from buffer 460 (PCM) (wherein, alternatively, delay information derived from the PCM buffer may have a predetermined amount). Based on the received information, the control logic 490 provides the scattering information 434 for scaling based on frames and scattering information 442 to the buffer 430 that removes the scattering and the scaling module 450. The control logic therefore generates frame-based scaling information 434 and scattering information 442 based on one or more audio content properties (e.g., such as whether there is an "inactive" frame for which it should be carried out the generation of comfortable noise,
It should be noted that the control logic 490 may implement some or all of the scratch buffer control functionality 100, the information 424 may correspond to the audio signal information 110, the control information 112 may correspond to the frame based scaling information 434, and the control information 114 may correspond to information 444 regarding scaling based on samples. It should also be noted that the time scaling module 450 may implement some or all functionalities of the time scaling module 200 (or vice versa), wherein the input audio signal 210 corresponds to the decoded 442 audio samples, and the time scaled version 212 of the audio input signal corresponds to time scaled samples 448.
[0086] It should further be noted that the audio signal decoder 400 corresponds to the decoder 300 of the audio signal, and thus the audio signal decoder 300 can implement some or all of the functionalities described with respect to the audio signal decoder 400 and vice versa. The out-of-synchronization buffer 320 corresponds to the de-synchronizing buffer 430, the decoder core 330 corresponds to the decoder core 440, and the temporal scaler module 340 corresponds to the time scaling module 450. Control 350 corresponds to control logic 490.
[0087] Some additional details regarding the functionality of the audio signal decoder 400 will be shown below. In particular, the proposed management of the synchronization buffer (JBM) will be described.
[0088] Described herein is a desynchronization buffer management (JBM) solution that can be used to deliver to the decoder 440 received packets 410 with frames containing coded speech or audio data while maintaining continuous reproduction. In a packet data connection, e.g. using voice over the Internet (VoIP) protocol, packets (e.g., packets 410) are typically subjected to variable transmission times and are lost during transmission, leading to out-of-sync of arrival times and the occurrence of missing packets in a receiver (e.g. a receiver comprising an audio signal decoder 400). In order to allow continuous delivery of the output signal without interruptions,
[0089] The following is an overview of the solution. For the described management of the desynchronization buffer, the encoded data contained in the received RTP packets (e.g., packets 410) are first depended (e.g. using deprancate module 420), and frames (e.g., frames 422) obtained with encoded data (e.g. voice included in the frame encoded using AMR-WB) are delivered to the buffer that removes the jitter (e.g., buffer 430 removing the out-of-sync). When new data encoded using pulse modulation (PCM data) is required for reproduction, it must be made available by the decoder (for example, by the decoder 440). For this purpose, frames (e.g., frames 432) are retrieved from a buffer that removes the synchronization (e.g., buffer 430 removing the out-of-sync). Thanks to the use of the buffer to remove the synchronization, it is possible to compensate for the arrival time fluctuations. In order to control the buffer depth, the time scale (TSM) modification is used (while modifying the time scale is also called short time scaling). Modification of the time scale may be performed based on coded frames (e.g. in buffer 430 removing out-of-sync) or in a separate module (e.g. in time scaling module 450), allowing more accurate matching of the PCM output signal (e.g., PCM output 448 or output signal) 412 PCM signal). Thanks to the use of the buffer to remove the synchronization, it is possible to compensate for the arrival time fluctuations. In order to control the buffer depth, the time scale (TSM) modification is used (while modifying the time scale is also called short time scaling). Modification of the time scale may be performed based on coded frames (e.g. in buffer 430 removing out-of-sync) or in a separate module (e.g. in time scaling module 450), allowing more accurate matching of the PCM output signal (e.g., PCM output 448 or output signal) 412 PCM signal). Thanks to the use of the buffer to remove the synchronization, it is possible to compensate for the arrival time fluctuations. In order to control the buffer depth, the time scale (TSM) modification is used (while modifying the time scale is also called short time scaling). Modification of the time scale may be performed based on coded frames (e.g. in buffer 430 removing out-of-sync) or in a separate module (e.g. in time scaling module 450), allowing more accurate matching of the PCM output signal (e.g., PCM output 448 or output signal) 412 PCM signal). the time scale (TSM) modification is used (while modifying the time scale is also called short time scaling). Modification of the time scale may be performed based on coded frames (e.g. in buffer 430 removing out-of-sync) or in a separate module (e.g. in time scaling module 450), allowing more accurate matching of the PCM output signal (e.g., PCM output 448 or output signal) 412 PCM signal). the time scale (TSM) modification is used (while modifying the time scale is also called short time scaling). Modification of the time scale may be performed based on coded frames (e.g. in buffer 430 removing out-of-sync) or in a separate module (e.g. in time scaling module 450), allowing more accurate matching of the PCM output signal (e.g., PCM output 448 or output signal) 412 PCM signal).
[0090] The concept described above is shown, for example, in Fig. 4, where a review of the management of the synchronization buffer is visible. In order to provide buffer depth control for the de-synchronization (e.g., buffer 430 removing disjoint) and / or the TSM (e.g. included in the time scaling module 450), control logic is used (e.g., control logic 490, which is assisted by delay estimation 470). and estimating 480 playback delay). It uses information about the target delay (e.g., information 472), playback delays (e.g., information 482) and whether discontinuous transmission (DTX) is currently used in combination with the generation of comfortable noise (CNG) (e.g., information 424).
5.4.1. Depaxiser module [0091] A depaketizer module 420 will be described below. The depaketizer module distributes the RTP packets 410 into single frames (access units) 422. It also calculates the RTP timestamp for all frames that are neither the only frame nor the first frame in the packet. . For example, the time stamp included in the RTP packet is assigned to the first frame. In the case of aggregation (i.e. in the case of RTP packets containing more than a single frame), the time stamp of the next frames is increased by a duration divided by the scale of the RTP time stamps. In addition to the RTP timestamp, each frame is also labeled with the system time in which the RTP packet has been received ("arrival time stamp"). As you can see, information about the RTP timestamp and information about the arrival time stamp can be provided e.g. to estimate the target delay 470. The depaketizer module also determines whether the frame is active or contains a silence insertion descriptor (SID). It should be noted that in non-active periods, only SID frames are received in some cases. Therefore, information 424 is provided to the control logic 490, which may, for example, include a SID tag.
5.4.2. Buffer removing out-of-sync. [0092] In a buffer 430 removing out-of-sync, frames 422 received via a network (e.g. TCP / IP networks) are stored until they are decoded (e.g. by the decoder 440). The frames 422 are queued according to the increasing values of the RTP timestamp to remove the change of order that could have occurred in the network. The frame at the beginning of the queue can be inserted into the decoder 440 and then removed (e.g. from the buffer 430 removing the out-of-sync). If the queue is empty or, in accordance with the difference between the timestamps of the frame at the beginning (queue) and the one previously read, the frame is missing,
[0093] In other words, the decoder 440 can be configured to generate a comfortable noise in the case of a signaling in the signal frame that a comfortable noise should be used, which is performed, for example, using the active "SID" tag. On the other hand, the decoder may also be configured to mask packet loss, e.g. by providing prediction (or extrapolated) audio samples in the event that the previous (last) frame was active (i.e., the generation of comfortable noise has been turned off). , while the synchronization buffer has been emptied (as a result of which the buffer 430 is provided by the buffer 430 removing the un-synchronized empty frame).
[0094] The de-synchronizing buffer 430 also supports frame-based time scaling by adding at the beginning (e.g., a synchronization buffer) an empty frame to achieve temporal stretching or skip a frame at the beginning (e.g., a synchronization buffer queue) to obtain time compression. In the case of inactive periods, the de-synchronizing buffer may behave as if the "NO_DATA" frames were added or suppressed.
5.4.3. Modifying the time scale (TSM) [0095] The following describes the modification of the time scale (TSM), which is also referred to herein briefly as a time scaling module or a time scaling module based on samples. In order to modify the time scale (called short time scaling), the packet-based modified WSOLA algorithm (overlap-summation based on waveform similarity) is used (see, for example [Lia01]) with built-in quality control. Some details can be seen, for example, in Fig. 9, which will be explained below. The level of time scaling depends on the signal: signals causing the creation of strong artifacts during scaling are detected by quality control, while low level signals that are close to silence, they are scaled to the greatest extent possible. Signals well suited for time scaling, such as periodic signals, are scaled using the internally generated offset. The offset is obtained on the basis of a similarity measure such as a normalized cross-correlation. In the case of overlap-adding (OLA), the end of the current frame (also referred to herein as the "second sample block") is moved (e.g. relative to the beginning of the current frame, which is also referred to herein as the "first sample block") to shorten or extend the frame.
[0096] As mentioned above, additional details regarding modifying the time scale (TSM) will be described below with reference to Fig. 9, which shows a modified WSOLA algorithm with quality control, as well as with reference to Figs. 10a, 10b and 11.
5.4.4. PCM Buffer [0097] The PCM buffer will be described below. The time scale modification module 450 changes the duration of the PCM frames provided by the decoder module using a time-varying scale. For example, with respect to the audio frame 432, it is possible for the decoder to provide 440 1024 samples (or 2048 samples). In contrast, with respect to the audio frame 432, the time scaling module 450 can provide a varying number of audio samples, which is related to time scaling based on samples. In contrast, in the sound card of the loudspeaker (or generally speaking in a sound-emitting device), a constant framing, e.g. of 20 ms, is normally expected.
[0098] Considering the entire chain, such PCM buffer 460 does not result in an additional delay. On the other hand, the delay is divided into a buffer 430 removing the synchronization and buffer 460 PCM. Nevertheless, the goal is to keep as few samples as possible stored in PCM buffer 460, because this increases the number of frames stored in buffer 430 removing out-of-sync, thus reducing the likelihood of late loss (when the decoder masks the missing frame that is received later).
[0099] The program pseudocode shown in Fig. 5 implements a PCM buffer level control algorithm. As can be seen in the program pseudocode shown in Fig. 5, the size of the soundcard frame ("soundCardFrameSize") is calculated based on the sampling frequency ("sampleRate"), whereby, for example, a frame duration of 20 ms is assumed. The number of samples per frame of the sound card is therefore known. The PCM buffer is then filled by decoding the 432 audio frames (also referred to as "accessUnit") until the number of samples in the PCM buffer ("pcmBuffer_nReadabteSamples ()") is no longer less than the number of samples per frame of the card soundcard ("soundCardFrameSize"). First, the frame (also referred to as "accessUnit") obtained (or requested) is from the staging buffer 430 as indicated by the reference numeral 510. Next, the "frame" of the audio samples is obtained by decoding the frame 432 requested from the desynchronization buffer, as indicated by reference numeral 512 A frame of decoded audio samples (e.g., designated 442) is obtained. Next, the frame of the decoded 442 audio samples undergoes modifying the time scale, whereby a "frame" of time scaled 448 audio samples is obtained, as indicated by the reference numeral 514. It should be noted that the frame of time scaled audio samples may contain a larger number of audio samples or a smaller number audio samples than the frame of the decoded 442 audio samples inserted into the time scaling module 450. Next,
[0100] This procedure is repeated until sufficient (time-scaled) audio samples are available in PCM buffer 460. As soon as there are enough (time-scaled) audio samples in the PCM buffer, a "frame" of time-scaled audio samples (having the frame length required by a sound reproducing device such as a sound card) is read from the PCM 460 buffer and sent to the sound reproducing device ( for example to a sound card), as indicated by reference numerals 520 and 522.
5.4.5. Estimating the target delay [0101] In the following, estimating the target delay that can be performed by the target delay estimating module 470 will be described. The target delay describes the desired buffering delay between the reproduction moment of the previous frame and the moment in which the frame could be received in the smallest delay of network transmission compared to all frames that are currently included in the history of the target delay estimating module 470. To estimate the target delay, two different modules are used to estimate synchronization, one module estimating long-term out-of-sync and one module estimating short-term out-of-sync.
Long-term decompression estimation module [0102] To calculate long-term out-of-sync, it is possible to use a FIFO data structure. The time range stored in the FIFO structure may differ from the number of saved positions if DTX (discontinuous transmission mode) is used. For this reason, the size of the FIFO structure is limited in two ways. It may contain at most 500 items (corresponding to 10 seconds at 50 packets per second) and the time range (difference between the RTP time stamps of the latest and oldest packet) of at most 10 seconds). If more items are needed, the oldest item is deleted. For each RTP packet received via the network, an entry is added to the FIFO structure. The item has three values: delay, offset and RTP time stamp. These values are calculated based on the time of receipt (represented for example by the time stamp received) and the RTP time stamp of the RTP packet, as is evident in the pseudo-code shown in Fig. 6.
[0103] As indicated by reference numbers 610 and 612, the time difference between the RTP time stamps of two packets (e.g., consecutive packets) is calculated (which leads to "rtpTimeDiff"), and the difference between the time stamps of receiving two packets is calculated. (for example, subsequent packages) (which leads to "rcvTimeDiff"). The timestamp is further transformed from the time base of the sending device to the time base of the receiving device as indicated by reference numeral 614 and leads to "rtpTimeTicks". Similarly, differences in RTP times (differences between RTP timestamps) are converted to the time scale / time base of the receiving device, as indicated by reference numeral 616, and leads to "rtpTimeDiff".
[0104] Next, the delay information ("delay") is updated based on the earlier delay information as indicated by reference numeral 618. For example, if the difference between the reception times (i.e. the difference between packet reception times) is greater than the difference in RTP times. (ie the difference between the times of sending packets), it can be concluded that the delay has increased. In addition, the offset time information ("offset") is computed, as indicated by the reference numeral 620, with the information regarding the offset time the difference between the receiving moment (i.e., the moment the packet has been received) and the moment at which the packet was sent (specified by the RTP timestamp transformed to the time scale of the receiver).
[0105] Next, some current information is recorded as "previous" information to be used in the next iteration, as indicated by reference numeral 624.
[0106] Long-time synchronization can be calculated as the difference between the maximum delay currently stored in the FIFO structure and the minimum delay value:
longTermJitter = longTermFifo_getMaxDelay () - longTermFifo_getMinDelay ();
Estimating short-time out-of-sync. [0107] The estimation of short-time out-synchronization will be described below. The estimation of short-time out-synchronization is carried out, for example, in two stages. In the first stage, the same disaggregation calculations are performed as those used for estimating long-term outgrowth, with the following modifications: the size of the FIFO structure window is limited to at most 50 positions, and the time range up to a maximum of 1 second. The resulting decoupling value is calculated as the difference between the 94% of the delay value currently stored in the FIFA structure (the three largest values are omitted) and the minimum delay value:
shortTermJitterTmp = shortTermFifo1_getPercentileDelay (94) shortTermFifo1_getMinDelay ();
[0108] In the second stage, different shifts between the short-term and long-term structure of the FIFO are compensated, which leads to the following result:
shortTermJitterTmp + = shortTermFifo1_getMinOffset ();
shortTermJitterTmp - = longTermFifo1_getMinOffset ();
[0109] The result is added to a different FIFO structure with a window size of at most 200 positions and a time range of at most four seconds. The maximum value stored in the FIFO structure is finally increased to the total multiple of the frame size and used as short-term out-of-sync:
shortTermFifo2_add (shortTermJitterTmp);
shortTermJitter = ceil (shortTermFifo2_getMax () / 20.f) * 20;
Estimating the target delay by combining the short-term / long-term decay estimation To calculate the target delay (e.g., target delay information 472) the long-term and short-term outgrowth estimates (e.g. referred to as "longTermJitter" and "shortTimeJitter" above) are combined with each other in different ways, depending on the current state. For active signals (or portions of the signal for which comfortable noise generation is not used), a range is used as the target delay (for example, determined by "targetMin" and "targetMax"). During DTX and when the DTX action is started, two different values are calculated as the target delay (e.g. "targetDtx" and "targetStartUp").
[0111] Details of how the different target delay values can be calculated are shown in Fig. 7. As indicated by reference numerals 710 and 712, the "targetMin" and "targetMax" values providing range assignment to active signals are calculated based on "shortTimeJitter" and long-term out-synchronization ("longTermJitter"). Calculation of the target delay during DTX (targetDtx & quot;) is indicated by numeral reference 714, and calculation of the target delay with respect to the start of operation (e.g. after DTX ("targetStartUp") is designated with the reference numeral 716.
5.4.6. Estimating the reproduction delay [0112] Hereinafter, estimating the reproduction delay that can be performed by the module 480 estimating the playback delay will be described. The playback delay determines the buffering delay between the reproduction of the previous frame and the moment in which the frame could be received in the smallest possible delay of transmission over the network as compared to all frames currently included in the target delay estimation history. It is calculated in milliseconds using the following formula:
playoutDelay = prevPlayoutOffset - longTermFifo_getMinOffset () + pcmBufferDelay;
[0113] The variable "prevPlayoutOffset" is recalculated each time the received frame is taken from the scramble removal module 430, which is done using the current system time expressed in milliseconds and the RTP timestamp transformed into milliseconds:
prevPlayoutOffset = sysTime - rtpTimestamp [0114] In order to prevent that the "prevPlayoutOffset" variable becomes out-of-date if the frame is not available, the variable is updated for frame-based time scaling. In the case of time-based stretching based on frames, the variable "prevPlayoutOffset" is increased by the duration of the frame, while in the case of time-based frame-based compression, the variable "prevPlayoutOffset" is reduced by the duration of the frame. The "pcmBufferDelay" variable describes the buffered time in the PCM buffer module.
5.4.7 Control logic [0115] The control logic (e.g. control logic 490) will be described in detail below. It should be noted, however, that the control logic 800 shown in Fig. 8 can be supplemented by any of the features and functionalities described with respect to the de-synchronizing buffer 100 and vice versa. It should further be noted that the control logic 800 may replace the control logic 490 shown in Fig. 4, but may optionally include additional features and functionalities. It is further not required that all the features and functionalities described above with reference to Fig. 4 are also included in the control logic800 shown in Fig. 8 and vice versa.
[0116] Fig. 8 is a flowchart of the control logic 800, which can of course also be implemented in hardware.
[0117] The control logic 800 includes a download 810 of the frame to be decoded. In other words, a frame to be decoded is selected and it is determined how this decoding should be performed. In check 814, a check is performed to see if the previous frame (e.g., the previous frame preceding the frame downloaded for decoding in step 810) was active or not. If it is found in check 814 that the previous frame was inactive, the first decision path (branch) 820 is selected which is used to adjust the inactive signal. Conversely, if it is found in check 814 that the previous frame was active, a second decision path (branch) 830 is selected which is used to adjust the active signal. The first decision path 820 includes determining the "blank" value performed in step 840, the value of 840 describing the difference between the playback delay and the target delay. The first decision path 820 further comprises making a decision 850 about the time scaling operation performed based on the gap value. The second decision path 830 includes selecting 860 time scaling based on whether the actual playback delay is included in the target delay range.
[0118] Hereinafter, additional details regarding the first decision path 820 and the second decision path 830 will be described.
[0119] In step 840 of the first decision path, check 842 checks whether the next frame is active. For example, check 842 may be to check whether the frame downloaded for decoding in step 810 is active or not.
Alternatively, check 842 may be to check whether the frame following the frame taken for decoding in step 810 is active or not. If it is found in check 842 that the next frame is not active or the next frame is not yet available, the variable "gap" assumes in step 844 the difference between the actual reproduction delay (determined by the variable "playoutDelay") and the target delay DTX (represented by the variable "targetDtx") described above in the chapter "Estimating the target delay". On the contrary, if it is found in check 840 that the next frame is active,
[0120] In step 850, first check whether the size of the variable "gap" is greater than the limit value (or the same). This is done in check 852. If it is determined that the size of the variable "gap" is less than the limit value (or the same), time scaling is not performed. Conversely, if it is found in check 852 that the size of the variable "gap" is greater than the limit value (or the same as the limit values, depending on the implementation), a decision is made that time scaling is required. Check 854 checks if the value of the variable "gap" is positive or negative (ie whether the variable "gap" is greater than zero or not). If it is determined that the value of the variable "interval" is not greater than zero (ie it is negative), the frame is inserted into the de-synchronizing buffer (frame-based time extension performed in step 856), and therefore frame-based time scaling is performed. This may for example be signaled by the frame-based scaling information 434. Conversely, if it is found that the value of the variable "spacing" is greater than zero, i.e. it is positive, the frame from the de-synchronizing buffer is skipped (time-based compression based on the frames carried out in step 856), and thus time scaling is performed in based on frames. This can be signaled by the frame-based scaling information 434. thus, time scaling based on frames is performed. This may for example be signaled by the frame-based scaling information 434. Conversely, if it is found that the value of the variable "spacing" is greater than zero, i.e. it is positive, the frame from the de-synchronizing buffer is skipped (time-based compression based on the frames carried out in step 856), and thus time scaling is performed in based on frames. This can be signaled by the frame-based scaling information 434. thus, time scaling based on frames is performed. This may for example be signaled by the frame-based scaling information 434. Conversely, if it is found that the value of the variable "spacing" is greater than zero, i.e. it is positive, the frame from the de-synchronizing buffer is skipped (time-based compression based on the frames carried out in step 856), and thus time scaling is performed in based on frames. This can be signaled by the frame-based scaling information 434. the frame from the de-synchronizing buffer is skipped (frame-based time compression performed in step 856), and therefore frame-based time scaling is performed. This can be signaled by the frame-based scaling information 434. the frame from the de-synchronizing buffer is skipped (frame-based time compression performed in step 856), and therefore frame-based time scaling is performed. This can be signaled by the frame-based scaling information 434.
[0121] In the following, the second decision branch 860 is described. In check 862, a check is performed to determine whether the playback delay is greater (or the same) than the maximum target (i.e. the upper limit of the target range) which is described, for example, by the variable & quot; targetMax ". If it is determined that the reproduction delay is greater than the maximum target value (or the same), the time scaling module 450 performs time compression (step 866, time-based compression based on samples using TSM), and thus time scaling is performed based on sample. This may for example be signaled by the scaling information 444 based on samples. If, however, check 862 is found, that the reproduction delay is less than the maximum target delay (or the same), check 864 is performed in which it checks if the playback delay is lower (or the same) than the minimum target value which is described, for example, by the variable "targetMin" . In the event that the reproduction delay is found to be less than the minimum target value (or the same), the time scaling module 450 performs temporal stretching (step 866, time stretching based on samples using TSM), and thus time scaling is performed based on sample. This may for example be signaled by the scaling information 444 based on samples. If, however, check 864 is found,
[0122] In summary, the control logic module (also referred to as the logic controlling the synchronization buffer management) shown in Fig. 8 compares the real delay (playback delay) with the desired delay (target delay). In the case of a significant difference, it initiates time scaling. During a comfortable noise (for example, when the SID tag is active), time scaling based on frames is initiated, which is carried out by the buffer module that removes the synchronization. During active periods, time scaling is initiated based on samples that are performed by the TSM module.
[0123] Fig. 12 shows an example of estimating target and playback delay. The cutoff 1210 of the graphic representation 1200 describes the time, and the elevation 1212 of the graphic representation 1200 describes the delay expressed in milliseconds. The "targetMin" and "targetMax" series create a range of the desired delay desired by the module estimating the target delay after disarming the windowed network. The "playoutDelay" latency usually remains in the range, but the adjustment may be slightly delayed due to the time scale being adapted to the signal.
[0124] Fig. 13 shows the operations associated with a time scale performed in the path shown in Fig. 12. The cutoff 1310 of the graphic representation 1300 describes the time expressed in seconds, and the ordinate 1312 describes time scaling expressed in milliseconds. In the 1300 graphic representation, positive values mean time stretching, while negative values mean time compression. During the impulse, both buffers are immediately emptied, and for stretching (plus 20 milliseconds at 35 seconds) one masked frame is inserted. For all other types of customization, it is possible to use the time scaling method based on samples that provide higher quality, which results in varying scales caused by the solution adapted to the signal.
[0125] In summary, the target delay is dynamically adjusted in the event of a desynchronization increase (as well as in the case of decreasing synchronization) in a certain window. When the target delay is increased or decreased, time scaling is usually performed, with the decision on the type of scaling being made in a manner adapted to the signal. In the event that the current frame (or previous frame) is active, time scaling based on samples is performed, where the actual time-based scaling delay based on the samples is adapted in a manner adapted to the signal to reduce the occurrence of artifacts. When using time scaling based on samples, there is usually no constant degree of time scaling.
5.8. Modification of the time scale shown in Fig. 9 [0126] In the following, referring to Fig. 9, details regarding modifying the time scale will be described. It should be noted that the modification of the time scale has been briefly described in chapter 5.4.3. The modification of the time scale, which can e.g. be performed by the time scaling module 150, will however be described in more detail below.
[0127] Fig. 9 is a flowchart of a modified WSOLA algorithm with quality control according to an embodiment of the present invention. It should be noted that the time scaling 900 shown in Fig. 9 can be supplemented by any of the features and functionalities described with respect to the time scaling module 200 shown in Fig. 2 and vice versa. It should further be noted that the time scaling 900 shown in Fig. 9 may correspond to the time scaling 340 module based on the samples shown in Fig. 3 and the time scaling module 450 depicted in Fig. 4. The time scaling 900 shown in Fig. 9 may further be replaced by time scaling 866 based on samples.
[0128] For time scaling (or a time scaling module or a time scale modifying module) 900, decoded samples 910 (audio) are input, e.g. in a form encoded using pulse code modulation (PCM). The decoded samples 910 may correspond to the decoded samples 442, the audio samples 332, or the audio input signal 210. Also, control information 912 is input to the time scaling module 900, which may correspond, for example, to information 444 on scaling based on samples. Control information 912 may, for example, describe a target scale and / or a minimum frame size (e.g., a minimum number of frame samples for 448 audio samples to be delivered to the 460 PCM buffer). The time scaling module 900 includes a switch (or dial) 920 that assumes, based on the scale information of the target decision, whether time compression should be performed, whether time stretching should be performed or no time scaling should be performed. For example, switching (or checking or selecting) 920 may be based on scaling information 444 based on samples received from control logic 490.
[0129] If it is determined from the target scale information that no time scaling should be performed, the received decoded samples 910 are sent in unmodified form to the output of the time scaling module 900. For example, the decoded samples 910 are sent in unmodified form to PCM buffer 460 as "time scaled" samples 448.
[0130] In the following, the processing will be described in the case of time-compression (which can be determined using check 920 based on the target scale information 912). In the case where time compression is desired, a calculation of 930 energy is carried out. During this 930 energy calculation, the energy of the sample block (e.g. a frame containing a given number of samples) is calculated. After calculating 930 energy, the selection (or switching or checking) is carried out 936. If it is determined that the energy value 932 provided by calculating 930 energy is greater (or the same) than the energy limit value (e.g. the energy limit Y), the first processing path 940 is selected comprising a tailored to the signal determining the degree of time scaling in time scaling based on samples. Conversely, if it is determined that the energy value 932 provided by calculating 930 energy is less (or the same) than the energy limit value (e.g., the energy limit Y), a second processing path 960 is selected, for which in time scaling in a constant degree of temporal scaling is used based on the samples. In the first processing path 940, in which the degree of time scaling is determined in a manner adapted to the signal, a similarity estimation 642 is made from the audio samples. In 642 similarity estimation, it is possible to include information 944 regarding the minimum frame size and may provide information 946 regarding the highest similarity (or regarding the location of the highest similarity). In other words, in 642 similarity estimation, it is possible to determine which position (e.g. which position of the samples in the sample block) is best suited for performing the overlap operation and summation with time compression.
The maximum similarity information 946 is sent to the quality control 950 in which it is calculated or estimated whether the overlap and aggregation operation using the maximum similarity information 946 will result in an audio quality (or the same) higher than the quality limit X (which can be constant or can be variable). If the quality control 950 determines that the quality of the result of the overlap and additive (or equivalent time-scaled version of the input audio signal that can be obtained by the overlap and add operation) will be smaller (or the same) than the quality limit X, time scaling is omitted, and the time scaling module 900 provides unscaled audio samples. On the contrary, if 950 quality control is found, that the quality of the result of the overlapping and aggregation operations using information 946 regarding the highest similarity (or regarding the location of greatest similarity) will be greater than the quality limit X or the same, an overlap operation 954 is performed, the offset used in the overlap and summation operation described is by information 946 regarding the highest similarity (or regarding the location of the highest similarity). The overlap and add operation thus provides a scaled block (or frame) of the audio samples. an overlap and additive operation 954 is performed, the offset used in the overlapping and summation operations is described by information 946 regarding the highest similarity (or regarding the location of the highest similarity). The overlap and add operation thus provides a scaled block (or frame) of the audio samples. an overlap and additive operation 954 is performed, the offset used in the overlapping and summation operations is described by information 946 regarding the highest similarity (or regarding the location of the highest similarity). The overlap and add operation thus provides a scaled block (or frame) of the audio samples.
[0131] The block (or frame) of time scaled audio samples 956 may, for example, correspond to time scaled samples 448. Similarly, a block (or frame) of un-scaled samples of 952 audio that is provided when a quality control 950 is found to be achievable the quality would be less than the quality limit X or the same, it may correspond to "time scaled" samples 448 (in which case, in fact, there is no time scaling).
[0132] Conversely, when it is found in check 936 that the energy of the block (or frame) of the input audio samples 910 is smaller than the energy limit Y (or the same), an overlap and add operation 962 is performed, the offset used in the overlap and add operation is determined by the minimum frame size (described by the minimum frame size information) and a block (or frame) of scaled 964 audio samples is obtained that may correspond to time scaled samples 448.
[0133] It should further be noted that the processing performed in the case of temporal stretching is analogous to the processing carried out during time compression, with modified similarity estimation, and overlapping and addition.
[0134] In conclusion, it should be noted that in the case of selecting time-compression or time stretching, three different cases are distinguished in a time-dependent or time-warped scaling based on the samples. If the energy of the block (or frame) of the input audio samples is a relatively low energy (e.g. less than the energy limit Y or the same), the time stretch or time compression operation by overlapping and summation is carried out with a constant time shift (i.e. with a constant rate) time compression or time stretching). On the contrary, if the energy of the block (or frame) of the input audio samples is greater than the energy limit Y (or the same), "Optimal" (sometimes also referred to as "proposed") the degree of time stretching or time compression is determined by estimating similarity (estimating 942 similarity). In the next quality control step, it is determined whether such an overlap and additive operation using a predetermined "optimal" degree of time stretching or time compression would allow obtaining sufficient quality. If it is found that sufficient quality can be achieved, the application and addition operation is carried out using a certain "optimal" degree of time stretching or time compression. On the contrary, if it is found that
[0135] Some further details will be described below regarding the time scaling-adapted quality that can be performed by the time scaling module 900 (or by the timer 200 or by the time scaling module 340 or by the time scaling module 450). Time scaling methods using overlap and summation (OLA) are widely available, but generally no adjustment of the time scaling results to the signal is performed. In the described solution, which can be used in the time scaling modules described herein, the degree of time scaling depends not only on the similarity obtained by estimating similarity (e.g. by estimating similarity of 942) position, which is considered optimal in the context of high quality time scaling, but also from the expected quality of the overlap and summation result (e.g. from overlapping and adding 954). In the time scaling module (e.g. in the time scaling module 900 or other temporal scaling modules described herein), two quality control steps have been introduced to decide whether time scaling would result in audible artifacts. In the case of the potential occurrence of artifacts, time scaling is postponed until they are less audible. whether time scaling would result in audible artifacts. In the case of the potential occurrence of artifacts, time scaling is postponed until they are less audible. whether time scaling would result in audible artifacts. In the case of the potential occurrence of artifacts, time scaling is postponed until they are less audible.
[0136] In the first quality control step, a targeted quality measure is calculated, which is done using the position p obtained using a similarity measure (e.g., 942 similarity estimation) as input data. In the case of a periodic signal, p is the basic frequency of the current frame. For the positions p, 2 * p, 3/2 * pi 1/2 * p, the normalized cross-correlation c () is calculated. It is expected that c (p) will be a positive value and c (1/2 * p) may be a positive or a negative value. In the case of harmonic signals, a positive c sign (2p) is also expected, and the c (3/2 * p) sign should be the same as the c sign (1/2 * p). This relationship can be used to create an intentional quality measure:
q = c (p) * c (2 * p) + c (3/2 * p) * c (1/2 * p).
[0137] The range of values q is [-2; 2]. The ideal harmonic signal would give a result of q = 2, while very dynamic signals and containing a wide bandwidth that could cause audible artifacts during time scaling result in a lower value. Due to the fact that time scaling is carried out frame by frame, the entire signal used to calculate c (2 * p) and c (3/2 * p) may not be available yet. The estimation can, however, also be carried out using the past samples. It is therefore possible to use c (-p) instead of c (2 * p), like using c (-1 / 2 * p) instead of c (3/2 * p).
[0138] In the second quality control step, the current value of the targeted quality measure q is compared with the dynamic minimum quality qMin value (which may correspond to the quality limit value X) to determine whether time scaling should be performed with respect to the current frame.
[0139] There are various reasons for using a dynamic minimum quality value: if q is of low value, because the signal has been rated as unreachable for a long time, the qMin value should be slowly decreased to ensure that the expected scaling will continue to be performed at that at the very moment with lower expected quality. On the other hand, signals with a high q value should not result in scaling multiple frames in a row, which could cause a decrease in quality in the context of long-term signal characteristics (e.g. rhythm).
[0140] To calculate the dynamic minimum value qMin of the quality (which can be for example equivalent to the quality limit X), the following formula is therefore used:
qMin = qMinInitial - (nNotScaled * 0,1) + (nScaled * 0.2) qMinInitial is a configuration value that provides optimization between specified quality and delay if the frame can be scaled with the desired quality, which value of 1 is a good compromise. nNotScaled is a frame counter that has not been scaled due to insufficient quality (q <qMin). nScaled is a frame counter that has been scaled for achieving the required quality (q> = qMin). The range of both counters is limited: their value is not reduced to negative values, nor is it increased above the marked value, which by default is 4 (for example).
[0141] The current frame is time scaled in position p, if q> = qMin, otherwise the time scaling is deferred to the next frame for which this condition is satisfied. The pseudocode shown in Fig. 11 illustrates quality control in time scaling.
[0142] As can be seen, the initial value of qMin is 1, wherein said initial value is designated as "qMinInitial" (see reference numeral 1110). Similarly, the maximum value of the nScaled counter (denoted as "variable qualityRise") assumes a value of 4 initially, as indicated by the reference numeral 1112. The maximum value of the counter nNotScaled is initially set to 4 (the variable "qualityRed"), see reference numeral 1114. Next, information is obtained about position p, which is done using a similarity measure and is indicated by the reference numeral 1116. Then the quality value q is calculated in relation to the position described by the position p value, in accordance with the formula indicated by numeral reference 1116. The qMin quality limit value is calculated from the qMinInitial variable as well as the nNotScaled and nScaled counter variables as numbered 1118. As you can see, the initial qMinInitial value of the qMin quality limit is reduced by a value proportional to the nNotScaled counter value and increased by a value proportional to the value of the nScaled counter. As it is visible, the maximum values of the counter variables nNotScaled and nScaled also determine the maximum increase in the limit value qMin and the maximum decrease in the limit value qMin. Then a check is made to see if the quality q value is greater than the qMin quality limit or the same as indicated by the reference numeral 1120. as well as the nNotScaled and nScaled counter variables, as indicated by reference numeral 1118. As can be seen, the initial value of qMinInitial regarding the qMin quality limit is reduced by a value proportional to the nNotScaled counter value and incremented by a value proportional to the nScaled counter value. As it is visible, the maximum values of the counter variables nNotScaled and nScaled also determine the maximum increase in the limit value qMin and the maximum decrease in the limit value qMin. Then a check is made to see if the quality q value is greater than the qMin quality limit or the same as indicated by the reference numeral 1120. as well as the nNotScaled and nScaled counter variables, as indicated by reference numeral 1118. As can be seen, the initial value of qMinInitial regarding the qMin quality limit is reduced by a value proportional to the nNotScaled counter value and incremented by a value proportional to the nScaled counter value. As it is visible, the maximum values of the counter variables nNotScaled and nScaled also determine the maximum increase in the limit value qMin and the maximum decrease in the limit value qMin. Then a check is made to see if the quality q value is greater than the qMin quality limit or the same as indicated by the reference numeral 1120. As can be seen, the initial value of qMinInitial regarding the qMin quality limit value is reduced by a value proportional to the nNotScaled counter value and incremented by a value proportional to the nScaled counter value. As it is visible, the maximum values of the counter variables nNotScaled and nScaled also determine the maximum increase in the limit value qMin and the maximum decrease in the limit value qMin. Then a check is made to see if the quality q value is greater than the qMin quality limit or the same as indicated by the reference numeral 1120. As can be seen, the initial value of qMinInitial regarding the qMin quality limit value is reduced by a value proportional to the nNotScaled counter value and incremented by a value proportional to the nScaled counter value. As it is visible, the maximum values of the counter variables nNotScaled and nScaled also determine the maximum increase in the limit value qMin and the maximum decrease in the limit value qMin. Then a check is made to see if the quality q value is greater than the qMin quality limit or the same as indicated by the reference numeral 1120. the initial qMinInitial value for the qMin quality limit value is reduced by a value proportional to the nNotScaled counter value and incremented by a value proportional to the nScaled counter value. As it is visible, the maximum values of the counter variables nNotScaled and nScaled also determine the maximum increase in the limit value qMin and the maximum decrease in the limit value qMin. Then a check is made to see if the quality q value is greater than the qMin quality limit or the same as indicated by the reference numeral 1120. the initial qMinInitial value for the qMin quality limit value is reduced by a value proportional to the nNotScaled counter value and incremented by a value proportional to the nScaled counter value. As it is visible, the maximum values of the counter variables nNotScaled and nScaled also determine the maximum increase in the limit value qMin and the maximum decrease in the limit value qMin. Then a check is made to see if the quality q value is greater than the qMin quality limit or the same as indicated by the reference numeral 1120.
[0143] If so, an overlap and add operation is performed as indicated by reference numeral 1122. In addition, the value of the counter variable nNotScaled is reduced, whereby it is ensured here that the value of said counter variable will not be negative. In addition, the value of the nScaled counter variable is increased, whereby it is ensured here that the nScaled value will not exceed the upper limit defined by the variable (or constant) qualityRise. Adjustment of counter variables is indicated by numerals 1124 and 1126.
[0144] In contrast, in the case of determining in a check marked with the reference numeral 1120 that the value in quality is lower than the limit value qMin, the overlapping and adding operations are skipped, the value of the counter variable nNotScaled is increased taking into account that the value The nNotScaled counter variable did not exceed the upper limit specified by the variable (or constant) qualityRed, and the value of the nScaled counter variable is reduced taking into account that the value of the counter variable nScaled did not become negative. The adjustment of counter variables in the case when the quality is insufficient is indicated by numerals 1128 and 1130.
5.9. The time counter shown in Figs. 10a and 10b. [0145] In the following, referring to Figs. 10 and 10b, a module adapted to a temporal scaling signal will be explained. Figs. 10 and 10b show a flow chart adapted to a time scaling signal. It should be noted that the time-scaling adapted to the signal, such as shown in Figs. 10a and 10b, can be used, for example, in a time scaling module 200, a time scaling module 340, a time scaling module 450 or a time scaling module 900.
[0146] The time scaling module 1000 shown in Figs. 10a and 10b includes energy calculation 1010 in which the energy of the frame (or part or block) of the audio samples is calculated. Energy calculation 1010 may, for example, correspond to energy calculation 930. Then, a check 1014 is performed in which it is checked whether the energy value obtained in the energy calculation 1010 is greater (or the same) than the energy limit value (which may be, for example, a constant energy limit value). If it is found in check 1014 that the energy value obtained in calculating energy 1010 is smaller (or the same) than the energy limit value, it can be assumed that it is possible to obtain sufficient quality using the overlap and add operations, and the overlap and add operation is performed in step 1018 with the maximum time shift (for maximum time scaling). Conversely, when it is found in check 1014 that the energy value obtained in the energy calculation 1010 is not less (or the same) than the energy limit, a search is performed for the best fit of the master segment in the search area, which is done using a similarity measure. The similarity measure can be, for example, a cross-correlation, a normalized cross-correlation, a function of the difference in mean values, or the sum of square errors. Here are some details about such a search for the best match, and the way that
[0147] In the following, reference is made to a graphic representation indicated by reference numeral 1040. The first representation 1042 represents a block (or frame) of samples that begins at time t1 and ends at time t2. As it can be seen, the block of samples beginning at the moment t1 and ending at the moment t2 can be divided logically into the first block of samples, which begins at the moment t1 and ends at the moment t3, and the second block of samples that begins at the moment t4 and ends at time t2.
The second block of samples is then temporally shifted with respect to the first sample block, as indicated by reference number 1044. As a result of the first time shift, the time-shifted second block of samples starts at, for example, at time t4 'and ends at time t2'. Between the moments t4 'and t3 there is thus a temporary overlap of the first block of samples and the time-shifted second block of samples. However, it can be seen that there is no good fit (i.e., there is not much similarity) between the first sample block and the time-shifted second sample block, e.g. in the overlap region between the t4 'and t3 moments (or in the aforementioned overlap region between the t4 moments) and t3). In other words, the timer can, for example, move the second block of samples, as indicated by reference numeral 1044, and determine the measure of similarity with respect to the overlap region (or with respect to a part of the overlap region) between the moments t4 'and t3. The time counter may also use an additional time shift with respect to the second sample block, as indicated by reference numeral 1046, whereby the (twice) time-shifted version of the second sample block starts at time t4 '' and ends at time t2 '' (where t2 ''> t2 '> t2, like t4' '> t4'> t4). The time counter may also specify (quantitative) similarity information representing the similarity between the first sample block and the time-shifted version of the second sample block, e.g. between t4 '' and t3 (or, for example, in a part located between the moments t4 '' and t3). The time counter therefore evaluates for which time offset the time-shifted version of the second block of samples the similarity in the area of overlap with the first sample block is maximum (or at least greater than the limit value). It is therefore possible to determine the time shift resulting in the "best fit", i.e. maximize (or at least obtain a large enough) similarity between the first sample block and the time-shifted version of the second sample block. If there is sufficient similarity between the first block of samples and the time shifted twice by the version of the second block of samples in the overlap region (e.g. between the times t4 '' and t3), then it can be expected with the level of reliability determined by the similarity measure used, that the overlap and summation operation ensuring overlap and summation of the first block of samples and the time-shifted version of the second block of samples will provide an audio signal that does not contain significant audible artifacts. It should further be noted that overlapping and summation between the first sample block and the time-shifted version of the second sample block results in a portion of the audio signal having a temporal stretch between the t1 and t2 "moments, which is not greater than the" original "audio signal stretching from t1 to t2.
[0148] In a similar manner, it is possible to achieve temporal compression, as will be explained with reference to the graphic representation indicated by reference numeral 1050. As can be seen, reference numeral 1052 indicates a primary block (or frame) of samples extending between moments t11 and t12. The primary sample block (or frame) can be divided, for example, into a first sample block that extends from the moment t11 to the instant t13, and a second sample block that extends from the moment t13 to the moment t12. The second block of samples is time-shifted to the left, as indicated by reference numeral 1054. As a result, (once) time-shifted version of the second block of samples starts at time t13 'and ends at time t12'. There is also a time overlap of the first block of samples and a time-shifted version of the second block of samples between the moments t13 'and t13. However, the time counter may specify (quantitative) similarity information representing the similarity between the first block of samples and the (once) time-shifted version of the second block of samples between the moments t13 'and t13 (or part of the time between the moments t13' and t13) and state that the similarity does not it is particularly good. The timer may also perform a second time shift of the second sample block to thereby obtain a time-shifted version of the second sample block, as indicated by reference numeral 1056, which starts at time t13 '' and ends at time t12 ''. There is therefore a time overlap of the first block of samples and (twice) time-shifted version of the second block of samples between the moments t13 '' and t13. Using the time scaling module, it can be concluded that the (quantitative) similarity information indicates a high similarity between the first sample block and the time-shifted version of the second sample block between the t13 '' and t13 moments. With the use of the time scaling module, it can therefore be concluded that the operation of overlapping and summation of the first block of samples and the time-shifted version of the second block of samples can be carried out with good quality and fewer audible artifacts (at least with the reliability provided by the similarity measure used).1058. The time-shifted version of the second block of samples may start at the moment t13 '' 'and end at the moment t12' ''. The time-shifted version of the second block of samples, however, may not show good similarity to the first block of samples in the overlap region between the moments t13 '' and t13, because the time shift was not appropriate. As a result, the timer can determine that the time-shifted version of the second block of samples provides the best fit (greatest similarity in the overlap region and / or around the overlap region and / or part of the overlap region) to the first block of samples. The time counter may thus perform overlapping and summation of the first block of samples and the time-shifted version of the second block of samples, if the additional quality check (which may be based on the second, more significant similarity) indicates obtaining sufficient quality. As a result of the overlap and add operations, a combined sample block extends from the moment t11 to the moment t12 '', which is shorter in time than the original sample block extending from the moment t11 to the moment t12. It is therefore possible to perform time compression. As a result of the overlap and add operations, a combined sample block extends from the moment t11 to the moment t12 '', which is shorter in time than the original sample block extending from the moment t11 to the moment t12. It is therefore possible to perform time compression. As a result of the overlap and add operations, a combined sample block extends from the moment t11 to the moment t12 '', which is shorter in time than the original sample block extending from the moment t11 to the moment t12. It is therefore possible to perform time compression.
[0149] It should be noted that the above functionalities, which have been described with reference to the graphical representations denoted by reference numerals 1040 and 1050, can be performed by searching 1030, which provides information about the location of the highest similarity obtained by searching for the best match (where information or value that describes the location of the highest similarity are also marked here as p). The similarity between the first block of samples and the time-shifted version of the second block of samples in the respective overlap regions can be determined using cross-correlation, using a normalized cross-correlation, using the function of difference in mean sizes or using the sum of square errors.
[0150] After determining the information regarding the location of the highest similarity (p), the matching quality 1060 is performed with respect to the identified location (p) of the highest similarity. These calculations may be carried out, for example, in the manner indicated in Fig. 11 by the reference numeral 1116. In other words, (quantitative) information on the quality of the match (which may for example be designated as q) may be calculated using a combination of four correlation values that can be obtained in relation to various time shifts (e.g., time shifts p, 2 * p, 3/2 * pi<sup>1</sup>/ 2 * p). It is therefore possible to obtain (quantitative) information (q) representing the quality of the match.
[0151] With reference to Fig. 10b, check 1064 is performed in which the quantitative information q representing q describing the quality of the match is compared to the qMin quality limit value. This check or comparison 1064 can provide an assessment of whether the match quality represented by the variable q is greater than the variable limit value qMin (or the same). If it is found in check 1064 that the match quality is sufficient (i.e., greater than a variable or similar quality limit value), an overlap and summation operation is applied (step 1068) using the highest similarity position (which is described for example by the variable p ). An overlap and summation operation is therefore performed, e.g. between the first block of samples and the time-shifted version of the second block of samples, which results in obtaining the "best fit" (ie the largest value of similarity information). Regarding the details, reference should be made, for example, to the explanations given with reference to the graphic representations 1040 and 1050. The application of overlapping and adding is also indicated by reference numeral 1122 in Fig. 11. In step 1072, a frame counter update is also performed. For example, the counter variable "nNotScaled" and the counter variable "nScaled" are updated, for example as described with reference to Fig. 11 and numerals 1124 and 1126. On the contrary, if it is found in check 1064 that the match quality is insufficient (for example, less than the variable limit value qmin (or the same)),
[0152] In summary, the time scaling module 1000, the functionality of which is described with reference to Figs. 10a and 10b showing the operation method, can perform time scaling based on samples using a quality control mechanism (steps 1060 to 1064).
5.10. The method shown in Fig. 14 [0153] Fig. 14 illustrates a method of controlling a method of controlling the delivery of a decoded audio content based on the input audio content. The method 1400 shown in Fig. 14 includes selecting time scaling based on frames or time scaling based on frames in a manner adapted to the signal.
[0154] It should further be noted that the method 1400 can be supplemented by any of the features and functionalities described herein, e.g. with respect to the control of the synchronization buffer.
5.11. The method shown in Fig. 15 [0155] Fig. 15 is a block diagram of a method 1500 for providing a time-scaled version of an input audio signal. The method includes calculating or estimating the quality of the time scaled version of the input audio signal that can be obtained by time scaling of the input audio signal. The method 1500 further comprises performing 1520 time scaling of the input audio signal based on calculating or estimating the quality of the time scaled version of the input audio signal that can be obtained by time scaling.
[0156] The method 1500 may be supplemented by any of the features and functionalities described herein, e.g. with respect to the time scaling module.
6. Summary [0157] In summary, the embodiments of the invention are a method for managing a desynchronization buffer and a speech and sound transmission device with high quality. The method and device can be used in conjunction with transmission codecs such as MPEG ELD, AMR-WB or codecs that will be developed in the future. In other words, embodiments of the invention provide a method for compensating for the un-synchronized reception in packet transmission.
[0158] Embodiments of the invention may be used, for example, in a technology called "3GPP EVS".
[0159] In the following, some aspects of an embodiment of the invention will be briefly described.
[0160] The solution for managing the out-of-sync buffer described herein is a system comprising a number of modules that are available and combined in the manner described above. It should further be noted that the aspects of the invention also apply to the properties of the modules themselves.
[0161] An important aspect of the present invention is a signal-adapted method for time scaling in adaptive management of a synchronization buffer. In the described solution, time scaling based on frames and time scaling based on samples are combined in the control logic, thanks to which a combination of the advantages of both methods was obtained. Available methods of time scaling are:
• Inserting / removing comfortable noise in DTX;
• Overlaying and summation (OLA) without correlation in the case of low signal energy (for example in the case of frames containing low signal energy);
• WSOLA algorithm for active signals;
• Inserting a masked frame to stretch with an empty desynchronization buffer.
[0162] The solution described here describes a mechanism for combining frame-based methods (inserting and removing comfortable noise and inserting masked frames for extension) with sample-based methods (WSOLA algorithm for active signals and non-synchronized overlap and adding (OLA) for signals) low energy). Figure 8 illustrates a control logic that selects the optimal technology for modifying the time scale according to an embodiment of the invention.
[0163] According to a further described aspect, a number of targets for adaptive management of the synchronization buffer are used. In the described solution, many optimization criteria are used when estimating the target delay used to calculate a single target reproduction delay. These criteria result in different objectives, primarily optimized for high quality or low latency.
[0164] Many of the purposes of calculating the target playback delay are:
• Quality: avoiding late loss (evaluation of the synchronization buffer);
• Delay: delay limit (evaluation of the synchronization buffer).
[0165] An (optional) aspect of the described solution is to optimize target delay estimation, thereby delaying the delay while avoiding late losses, and also keeping the small reserve in synchronization buffer to increase the likelihood of interpolation that is intended to allow high-quality decoding in the decoder. error masking.
[0166] Another (optional) aspect is recovering during masking of TCX with respect to late frames. For most current solutions, frame management of late frame buffer is rejected. Mechanisms for using late frames in ACELP-based decoders have been described [Lef03]. According to an aspect, such a mechanism is also used for frames other than ACELP frames, e.g. frames coded in the frequency domain such as TCX, which is intended to generally facilitate recovery of the decoder state. Frames that were received late and were already masked are still delivered to the decoder in order to improve recovery of the decoder state.
[0167] Another important aspect of the present invention is the quality-adjusted time scaling which has been described above.
[0168] Still summarizing, embodiments of the present invention provide a complete solution for the management of the synchronization buffer that can be used to provide the user with a better experience in the case of packet transmission. It has been observed that the presented solutions work better than any known solutions for managing the synchronization buffer known to the inventors.
7. Alternative implementations [0169] Certain aspects have been described in the context of the device, however, it is evident that these aspects also represent a description of the corresponding method, wherein the block or device corresponds to a method step or process properties of the method step. Analogously, the aspects described in the context of the method step also represent a description of the corresponding block, element or property of the respective device. Some or all of the method steps may be performed by a hardware device (or using it), e.g. a microprocessor, a programmable computer or an electronic circuit. In some embodiments of the invention, some or the most important steps of the method may be performed using such a device.
[0170] The encoded audio signal according to the invention may be recorded on a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium, e.g. a network
Internet.
[0171] Depending on some implementation requirements, embodiments of the invention may be implemented in hardware or in software. The implementation can also be carried out using a digital data carrier, in particular a floppy disk, DVD disc, Blue-Ray disc, CD, ROM memory, PROM memory, EPROM memory, EEPROM memory or FLASH memory, on which readable electronic signals are stored controls that are used (or are able to be used) by a programmable computer system, which ensures that the corresponding method is performed. The digital storage medium can thus provide the ability to read it using a computer.
[0172] Some embodiments of the invention are a data carrier including readable electronic control signals that can be used by a programmable computer system, which ensures carrying out one of the methods described herein.
[0173] Embodiments of the invention may generally be implemented as a computer program product comprising program code, wherein the program code allows one of the methods described herein to be performed when the computer program product is executed using a computer. The program code can be stored, for example, on a medium that can be machine readable.
[0174] Other embodiments are a computer program that allows one of the methods described herein to be performed and stored on a medium that can be machine readable.
[0175] In other words, the embodiment of the method according to the invention is thus a computer program comprising a program code, the program code making it possible to carry out one of the methods described herein, when the computer program is executed using a computer.
[0176] A further embodiment of the methods according to the invention is thus a data carrier (or a digital storage medium or a medium that can be read using a computer) comprising a computer program stored on it for carrying out one of the methods described herein. The data medium, the digital storage medium or the stored medium is usually real and / or non-volatile.
[0177] A further embodiment of the method according to the invention is thus a data stream or a sequence of signals representing a computer program to perform one of the methods described herein. The data stream or the sequence of signals can be formatted, for example, in such a way that they can be transmitted over a telecommunications connection, e.g. via the Internet.
[0178] A further embodiment is a processing device, e.g. a computer or programmable logic, configured or adapted to perform one of the methods described herein.
[0179] A further embodiment is a computer on which a computer program is installed to perform one of the methods described herein.
[0180] A further embodiment is a device or system configured to provide transmission to the receiver (e.g., electronically or optically) of a computer program to perform one of the methods described herein. For example, the receiver may be in the form of a computer, a mobile device, memory, etc. The device or system may for example be a file server for transferring a computer program to a receiver.
[0181] In some embodiments of the invention, some or all of the functionality of the described methods may be provided by a programmable logic (e.g., a user-programmed logic). In some embodiments of the invention, the user-programmed logic can cooperate with a microprocessor to perform one of the methods described herein. Generally speaking, the methods are preferably carried out using a hardware device.
[0182] The device described herein may be implemented using a hardware device, using a computer or a combination of a hardware device and a computer.
[0183] The methods described herein may be performed using a hardware device, using a computer or a combination of a hardware device and a computer [0184] The embodiments described above are merely illustrative of the concept of the present invention. It is to be understood that those skilled in the art will notice modifications and variations of the solutions and details described herein. The invention is therefore only limited by the scope of the appended claims, and not by the specific details set out to provide a description and explanation of the embodiment.
References [0185] [Lia01] YJ Liang, N. Faerber, B. Girod: "Adaptive playout scheduling using time-scale modification in packet voice communications", 2001 [Lef03] P. Gournay, F. Rousseau, R. Lefebvre: " Improved packet loss recovery using 15 late frames for predictionbased speech coders, 2003 [AHEVS-044]: FRAUNHOFER GESELLSCHAFT: "On Jitter Buffer Management in the Design Constraints", 3GPP DRAFT; AHEVS-044, 3RD GENERATION PARTNERSHIP PROJECT (3GPP), MOBILE COMPETENCE CENTER; 650, ROUTE DES LUCIOLES; F-06921 SOPHIA20 ANTIPOLIS CEDEX; FRANCE, vol. SA WG4, no. San Diego; May 10 201, XP050527151, disclosing control of a synchronization buffer to control the delivery of audio content. This document discloses an adaptation module, which takes into account the buffer status and input data from the network analysis module to control time scaling. Time scaling can be performed as time-based frame-based time scaling or time-based scaling based on samples.
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung EV, Germany Plenipotentiary:
EP 3 011 692 B1 Z - 15937/17
93 members in 19 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 13173159 | European Patent Office (EPO) | A | |
| 13173159 | European Patent Office (EPO) | A | |
| 14167061 | European Patent Office (EPO) | A | |
| 14167061 | European Patent Office (EPO) | A | |
| 14731262 | European Patent Office (EPO) | A | |
| 13173159 | – | – | – |
| 14167061 | – | – | – |
| 147312623 | – | – | – |
| EP20130173159 | – | – | – |
| EP20140167061 | – | – | – |
| EP20140731262 | – | – | – |
Members93
| Document | Office | Kind | |
|---|---|---|---|
| CA2916121A1 | Canada | A1 | |
| CA2916126A1 | Canada | A1 | |
| CA2964362A1 | Canada | A1 | |
| CA2964368A1 | Canada | A1 | |
| WO2014202647A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014202672A2 | World Intellectual Property Organization (WIPO) | A2 | |
| TW201517025A | Taiwan Province of China | A | |
| TW201517026A | Taiwan Province of China | A | |
| WO2014202672A3 | World Intellectual Property Organization (WIPO) | A3 | |
| AR096688A1 | Argentina | A1 | |
| AR096690A1 | Argentina | A1 | |
| SG11201510459YA | Singapore | A | |
| SG11201510501YA | Singapore | A | |
| AU2014283256A1 | Australia | A1 | |
| AU2014283320A1 | Australia | A1 | |
| KR20160021886A | Republic of Korea | A | |
| KR20160023830A | Republic of Korea | A | |
| CN105474313A | China | A | |
| MX2015017364A | Mexico | A | |
| MX2015017831A | Mexico | A | |
| CN105518778A | China | A | |
| EP3011564A2 | European Patent Office (EPO) | A2 | |
| EP3011692A1 | European Patent Office (EPO) | A1 | |
| US2016171990A1 | United States of America | A1 | |
| US2016180857A1 | United States of America | A1 | |
| JP2016527540A | Japan | A | |
| AU2014283320B2 | Australia | B2 | |
| JP2016529536A | Japan | A | |
| TWI581257B | Taiwan Province of China | B | |
| TWI582759B | Taiwan Province of China | B | |
| EP3011692B1 | European Patent Office (EPO) | B1 | |
| BR112015032174A2 | Brazil | A2 | |
| RU2016101339A | Russian Federation | A | |
| RU2016101580A | Russian Federation | A | |
| AU2017204613A1 | Australia | A1 | |
| HK1223727A | Hong Kong, China | A | |
| HK1223727A1 | Hong Kong, China | A1 | |
| HK1224447A | Hong Kong, China | A | |
| HK1224447A1 | Hong Kong, China | A1 | |
| AU2014283256B2 | Australia | B2 | |
| PT3011692T | Portugal | T | |
| ES2642352T3 | Spain | T3 | |
| PL3011692T3This record | Poland | T3 | |
| MX352748B | Mexico | B | |
| JP6251464B2 | Japan | B2 | |
| SG10201708531PA | Singapore | A | |
| EP3011564B1 | European Patent Office (EPO) | B1 | |
| JP6317436B2 | Japan | B2 | |
| MX355850B | Mexico | B | |
| PT3011564T | Portugal | T | |
| ES2667823T3 | Spain | T3 | |
| EP3321934A1 | European Patent Office (EPO) | A1 | |
| EP3321935A1 | European Patent Office (EPO) | A1 | |
| US9997167B2 | United States of America | B2 | |
| US2018190302A1 | United States of America | A1 | |
| RU2662683C2 | Russian Federation | C2 | |
| PL3011564T3 | Poland | T3 | |
| RU2663361C2 | Russian Federation | C2 | |
| CA2916121C | Canada | C | |
| US10204640B2 | United States of America | B2 | |
| AU2017204613B2 | Australia | B2 | |
| KR101952192B1 | Republic of Korea | B1 | |
| KR101953613B1 | Republic of Korea | B1 | |
| US2019147901A1 | United States of America | A1 | |
| EP3321935B1 | European Patent Office (EPO) | B1 | |
| CA2916126C | Canada | C | |
| HK1255499A | Hong Kong, China | A | |
| HK1255499A1 | Hong Kong, China | A1 | |
| MY170699A | Malaysia | A | |
| CN105474313B | China | B | |
| CN110211603A | China | A | |
| PT3321935T | Portugal | T | |
| CN105518778B | China | B | |
| MY171256A | Malaysia | A | |
| PL3321935T3 | Poland | T3 | |
| ES2739481T3 | Spain | T3 | |
| CA2964362C | Canada | C | |
| CA2964368C | Canada | C | |
| BR112015031825A2 | Brazil | A2 | |
| US10714106B2 | United States of America | B2 | |
| HK1255429B | Hong Kong, China | B | |
| US2020321014A1 | United States of America | A1 | |
| BR112015032174B1 | Brazil | B1 | |
| US10984817B2 | United States of America | B2 | |
| US2021233553A1 | United States of America | A1 | |
| BR112015031825B1 | Brazil | B1 | |
| US11580997B2 | United States of America | B2 | |
| CN110211603B | China | B | |
| EP3321934B1 | European Patent Office (EPO) | B1 | |
| EP3321934C0 | European Patent Office (EPO) | C0 | |
| US12020721B2 | United States of America | B2 | |
| PL3321934T3 | Poland | T3 | |
| ES2979208T3 | Spain | T3 |
Numbers
- Publication
- 3011692
- Publication, DOCDB
- 3011692
- Publication, EPODOC
- PL3011692T
- Application
- 14731262
- Application, DOCDB
- 14731262
- Application, EPODOC
- PL20140731262T
Titles2
- English
- JITTER BUFFER CONTROL, AUDIO DECODER, METHOD AND COMPUTER PROGRAM
- Polish
- Sterowanie buforem rozsynchronizowania, dekoder sygnału audio, sposób i program komputerowy
Classification
- CPC, 7
- G10L19/012
- G10L19/022
- H04J3/0632
- G10L21/04
- H04J3/0664
- H04J3/06
- G10L19/04
- IPC, 5
- H04J3 06
- G10L19 012
- G10L19 04
- G10L21 04
- H04L49 9023