Method and apparatus for an adaptive de-jitter buffer
35 claims: 6 independent, 29 dependent
- 1Zastrzeżenia patentowe 1. Urządzenie zawierające:jednostkę pamięci (256) skonfigurowaną do przechowywania pakietów danych;oraz pierwszy kontroler (254) skonfigurowany do porównywania określonej liczby pakietów przechowywanych w jednostce pamięci z pierwszym progiem kompresji czasu oraz pierwszym progiem rozciągania czasu dla jednostki pamięci, przy czym pierwszy kontroler jest ponadto przystosowany do generowania sygnału sterowania 53/59P27611PL00 dopasowaniem czasu wskazującego kompresję, jeśli liczba przechowywanych pakietów przekracza pierwszy próg kompresji czasu oraz do generowania sygnału sterowania dopasowaniem czasu wskazującego rozciąganie, jeśli liczba przechowywanych pakietów przekracza pierwszy próg rozciągania czasu, przy czym pierwszy kontroler jest skonfigurowany do identyfikowania początku i końca strumienia mowy, przy czym pierwszy kontroler jest skonfigurowany do identyfikowania końcowej części strumienia mowy (510) i dokonania kompresji co najmniej jednego pakietu w końcowej części strumienia mowy.
- 2Urządzenie według zastrzeżenia 1, w którym pierwszy kontroler jest ponadto przystosowany do porównywania liczby przechowywanych pakietów ze zbiorem wartości progowych kompresji czasu oraz ze zbiorem wartości progowych rozciągania czasu, przy czym każdy z tych zbiorów, zbioru wartości progowych kompresji czasu i zbioru wartości progowych rozciągania czasu, odpowiada unikalnym wartościom procentowym docelowej długości opóźnienia dla jednostki pamięci.
- 3Urządzenie według zastrzeżenia 1, w którym pierwszy kontroler jest ponadto skonfigurowany do generowania sygnału sterującego dopasowaniem czasu dla rozciągania, jeśli następny sekwencyjny pakiet jest odbierany po upływie przewidywanego czasu odtwarzania dla następnego sekwencyjnego pakietu.
- 4Urządzenie według zastrzeżenia 3, w którym pierwszy kontroler jest ponadto skonfigurowany do uśredniania stanu jednostki pamięci w oknie czasowym przed dokonaniem porównania liczby przechowywanych pakietów z progami dopasowania czasu.
- 5Urządzenie według zastrzeżenia 4, w którym pierwszy kontroler jest ponadto skonfigurowany do filtrowania liczby 53/59P27611PL00 pakietów przechowywanych w jednostce pamięci, w oknie czasowym.
- 6Urządzenie według zastrzeżenia 5, w którym pierwszy kontroler jest ponadto skonfigurowany do określania docelowej długości opóźnienia, a także określania okna czasowego jako funkcji docelowej długości opóźnienia.
- 7Urządzenie według zastrzeżenia 6, w którym pierwszy kontroler jest ponadto skonfigurowany do wyznaczenia docelowej długości opóźnienia jako docelowej liczby pakietów, które mają być przechowywane w jednostce pamięci.
- 8Urządzenie według zastrzeżenia 4, w którym pierwszy kontroler jest ponadto skonfigurowany do porównywania uśrednionej liczby pakietów przechowywanych w jednostce pamięci z progami dopasowania czasu.
- 9Urządzenie według zastrzeżenia 1, w którym pierwszy kontroler jest ponadto skonfigurowany do generowania sygnału sterującego dopasowaniem czasu, który jest wielostanowym sygnałem sterującym.
- 10Urządzenie według zastrzeżenia 9, w którym pierwszy kontroler jest ponadto skonfigurowany do wyznaczania docelowej długości opóźnienia dla jednostki pamięci, przy czym jednostką pamięci jest adaptacyjny bufor eliminujący jitter, zaś docelową długością opóźnienia jest docelowa długość bufora eliminującego jitter.
- 11Urządzenie według zastrzeżenia 1, w którym pierwszy kontroler jest ponadto przystosowany do inicjowania kompresji co najmniej jednego pakietu, gdy liczba przechowywanych pakietów przekracza docelową długość opóźnienia. 53/59P27611PL00
- 12Urządzenie według zastrzeżenia 11, w którym pierwszy kontroler jest ponadto skonfigurowany do utrzymania danej zawartości procentowej niedomiarów w wyniku opóźnionych pakietów.
- 13Urządzenie według zastrzeżenia 12, w którym pierwszy kontroler jest ponadto skonfigurowany do obliczania docelowej długości opóźnienia jako:If (PER d ei ay TARGET_VALUE) then DEJITTER_DELAY = DEJITTER_DELAY - CONSTANT;If (PERdeiay TARGET VALUE PERdeiay = last PERdeiay) then DEJITTER_DELAY = DEJITTER_DELAY + CONSTANT;Set DEJITTER_DELAY = ΜΑΧ(MIN_JITTER,DEJITTER_DELAY);oraz DEJITTER_DELAY = MIN(MAX_JITTER,DEJITTER_DELAY), gdzie PER de iay jest współczynnikiem niedomiarów wynikających z opóźnionych pakietów, TARGET_VALUE jest docelowym współczynnikiem opóźnionych pakietów, DEJITTER_DELAY jest docelową długością opóźnienia adaptacyjnego bufora eliminującego jitter, CONSTANT jest z góry określoną wartością, zaś MAX_JITTER i MIN_JITTER są z góry określonymi wartościami oznaczającymi odpowiednio maksymalną i minimalną docelową długość opóźnień.
- 14Urządzenie według zastrzeżenia 13, w którym pierwszy kontroler jest skonfigurowany do obliczenia wartości PER de iay j ako:PER delay = PER_CONSTANT x PER delay + (1 - PER_CONSTANT) * Current_PER delay gdzie PER_CONSTANT jest stałą czasową filtra stosowanego do oszacowania wartości PER de iay· 53/59P27611PL00
- 15Urządzenie według zastrzeżenia 14, w którym pierwszy kontroler zawiera:jednostkę obliczania błędu pakietu, skonfigurowaną do obliczania wartości Current_PER de i ay jako współczynnika opóźnionych pakietów, przy czym opóźnione pakiety są odbierane po upływie przewidywanego czasu odtwarzania, a także Current_PER de iay.
- 16Urządzenie według zastrzeżenia 15, w którym jednostka obliczania błędu pakietu jest skonfigurowana do obliczania wartości Current_PER de i ay jako stosunku opóźnionych pakietów do całkowitej liczby odebranych pakietów, wliczając w to opóźnione pakiety, mierzone od ostatniej aktualizacji PER de i ay do bieżącej aktualizacji i która jest obliczana jako:Current PER delcv Liczba niedomiarów opóźnienia od ostatniej aktualizacji Liczba pakietów odebranych od ostatniej aktualizacji
- 17Urządzenie według zastrzeżenia 16, w którym pierwszy kontroler jest skonfigurowany do identyfikowania pierwszej części odebranych pakietów, przy czym pierwsza część odpowiada strumieniowi mowy, zaś strumień mowy zawiera wiele sekwencyjnycń pakietów.
- 18Urządzenie według zastrzeżenia 17, w którym pierwszy kontroler jest skonfigurowany do identyfikowania pierwszej części poprzez kodowanie pierwszej części.
- 19Urządzenie według zastrzeżenia 18, w którym pierwszy kontroler jest skonfigurowany do określenia przewidywanego czasu odtwarzania dla pierwszego pakietu strumienia mowy, a także do zainicjowania odtwarzania pierwszego pakietu strumienia mowy przed przewidywanym czasem odtwarzania. 53/59P27611PL00
- 20Urządzenie według zastrzeżenia 19, przy czym pierwszy kontroler jest ponadto skonfigurowany do inicjowania rozciągania kolejnych pakietów po odtworzeniu pierwszego pakietu.
- 21Urządzenie według zastrzeżenia 11, w którym pierwszy kontroler jest skonfigurowany do identyfikowania końcowej części strumienia mowy przez szybkość kodowania odbieranych pakietów.
- 22Urządzenie według zastrzeżenia 21, w którym pierwszy kontroler jest skonfigurowany do identyfikowania końcowej części strumienia mowy za pośrednictwem wskaźnika ciszy.
- 23Urządzenie według zastrzeżenia 22, w którym pierwszy kontroler jest skonfigurowany do identyfikowania końcowej części strumienia mowy za pośrednictwem wskaźnika końca strumienia mowy.
- 24Sposób przetwarzania pakietowych danych obejmujący:gromadzenie pakietów danych w jednostce pamięci;określanie docelowej długości (100) opóźnienia dla jednostki pamięci;dokonanie oceny stanu jednostki pamięci pod względem docelowej długości opóźnienia, gdzie stan jednostki pamięci jest miarą danych przechowywanych w jednostce pamięci;inicjowanie dopasowania czasu co najmniej jednego pakietu z jednostki pamięci, jeśli stan pamięci przekracza docelową długość opóźnienia;identyfikowanie początku i końca strumienia mowy;identyfikowanie końcowej części strumienia mowy (510);a także dokonywanie kompresji co najmniej jednego pakietu w końcowej części strumienia mowy. 53/59P27611PL00
- 25Sposób według zastrzeżenia 24, znamienny tym, że ponadto obejmuj e:obliczanie docelowej długości opóźnienia w następujący sposób: If (PERdelay TARGET_VALUE) (102) then DEJITTER_DELAY = DEJITTER_DELAY - CONSTANT (104);If (PER d eiay TARGET_VALUE PER de iay = last_PER de iay) (103) then DEJITTER_DELAY = DEJITTER_DELAY + CONSTANT (108);Set DEJITTER_DELAY = ΜΑΧ(MIN_JITTER,DEJITTER_DELAY) (110);oraz DEJITTER_DELAY = MIN(MAX_JITTER,DEJITTER_DELAY) (112), gdzie PER de i ay jest współczynnikiem niedomiarów wynikających z opóźnionych pakietów, TARGET_VALUE jest docelowym współczynnikiem opóźnionych pakietów, DEJITTER_DELAY jest docelową długością opóźnienia adaptacyjnego bufora eliminującego jitter, CONSTANT jest z góry określoną wartością, zaś MAX_JITTER i MIN_JITTER są z góry określonymi wartościami oznaczającymi odpowiednio maksymalną i minimalną docelową długość opóźnienia.
- 26Sposób według zastrzeżenia 25, obejmujący ponadto:generowanie sygnału sterującego dopasowaniem czasu;odbieranie wielu sekwencyjnych pakietów;a także dodawanie z nakładaniem segmentów w odpowiedzi na sygnał sterujący dopasowaniem czasu.
- 27Sposób według zastrzeżenia 26, w którym dodawanie z nakładaniem obejmuje łączenie co najmniej dwóch spośród wielu segmentów zgodnie z następującym wyrażeniem:a) OutSegmeHt[i] = b) OutSegntenl[i\ = (Segmenll(j) * (WindowSize - i) + (Segment2(i) * i) WmdowSize (Segmeni2(f) * (WindowSize - i) + (SegmenlAfi) * i) WindowSize 53/59Ρ27611PL00 i = 0. WindowSize — 1 WindowSize = R WindowSize gdzie OutSegment jest segmentem powstałym z procesu dodawania z nakładaniem, Segmentl i Segment2 są segmentami, które mają zostać dodane z nakładaniem, WindowSize odpowiada pierwszemu segmentowi, zaś RWindowSize odpowiada drugiemu segmentowi.
- 28Sposób według zastrzeżenia 27, w którym dodawanie z nakładanie obejmuje ponadto:identyfikowanie części o maksymalnej korelacji między pierwszym segmentem i drugim segmentem.
- 29Sposób według zastrzeżenia 28, w którym identyfikowanie części o maksymalnej korelacji między pierwszym segmentem a drugim segmentem obejmuje ponadto:identyfikowanie części o maksymalnej korelacji przez obliczenie maksymalnej korelacji jako: ^[(x(z) - mx) x (y(i -d)- my)] Corr(d) - - . = jZ- wa ') A2 my c 2 gdzie x oznacza pierwszy segment, y oznacza drugi segment mowy, m oznacza okno korelacji, i jest wartością indeksu, zaś d oznacza część korelacji.
- 30Sposób według zastrzeżenia 24, obejmujący ponadto:dopasowywanie czasu wielu sekwencyjnych pakietów;zatrzymywanie dopasowania czasu dla co najmniej jednego sekwencyjnego pakietu, przy czym co najmniej jeden sekwencyjny pakiet jest kolejnym pakietem za wspomnianymi wieloma sekwencyjnymi pakietami;a także uaktywnianie dopasowania czasu po wspomnianym co najmniej jednym sekwencyjnym pakiecie. 53/59P27611PL00
- 31Sposób według zastrzeżenia 24, obejmujący ponadto:obliczanie stopnia dopasowania czasu, gdzie stopień dopasowania czasu jest liczbą pakietów poddanych dopasowaniu czasu w oknie czasowym;a także inicjowanie dopasowania czasu dla pakietów w funkcji stopnia dopasowania czasu.
- 32Odczytywalny przez komputer nośnik pamięci zawierający zestaw instrukcji, przy czym zestaw instrukcji zawiera:procedurę wejściową służącą do zapisywania pakietów danych w jednostce pamięci;procedurę obliczania docelowej długości opóźnienia służącą do wyznaczania docelowej długości opóźnienia dla jednostki pamięci;pierwszą procedurę służącą do oceniania stanu jednostki pamięci pod względem docelowej długości opóźnienia, przy czym stan jednostki pamięci stanowi miarę danych przechowywanych w jednostce pamięci;drugą procedurę inicjującą dopasowanie czasu co najmniej jednego pakietu z jednostki pamięci, jeśli stan jednostki pamięci przekracza docelową długość opóźnienia;identyfikowanie początku i końca strumienia mowy;identyfikowanie końcowej części strumienia mowy (510);a także dokonanie kompresji co najmniej jednego pakietu w końcowej części strumienia mowy. Qualcomm Incorporated Pełnomocnik: 53/59P27611PL00 1/35 ω π ω 53/59P27611PL00 2/35 53/59P27611PL00 3/35 Następny ---------do odtworzenia Pakiet wokodera 53/59P27611PL00 4/35 SYGNAŁY FIG. 3 53/59P27611PL00 5/35 53/59P27611PL00 6/35 SYGNAŁY FIG. 5 53/59P27611PL00 7/35 100 ( START 53/59P27611PL00 8/35 ZMIANA OPÓŹNIENIA OPÓŹNIENIE' 1 OCZEKUJE OCZEKUJE OCZEKUJE OCZEKUJE OCZEKUJE (D init ) PKT I PKT 2 PKT 3 PKT 4 PKT 4 LUB PKT 5 FIG. 7C 53/59P27611PL00 9/35 53/59P27611PL00 10/35 FIG. 8B 53/59P27611PL00 11/35 SYGNAŁY FIG. 9 53/59P27611PL00 12/35 ω Μ ω (40ms + t 2 ) 53/59P27611PL00 13/35 250 53/59P27611PL00 14/35 53/59P27611PL00 100 53/59P27611PL00 16/35 SYGNAŁY LU z. $ CŁ FIG. 15 O 101 53/59P27611PL00 AMPLITUDA SYGNAŁU (WOLTY) 102 53/59P27611PL00 18/35 SEGMENT MOWY SKOMPRESOWANY SEGMENT MOWY 103 53/59P27611PL00 19/35 FIG. 18A 53/59P27611PL00 104 20/35 105 53/59P27611PL00 21/35 106 53/59P27611PL00 PAKIETY ODEBRANE POZA KOLEJNOŚCIĄ FIG. 19 107 53/59P27611PL00 23/35 UŻYTK0WNIK1 UŻYTK0WNIK2 108 53/59P27611PL00 24/35 START KONIEC 109 53/59P27611PL00 25/35 START 110 53/59P27611PL00 26/35 111 53/59P27611PL00 27/35 112 53/59P27611PL00 28/35 ZMIENNE OPÓŹNIENIE - RÓWNOMIERNE I UJ Z UJ z 'NI Ό Cl O UJ Z z UJ Ξ N LU N O QC CO O Q LU ki z O -CO O Z V, ε c FIG. 26 113 53/59P27611PL00 29/35 700 FIG. 27 114 53/59P27611PL00 STEROWANIE rDOPASOWANIEM CZASUn 115 53/59P27611PL00 31/35 116 53/59P27611PL00 32/35 © O fc START N DC LU □a o CL KONIEC 117 53/59P27611PL00
- 3333/35 118 53/59P27611PL00
- 3434/35 FIG. 32 119 53/59P27611PL00
- 3535/35 LU
Independent claims35
283 paragraphs in 24 sections, as filed
Description
Priority claim under 35 USC §119
[0001] This patent application claims priority of US Provisional Application No. 60 / 606,036 entitled "Adaptive De-Jitter Buffer For Voice Over IP for Packet Switched Communications, filed August 30, 2004, the rights of which have been acquired by the present applicant.
BACKGROUND
Technical field
The present invention relates to wireless communication systems, and more particularly to adaptive de-jitter buffer for Internet Protocol Voice (VoIP) transmission for packet-switched communication. The invention is applicable to any system where packets may be lost.
Background
[0003] In a communication system, packet delay between endpoints can be defined as the time from being generated at a source until the packet reaches its destination. In a packet switched communication system, the delay of packets traveling from source to destination may vary according to various operating conditions, including, but not limited to, channel conditions and network load. Channel conditions refer to the quality of the wireless link. Factors that determine the quality of a wireless link include signal strength, cell phone speed, and / or physical obstructions.
[0004] The end-to-end delay includes the delay introduced into the network and the various elements that the packet passes through. Many factors contribute to the delay between endpoints. Latency variability
53 / 59P27611PL00 between endpoints is referred to as jitter. Jitter can cause packets to be picked up when they are no longer useful. For example, in a low latency application such as voice transmission, if a packet is received too late, it may be rejected by the receiver. Such conditions lead to a deterioration in the quality of communication.
[0005] WO 00/24144 A (TIERNAN COMMUNICATIONS, INC.) Issued on April 27, 2000 (2000-04-27) describes a phase-locked loop (PLL) for packet synchronization that comes with delay variations due to the asynchronous transfer mode. (ATM) or asynchronous multiplexing. The phase-locked loop includes a recovery circuit for controlling the clock timing and a clock timing recovery unit that monitors the PLL buffer level.
[0006] US 2004/156397 A1 (HEIKKINEN ARI ET AL) issued on August 12, 2004 (2004-08-12) discloses a device that processes packaged and encoded speech data to be audible to a listener. The device includes a speech decoder to perform a "time warping" operation to extend or shorten the duration of a speech frame.
[0007] US-B1-6 496 794 (KLEIDER JOHN ERIC ET AL) issued on December 17, 2002 (2002-12-17) discloses a communication system including a variable size / rate buffer, a speech buffer, a buffer control block, and a uniform rate variation module that correlates pre-encoded data at different rates, and cuts or concatenates and adjusts the speech data.
[0008] US 2004/120309 A1 (KURITU ANTTI ET AL) of June 24, 2004 (2004-6-24) describes a method for changing the jitter buffer size in a communication system including a packet network for buffering received packets.
53 / 59P27611PL00 containing audio data to be able to compensate for varying packet received delays. There it is also proposed to establish that the current jitter size is changed.
[0009] In the article by E. Moulines and W. Verhelst entitled "Time-Domain and Frequency-Domain Techniques for Prosodic Modification of Speech in Speech Coding and Synthesis, pages 519-555, ΧΡ002366713, discloses a model of speech formation, modification of the time scale and pitch scale, the short time Fourier transform (STFT) as a representation of the time- frequency for analysis, modification and synthesis of signals slowly changing in time. Also, STFT analysis and synthesis properties are defined as quantities varying with time.
BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Fig. 1 is a block diagram of a prior art communication system in which the Access Terminal includes a de-jitter buffer.
[0011] Fig. 2 shows a prior art de-jitter buffer.
[0012] Fig. 3 is a timing diagram illustrating the transmission, reception, and recovery of packets causing "underflow.
[0013] Figures 4A and 4B are two scenarios showing timing charts to illustrate the calculation of optimal de-jitter buffer lengths.
[0014] Fig. 5 is a timing diagram illustrating the course of "underflow due to delayed packets.
[0015] Fig. 6 is a flowchart illustrating the calculation of a target de-jitter buffer length.
[0016] Fig. 7A is a timing diagram illustrating packet transmission in a first scenario.
53 / 59P27611PL00
[0017] Fig. 7B is a timing diagram illustrating the receipt of packets without de-jitter adaptation.
[0018] Fig. 7C is a timing diagram illustrating the receipt of packets with de-jitter adaptation where the receiver can receive a packet following the expected time for the packet.
[0019] FIG. 8A is a flowchart illustrating one example of an implicit buffer adaptation that allows a receiver to receive a packet after expected time for the packet.
[0020] Fig. 8B is an operating mode state diagram for the adaptive de-jitter buffer.
[0021] Fig. 9 is a timing diagram illustrating the application of de-jitter adaptation to another example.
[0022] FIG. 10 is a diagram illustrating the transmission of voice information on speech streams according to one example, and the de-jitter delay is not sufficient to avoid data collision.
[0023] Fig. 11 is a block diagram of a communication system including an adaptive elimination buffer.
[0024] Fig. 12 is a block diagram of a receiver portion including an adaptive de-jitter buffer and a time warping unit.
[0025] Fig. 13A is one example of an adaptive de-jitter buffer including compression and stretch thresholds.
[0026] Fig. 13B is one example of an adaptive de-jitter buffer including multiple compression and stretch thresholds.
[0027] FIG. 14 is a timing diagram illustrating a time warping when receiving packets having different delays.
53 / 59P27611PL00
[0028] Figure 15 is a timing diagram illustrating examples of: i) compression of the silence portion of a speech segment; and ii) extending the silence portion of the speech segment.
[0029] Fig. 16 is a timing diagram illustrating a speech signal in which portions of the speech signal may repeat.
[0030] Fig. 17A is a diagram illustrating a speech segment where multiple PGM samples have been identified in a reference window for an addoverlap, referred to as RWindowSize, and a target or desired segment size, referred to as Segment, has been identified.
[0031] Fig. 17B is a diagram illustrating an application of an add-overlap operation to compress a speech segment in accordance with one example.
[0032] Fig. 18A is a diagram illustrating a plurality of speech segments where the number of PGM samples has been identified in a reference window for an add overlap operation, referred to as RWindowSize, and where a target or desired segment size, referred to as Segment in preparation for stretching, has been identified. current speech segment.
[0033] Fig. 18B is a diagram illustrating the use of an add-overlap operation to stretch a speech sample according to one example.
[0034] Fig. 18C is a diagram illustrating the application of a speech sample stretching operation in accordance with an alternate example.
[0035] FIG. 19 is a diagram illustrating how packets are stretched to allow for the arrival of delayed packets and packets that arrive out of order as in Hybrid ARQ retransmission.
[0036] Fig. 20 is a diagram illustrating the timing of a conversation between two users.
53 / 59P27611PL00
[0037] Fig. 21 is a flowchart illustrating an improvement at the start of a speech stream according to one example.
[0038] Fig. 22 is a diagram illustrating an improvement at the start of a speech stream according to an alternative example.
[0039] Fig. 23 is a diagram illustrating an improvement at the end of speech streams.
[0040] Fig. 24 is a flowchart illustrating an improvement at the end of a speech stream according to one example.
[0041] Fig. 25 is a diagram illustrating the operation of a prior art de-jitter buffer and decoder circuit in which the de-jitter buffer delivers packets to a decoder at regular intervals.
[0042] Fig. 26 is a diagram illustrating the operation of an adaptive de-jitter buffer and a decoder according to one example in which an adaptive de-jitter buffer delivers packets to a decoder at unequal time intervals.
[0043] Fig. 27 is a block diagram illustrating an Access Terminal (AT) according to one example, including an adaptive de-jitter buffer and a time warping control unit.
[0044] Fig. 28 shows a receiver portion including adaptive de-jitter buffer adapted to time warped packets according to one example.
[0045] Fig. 29 shows an alternate example of a receiver including an adaptive de-jitter buffer and adapted to time warped packets according to another example.
[0046] FIG. 30 is a flowchart illustrating one example of a scheduler at a decoder in one example of a receiver including an adaptive buffer.
Time warped packets according to one example. 53 / 59P27611PL00 de-jitter and adapted to time warped packets.
[0047] Fig. 31 is a flowchart illustrating a scheduler in an audio interface unit in one example of a receiver.
[0048] Fig. 32 shows a time warp unit where the scheduling is computed outside of the decoder.
[0049] Fig. 33 shows a time warp unit in which the scheduling is calculated at the decoder.
DETAILED DESCRIPTION
[0050] In packet switched systems, data is formed as packets and routed through the network. Each packet is sent to its destination on the network based on an assigned address contained within the packet, usually in the header. End-to-end packet latency, or the time it takes for an intra-network packet to arrive from the first user or "sender to a second user, or" recipient, varies depending on, among other things, channel conditions, network load, system quality of service (QoS). and other streams competing for resources. It should be noted that, for clarity, the following description describes spread spectrum communication systems supporting packet communication, including code division multiple access (CDMA), orthogonal frequency division multiple access (OFDMA), wideband code division multiple access (W-CDMA) systems. ), Global System for Mobile Communications (GSM), systems supporting IEEE standards such as 802.11 (A, B, G), 802.16, and others, but not limited to.
[0051] In a wireless communication system, each packet may experience a different source-to-destination transmission delay than other packets belonging to the same flow. This delay variation is known as "jitter."
53 / 59P27611PL00
Jitter causes additional complications for the receiver side application. If the receiver does not apply jitter corrections, then the received message will be garbled after packet recovery. Some systems correct the jitter when reconstructing messages from received packets. Such systems include a de-jitter buffer that adds latency, referred to as de-jitter buffer delay. When the de-jitter buffer applies a constant high de-jitter buffer delay, it can accommodate a large amount of jitter as packets arrive; however, this solution is inefficient since the low latency packets are also processed using the high delay de-jitter buffer, even though the packets might have been processed earlier. This leads to longer end-to-end delay for packet cuttings than would be achievable with a lower de-jitter buffer delay.
[0052] To prevent this effect, VoIP systems including de-jitter buffers may try to adapt to changes in packet delay. For example, the de-jitter buffer can detect changes in packet delay as a result of analyzing packet arrival statistics. Many de-jitter buffer implementations do not adjust their delay at all and are configured to have a conservatively long delay. In this case, the de-jitter buffer may add excessive delay to the packets causing the user experience to be suboptimal.
[0053] The following description describes an adaptive de-jitter buffer that adapts to changes in packet delay by changing its de-jitter buffer delay. This de-jitter buffer uses speech time warping to correct its
Variable delay packet tracking capability. The following description may apply to packet communication such as, for example, periodic data communication, low latency communication, sequential data processing, or designated playback speed. In particular, the following description relates in detail to voice communication where data or speech and silence is initiated at a source and transmitted to a destination for opening. The original data is packaged and encoded using a known encoding scheme. At the receiver, an encoding scheme is defined for each data packet. In speech communication, for example, the type of speech coding is different from the type of silence coding. This allows the communication system to take advantage of the periodic nature of speech that contains bits of silence. For speech communication, data is sequential and speech content may be repetitive. Packet speech requires low latency because voice participants do not want to hear delays, but the communication quality allows only limited latency. The speech contained in packets can reach the recipient via different paths, however, upon arrival, the packets are recompiled to their original sequence. Therefore, the received speech contained in the packets is played back sequentially. If a packet is lost in air transmission or during physical layer processing, the packet is not reconstructed, but the receiver must estimate or guess the contents of that packet. Moreover, the reproduction rate of speech communication has a predetermined playback speed or range. If the playback is outside this range, the quality on the receiver side deteriorates. The use in speech communication is an exemplary application of the present invention. Other applications may include video communication, game communication, or other types of communication having similar properties, specifications, and / or requirements
53 / 59P27611PL00 for speech communication. For example, video communication may require speeding up or slowing down the playback. The present description may be desirable for this kind of application. The adaptive de-jitter buffer of the present invention may allow a receiver to achieve a quality of service defined by the system jitter requirements. The adaptive de-jitter buffer adjusts the target de-jitter buffer length, for example, the amount of data stored in the de-jitter buffer, to the timing and the amount of data received in the adaptive de-jitter buffer. In addition, the adaptive de-jitter buffer uses the state or size of the de-jitter buffer, such as measuring data stored in the adaptive de-jitter buffer, to determine when time warping is beneficial for processing and playing back received data. For example, when data arrives at the adaptive de-jitter buffer at a slow rate, the adaptive de-jitter buffer provides this information to the time warping unit, allowing the unit to stretch the received packets. If the data stored in the adaptive de-jitter buffer exceeds a threshold, the adaptive de-jitter buffer alerts the time warping unit to compress the packets so as to efficiently follow the incoming data. It should be noted that the time warping is within limits that can be determined by the application and type of communication. For example, in speech communication, time warping should not compress the speech, that is, increase the pitch so that the listener is unable to understand the message. Likewise, time warping should not extend speech beyond this range. Ideally, the timing range is defined so that the listener experiences little or no discomfort.
53 / 59P27611PL00
Communication system
Fig. 1 is a block diagram illustrating a digital communication system 50. Two access terminals (ATs) 52 and 82 communicate via a base station (BS) 70. Within the AT 52, the transmitting processing unit 64 transmits the voice data to the encoder. 60, which encodes and packets the voice data and sends the packet data to a lower layer processing unit 58. For transmission, the data is then sent to the BS 70. The BS 70 processes the received data and transmits the data to the AT 82, where it is received by the lower layer processing unit 88. The data is then delivered to a de-jitter buffer 86, which stores the data so as to hide or reduce the effect of jitter. Data is sent from de-jitter buffer 86 to decoder 84 and on to receive processing unit 92.
[0055] For transmission from an AT 82, data / voice is provided from a transmitting processing unit 94 to an encoder 90. The lower layer processing unit 88 processes the data for transmission to a BS 70. To receive data from a BS 70 at the AT 52, the data is received at a lower layer processing unit 58. The data packets are then sent to de-jitter buffer 56, where they are stored until the required buffer length or delay is reached. When this length or delay is reached, the de-jitter buffer 56 begins sending data to decoder 54. The decoder 54 converts the data contained in the packets into voice data packets and sends these packets to the receiving processing unit 62. In the present example, the operation of the AT 52 is analogous to that of the AT 82.
Anti-jitter buffer
53 / 59P27611PL00
[0056] A memory or de-jitter buffer is used in AT terminals as described above to hide the effect of jitter. In one example, an adaptive de-jitter buffer is used for packet switched communication, such as for example VoIP communication. The de-jitter buffer has an adaptive buffer memory and uses speech time warping to improve its ability to track variable delay and jitter. In this example, the de-jitter buffer processing is coordinated with the operation of the decoder, the de-jitter buffer identifying the possibility or need for packet time warping and instructing the decoder to time warp the packets. The decoder performs time warping of the packets by compressing or stretching the packets according to instructions issued by the de-jitter buffer.
[0057] Fig. 2 is one example of a de-jitter buffer. Incoming encoded packets are accumulated and saved in a buffer. In one example, the buffer is a first in, first out (FIFO) buffer, in which data is received in a specific order and processed in the same order; the first data processed is the first data received. In another example, the de-jitter buffer is an ordered list that keeps track of which packet should be processed next. The adaptive de-jitter buffer may be a unit of memory where the state of the de-jitter buffer is a measure of data (or number of packets) stored in the adaptive de-jitter buffer. The data processed by the de-jitter buffer may be sent therefrom to a decoder or other device. The encoded data may correspond to a fixed amount of speech data, for example 20 milliseconds, corresponding to 160 samples of speech data, at a sampling rate of 8 kHz. In one example of the present invention, the number of samples produced
By a decoder having the time warping capability, the decoder can change based on whether a given packet has been matched or not. When the de-jitter buffer instructs the decoder / time warrior to stretch the packet, the decoder / time warrior can produce more than 160 samples. On the other hand, when the de-jitter buffer instructs the decoder / time warrior to compress the packet, the decoder / time warrior may produce less than 160 samples. It should be noted that the alternative systems may have different reproduction schemes, e.g. with voice coding with a frame length other than 20 ms.
[0058] Packets arriving in the de-jitter buffer may not arrive at regular intervals. Therefore, one of the design purposes of the de-jitter buffer is to match irregularities in the incoming data. In one example of the present invention, the de-jitter buffer has a target de-jitter buffer length. Target de-jitter buffer length refers to the amount of data required to be stored in the de-jitter buffer before the first packet is played back. In another example, the target de-jitter buffer length can refer to the amount of time the first packet in the de-jitter buffer must be delayed before starting playback. The target length of the de-jitter buffer is shown in Fig. 2. By accumulating a sufficient number of packets in the de-jitter buffer before starting packet playback, the de-jitter buffer is able to play consecutive packets at regular intervals minimizing the risk of packet exhaustion. Fig. 2 shows a de-jitter buffer, in which a vocoder packet first received in the de-jitter buffer is the next packet scheduled for output from the de-jitter buffer. The de-jitter buffer contains a sufficient number of packets
53 / 59P27611PL00 to obtain the required de-jitter buffer delay. In this way, the de-jitter buffer smoothes the jitter experienced by the packets and hides the variation in the time of packet arrival at the receiver.
[0059] Fig. 3 shows timing of transmission, reception and playback for packets in various scenarios. The first packet, PKT 1, is transmitted at time ti and is played back upon receipt at time ti. Subsequent packets, PKT 2, PKT 3, and PKT 4, are transmitted at 20 ms intervals after PKT 1. In the absence of time warping, the decoders reconstruct the packets at regular intervals (e.g., 20 ms) from the moment of playing the first packet. For example, if the decoder plays packets at regular intervals of 20 ms, the first packet received is played at time ti, and subsequent packets will be played 20 ms after ti, 40 ms after ti, 60 ms after ti and so on. As shown in Fig. 3, the projected playback time (without de-jitter delay) of PKT 2 is t2 = ti + 20 ms. PKT 2 is received before its expected playback time, t2- On the other hand, packet 3 is received after its expected playback time ta = t2 + 20 ms. This condition is referred to as underflow. An underflow occurs when the reproducing device is ready to reproduce a packet but the packet is not yet present in the de-jitter buffer. The underflow usually causes the decoder to produce blotted spots and degrade the playback quality.
[0060] Fig. 3 additionally shows a second scenario in which the de-jitter buffer introduces a delay tdjb before starting playback of the first packet. In this scenario, the de-jitter buffer delay is added to allow the reproducing device to receive packets (or samples) every 20 ms. In this scenario, even though PKT 3 is received after the expected expiration
This addition of the de-jitter buffer delay allows the recovery of PKT 3 20 ms after recovery of PKT 2.
[0061] PKT 1 is sent at time t<sub>0</sub>, received at time ti and instead of being played back at time ti as has been done so far, is now played back at time tl + tdjb = ti '. The reproducing apparatus recreates PKT 2 at a predetermined time interval, e.g. 20 ms, after PKT 1 or at time t<sub>2</sub>'= ti + tdjb + 20 ms = t<sub>2</sub> + t<sub>d</sub>jb, and PKT 3 at time t<sub>3</sub>'= t<sub>3</sub> + t<sub>d</sub>jb- Playback delay ot<sub>d</sub>jb allows the third packet to be played back without underflow. So, as shown in Fig. 3, introducing a de-jitter buffer delay can reduce underflow and prevent speech quality degradation.
[0062] Speech consists of periods of speech streams and periods of silence. Stretching / compressing silence periods has little or no effect on speech quality. This allows the de-jitter buffer to delay plays of the first packet differently for each talkspurt.
[0063] Figs. 4A and 4B are transmission and reception timings for the different speech streams. Note that the amount of the de-jitter buffer delay is specified to prevent underflow. This is referred to as the "optimal de-jitter buffer delay." The optimal de-jitter buffer delay is related to the target de-jitter buffer length. In other words, the target length of the de-jitter buffer is determined to allow sufficient data to be stored in the buffer such that packets are played back according to the parameters of the reproduction apparatus. The optimal de-jitter delay may be determined by the greatest end-to-end delay experienced by the system. Alternatively, the optimal de-jitter buffer delay may be based on the average
The delay experienced by the system. Other methods for determining the optimal de-jitter buffer delay specific to a given criteria or system design may also be used. In addition, the target de-jitter buffer length is determined to cause optimal de-jitter delay, and therefore the target de-jitter buffer length can be calculated from the forwarding rate of received packets, Packet Error Rate (PER), or other performance statistics.
[0064] Figs. 4A and 4B are optimal de-jitter buffer delays for two examples. As illustrated, the time interval between transmission and reception of the sequential packets varies with time. Since PKT 3 has the longest delay from transmission to reception, this difference is used to determine the optimal delay for de-jitter processing.
[0065] Using a de-jitter buffer with a target de-jitter buffer length can help avoid at least some of the under-measurement conditions . Referring again to Fig. 3, the second scenario removes the underflow (occurs when the decoder was waiting for a packet and the playback device was ready to recover the packet, but there were no packets in the accumulator buffer). Here, PKT 2 is reconstructed after a specified time, 20 ms, after time ti, with ti being the playback time of PKT 1. Although PKT 3 is scheduled or scheduled to be played at ta, PKT 3 is not received, until after a while ta. In other words, the reproducing device is ready to play PKT 3, but this packet is not present in the memory buffer. Since PKT 3 is not available for timed playback and is not played,
A large amount of jitter and underflow with respect to packet PKT 3 arises. PKT 4 is played back at time t<sub>4</sub>PKT 4 is expected to play back the expected time. Note that the time t<sub>4</sub> is calculated from the time ta. Since each packet may contain more than one voice packet, underflow loss of packets reduces voice quality.
[0066] Another scenario considered relates to a course of "underflow due to delayed packets as shown in Fig. 5, where transmission, reception and predicted playback time waveforms are shown. In this scenario, each packet is received shortly after its anticipated recovery time. For example, the projected playback time for PKT 50 is t<sub>0</sub>but PKT 50 is not received until ti 'after ti. A next packet, 51, is anticipated at ti but is not received until ti' after ti. This causes a series of underflows leading to a high percentage of "delay underflows, underflows due to the delayed packet, and thus leads to longer endpoint delays.
[0067] Of course, a de-jitter buffer that delays playback by a large amount of time will be effective in keeping the underflow to a minimum. However, this kind of de-jitter buffer introduces a large de-jitter buffer delay to the packet delay between endpoints. High packet latency between endpoints can lead to difficulties in keeping the conversation going. Delays greater than 100 ms may cause the listener to think that the speaker has not finished speaking. Therefore, good quality ideally takes into account both underflow avoidance and end-to-end delay reduction. There is a problem where solving one problem
Another problem may be magnified. In other words, lower latency generally causes more underflow, and vice versa. Therefore, there is a need to balance these two competing goals. In particular, there is a need for the de-jitter buffer to track and avoid underflow while reducing end-to-end delay.
The target length of the de-jitter buffer
[0068] The design goal of the adaptive de-jitter buffer is to allow the system to achieve a specific "voice packet underflow rate" while still achieving low end-to-end delay. Since perceived quality is a function of the underflow percentage, the ability to obtain a specific underflow percentage allows you to control the voice quality. Packet underflow in the de-jitter buffer can occur when a packet is missing. A package may be missing if it is lost or delayed. A lost packet causes underflow if it is discarded before reaching the receiver, such as when discarded at a specific location in the access network, such as the physical layer or forward link scheduler. In this scenario, the underflow cannot be corrected by a de-jitter buffer delay as the given packet will never reach the de-jitter buffer. Alternatively, underflow may occur as a result of a packet being delayed that arrives after its recovery time. In addition to tracking underflow due to packet delayed, the adaptive de-jitter buffer can also track underflow due to packet loss.
[0069] The number of underflows due to a packet delayed may be controlled at the expense of underflows for the de-jitter buffer delay. Value representing the target
The percentage of underflow due to delayed packets is referred to as the "underflow target amount." This value is a target value for de-jitter buffer operation and is chosen to keep the inter-endpoint delay within a reasonable limit. In one instance, a value of 1% (0.01) may be used as a "target underflow" in one case. Another example uses a value of 0.5% (0.005). The de-jitter buffer delay can be adjusted to achieve the "target amount of underflow".
[0070] In one example of the present invention, a filtered percentage of underflow due to delayed packets (referred to as "delay underflow" hereafter) may be used to adjust the de-jitter buffer delay. At the end of each silence period (or at the start of each speech stream), the de-jitter buffer delay is updated as shown in FIG. 6. As shown in FIG. 6, the algorithm is as follows:
1) If (PER<sub>d</sub>ei<sub>ay</sub> <TARGET_VALUE) then
DEJITTER_DELAY = DEJITTER_DELAY - CONSTANT;
2) If (PER<sub>d</sub>ei<sub>ay</sub> > TARGET_VALUE && PER<sub>d</sub>eia<sub>Y</sub> > = t_PER forest<sub>d</sub>eia<sub>Y</sub>) then
DEJITTER_DELAY = DEJITTER_DELAY + CONSTANT;
3) Set DEJITTER_DELAY = ΜΑΧ (MIN_JITTER, DEJITTER_DELAY); and
4) DEJITTER_DELAY = MIN (MAX_JITTER, DEJITTER_DELAY). (1)
[0071] In this example, the initial de-jitter buffer delay may be set to a fixed value, such as 40 ms. TARGET_VALUE is the target value for "latency underflows (for example, 1%). PER<sub>de</sub>and a<sub>Y</sub> is a filtered packet delay underflow value, where the filter parameters allow the value of TARGET_VALUE to be reached. t_PER forest<sub>de</sub>and a<sub>Y</sub> is the PER value<sub>de</sub>and a<sub>Y</sub> in
53 / 59P27611PL00 of previous de-jitter delay update. DEJITTER_DELAY is the target de-jitter buffer length as defined above. In this example, the CONSTANT parameter is 20 ms. MIN_JITTER and MAX_JITTER are the minimum and maximum de-jitter buffer delay values, respectively; in one example, the values are 20 ms and 80 ms, respectively. The MIN_JITTER and MAX_JITTER values can be estimated from system simulation. The values (MIN_JITTER, MAX_JITTER, CONSTANT) can be optimized depending on the communication system where the jitter suppression buffer is used.
[0072] The PER value<sub>de</sub>and<sub>ay</sub> may be updated at the end of each silence period or at the beginning of each talkspurt, with PER<sub>de</sub>and<sub>ay</sub> is calculated as follows:
PER<sub>delay</sub> = PER_CONSTANT χ PER<sub>delay</sub> + (1 - PER_CONSTANT) χCurrent_PER<sub>delay</sub> (2)
[0073] PER_CONSTANT is the filter time constant used to estimate the PER value<sub>de</sub>and<sub>ay</sub>. The value of this constant determines the filter memory and achieves the value of TARGET_VALUE. Current_PER<sub>de</sub>and<sub>ay</sub> is the coefficient of the "latency underruns observed between the last update of the PER value<sub>de</sub>and<sub>ay</sub> and the current update.
[0074] The value of Current_PER<sub>de</sub>and<sub>ay</sub> is defined as the ratio of the number of under-delay packets to the total number of packets received between the last PER update<sub>de</sub>and<sub>ay</sub> and the current update.
_ Number of latency underflows since last Wlliiclll 11α1 \ .ι i \ O update <sup>_</sup> Number of packets received since the last update
[0075] With reference to FIG. 6, a de-jitter buffer delay calculation and update process 100 begins in step 101 by initializing the value of the de-jitter delay.
53 / 59P27611PL00
DEJITTER_DELAY. In step 102, the value of PER<sub>de</sub>and<sub>ay</sub> is compared to the value of TARGET_VALUE. If the value of PER<sub>de</sub>and<sub>ay</sub> is less than TARGET_VALUE, then the value of CONSTANT is subtracted from the value of DEJITTER_DELAY in step 104. If the value of PER<sub>de</sub>and<sub>ay</sub> is greater than TARGET_VALUE in step 102 and also PER<sub>de</sub>and<sub>ay</sub> is greater than TARGET_VALUE and greater than or equal to the LAST_PERDELAY value in step 103, and is not less than the last PER value<sub>d</sub>ei<sub>ay</sub> in step 102, then the process reaches the decision step 108. The value of DEJITTER_DELAY is set and the value of DEJITTER_DELAY plus the value of CONSTANT in step 108. Continuing from step 103, if PER<sub>de</sub>and<sub>ay</sub> is not greater than TARGET_VALUE and not greater than or equal to LAST_PERDELAY, process proceeds to step 110. Continuing from step 104, DEJITTER_DELAY is set to the maximum of MIN_JITTER and DEJITTER_DELAY in step 110. From step 110, the process proceeds to step 112 to setting the DEJITTER_DELAY parameter to a value equal to the minimum of MAX_JITTER and DEJITTER_DELAY in step 112.
Delay Tracking [0076] The de-jitter buffer may enter a mode that tracks the delay (instead of tracking the underflow rate). The tracked delay may be the end-to-end delay or the de-jitter delay. In one instance, the de-jitter buffer enters the "follow delay" mode when the target underflow rate can easily be met. That is, the de-jitter buffer is capable of achieving a lower underflow rate than the target underflow rate for a period of time. This period of time can be anything from several hundred milliseconds to several seconds.
[0077] In this mode, the de-jitter buffer has the target delay value. This is similar to the underflow target described above. Equation (1) above may be
Used to select the target value, and the underflow factor may be used in an analogous manner to calculate the Target Delay value. When the de-jitter buffer enters this mode in which it selects a target value for the Target Delay, this may allow it to reduce its target underflow ratio as long as the Target Delay value is kept.
Implicit buffer adaptation
[0078] In some situations, the decoder may wait to play back a packet that has not yet been received. This situation is shown in FIG. 5, where PKT 50 is the projected playback time of PKT 50 is t<sub>0</sub>but PKT 50 is received after this time. Similarly, PKT 51 is received after its anticipated playback time t<sub>2</sub>PKT 52 is received after its anticipated playback time t<sub>2</sub> and so on. It should be noted that packets arrive fairly regularly, but since PKT 50 was received slightly after its anticipated playback time had elapsed, this caused all subsequent packets to also expire their playback times. On the other hand, if the decoder could erase at time t and still play back packet 50 at time t<sub>2</sub>this would allow all packets to meet their recovery times. By playing PKT 50 after an erasure instead of playing PKT 50, the length of the de-jitter buffer is effectively adapted.
[0079] It should be noted that reconstructing PKT 50 after it has been cleared may cause discontinuities that can be removed by applying the phase matching technique described in co-pending application number 11 / 192,231 entitled "PHASE MATCHING IN VOCODERS, filed on July 7, 2005.
As shown in Fig. 7A, gaps, such as a slot, may be present in packet reception.
Time between packets PKT 3 and PKT 4. Packet arrival delay may be different for each packet. The de-jitter buffer may respond immediately by adjusting to compensate for this delay. As shown, PKT 1, PKT 2 and PKT 3 are received at times ti, t, respectively.<sub>2</sub> and this one. At the moment t<sub>4</sub> PKT 4 is expected to receive, but PKT 4 has not yet arrived. In Fig. 7A, it is assumed that packet reception is expected every 20 ms. In this illustration, PKT 2 is received 20 ms after PKT 1, and PKT 3 is received 40 ms after PKT 1. Receipt of PKT 4 is expected 60 ms after PKT 1, but does not arrive earlier than 80 ms after PKT 1. ITEM 1.
[0081] In Fig. 7B, an initial delay is introduced into the de-jitter buffer before playing back the first PKT 1 packet received. Here, the initial delay is Dinit. In this case, PKT 1 will be played back by the buffer at time D<sub>in</sub>it, PKT 2 in time D.<sub>in</sub>it + 20 ms, PKT 3 in time D<sub>in</sub>it + 40 ms and so on. In Fig. 7B, when PKT 4 does not arrive in the expected time, D<sub>in</sub>it + 60ms, de-jitter buffer can restore the erase. At the next moment of packet playback, the de-jitter buffer will attempt to restore PKT 4. If PKT 4 still has not arrived, then at<sub>in</sub>it + 80 ms another erasure may be sent. Erases will continue to be played until PKT 4 arrives in the de-jitter buffer. After PKT 4 arrives in the de-jitter buffer, PKT 4 is played back. Such processing causes a delay as no other packets are played until PKT 4 is received. When the system is unable to return to normal operation, that is, never receives PKT 4, the system may reset the process, allowing the recovery of packets following PKT 4 without recovering the PKT 4 packet.
The delay between the de-jitter buffer endpoints may increase as the sending of erasures continues for a long period of time prior to the arrival of PKT 4.
[0082] In contrast, according to the example shown in Fig. 7C, if the packet does not arrive or reception of the packet is delayed, the erasure is reconstructed at the expected playback time of PKT 4. This is similar to the scenario described with reference to Fig. 7B above, where the system waited for PKT 4. At the next playback time, if PKT 4 still has not arrived but PKT 5 has arrived, PKT 5 is restored. For additional illustration, suppose the reception of PKT 4 is delayed and the de-jitter buffer waits to receive PKT 4 at time D<sub>lnlt</sub> + 80 ms. When PKT 4 is delayed, the erasure is recreated. During D<sub>lnlt</sub> + 100 ms, if PKT 4 still has not arrived, PKT 5 is reconstructed instead of playing another erase. In this second scenario, delay adjustments are performed immediately and excessive delays between endpoints in the communication network are avoided. This process can be referred to as IBA because the size of the data that is being averaged in the buffer before playback increases and decreases depending on the reception of the data.
[0083] Implicit Buffer Adaptation (IBA) process 200 is illustrated with the flowchart of Fig. 8A. Process 200 may be implemented in a controller in an adaptive de-jitter buffer, such as output controller 760 or de-jitter buffer controller 756. Process 200 may reside in other portions within the system supporting the adaptive de-jitter buffer. At step 202, a request to provide the next packet for playback is received in the adaptive de-jitter buffer. The next packet is identified as a packet with index i in sequence, in particular PKT [i]. In step 204, if
Implicit Buffer Adaptation (IBA) mode is enabled, the process proceeds to step 206 to operate according to the IBA mode; and if the IBA mode is off, the process continues and proceeds to step 226 for processing without IBA mode.
[0084] If PKT [i] is received in step 206, then the adaptive de-jitter buffer provides PKT [i] for playback in step 208. Mode 208 is turned off in step 210 and the index i is incremented, i.e. (i = i + 1). Moreover, if PKT [i] is not received in step 206 and if PKT [i + 1] is received in step 214, the process proceeds to step 216 for reproducing PKT [i + 1]. In step 218, the IBA mode is turned off and the index i is incremented twice, i.e. (i = i + 2) in step 220.
[0085] If in step 214 the PKT [i] and PKT [i + 1] packets are not received, then the controller initiates the erasure recovery in step 222; and the index i is incremented in step 224. It should be noted that in the present example, in IBA mode, the controller checks up to two (2) packets in response to a request for the next packet, for example received in step 202. This solution efficiently realizes the packet window, inside which the controller looks for received packets. In alternative examples, a different window size may be performed, for example, searching for three (3) packets that would be sequentially numbered i, i + 1, 1 + 2 in this example.
[0086] Returning to step 204, if IBA mode is not enabled, processing continues with step 226 to determine if PKT [i] is received. If received, PKT [i] is delivered to be restored in step 228 and the index i is incremented in step 230. If PKT [i] is not received in step 226, the adaptive de-jitter buffer provides erasure to be restored in step 232. . IBA mode is enabled when PKT [i] has not been received and the erasure has been restored instead.
53 / 59P27611PL00
[0087] Fig. 8B is a state diagram related to the IBA mode. In normal mode 242, if the adaptive de-jitter buffer supplies PKT [i] for playback, the controller remains in the normal mode. The controller transitions from the normal mode 242 to the IBA mode 240 when the erasure is played back. After entering the IBA 240 mode, the controller remains in that mode when the erase is opened. The controller transitions from the IBA mode 240 to the normal mode 242 when retrieving PKT [i] or PKT [i + 1].
[0088] Fig. 9 shows one example of a de-jitter buffer implementing the IBA mode as shown in Figs. 8A and 8B. In this illustration, the playback device requests samples to be played back from the decoder. The decoder then requests a sufficient number of packets from the de-jitter buffer to allow uninterrupted playback by the playback device. In this illustration, the packets carry voice communication and the playback device plays a sample every 20 ms. Alternative systems may deliver packet data from the de-jitter buffer to the playback device through other configurations, and the packet data may be other than voice communication.
[0089] The de-jitter buffer is shown in Fig. 9 as a packet stack. In this illustration, the buffer first receives PKT 49 and then receives PKT 50, PKT 51, PKT 52, PKT 53, and so on. The package number in this illustration relates to the sequence of packages. However, on a packetized system, there is no guarantee that packets will be received in this order. For clarity, in this illustration, packets are received in the same numerical order as they were transmitted, which is also the playback order. For purposes of illustration in FIG. 9, consecutively received packets are placed on top of previously received packets in a de-jitter buffer;
For example, PKT 49 is placed on top of PKT 50, PKT 51 is placed on top of PKT 50, and so on. The packet at the bottom of the stack in the de-jitter buffer is the first to be sent to the playback device. Also, note that the target de-jitter buffer length is not shown in this illustration.
[0090] Fig. 9 is a diagram of packet reception, anticipated packet reception time, and packet recovery time versus time. The updated buffer status is illustrated each time a packet is received. For example, PKT 49 is received at time t<sub>0</sub>wherein PKT 49 is scheduled to be played back at time ti. The state of the buffer on receipt of PKT 49 is illustrated at the top of the graph above time t<sub>0</sub>that is, the time of receipt of PKT 49. The time of receipt of each packet received in the de-jitter buffer is shown as RECEIVED. The ANTICIPATED PLAYBACK time is plotted directly below the RECEIVED time. The playing times are identified as PLAYBACK.
[0091] In this example, initially the next packet to be played back is PKT 49 which is scheduled to be played back at time t. The next sequential packet is expected at time ti, and so on. The first packet, PKT 49, is received before the scheduled playback time, to- Therefore, PKT 49 is played back at time as expected. The next packet of PKT 50 is predicted at time ti. However, reception of PKT 50 is delayed and the erasure is sent to the reproduction apparatus instead of PKT 50. The delay of PKT 50 causes an underflow as previously described. PKT 50 is received after the scheduled reproduction time, ti, i before the next scheduled reproduction time t<sub>2</sub>. Upon receipt, PKT 50 is stored in a de-jitter buffer. Therefore, when it is received
The next packet recovery request at this time, the system searches for the sequential packet with the lowest number in the de-jitter buffer; and PKT 50 is delivered to the reproduction apparatus to be played back at a time. Note that using the IBA mode, even though PKT 50 is not received in time to be played back as predicted, PKT 50 is played back later and the rest the sequence is restarted from that point. As illustrated, successive packets PKT 51, PKT 52, and so on are received and played back in time to avoid further erasure.
[0092] While IBA mode may appear to increase packet latency between endpoints, this is not actually the case. Since the IBA mode leads to fewer underflows, a lower de-jitter buffer value is kept as estimated from Equation 1 above. Therefore, the overall effect of the IBA mode may be to decrease the average overall packet delay between endpoints.
[0093] The IBA mode can improve the processing of communication containing speech streams. Speech stream refers to the part of voice communication where speech communication includes portions of speech and silence consistent with normal speech patterns. In speech processing, the vocoder produces one type of packet for speech and another type of packet for silence. Speech packets are encoded at one encoding rate and silence is encoded at a different encoding rate. When coded packets are received in the de-jitter buffer, the de-jitter buffer identifies the packet type from the code rate. The de-jitter buffer assumes that the speech frame is part of a speech stream. The first frame containing no silence is the start of the speech stream. The speech stream ends when a silence packet is received. Not in discontinuous transmission
All silence packets are transmitted as the receiver may implement simulated noise to include silence portions in communication. In continuous transmission, all silence packets are transmitted and received. In one example, the de-jitter buffer adjusts the length of the de-jitter buffer in accordance with the type of packets received. In other words, the system may decide to reduce the de-jitter buffer length required for the silence portions in communication. It should be noted that the IBA methods can be applied to any communication in which playback is performed according to a predetermined timing scheme, such as a constant rate and etc.
Time alignment
[0094] A speech stream is generally made up of a plurality of data packets. In one example, playback of the first speech stream packet may be delayed by a length equal to the de-jitter buffer delay. The de-jitter buffer delay may be determined in various manners. In one scenario, the de-jitter buffer delay may be computed from an algorithm such as Equation 1 above. In another scenario, the de-jitter delay may be the time it takes to receive voice data equal to the de-jitter delay length. Alternatively, the de-jitter buffer delay may be selected as the smaller of the above-mentioned values. In this example, suppose the de-jitter delay is calculated as 60 ms using Equation 1 and a first speech stream packet is received at a first time t. When the next speech stream packet is received 50 ms after the first packet, the adaptive de-jitter buffer data is 60 ms de-jitter delay. In other words, the time from receiving the packet in the adaptive buffer
53 / 59P27611PL00 de-jitter to be reproduced is 60 ms. Note that the target adaptive de-jitter buffer length may be set to obtain a delay of 60 ms. This kind of calculation determines how many packets are to be kept in order to obtain the delay time. [0095] The adaptive de-jitter buffer monitors data entry and exit from the buffer and adjusts the buffer output to maintain a target buffer delay length, i.e., data amount, to achieve the target delay time. When the de-jitter buffer sends the first packet of the speech stream to be played, the delay is equal to Δ, where A = MIN (de-jitter delay, time used to receive the voice data equal to the de-jitter delay). Successive speech stream packets are delayed by ∆ plus the time it takes to recover previous packets. Thus, the de-jitter buffer delay of successive packets of the same talkspurt is implicitly determined after the de-jitter buffer delay for the first packet is defined. In practice, this definition of the de-jitter buffer delay may require additional consideration to accommodate situations such as those illustrated in FIG. 10.
[0096] Fig. 10 shows the transmission of voice information on speech streams. Speech stream 150 is received at time t and speech stream 154 is received at time t. There is a 20 ms silence period 152 between speech stream 150 and speech stream 154. The adaptive de-jitter buffer may save received data and determine delays for the playback of each speech stream. In this example, speech stream 150 is received in the adaptive de-jitter buffer at time t, with the adaptive de-jitter buffer delay time calculated as 80 ms. The de-jitter buffer delay is added to the reception time, resulting in time
53 / 59P27611PL00 reproduction. Thus, speech stream 150 is delayed 80 ms before being played back by the adaptive de-jitter buffer. Speech stream 150 begins playback at ti, with ti = to + 80 ms or 80 ms after receipt of speech stream 150, and ends at c when t<sub>4</sub>. Using an algorithm such as Equation 1 to calculate the target de-jitter buffer length as above, the de-jitter buffer delay applied to speech stream 154 is 40 ms. That is, the first packet of talkspurt 154 is to be played back at time t<sub>3</sub>, with t<sub>3</sub> = t<sub>2</sub> + 40 ms or 40 ms after receipt of speech stream 154. Playing back packet 154 at time t<sub>3</sub> however, it conflicts with playback of the last packet of speech stream 150, which terminates playback at time t<sub>4</sub>. Therefore, the calculated de-jitter delay of 40 ms (for packet 154) does not allow enough time to complete playback of speech stream 150. To avoid this kind of conflict and allow both packets to be properly restored, the first packet of speech stream 154 should be played after recovery. the last packet of speech stream 150 with a silence period in between. In this example, speech stream 150 and speech stream 154 overlap by cs<sub>3</sub> to cńwili vol<sub>4</sub>. Therefore, the reproduction method in this scenario is not desirable. In order to prevent packet playback overlap as described above, there is a need to detect when the last packet of the previous talkspurt is played back. Thus, the de-jitter buffer delay calculation may consider the timing of the recovery of previously reconstructed packets, so as to avoid overlap or conflict.
[0097] As described above, in one example, the de-jitter buffer delay is computed or updated at the start of the speech stream. Constrain the update of the de-jitter buffer delay to the beginning of the stream
The speech can be limiting, however, as speech streams often differ in length and operating conditions may also change during the speech stream. Consider the example of Fig. 10. You may need to update the de-jitter buffer delay during the speech stream.
[0098] It should be noted that it is desirable to control the data stream output from the adaptive de-jitter buffer to maintain a target delay length. Thus, if the adaptive de-jitter buffer receives data with varying delays, the data output from the adaptive de-jitter buffer is adjusted to allow the buffer to fill with sufficient data to achieve the target length of the adaptive de-jitter buffer. Time warping may be used to stretch packets when the adaptive de-jitter buffer receives an insufficient number of packets to maintain the target delay length. Similarly, time warping can be used to compress packets when the adaptive de-jitter buffer receives too many packets and keeps packets above the target delay length. The adaptive de-jitter buffer may work in coordination with the decoder to time warp the packets as described herein.
[0099] Figure 11 is a block diagram of a system including two receivers communicating via a network element. These receivers are the AT terminal 252 and the AT terminal 282; as illustrated, the AT 252 and the AT 282 are adapted to communicate via the BS 270. At the AT 252, the transmitting processing unit 264 transmits the voice data to an encoder 260 which digitizes the voice data and outputs the packet data to a lower layer processing unit 258. Packages are then
Sent to BS 270. When the AT 252 receives data from BS 270, the data is first processed in a lower layer processing unit 258 from which data packets are delivered to adaptive de-jitter buffer 256. Received packets are stored in the adaptive de-jitter buffer 256 until the target de-jitter buffer length is reached. Upon reaching the target de-jitter buffer length, the adaptive de-jitter buffer 256 sends the data to decoder 254. In the illustrated example, compression and stretching as part of the time warping implementation may be performed at decoder 254 which converts packet data to voice data and outputs voice data to the receiving processing unit 262. In another example of this invention, compression and time stretching (time warping) may be performed in an adaptive de-jitter buffer by a controller (not shown). Operation of the AT 282 is similar to that of the AT 252. The AT 282 transmits the data on the path from the transmitting processing unit 294 to the encoder 290, to the lower layer processing unit 288, and ultimately to the base station BS 270. An AT 282 receives data on a path from a lower layer processing unit 288 to an adaptive de-jitter buffer 286 to a decoder 284 to a receive processing unit 292. Further processing is not illustrated but may affect the playback of data such as voice and may include processing audio, screen displays and the like.
[0100] The de-jitter buffer equations in equation (1) calculate the de-jitter buffer delay at the beginning of the talkspurt. The de-jitter delay may represent a specific number of packets, e.g. as determined by the speech streams, or may represent an expected time equivalent to playback
53 / 59P27611PL00 of data such as voice data. It should be noted here that the de-jitter buffer has a target size and this determines the amount of data that the de-jitter buffer expects to be stored at all time points.
[0101] Variation in packet delay due to channel conditions and other operating conditions can lead to differences in packet arrival time at the adaptive de-jitter buffer. Consequently, the amount of data (number of packets) in the adaptive de-jitter buffer may be less than or greater than the calculated de-jitter buffer delay value, DEJITTER_DELAY. For example, packets may arrive at the de-jitter buffer at a slower or faster rate than they were initially generated at the encoder. When packets arrive at the de-jitter buffer at a slower rate than expected, the de-jitter buffer may run out as the incoming packets do not complete the outgoing packets at the same rate. Alternatively, if packets arrive at a rate greater than the generation rate at the encoder, the de-jitter buffer may begin to increase in size as packets do not exit the de-jitter buffer as soon as they arrive. The former condition may lead to underflow, while the latter condition may cause large end-to-end delays due to higher buffering times in the de-jitter buffer. The second condition is important because if the delay between endpoints of the packet data system is decreasing (the AT moves to a less busy area or the user moves to an area with better channel quality), it is desirable to implement this delay reduction to speech reproduction. The end-to-end delay is an important factor in speech quality and any reduction in playback delay is seen as an increase in the quality of the conversation or speech.
53 / 59P27611PL00
[0102] To correct for a discrepancy in the de-jitter buffer between the value of DEJITTER_DELAY and the amount of data actually present in the de-jitter buffer, one example of the de-jitter buffer uses time warping. Time warping relates to stretching or compressing the duration of the speech packet. The de-jitter buffer performs time warping by stretching speech packets when the adaptive de-jitter buffer begins to run out and compressing speech packets when the adaptive de-jitter buffer becomes larger than DEJITTER_DELAY. The adaptive de-jitter buffer may work in coordination with the decoder to time warp the packets. Time warping provides a substantial improvement in speech quality without increasing the end-to-end delay.
[0103] Fig. 12 is a block diagram of an example of an adaptive de-jitter buffer performing time warping. The physical layer processing unit 302 provides data to the data stack 304. The data stack 304 forwards the packets to the adaptive de-jitter buffer and control unit 306. The forward link bearer access (MAC) processing unit 300 provides the de-jitter buffer processing unit 306 with a handoff indication. The MAC layer implements the protocols for receiving and sending data in the physical layer, i.e. via radio communication. The MAC layer may contain information related to security, encryption, authentication, as well as connection information. In an IS856-enabled system, the MAC layer contains rules governing the operation of the Control Channel, Access Channel as well as Forward and Reverse Traffic Channels. The target length estimator 314 provides a target buffer length
The de-jitter into the de-jitter buffer using the calculations in Equation 1. Input to the target length estimator 314 includes packet arrival information and a current packet error rate (PER). It should be noted that different configurations may include a target length estimate 314 within the adaptive de-jitter buffer and control unit 306.
[0104] In one example, the adaptive de-jitter buffer and control unit 306 further include a reproduction control that controls the rate of delivery of data to be played back. From the adaptive de-jitter buffer and control unit 306, packets are sent to a Discontinuous Transmission (DTX) unit 308, where the DTX unit 308 provides background noise information to decoder 310 when speech data is not received. It should be noted that the packets provided by the adaptive de-jitter buffer and control unit 306 are decodable and may be referred to as a vocoder packet. Decoder 310 decodes the packets and provides Pulse Code Modulation (PGM) speech samples to a time warping unit 312. In alternative examples, time warping unit 312 may be implemented within decoder 310. A time warping unit 312 receives the time warping indicator from the adaptive de-jitter buffer and control unit 306. The time warping indicator may be a control signal, an instruction signal, or a flag. In one example, the time warp indicator may be a multi-state indicator having, for example, compressed, stretched, and no time warping states. Different values may be used for different levels of compression and / or different levels of stretching. In one example, the time warp indicator instructs the time warping unit 312 to stretch or compress data. The time alignment indicator indicates stretching, compression or no alignment. Indicator
The time warping unit may be considered to be an operation initiating control signal at the time warping unit 312. The time warp indicator may be a message specifying how packets should be stretched or compressed. The time warping indicator can identify packets to be time warped as well as what operation to perform, stretching or compression. Moreover, the time warping indicator may provide a selection of options for the time warping unit 312. During the silence period, the DTX module modifies the erase stream provided by the de-jitter buffer to a stream of erasure frames and silent frames that are used by the decoder to reconstruct a more accurate and higher quality background noise. In an alternate example, the time warping indicator turns time warping on and off. In yet another example, the index identifies the amount of compression and stretch used in reproduction. The time warping unit 312 may modify the samples from the decoder and provides the samples to the audio processing circuit 316, which may include an interface and conversion unit as well as a driver and an audio loudspeaker.
[0105] While the time warp indicator identifies situations when compression or stretching should be performed, there is a need to know to what extent the time warping should be applied to a given packet. In one example, the time warping value is fixed and the packets are time warped according to the speech cycle or pitch.
[0106] In one embodiment, the time warp index is communicated as a percentage of the target stretch or compression level. In other words, the time warp index tells you to compress by a given percentage or stretch by a given percentage.
53 / 59P27611PL00
[0107] In one scenario, it may be necessary to recognize a known property of the incoming data. For example, an encoder can predict data of known pitch or for example having specific length properties. In this situation, since a particular property is envisaged, it would not be desirable to modify the received data with time warping. For example, an encoder can expect incoming data to have a particular pitch length. However, if time warping is on, the length of the tone may be modified by time warping. Therefore, in this scenario, time warping should not be enabled. Tone-based communication includes, but is not limited to, Teletype / Communication Device for the Deaf (TTY / TDD) information, applications that use keyboard input, or other applications that use tone-based communications. In this type of communication, the length of the tone carries information, and therefore modifying the pitch or length of the tone, such as compression or playback stretching, may result in the loss of this information. In TTY, TDD, and other applications that enable reception by hearing impaired people, the decoder also provides its in-band processing status for this type of communication. This indication is used to mask the time warp indications provided by the de-jitter buffer. If the decoder processes packets with TTY / TDD information, time warping should be turned off. This can be done in two ways: by providing the TTY / TDD state to the de-jitter buffer controller, or by providing the TTY / TDD state to the time warping unit. If the TTY / TDD state of the decoder is provided to the de-jitter buffer controller, the controller should not indicate any stretch or compression indication when the vocoder indicates TTY / TDD processing. If the decoder TTY / TDD status is
Provided to the time warping unit, this acts as a filter and the time warping unit does not operate in the time warping indications if the decoder is processing the TTY / TDD information.
[0108] In the system shown in FIG. 12, the adaptive de-jitter buffer and control unit 306 monitors the rate of the incoming data and generates a time warping indicator when too many or too few packets are available or buffered. The adaptive de-jitter buffer and control unit 306 determines when to time warp and what action should be taken. In fig. 13A illustrates the operation of one example of adaptive de-jitter buffer making a time warping decision using compression and stretching thresholds. The de-jitter buffer stores packets that may have arrived at irregular intervals. Target de-jitter buffer length estimator 314 generates a target de-jitter buffer length; the target de-jitter buffer length is input to the de-jitter buffer. In practice, the adaptive de-jitter buffer and control unit 306 uses the de-jitter buffer length value to make control decisions about the operation of the de-jitter buffer and to control playback. The compression and stretching thresholds indicate when compression and stretching are enabled, respectively. These thresholds may be specified as a fraction of the target de-jitter buffer length.
[0109] As shown in Fig. 13A, a target de-jitter buffer length is given as L.<sub>Targe</sub>t · The compression threshold is given as T.<sub>What</sub>mpress / and the stretch threshold is given as T.<sub>Expand</sub>. When the length of the de-jitter buffer is increased above the compression threshold, T<sub>What</sub>mpress / then the de-jitter buffer indicates to the decoder that the packets should be compressed.
53 / 59P27611PL00
[0110] Similarly, when the length of the de-jitter buffer decreases below the stretching threshold, T<sub>Expand</sub>, the de-jitter buffer indicates to the decoder that packets should be stretched and effectively reconstructed at a lower rate.
[0111] The operating point between the stretch and compression thresholds avoids underflow as well as excessive delay increments between endpoints. Therefore, the target operating point lies between the T thresholds<sub>What</sub>m<sub>P.</sub>re<sub>SS</sub> and T.<sub>E.</sub>x<sub>bye</sub>n / a In one example, the stretch and compression threshold values are set to 50% and 100% of the de-jitter target buffer value. Although in one example the time warping may be performed inside the decoder, in alternative examples the function may be performed outside the decoder, e.g. after the decoding step. However, it may be simpler to implement the time warping of the signal before synthesizing the signal. If such time warping methods were to be used after decoding the signal, the pitch period of the signal should be estimated.
[0112] In some scenarios, the length of the de-jitter buffer may be longer, such as in a W-CDMA system. The time warp threshold generator may generate a plurality of compression and stretching thresholds. These thresholds can be calculated in response to working conditions. Fig. 13B shows multi-level thresholds. T.<sub>C.</sub>i is the first compression threshold, T.<sub>C.</sub>2 is the second compression threshold, and T.<sub>C.</sub>3 is the third compression threshold. The values of T are also illustrated<sub>Ei</sub>, T<sub>E2</sub>, T<sub>E3</sub> representing three different stretch threshold values. These thresholds can be based on the percentage of time warping (how many packets are time warped), on compressed packets, on percentage of stretched packets, or on the ratio of the two values. The number of thresholds can be changed in
As needed, in other words more or fewer thresholds may be needed. Each of these thresholds is associated with a different degree of compression or stretching, for example, for systems requiring high granularity, more thresholds may be used, and for lower granularity, fewer thresholds may be used. T.<sub>E2</sub>, T<sub>E2</sub> and T.<sub>E3</sub> etc. may be a function of the target delay length. The threshold can be changed by tracking the delay underflow and based on error statistics such as the PER parameter.
[0113] Fig. 14 shows packet recovery with and without time warping. In Fig. 14, PKT 1 is transmitted at time t<sub>2</sub>, PKT 2 is sent at time t<sub>2</sub> and so on. Packets arrive at the receiver as indicated with PKT 1 arriving at time t<sub>2</sub>'and PKT 2 arrives at t<sub>2</sub>'' For each packet, the play time without applying time warping is given as PLAYBACK WITHOUT MATCHING. In contrast, the time warping playback time is given as MATCH PLAY. Since the present example deals with real-time data such as speech communication, the anticipated packet recovery time is at predetermined intervals. During playback, ideally each packet arrives before the estimated playback time. If a packet arrives too late for recovery within the estimated time, recovery quality can be affected.
[0114] PKT 1 and 2 are received on time and are played back without time warping. PKT 3 and PKT 4 are received at the same time tU. The reception time for both packets is satisfactory as each packet is received before the respective predicted playback time, tU 'for PKT 3 and ts' for PKT 4. Packets 3 and 4 are rendered timely without applying time warping. Problem
Appears when, at time te ', packet PKT 5 is received after predicted reproduction time has elapsed. Instead of PKT 5, the erase at the estimated playback time is recreated. PKT 5 arrives later after the erasure has started playing.
[0115] In the first scenario without time warping, PKT 5 is discarded and PKT 6 is received and played back at the next scheduled reproduction time. Note that in this case PKT 6 was received in time for playback. In the second scenario, if PKT 5 and all packets following PKT 5 are delayed, then each packet may arrive too late for anticipated reproduction and result in a sequence of erasures. In both of these scenarios, information is lost, i.e., PKT 5 is discarded in the first scenario, and PKT 5 and subsequent packets are discarded in the second scenario.
[0116] Alternatively, the use of the IBA technique allows PKT 5 to be reproduced at the next predicted reproduction time while continuing playback of subsequent packets from that point. The IBA technique prevents data loss but delays the packet stream.
[0117] Such out-of-time reproduction can increase the overall delay between endpoints in a communication system. As shown in Fig. 14, inter-packet delays can cause loss of information or delays in playback.
[0118] By implementing time warping when PKT 5 arrives after its estimated reproduction time, packets are stretched and puncturing can be avoided. For example, stretching PKT 4 may result in playback in 23 ms instead of 20 ms. PKT 5 is played when it is received. This happens faster than it would be played if an erase was sent instead (such as
53 / 59P27611PL00 is illustrated in one alternative for no time warping but with IBA as described in Fig. 14). Stretching PKT 4 instead of sending an erasure results in less loss of playback quality. Time warping therefore provides better overall reproduction quality as well as reduced latency. As illustrated in Fig. 14, packets following PKT 5 are reconstructed earlier using time warping than without using a time warping technique. In this particular example, PKT 7 is played back at time t<sub>9</sub>when time warping is applied, which is earlier than without using time warping.
[0119] One application of time warping is to improve the reproduction quality taking into account changes in operating conditions as well as changes in properties of the information transmitted in speech transmission. As speech properties change, for the speech stream and silence periods, the target de-jitter delay length and the compression and stretching thresholds for each data type may be different.
[0120] Fig. 15 shows examples of "silence compression and" silence stretching due to delay differences due to jitter elimination between one speech stream and another. In Fig. 15, shaded areas 120, 124, and 128 represent speech streams and unshaded areas 122 and 126 represent silence periods of the received information. The received speech stream 120 begins at ti and ends at ti<sub>2</sub>. At the receiver, a de-jitter delay is introduced, and therefore playback of speech stream 120 begins at time t<sub>2</sub>'. The de-jitter buffer delay is identified as the difference between time t<sub>2</sub>'a moment ti. The received silence period 122 begins at time t<sub>2</sub> and ends in cńwili vol<sub>3</sub>. The silence period 122 is compressed and recreated as a silence period from 132 to
53 / 59P27611PL00 of the moment t<sub>2</sub>'until t<sub>2</sub>'that is shorter than the initial duration of the received silence period 122. The stream of speech 124 begins at t<sub>2</sub> and ends in cńwili vol<sub>4</sub> in the source. Speech stream 104 is played back at the receiver from time t<sub>2</sub>'until at t<sub>4</sub>'. The period of silence 126 (time from t<sub>4</sub> to vol<sub>2</sub>) is stretched at the receiver on playback as a silence period 136, the period (t<sub>5</sub>'- vol<sub>4</sub>') is longer than the period (t<sub>5</sub> - vol<sub>4</sub>). The silence period can be compressed when the de-jitter buffer needs to recover packets sooner, and stretched when the de-jitter buffer needs to delay packet playback. In one example, compressing or stretching the periods of silence slightly degrades voice quality. Thus, adaptive de-jitter delays may be achieved without degrading voice quality. In the example of Fig. 15 The adaptive de-jitter buffer compresses and stretches the silence periods identified and controlled by the adaptive de-jitter buffer.
[0121] It should be noted that the term time warping as used herein refers to adaptive reproduction control in response to the arrival time and the length of data received. Time warping may be performed using data compression for playback, data stretching for playback, or by using both compression and data stretching for playback. In one example, a threshold is used to initiate compression. In another example, the threshold is used to trigger the stretching. In yet another example, two triggers are used: one for compression and one for expansion. In still other examples, multiple trigger signals may be used to indicate different levels of time warping, e.g., fast playback at different rates.
[0122] Time warping may also be performed inside a decoder. Techniques to perform time warping in
53 / 59P27611PL00 decoder are described in co-pending Patent Application No. 11 / 123,467, entitled "Time Warping frames inside the Vocoder by Modifying the Residual, filed May 5, 2005.
[0123] In one example, the time warping comprises a method "merging speech segments. Merging the speech segments includes comparing the speech samples in at least two consecutive speech segments and, if a correlation is found between the compared segments, creating a single segment from the at least two consecutive segments. Speech merging is performed while trying to maintain speech quality. Preserving speech quality and minimizing the introduction of artifacts such as sounds that degrade the quality for the user, including "tapping and popping", into the output speech is achieved by carefully selecting the segments to merge. The selection of speech segments is based on the similarity or correlation of the segments. The greater the similarity of the speech segments, the better the resulting speech quality and the lower the likelihood of introducing speech artifacts.
[0124] Fig. 16 shows a speech signal plotted against time. The vertical axis represents the signal amplitude and the horizontal axis represents time. It should be noted that the speech signal has a very characteristic course, with parts of the speech signal repeating over time. In this example, the speech signal includes a first segment from ti to t<sub>2</sub>which repeats as the second segment in time from t<sub>2</sub> to vol<sub>2</sub>. Upon finding such a repeat, one or more segments, such as a segment from time t<sub>2</sub> to vol<sub>2</sub>, can be eliminated with little or no effect on sample reproduction quality.
[0125] In one example, Equation 4, as shown below, may be used to find a relationship between two speech segments. Correlation is a measure of the degree of relationship
53 / 59Ρ27611PL00 between two segments. Equation 4 provides the absolute and limited correlation coefficient (-1 to +1) as a measure of the degree of a relationship, with a small negative number reflecting a weaker relationship, i.e., a smaller correlation, than a large positive number, which reflects a stronger relationship, i.e. a stronger correlation. If the application of Equation 4 indicates "high similarity, then time warping is performed." If the application of Equation 4 indicates little similarity, artifacts may be present in the merged speech segment. The correlation is given by the expression:
Corr (d) = ^ [(λ (0 - mx) x (y (i -d) - my)]
<img file="PL1787290T3_D0001.tif" />
<img file="PL1787290T3_D0002.tif" />
(4)
[0126] In Equation 4, x and y represent two speech segments, m is the window where the correlation between the two segments is calculated, d is the correlation part, and i is the index. If applying equation 4 indicates segments that can be merged without introducing artifacts, the merging can be done using the "add overlap" technique. The "add-overlap" technique combines the compared segments and produces one speech segment from two different speech segments. A combination using the "add-overlap" technique can be based on an equation such as, for example, Equation 5:
OutSegment [z] = (SegmentUj) * (WindowSize - i) + (Segment2 (i) * i) WindowSize (Segment2 (i) * (WindowSize -1) + (SegmenLoj * ij
WindowSize (5) and - 0. WindowSize -1 WindowSize = R WindowSize
[0127] The resulting samples may be Modulation samples
Pulse Code (PGM). Each PGM sample has a predetermined format defining the bit length and format of the PGM sample. For example, a signed 16 bit number may be a format
53 / 59P27611PL00 representing a PGM sample. The add-overlap technique obtained by using Equation 5 includes weighting to obtain a smooth transition between the first segment 1 PGM sample and the last segment 2 PGM sample. In Equation 5, "RWindowSize is the number of PGM samples in the reference window and" OutSegment is the resulting segment size. " obtained by the add-overlap technique. The value of "WindowSize is the size of the reference window and the value of" Segment is the target size of the segment. " These variables are determined depending on the sampling rate, the content of the frequency in speech and the desired trade-off between quality and computational complexity.
[0128] The add-overlap technique described above is illustrated in Figs. 17A and 17B. Fig. 17A shows a speech segment made up of 160 PGM samples. In this example, the RWindowSize variable is represented by PGM samples 0-47. In other words, PGM samples 0-47 correspond to the number of samples in the reference window of WindowSize. The Segment variable corresponds to the size of the target search area and is represented by PGM samples 10-104. In this example, PGM samples 0-47 are compared with samples 10-104, one PGM sample at a time, to find the best correlation between the reference samples and the search target area. The position in the search area where the maximum correlation was found is referred to as the "offset. At the offset point, the RWindowSize variable can be combined with the part of the Segment variable corresponding to the size of RWindowSize. The speech segment corresponding to PGM 104 - 160 remains intact.
[0129] In Fig. 17B, the first RWindowSize samples of the speech segment are compared to successive portions of the speech segment, one PGM sample at a time. The location where the maximum correlation between RWindowSize and the corresponding sample length within the target search area is found
53 / 59P27611PL00 (Segment) is designated as "offset. The offset length is the distance from the start of the speech segment to the point of maximum correlation between RWindowSize and Segment. Once the maximum correlation is found, RWindowSize and the appropriate Segment length (at the offset point) are merged. In other words, the add-overlap technique is performed by adding RWindowSize to a portion of a Segment of the same length. This is done at the offset point as shown in the illustration. The rest of the samples are copied from the original segment as illustrated. The resulting speech segment consists of the remaining samples copied as-is from the original speech segment, attached to the merged segment as illustrated. The resulting packet is shorter than the original segment by the shift length. This process is known as speech compression. The smaller the speech segment is compressed, the less likely it is that the user will notice any deterioration in quality.
[0130] Speech stretching is performed when the de-jitter buffer contains a small number of voice packets. The underflow is more likely to occur if the de-jitter buffer contains a small number of packets. When an underflow occurs, the de-jitter buffer may provide erasure to the decoder. However, this leads to a deterioration in voice quality. To prevent such degradation of voice quality, playback of the last few packets in the de-jitter buffer may be delayed. This is achieved by stretching the packets.
[0131] Speech distension may be achieved by repeating multiple PGMs of a speech segment. Repeating multiple PGM samples while avoiding artifacts or pitch flattening is achieved by processing more samples
Speech PGM than when speech compression is performed over time. For example, the number of PGM samples used for
Speech extension implementations may be twice the number of PGM samples used in speech time compression. Additional PGM samples may be obtained from the previous playback speech packet.
[0132] Fig. 18A shows one example of speech extension where each speech packet or segment is 160 PGM samples in length and "pre-stretched speech segment is generated." In this example, two speech segments are compared: "current speech segment and" previous speech segment. The first RWindowSize PGM samples of the current speech segment are selected as reference samples. These RWindowSize samples are compared with the Segment variable of the previous speech packet, where the maximum correlation point (or offset) is determined. The RWindowSize of PGM samples is added overlapping to the appropriate Segment size within the previous packet at the offset point. The pre-stretched speech segment is created by copying and appending the rest of the samples from the previous speech segment to the segment added with the overlay as shown in Fig. 18A. The length of the stretched speech segment is then the sum of the length of the pre-stretched speech segment and the length of the current speech segment as illustrated in Fig. 18A. In this example, the PGM samples are offset from the start of the speech segment.
[0133] In another example, the current packet or speech sample is stretched as shown in Fig. 18B. The reference samples, RWindowSize, are positioned at the beginning of the current speech segment. RWindowSize is compared to the rest of the current speech packet until a point of maximum correlation (offset) is located. Reference samples are added overlapping to the corresponding PGM samples found to have the maximum correlation within the current speech segment. The extended speech segment is then created by copying the PGM samples,
Starting from the beginning of the packet to the offset point, appending the added overlapping segment to it, and copying and appending the remaining PGM samples, unmodified, from the current packet. The length of the extended speech segment is equal to the sum of the offset plus the length of the original packet.
[0134] In another example, speech is stretched as shown in Fig. 18C, where a window of RWindowSize is contained within the current packet or speech segment and is not at the beginning of the packet. Roffset is the length of the speech segment corresponding to the distance between the start of the current packet and the point where RWindowSize begins. RWindowSize is added overlapping to the corresponding PGM sample size in the current packet at the point of maximum correlation. An extended speech segment is then created by copying the PGM samples starting at the beginning of the original or current packet and ending at the offset point and appending the added overlapping segment and the rest of the PGM samples from the original packet. The length of the resulting stretched speech segment is the sum of the original packet length and the offset minus the Roffset samples, that is, the number of PGM samples in Roffset, as defined above.
Filtered time warp thresholds
In order to avoid oscillating compression and stretching decisions when the number of packets retained in the adaptive de-jitter buffer changes rapidly, the variables used to evaluate the state of the adaptive de-jitter buffer, that is, the number of packets stored in the adaptive de-jitter buffer, filters in one example these kinds of variables appear in the sampling window. Adaptive de-jitter buffer status can refer to the number of packets stored in the adaptive de-jitter buffer or any
Variables used to evaluate the data stored in the adaptive de-jitter buffer. In a system supporting packet data delivery, IS-856, referred to as lxEV-DO, the delivery of packets to a given receiver is time-multiplexed on the forward link, the receiver may receive several packets simultaneously, after which no packets are received for some time. This results in the receipt of packet data at the adaptive de-jitter buffer at the receiver. The received data is effectively "packaged" where there may be two or more packets arriving closely together in time. Such packet forming can easily lead to an oscillation between packet stretching and compression where the adaptive de-jitter buffer issues time warping instructions in response to received data rate and buffer status. For example, consider an example where the computed de-jitter buffer value (delay or length) is 40 ms at the start of the speech stream. Thereafter, the load on the de-jitter buffer drops below the stretching threshold, causing a decision to stretch the data packet. Immediately after playing this package, a bundle of three packages arrives; the incoming data fill the size of the de-jitter buffer such that the compression threshold is exceeded. This will compress the packages. Since the arrival of the packet packet may be preceded by a certain period in which no packets are received, the de-jitter buffer may exhaust again, causing packets to stretch. This kind of switching between stretching and compression can cause a large percentage of packets to be time warped. This is undesirable as we would like to limit the percentage of packets in which the contained information has been modified due to time warping to a small value.
53 / 59P27611PL00
[0136] In one example, such oscillation is avoided by mitigating the effect that packetizing may have on adaptive de-jitter buffer control, time warping and data recovery. This example uses average values to determine when to perform time warping. These average values are computed by filtering the variables used in this type of calculation. In one example, the compression and stretching thresholds are determined by filtering or averaging the size of the de-jitter buffer. Note that the size of the buffer refers to the current state of the buffer.
[0137] Comparing the filtered buffer size value with the stretching threshold may result in more underflow as some packets that would be stretched with the unfiltered value are not stretched with the filtered value. On the other hand, comparing the filtered value with the compression threshold can serve to dampen most of the oscillations (or switch between time warping controls) with minimal or no negative influence. Therefore, the compression and stretch thresholds may be treated differently.
[0138] In one example, the instantaneous adaptive de-jitter buffer size value is checked for a stretching threshold. In contrast, the filtered de-jitter buffer value is checked against a compression threshold. One configuration uses an infinite impulse response (IIR) filter to determine the average size of the adaptive de-jitter buffer, wherein the adaptive de-jitter buffer has a filtered value that may be recalculated periodically, for example every 60 ms. The filter time constant can be obtained from the package building statistics and
An example value for 1xEV-DO Rev A could be 60 milliseconds. The packet statistics are used to derive the filter time constant as they strongly correlate with how the instantaneous de-jitter buffer size oscillates during operation.
Stretching as a result of a lost package
[0139] As noted above, the adaptive de-jitter buffer and various adaptive de-jitter buffer control and time warping control methods for received data can be adapted to specific system properties and operating conditions. For communication systems implementing a repetition request scheme to improve performance, such as a Hybrid Automatic Retry Request (H-ARQ) scheme, such repetition processing has the effect of stretching the speech packet. In particular, the H-ARQ scheme can cause packets to arrive with a change of order (i.e., in a different order). Consider Fig. 19 illustrating a de-jitter buffer of a length and a stretch threshold T<sub>Expand</sub>, specified as 50% of the target de-jitter buffer length. The current packet being played back has sequence number 20, PKT 20. The de-jitter buffer contains three packets with sequence numbers 21, 23 and 24, labeled PKT 21, PKT 23 and PKT 24, respectively. When the reproduction apparatus requests the next packet after PKT 20 has been reproduced, the stretching threshold is not reached because the de-jitter buffer contains enough packets to maintain the buffer length greater than 50% of the calculated de-jitter buffer length. Therefore, PKT 21 is not stretched in the present example. This can cause an underflow if PKT 22 does not arrive before PKT 21 has finished playing since the packets are played sequentially and therefore
The playback apparatus cannot restore PKT 23 before PKT 22. Even if the stretch threshold has not been exceeded, one example predicts a discontinuity in received packets and chooses to stretch PKT 21 to give more time for the arrival of PKT 22. W thus, stretching PKT 21 can avoid packet skipping and erasure. Thus, the packet can be stretched even if the length of the de-jitter buffer is greater than the stretch threshold T<sub>Expand</sub>.
[0140] Conditions under which packets are to be stretched can be enhanced. As described above, a packet may be stretched if the size of the de-jitter buffer is smaller than the stretching threshold. In another scenario, a packet may be stretched if no packet with the sequential sequence number is present in the de-jitter buffer.
[0141] As mentioned before, the de-jitter buffer delay may be computed at the beginning of the speech stream. Since network conditions, including, but not limited to, channel conditions and load conditions, may change over the course of the speech stream, especially over the course of a long speech stream, one example is configured to change the buffer delay. de-jitter during the duration of the speech stream. Thus, the de-jitter buffer equations given above may be periodically recalculated every CHANGE_JITTER_TIME seconds for the duration of the speech stream. Alternatively, the variables may be recalculated when the event that causes them occurs, such as a significant change in operating conditions, load, air interface indication, or other event. In one example, the CHANGE_JITTER_TIME value may be set to 0.2 seconds (200 ms).
[0142] The time warping thresholds, that is, the compression and extension thresholds, can provide an indication of how to vary
53 / 59P27611PL00 values during the speech stream. Normal operation applies to receiver operation when the adaptive de-jitter buffer state is between the compression and stretching thresholds and about the target de-jitter buffer length. Each threshold acts as a trigger. When the threshold is reached or exceeded, packets present in the adaptive de-jitter buffer may be stretched or compressed depending on the threshold. The adaptive de-jitter buffer size may continue to stretch or contract as packets are received. This constant resizing of the adaptive de-jitter buffer indicates that there may be continual approach to the stretch and compression thresholds during communication. In general, the system attempts to maintain the size of the adaptive de-jitter buffer between the stretch and compression thresholds, which is considered a stable state. In a steady state, the size of the adaptive de-jitter buffer is not changed; and a change in packet reception and thus a change in the adaptive de-jitter buffer size may automatically cause the compression / stretching threshold to be reached and the packets to compress / stretch accordingly, until the new adaptive de-jitter buffer delay is reached. In this scenario, the target delay length of the adaptive de-jitter buffer is updated according to the CHANGE_JITTER_TIME value. The actual de-jitter buffer size need not necessarily be calculated, since the de-jitter buffer size automatically changes on trigger as a result of reaching any of the time warp stretch / compression thresholds. In one example, the CHANGE_JITTER_TIME value may be set to 0.2 seconds (200 ms).
PRE-ADJUSTMENT OF TIME ON CALL TRANSFER
53 / 59P27611PL00
[0143] Handoff events are typically accompanied by a loss of coverage for a short period of time. As a handoff is approaching, the AT may experience poor channel conditions and increased packet delay. In one example, handoff conditions are handled in a special way by applying time warping on speech packets. Once the AT decides to hand over the call to the new base station, this information can be used to control the de-jitter buffer. Upon receipt of such handoff signal, the AT enters a "pre-matching mode as shown in pre-matching mode 244 of Fig. 8B." In this mode, the AT terminal stretches packets until either of two conditions is met. On the first condition, the de-jitter buffer continues to accumulate packets, and the accumulation increment results in a de-jitter buffer size of PRE_WARPING_EXPANSION. In other words, packet stretching is performed until the value of PRE_WARPING_EXPANSION is reached. Alternatively, for the second condition, the time period of WARPING_TIME has been reached. The timer starts upon receipt of a handover signal or an idle indicator; the timer stops at the moment WARPING_TIME. Upon meeting either of these two conditions, the AT exits the time warping mode. During pre-warping mode, no packets are compressed unless the End_Talkspurt condition (described later) is met, as the de-jitter buffer wants to accumulate enough packets to send them at regular intervals to the playback device. In an example where packets are expected at regular intervals, such as 20 ms, the PRE_WARPING_EXPANSION value may be set to 40 ms and the WARPING_TIME value may be 100 slots (166 ms).
53 / 59P27611PL00
[0144] Call handovers are only one form of downtime events. The de-jitter buffer may implement call forwarding or some other type of downtime. The information required for this is how much overhead is required to handle downtime (PRE_WARPING_EXPANSION) and how long the de-jitter buffer will be in this downtime avoidance mode (WARPING_TIME).
COUNTING DELAY INFLUENCES
[0145] Since the adaptive de-jitter buffer equations given above are designed to achieve a target delay underflow percentage, it is desirable to measure the number of delay underflows accurately. When an underflow occurs, it is not known whether the underflow was caused by packet delay or by dropping a packet somewhere on the network, that is, in the transmission path. Therefore, there is a need to precisely account for the type of underflow.
[0146] In one example, for communication using RTP / UDP / IP, each packet includes an RTP sequence number. Sequence numbers are used to arrange the received packets in the order in which they were transmitted. When an underflow occurs, the RTP sequence number of the causing underflow may be stored in a memory, such as a memory matrix. If a packet with an identified sequence number arrives later, this underflow is counted as "delay underflow."
[0147] "The delay underflow factor is the ratio of the number of underflows to the total number of packets received. The underflow number and the number of packets received are set to zero each time the de-jitter buffer equations are updated.
Streamline the start or end of a speech stream
53 / 59P27611PL00
[0148] Consider Fig. 20, which illustrates the timing of a conversation between two users. In this graph, the vertical axes represent time. Each user transmits speech streams and silence periods which are then received by the other user. For clarity, the shaded block segments 400 and 410 represent User 1 speech streams (speech segments). The non-shaded block segment 405 represents User 2 speech streams. The areas outside of the speech streams on the time waveform represent periods of time when the users are not talking but may be listening to the other user or receiving a period of silence. Segment 400 is played for User 2. After segment 400 is played for User 2, User 2 waits for a short period of time before speaking. The start of User 2's first speech segment 405 is then heard by User 1. The Conversational Round Trip Delay (RTD) as perceived by User 1 is the period of time between when User 1 stops speaking and when User 1 hears the start of User 2's speech segment. A conversational RTD is not a one-way delay between endpoints, but is user-specific and relevant to users. For example, if the conversational RTD is too great for User 1, it will prompt User 1 to start speaking again without waiting for User 2's speech segment to be played. This breaks the conversation stream and is seen as a deterioration in the quality of the conversation.
[0149] The conversational RTD delay experienced by
User 1 can be changed in various ways. In one example, the time that the speech segment of User 1 is played back to User 2 may be changed. In the second example, the time during which User 1 is played back to User 2's speech segment is changed. Note that
That delays of the speech stream start and end times alone affect the voice quality of the conversation. The design goal is to further reduce delays at the beginning and end of the speech stream.
[0150] In one example, the goal is to improve the start of the speech stream. This improvement may be achieved by manipulating the first packet of User 1's speech stream so that the listener User 2 receives the packet faster than the default implementation of adaptive de-jitter buffer delay. The delay applied to the packet in the adaptive de-jitter buffer may be the default adaptive de-jitter buffer delay, a computed value, or a value chosen so that the listener receives the packet at a given time. In one example, the timing of the first speech stream packet is changed by recalculating the adaptive de-jitter buffer delay at the start of each received speech stream. When the adaptive de-jitter delay applied to a first speech stream packet is reduced, the first packet is sent to the listener. As the applied delay is increased, the first packet is received by the listener later. The default de-jitter buffer delay for the first packet may be less than the calculated de-jitter buffer delay, and vice versa. In the illustrated example, the de-jitter delay of the first packet of each talkspurt is limited by a value specified as MAX_BEGINNING_DELAY, which can be measured in seconds. This value may be a recalculated de-jitter delay value or a delay designed for the listener to receive the packet within the specified time. The MAX_BEGINNING_DELAY value may be less than the actual calculated buffer delay value
53 / 59P2761 de-jitter 1PL00. When the MAX_BEGINNING_DELAY value is less than the calculated de-jitter delay and is applied to the first speech stream packet, subsequent speech stream packets will be automatically stretched. Packets are automatically stretched from one packet to the next because the de-jitter buffer cannot receive packets at the same rate as it retrieves packets. When the de-jitter buffer recovers the packets, the de-jitter buffer reduces in size and approaches the stretching threshold. After the stretching threshold is reached, stretching is triggered and subsequent speech stream packets are stretched until the de-jitter buffer has received enough incoming packets to exceed the stretching threshold. As a result of using the MAX_BEGINNING_DELAY value, the first packet of the speech stream is received by the listener faster and the subsequent packets are stretched. The listener is pleased to receive the initial packet faster. Improving the start of a speech stream has a tendency to slightly increase the number of underflow; however, a proper value of MAX_BEGINNING_DELAY alleviates this effect. In one example, the MAX_BEGINNING_DELAY value is calculated as a fraction of the actual jitter suppression target; for example, a MAX_BEGINNING_DELAY value of 0.7 of the JITTER BUFFER LENGTH ELIMINATING TARGET value may lead to a slight increase in the underflow number. In another example, the MAX_BEGINNING_DELAY value may be a fixed value such as 40 ms, leading to a slight increase in underflow, for example in a system supporting 1xEV-DO Rev A.
[0151] Stretching consecutively in a speech stream does not degrade overall voice quality. This is illustrated in Fig. 20 where User 2 receives the first speech stream packet from User 1 and the initial or "one-way delay is limited to the value of T."<sub>for</sub>. According to what
Illustrated, speech segment 400 is received by User 2 without any stretching or compression, however, speech segment 405 is compressed on reception at User 1.
[0152] Fig. 21 is a flowchart illustrating an improvement to the start of a speech stream. First, in step 510, it is determined whether the system is in silence mode. The silence mode may correspond to a period of silence between speech streams, or a period of time when packets are not received by the de-jitter buffer. If the system is not in silent mode, the process terminates. If the system is in silent mode, in step 520, an estimation of a target de-jitter buffer length is made. Then, in step 530, it is determined if the system is improved. The improvement, according to one example, indicates that the calculated target adaptive de-jitter buffer length is greater than a given value, which in one example is given as an improvement factor, for example, MAX_BEGINNING_DELAY; the system waits for a period of time equal to the improvement factor or a fraction of the target length to start reproducing, in step 540. If the system is not upgraded, then the system waits for a new target value to begin reproducing, in step 550. The new target value may be the computed target de-jitter buffer length or the maximum de-jitter buffer length.
[0153] Fig. 22 also shows an improvement to the start of a speech stream. Process 580 is illustrated starting with identifying a speech stream. Two scenarios are considered: i) with time warping and also ii) without time warping. In this example, 20 ms speech packets are used. Speech packets of any length can be used. In this case, the adaptive de-jitter buffer waits 120 ms before starting packet playback. This value is
The target adaptive de-jitter buffer length is received from the adaptive de-jitter target length estimator in step 582. In the present example, 120 ms is equivalent to receiving six (6) packets of 20 ms each without time warping. If no time warping is used at 584, these six (6) packets are delivered in 120 ms. Therefore, in the first scenario, the de-jitter buffer will start playing after receiving six packets. This is equivalent to a 120 ms delay. In the second scenario, using time warping, the de-jitter buffer may stretch the first four (4) packets received and start packet playback after receiving four (4) packets. Thus, even though the de-jitter delay of 80 ms in this case is less than the estimated 120 ms de-jitter buffer delay, potential underflow can be avoided by stretching the first few packets. In other words, packet playback may start faster with time warping than without time warping. Time warping can then be used to improve the start of a speech stream without affecting the underflow number.
[0154] In another example, the end of the speech stream may be improved. This is done by compressing the last few packets, thereby reducing end-to-end delay. In other words, the delay at the end of the speech stream is reduced and the second user hears back from the first user faster. The enhancement at the end of the speech stream is shown in Fig. 23. Here, the 1/8 code rate packet indicates the end of the speech stream. It differs from full code rate (rate 1), half code rate (rate U) and quarter code rate (rate M) packets that can be used for voice transmission. For transmission in
Packets with other code rates may also be used during periods of silence or at the end of a speech stream. The implementation of 1/8 code rate packets as silence indicator packets in voice communications is described in more detail in concurrent US application number 11 / 123,478, with priority date February 1, 2005, entitled "METHOD FOR DISCONTINUOUS TRANSMISSION AND ACCURATE REPRODUCTION OF BACKGROUND NOISE INFORMATION (A method for discontinuously transmitting and accurately reproducing background noise information).
[0155] As shown in Fig. 23, without time warping, packets N to N + 4 are reconstructed in 100 ms. By compressing the last few packets of the speech stream, the same N to N + 4 packets can be played in 70 ms instead of 100 ms. When time compression is implemented, the quality of speech may be slightly or not deteriorated at all. The enhancement at the end of the speech stream presupposes that the receiver has the knowledge to identify the end of the speech stream and to predict when it is approaching that end.
[0156] When sending according to the Real Time Transmission (RTP) protocol, the "end of talkspurt" indicator may be set in the last packet of each talkspurt. When a packet is delivered for playback, packets in the de-jitter buffer are checked for the "end of speech stream" indicator. If this indicator is set in one of the packets and there are no missing sequence numbers between the current packet delivered for playback and the end of speech stream packet, the packet delivered for playback is compressed as well as all future packets of the current speech stream.
[0157] In another example, the system goes to silence while in speech stream mode and either 1/8 rate packet is delivered to the playback device.
Or packet with a bit set Silence Indicator Description (SID). 1/8 bitrate packet can be detected by checking its size. The SID bit is carried in the RTP header. The system enters speech stream mode when in silent mode and a packet is delivered for playback that is neither a 1/8 rate packet nor has the SID bit set. It should be noted that in one example, the adaptive de-jitter buffering methods presented may be performed while the system is in the talkspurt state, and may be ignored during a silence period.
[0158] It should be noted that this method can correctly discard duplicate packets that are arriving late. If a duplicate packet arrives, it will simply be discarded since the first occurrence of that packet was restored in the correct time and its sequence has not been recorded in a table containing potential "delay gaps."
[0159] When sending voice packets according to the RTP protocol in one example, the end-of-speech indicator may be set in the last packet of each talkspurt. When a packet is delivered for playback, the packets in the de-jitter buffer are checked for the "end of speech stream" indicator. If this indicator is set in one of the packets and there are no missing sequence numbers between the current packet delivered for playback and the end of talkspurt packet, the packet delivered for playback is compressed as well as all future packets of the current talkspurt.
[0160] A flowchart illustrating the end enhancement of a speech stream according to one example is shown in Fig.
24. A new packet starts in step 600. In step 605, if the length of the de-jitter buffer is greater than or equal to a compression threshold, a compression indication is generated in step 635, i to a new packet in step 600,
A tail piece is provided. In step 605, if the de-jitter buffer is not greater than or equal to the compression threshold, it is determined in step 610 as to whether the length of the de-jitter buffer is less than or equal to the stretching threshold. If there is, in step 615 it is determined whether the tail is equal to the packet coding rate which may represent a silence period or end of a speech stream. In one example, a contiguous chain of 1/8 code rate packets may be sent at constant intervals, for example, 20 ms, during a silence period or at the end of a speech stream. As shown in Fig. 24, if it is determined in step 615 that the tail is not equal to a packet by 1/8 code rate, the segment is stretched in step 620 and returns to a new packet in step 600. In step 6, 625, it determines whether the tail is 1/8 the code rate. If in step 625 the trailing portion is 1/8 of the code rate, in step 635 a compression rate is generated. If it is not 1/8 of the code rate then playback proceeds normally in step 630 without applying time warping.
TIME QUALITY OPTIMIZER
[0161] When a number of consecutive packets are compressed (or stretched), this can noticeably speed up (or slow down) the sound and cause quality degradation. This deterioration can be avoided by spacing the time warped packets away from each other, i.e., a few untimed packets follow the time warped packet before another packet is time warped.
[0162] If the above matched packet splitting is applied to stretching, this can cause some packets that would otherwise be stretched not to be stretched. This can lead to underflow,
As packet stretching is performed when the packets in the de-jitter buffer are low. Thus, in one example, the above matched packet splitting may be applied to compressed packets, i.e., the compressed packet may be followed by several uncompressed packets before another packet is compressed. Typically the number of packets that should not be compressed between two compressed packets can be 2 to 3.
Set of conditions for activating time warping
[0163] A number of conditions for enabling time warping (stretching / compressing) of voice packets are described herein. The following is a combined set of rules (in pseudo-code) to determine whether a given packet should be compressed, stretched, or neither.
[0164]
If (in the Initial Fitting Phase (Call Transfer Detected) and Speech End Not Detected) and not reached DEJITTER_TARGET + PRE_WARPING_EXPANSI0N) Stretch Pack End If Otherwise If (End Speech End Detected) If (End Speech End Is Detected) Compression) Compress End If Otherwise If (Triggered Tension Threshold or No Next Packet Queued)
53 / 59P27611PL00
Stretch End If End If End If.
[0165] Fig. 25 shows an implementation of a conventional de-jitter buffer combined with a decoder function. In Fig. 25, packets are expected to arrive at the de-jitter buffer at 20 ms intervals. It is observed in this example that the packets arrive at irregular intervals, that is, with a jitter effect. The de-jitter buffer accumulates packets until a certain length of the de-jitter buffer is reached such that the de-jitter buffer is not exhausted after it starts sending packets at regular intervals such as 20 ms. With the de-jitter buffer length required, the de-jitter buffer starts playing back packets at regular 20 ms intervals. The decoder receives these packets at regular intervals and converts each packet to 20 ms of voice for each packet. In other examples, other time intervals may be selected.
[0166] Fig. 26 shows an example of adaptive de-jitter for time warping for comparison. In this case, packets arrive at the adaptive de-jitter buffer at irregular intervals. In this case, however, the target de-jitter buffer length is much shorter. This is because time warping allows packets to be stretched if the de-jitter buffer begins to run out, allowing time to replenish the adaptive de-jitter buffer. The decoder may stretch packets if the adaptive de-jitter buffer begins to run out and compress packets if the adaptive de-jitter buffer begins to accumulate too many packets. It is observed that uneven delivery of voice packets is a signal
Decoder and time warping unit from adaptive de-jitter buffer. These packets may arrive at irregular intervals as the decoder uses time warping to convert each packet to a different length voice packet depending on the arrival time of the original packet. For example, in this case, the decoder converts each packet into 15-35 ms voice chunks per packet. Since packets can be played back faster due to time warping, the required buffer size is smaller, resulting in lower network latency.
[0167] Fig. 27 is a block diagram illustrating an AT terminal according to one example. Adaptive de-jitter buffer 706, time warp control unit 718, receive circuit 714, processor control 722, memory 710, transmission circuit 712, decoder 708, H-ARQ control unit 720, encoder 716, speech processing unit 724, speech stream ID 726 the error correction circuit 704 may be coupled together as shown in the previous embodiments. Moreover, they can be connected via the communication bus 702 shown in Fig. 27.
[0168] Fig. 28 shows packet processing in one example where packets are received via de-jitter buffer and finally played back by a loudspeaker. As shown, packets are received in a de-jitter buffer. The de-jitter buffer sends packets and time warping information to the decoder after packet requests arrive from the decoder. The decoder sends samples to the output driver upon requests from the output driver.
[0169] An input controller within the de-jitter buffer keeps track of the incoming packets and indicates if there is any error in the incoming packets. Buffer
De-jitter can receive packets that have sequence numbers. For example, the error can be detected by the input controller when the incoming packet has a sequence number that is smaller than the sequence number of the previous packet. A classifying unit located inside the input controller in Fig. 28 classifies the incoming packets. The various categories determined by the classifying unit may fall into categories such as "good packets," delayed packets, "bad packets, and the like. In addition, the control unit may compare the packets and send this information to the elimination buffer controller.
[0170] The de-jitter buffer controller shown in Fig. 28 receives bi-directional input from an input controller and an output de-jitter buffer controller. The de-jitter buffer controller receives data from the input controller, the data indicating characteristics of the incoming data, such as the number of good packets received, the number of bad packets received, and the like. The de-jitter buffer may use this information to determine when the de-jitter buffer should be decreased or increased, which may result in sending a signal to a time warping controller to perform compression or stretching. A packet error rate (PER) unit within a de-jitter buffer controller unit calculates the PER delay. The output de-jitter buffer controller requests packets from the de-jitter buffer. The output de-jitter buffer controller unit may also indicate which packet was most recently played.
[0171] The decoder sends packet requests to the de-jitter buffer and receives packets from the de-jitter buffer after such requests. A time warping controller unit inside the decoder receives the control information
Time warping from the output de-jitter buffer controller. Time warp control information indicates whether packets are to be compressed, stretched, or left unchanged. Packets received by the decoder are decoded and converted into speech samples; and upon request from a buffer within an output driver, the samples are sent to that output driver. Sample requests from the output driver are received by the output controller inside the decoder.
Phase matching
[0172] As previously noted, receiving a packet after its estimated reproduction time may result in recovery of erasures instead of a delayed packet. The receipt of wipes or missing packets in the adaptive de-jitter buffer can result in discontinuities in decoded speech. When the adaptive de-jitter buffer recognizes potential discontinuities, it may request the decoder to perform phase matching. As illustrated in Fig. 28, the adaptive de-jitter buffer 750 may include a phase match controller that receives an input from an output controller 760. Phase warping control information is sent to a phase adjusting unit, which may be located at decoder 762. In one example, phase warping control information may be include the information "phase shift and" run length. " The phase shift is the difference between the number of packets that the decoder has decoded and the number of packets encoded by the encoder. The run length refers to the number of consecutive erasures the decoder has decoded immediately before decoding the current packet.
[0173] In one example, both phase warping and time warping are implemented at the decoder using common control code or software.
53 / 59P27611PL00
In one example, the decoder implements waveform interpolation, with:
a) if no time warping or phase warping is used, vocoding is performed using the variable waveform_interpolation with 160 samples;
b) if time warping but not phase warping is used, vocoding is performed using a waveform_interpolation_decoding with a number of samples of (160 + -N * Pitch Period), where N can be 1 or 2;
c) if no time warping is applied, but phase warping is used, vocoding is performed using the variable waveform_interpolation_decoding with a sample number of (160 - Δ), where Δ is the amount of phase warping;
d) if both time warping and phase warping are used, vocoding is performed using the variable waveform_interpolation_decoding with a number of samples of (160 - Δ + -N * Pitch Period), where Δ is the amount of phase warping;
[0174] A clock input to the output driver determines how often data is requested by a buffer within the output driver. It is the main clock in the system and can be implemented in many different ways. The dominant system clock can be derived from the PGM sample rate. For example, if narrowband speech is being communicated, the system plays back 8,000 PGM samples per second (8 kHz). This clock can clock the rest of the system. One approach is to allow the audio interface 770 to request more samples from the decoder as they are needed. Another approach is to let the decoder / time warping run independently and because this module has information about how many samples
53 / 59P27611PL00
The PGM has been delivered previously and also knows when to re-deliver more samples.
[0175] The scheduler may be located at the decoder 762 or at the audio interface and control unit 810. When located at the audio interface control unit 810, the scheduler bases the next packet request on the number of PGM samples received. When the scheduler is located in the decoder, the scheduler may request packets every t ms. For example, the decoder scheduler may request packets from the adaptive de-jitter buffer 750 every 2 ms. If no time warping is enabled in the decoder, or if the time warping unit is not located at decoder 762, the scheduler sends a sample set to the audio interface and control unit 770 corresponding to the exact number of samples in 1 packet. For example, when audio interface unit 770 requests samples every 2 ms, output decoder controller 766 sends 16 PGM samples (1 packet corresponds to 20 ms of 160 speech data samples at 8 kHz sampling rate). In other words, when the time warp controller is outside the decoder, the decoder output is a normal sample conversion packet. Audio interface unit 770 converts the number of samples into the number of samples it would receive if the decoder performed time warping.
[0176] In another scenario, when the time warp controller is located inside the decoder and time warping is on, the decoder may output fewer samples in the compression mode; and in the stretching mode, the decoder can output a larger number of samples.
[0177] Fig. 30 further shows a scenario where the scheduling function is performed by a decoder. In stage
902, the decoder sends a packet request from the de-jitter buffer. The packet is received in step 904. In step 906, this packet is converted to "N samples." “N samples generated are delivered to the unit
In step 908, in step 910 the next packet request is scheduled as a function of N.
[0178] Fig. 31 shows scheduling beyond decoder in the audio interface and control unit. The audio interface unit first requests the set of PGM samples in step 1002. The requested PGM samples are received in step 1004, and in step 1006 the next packet request is scheduled as a function of N.
[0179] The time warp indicator may be part of an adaptive de-jitter buffer instruction such as, for example, a time warp failure indicator. Fig. 32 shows a time warping unit where the scheduling is computed outside of the decoder, for example in an audio interface and a control unit. The packet type, time warp index, and amount of warp to be performed are input into the time warping unit.
[0180] Fig. 33 shows a time warp unit where the scheduling is calculated per time warp at the decoder. The time warp unit input includes a packet type, a time warp indicator, and a warp amount to be performed. The amount of warping and the inclusion is an input to the quality optimization unit of the time warp unit. The output is time warping information.
[0181] Although specific examples of the present invention have been described herein, those skilled in the art can devise variations of the present invention without departing from the spirit of the invention. For example, the content presented here relates to circuit-switched network elements, but may very well refer to packet-switched network elements. Also, the contents of this specification are not limited to pairs of authentication triplets (Authenticate Triple Octrs), but may also be applied to a single triplet.
53 / 59P27611PL00 including two SRES values (one in the regular format and one in the newer format disclosed here).
[0182] It will be appreciated by those skilled in the art that information and signals may be represented using any of a wide variety of technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips as may be referred to in the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof. .
[0183] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithms described in connection with the examples disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. In order to clearly illustrate this interchangeability of hardware and software, various illustrative elements, blocks, modules, circuits, methods, and algorithms have been described above generally in terms of their functionality. Whether this functionality is implemented on a hardware or software platform depends on the specific application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in various ways for each particular application, but such implementation decisions should not be construed as a departure from the scope of the present invention.
[0184] The various illustrative logic blocks, modules, and circuits described in connection with the examples disclosed herein may be implemented or executed in a general purpose processor, digital signal processor (DSP), specialized integrated circuit (ASIC), directly programmable gate array (ERGA ) or other programmable logic device, discrete logic of gates or transistors, discrete hardware components, or
53 / 59P27611PL00 of any combination of the foregoing, designed to perform the functions described herein. The general purpose processor may be a microprocessor, but alternatively may be a conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of counting devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with the DSP core, or any other configuration.
[0185] The methods and algorithms described in connection with the examples disclosed herein may be implemented directly in a hardware platform, in a software module executed by a processor, or a combination of both. The program module may reside in RAM, flash memory, ROM, EPROM, EEPROM, registers, hard disk, removable disk, CD-ROM, or any other medium known in the art. The medium can be connected to the processor, so that the processor can read and write information on the medium. Alternatively, the data carrier may be integrated with the processor. The processor and the medium may be located on a specialized ASIC.
Contents24
55 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55
62 members in 17 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 60603604 | United States of America | P | |
| 60603604 | United States of America | P | |
| 05794150 | European Patent Office (EPO) | A | |
| 2005030894 | United States of America | W | |
| 2005030894 | United States of America | W | |
| EP20050794150 | – | – | – |
| US20040606036P | – | – | – |
| WO2005US30894 | – | – | – |
Members62
| Document | Office | Kind | |
|---|---|---|---|
| US2006045138A1 | United States of America | A1 | |
| US2006045139A1 | United States of America | A1 | |
| CA2578737A1 | Canada | A1 | |
| CA2691589A1 | Canada | A1 | |
| CA2691762A1 | Canada | A1 | |
| CA2691959A1 | Canada | A1 | |
| US2006050743A1 | United States of America | A1 | |
| WO2006026635A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2006056383A1 | United States of America | A1 | |
| WO2006026635A3 | World Intellectual Property Organization (WIPO) | A3 | |
| TW200629816A | Taiwan Province of China | A | |
| MX2007002483A | Mexico | A | |
| EP1787290A2 | European Patent Office (EPO) | A2 | |
| KR20070065876A | Republic of Korea | A | |
| CN101048813A | China | A | |
| JP2008512062A | Japan | A | |
| BRPI0514801A | Brazil | A | |
| KR20090026818A | Republic of Korea | A | |
| KR20090028640A | Republic of Korea | A | |
| KR20090028641A | Republic of Korea | A | |
| KR100938032B1 | Republic of Korea | B1 | |
| KR100938034B1 | Republic of Korea | B1 | |
| EP2189978A1 | European Patent Office (EPO) | A1 | |
| KR100964436B1 | Republic of Korea | B1 | |
| KR100964437B1 | Republic of Korea | B1 | |
| JP2010136346A | Japan | A | |
| EP2200024A1 | European Patent Office (EPO) | A1 | |
| EP2204796A1 | European Patent Office (EPO) | A1 | |
| CA2578737C | Canada | C | |
| JP2010226744A | Japan | A | |
| US7817677B2 | United States of America | B2 | |
| CN101867522A | China | A | |
| CN101873266A | China | A | |
| CN101873267A | China | A | |
| US7826441B2 | United States of America | B2 | |
| US7830900B2 | United States of America | B2 | |
| EP1787290B1 | European Patent Office (EPO) | B1 | |
| AT488838T | Austria | T | |
| ATE488838T1 | Austria | T1 | |
| DE602005024825D1 | Germany | D1 | |
| ES2355039T3 | Spain | T3 | |
| PL1787290T3This record | Poland | T3 | |
| CA2691762C | Canada | C | |
| JP4933605B2 | Japan | B2 | |
| CN101048813B | China | B | |
| CN101873267B | China | B | |
| CN102779517A | China | A | |
| US8331385B2 | United States of America | B2 | |
| JP2013031222A | Japan | A | |
| EP2200024B1 | European Patent Office (EPO) | B1 | |
| ES2405750T3 | Spain | T3 | |
| DK2200024T3 | Denmark | T3 | |
| PT2200024E | Portugal | E | |
| CA2691959C | Canada | C | |
| PL2200024T3 | Poland | T3 | |
| MY149811A | Malaysia | A | |
| JP5389729B2 | Japan | B2 | |
| JP5591897B2 | Japan | B2 | |
| TWI454101B | Taiwan Province of China | B | |
| CN101873266B | China | B | |
| EP2204796B1 | European Patent Office (EPO) | B1 | |
| BRPI0514801B1 | Brazil | B1 |
Numbers
- Publication, DOCDB
- 1787290
- Publication, EPODOC
- PL1787290T
- Application
- 794150
- Application, DOCDB
- 05794150
- Application, EPODOC
- PL20050794150T
Titles2
- English
- METHOD AND APPARATUS FOR AN ADAPTIVE DE-JITTER BUFFER
- Polish
- Sposób i urządzenie dla adaptacyjnego bufora eliminującego jitter
Classification
- CPC, 14
- H04L12/66
- H04L12/00
- G10L19/005
- H04J3/0632
- H04L65/80
- H04L47/29
- H04L47/30
- H04L47/28
- H04L47/2416
- H04L49/9094
- H04L65/1101
- H04L65/752
- H04L47/10
- H04L49/90
- IPC, 5
- G10L19 00
- H04L49 9023
- H04L12 56
- H04L47 2416
- H04L47 30
