Methods for changing the size of a jitter buffer and for time alignment, communications system, receiving end, and transcoder
5 claims: 1 independent, 4 dependent
- 1Method for carrying out a time alignment in a transcoder of a radio communications system, which time alignment is used for decreasing a buffering delay, said buffering delay resulting from buffering speech data encoded by said transcoder before transmitting said speech data over a radio interface of said radio communications system in order to compensate for a phase shift in a framing of said speech data in said transcoder and at said radio interface, the method comprising:- determining whether a time alignment has to be carried out;and - in case it was determined that a time alignment has to be carried out, condensing speech data for achieving the required time alignment by discarding at least one frame of speech data, wherein gain parameters and Linear Predictive Coding coefficients of frames of speech data surrounding the at least one discarded frame are modified to smoothly combine the frames surrounding the at least one discarded frame.
104 paragraphs in 5 sections, as filed
FIELD OF THE INVENTION
0001The invention relates to a method for carrying out a time alignment in a transcoder of a radio communications system, which time alignment is used for decreasing a buffering delay , said buffering delay resulting from buffering speech data encoded by said transcoder before transmitting said speech data over a radio interface of said radio communications system in order to compensate for a phase shift in a framing of said speech data in said transcoder and at said radio interface. Further, the invention relates to such a radio communications system and to a transcoder for a radio communications system and to an apparatus.
BACKGROUND OF THE INVENTION
0002An example of a packet network is a voice over IP (VoIP) network.
0003IP telephony or voice over IP (VoIP) enables users to transmit audio signals like voice over the Internet Protocol. Sending voice over the internet is done by inserting speech samples or compressed speech into packets. The packets are then routed independently from each other to their destination according to the IP-address included in each packet.
0004One drawback in IP telephony is the availability and performance of networks. Although the local networks might be stable and predictable, the Internet is often congested and there are no guarantees that packets are not lost or significantly delayed. Lost packets and long delays have an immediate effect on speech quality, reciprocity and the pace of conversation.
0005Because of the independent routing of the packets, the packets moreover take variable times to go through the network. The variation in packet arrival times is called jitter. To play out the voice in the receiving end correctly, though, the packets must be in the order of transmission and equally spaced. To achieve this requirement a jitter buffer can be employed. The jitter buffer can be located before or after a decoder used at the receiving end for decoding the speech which was encoded for transmission. In the jitter buffer, the right order of packets can then be assured by checking sequence numbers contained in the packets. Equally contained timestamps can further be used to determine the jitter level in the network and for compensating for the jitter in play out.
0006The size of the jitter buffer, however, has a contrary effect on the number of packets that are lost and on the end-to-end delay. If the jitter buffer is very small, many packets are lost because they have arrived after their playout point. On the other hand, if the jitter buffer is very large an excessive end-to-end delay appears. Both, packet loss and end-to-end delay, have an effect on speech quality. Therefore, the size of the jitter buffer has to result in an acceptable value for both, packet loss and delay. Since both can vary in time, adaptive jitter buffers have to be employed in order to be able to continuously guarantee a good compromise for the two factors. The size of an adaptive jitter buffer can be changed based on measured delays of received speech packets and measured delay variances between received speech packets.
0007Known methods adjust the jitter buffer size in the beginning of a talkspurt. At the beginning of a talkspurt and therefore at the end of a pause in speech, the played out speech is not affected by the adjustment of the jitter buffer size. This means, however, that an adjustment has to be delayed until a beginning of a talkspurt occurs and that a voice activity detector (VAD) is needed. Such methods are described e.g. in "<nplcit id="ncit0001" npl-type="s"><text>An algorithm for playout of packet voice based on adaptive adjustment of talkspurt silence periods", LCN '99, Conference on Local Computer Networks, 1999, Pages 224 - 231, by J. Pinto and K.J. Christensen</text></nplcit>, and in "<nplcit id="ncit0002" npl-type="s"><text>Adaptive playout mechanisms for packetized audio applications in wide-area networks", INFOCOM '94, 13th Proceedings IEEE Networking for Global Communications, 1994, Pages 680 - 688, vol.2, by R. Ramjee, J. Kurose, D. Towsley and H. Schulzrinne</text></nplcit>.
0008A similar problem with jitter buffers can arise e.g. in voice over ATM networks.
0009A similar problem can moreover arise during time alignment in GSM (global system for mobile communications) or 3G (third generation) systems. In radio communications systems like GSM or a 3G system, the air interface requires a tight synchronization between uplink and downlink transmission. However, at the start of call or after a handover, the initial phase shift between uplink and downlink framing in a transcoder used on the network side for encoding data for downlink transmissions and decoding data from uplink transmissions is different from the corresponding phase shift at the radio interface. This phase shift can also be seen in the phase shift only of the downlink framing in the transcoder and at a radio interface of the radio communications system. Therefore, a downlink buffering is needed to achieve a correct synchronization for the air interface, which buffer is included in GSM in a base station and in 3G networks in a radio network controller (RNC) of the communications system. The buffering leads to an additional delay of up to one speech frame in the base station in downlink direction. To minimize this buffering delay, a time alignment procedure can be utilized on the network side. The time alignment is used to align the phase shift in the framing of the transcoder and thus to minimizing the buffering delay after a call set-up or handover. During the time alignment, the base station or radio network controller requests the transcoder to carry out a desired time alignment. In the time alignment, the transmission time instant of an encoded speech frame and the following frames need to be advanced or delayed. Thereby the window (one speech frame) of input buffer of linear samples before the encoder has to be slided in to the desired direction by the amount of samples requested by the base station. Presently, a time alignment is carried out by dropping or repeating speech samples, which leads to a deterioration of the speech quality.
0010<patcit id="pcit0001" dnum="US5664044A"><text>US patent 5,664,044</text></patcit> aims at providing a system and method for allowing user-controlled, variable-speed synchronized playback of an existing, digitally-recorded audio/video presentation. The digital data stream is a multiplexed, compressed audio/video data stream such as that specified in the MPEG standard. In one embodiment, the user directly controls the rate of the video playback by setting a video scaling factor. The length of time required to play back an audio frame is then adjusted automatically using the time domain harmonic scaling method so that it approximately matches the length of time a video frame is displayed. The number of frames of compressed digital audio in an audio buffer is monitored and the time domain harmonic scaling factor is adjusted continuously during playback to ensure that the audio buffer does not overflow or underflow. An underflow or overflow condition in the audio buffer would eventually cause a loss of synchronization between the audio and video.
SUMMARY OF THE INVENTION
0011It is an object of the invention to improve the time alignment in radio communications systems.
0012This object is reached on the one hand with a method for carrying out a time alignment in a transcoder of a radio communications system, which time alignment is used for decreasing a buffering delay, said buffering delay resulting from buffering speech data encoded by said transcoder before transmitting said speech data over a radio interface of said radio communications system in order to compensate for a phase shift in a framing of said speech data by said transcoder and by said radio interface. First, it is determined whether a time alignment has to be carried out.
0013In case it was determined that a time alignment has to be carried out, speech data is condensed for achieving the required time alignment by discarding at least one frame of speech data. Gain parameters and Linear Predictive Coding (LPC) coefficients of frames of speech data surrounding the at least one discarded frame are moreover modified to smoothly combine the frames surrounding the at least one discarded frame.
0014The object of the invention is reached on the other hand with a radio communications system comprising at least one radio interface for transmitting encoded speech data and at least one transcoder. Said transcoder includes at least one encoder for encoding speech data to be used for a transmission via said radio interface. The transcoder further includes processing means for carrying out a time alignment on encoded speech samples according to the proposed method. The radio communications system moreover comprises buffering means arranged between said radio interface and said transcoder for buffering speech data encoded by said transcoder before transmitting said encoded speech data via said radio interface in order to compensate for a phase shift in a framing of said speech data by said transcoder and by said radio interface. Finally, the radio communications system comprises processing means for determining whether and to which extend the speech samples encoded by said encoder have to be time aligned before transmission in order to minimize a buffering delay for encoded speech data resulting from a buffering by said buffering means. The object of the invention is equally reached with such a transcoder for a radio communications system.
0015The invention proceeds from the idea that the time alignment in a transcoder of a radio communications system could be achieved with less effect on the encoded speech samples, if it is not carried out by simply dropping or repeating speech samples, but rather by compensating for the time alignment in a more sophisticated way that results in less effect on the quality of the speech data. The proposed compensation of a time alignment ensures only smooth transitions within the aligned speech data. It is thus an advantage of the invention that it enables in a simple way an improved time alignment.
0016It is to be noted that strictly speaking the mentioned phase shift relates to the time difference of sending and receiving the first data bit of a frame on uplink vs. downlink, i.e. how data frames aligned in time at an observation point in different transmission directions. For GSM, e.g., initially this time difference is not equal between air and abis interfaces. After the time alignment, the time difference should be almost equal, i.e. minimal buffering.
0017It becomes apparent that the invention is based on changing the amount of currently available audio data based on this existing audio data such that a necessary change can be achieved without severe deterioration of the audio data during ongoing transmission.
0018The invention can be employed in particular, though not exclusively, in a Media Gateway as well as in GSM and 3G time alignments.
BRIEF DESCRIPTION OF THE FIGURES
0019In the following, the invention is explained in more detail with reference to drawings, of which <dl id="dl0001" compact="compact"><dt>Fig. 1</dt><dd>illustrates the principle of three approaches for changing the jitter buffer size;</dd><dt>Fig. 2</dt><dd>shows diagrams illustrating an increase in jitter buffer size according to the first approach based on a method for bad frame handling;</dd><dt>Fig. 3</dt><dd>shows diagrams illustrating a decrease in jitter buffer size according to the first approach ;</dd><dt>Fig. 4</dt><dd>shows diagrams illustrating the principle of a time domain time scaling according to a secondapproach;</dd><dt>Fig. 5</dt><dd>is a flow chart of the second approach;</dd><dt>Fig. 6</dt><dd>shows diagrams further illustrating the second approach;</dd><dt>Fig. 7</dt><dd>shows diagrams illustrating the principle of a third approachfrequency domain time scaling;</dd><dt>Fig. 8</dt><dd>is a flow chart of the third approach;</dd><dt>Fig. 9</dt><dd>shows jitter buffer signals before and after time scaling according to the third approach;</dd><dt>Fig. 10</dt><dd>is a flow chart illustrating a fourth approach changing a jitter buffer size in the parametric domain;</dd><dt>Fig. 11</dt><dd>schematically shows a part of a first system;</dd><dt>Fig. 12</dt><dd>schematically shows a part of a second system; and</dd><dt>Fig. 13</dt><dd>schematically shows a part of a third system; and</dd><dt>Fig. 14</dt><dd>schematically shows a communications system in which a time alignment according to the invention can be employed.</dd></dl>
DETAILED DESCRIPTION OF THE INVENTION
0020On the left hand side of the <figref idref="f0001">figure 1</figref>, an increase of a packet stream is shown, while on the left hand side, a decrease of a packet stream is shown. The upper part of the figure shows for both cases original streams, the middle part for both cases streams treated according to a first approach and the lower part for both cases streams treated according to a second or third approach.
0021In the upper left part of <figref idref="f0001">figure 1</figref>, a first packet stream with eight original packets 1 to 8 including speech data is indicated. This packet stream is contained in a jitter buffer of a receiving end in a voice over IP network before an increase of the jitter buffer size becomes necessary. In the upper right part of <figref idref="f0001">figure 1</figref>, a second packet stream with nine packets 9 to 17 including speech data is indicated. This packet streams is contained in a jitter buffer of a receiving end in a voice over IP network before a decrease of the jitter buffer size becomes necessary.
0022On the left hand side in the middle of <figref idref="f0001">figure 1</figref>, the first packet stream is shown after an increase of the jitter buffer size. The jitter buffer size was increased by providing an empty space of the length of one packet between the original packet 4 and the original packet 5 of the packet stream in the jitter buffer. This empty space is filled by a packet 18 generated according to a bad frame handling BFH as defined in the above mentioned ITU-T G.711 codec, the empty space simply being considered as lost packet. The size of the original stream is thus expanded by the length of one packet.
0023On the right hand side in the middle of <figref idref="f0001">figure 1</figref>, in contrast, the second packet stream is shown after a decrease of the jitter buffer size. It is realized in the first described approach by overlapping two consecutive packets, in this example, the original packet 12 and the original packet 13 of the second packet stream. The overlapping reduces the number of speech samples contained in the jitter buffer in the length of one packet, the size of which can thus be reduced by the length of one packet.
0024On the left hand side at the bottom of <figref idref="f0001">figure 1</figref>, the first packet stream is shown again after an increase of the jitter buffer size which resulted in an empty space of the length of one packet between the original packets 4 and 5 of the first packet stream. This time, however, the original packets 4 and 5 were time scaled in the time domain or in the frequency domain according to the second or third approach in order to fill the resulting empty space. That means the data of original packets 4 and 5 was expanded to fill the space of three instead of two packets. The size of the original stream was thus expanded by the length of one packet.
0025On the right hand side at the bottom of <figref idref="f0001">figure 1</figref>, finally, the second packet stream is shown again after a decrease of the jitter buffer size. The corresponding decrease of the data stream was realized according to the second or third approach by time scaling the data of three original packets to the length of two packets. In the presented example, the data of the original packets 12 to 14 of the second packet stream were condensed to the length of two packets. The size of the original stream was thus reduced by the length of one packet.
0026The increase and decrease of the jitter buffer size according to the first approach will now be explained in detail with reference to <figref idref="f0002">figures 2</figref> and <figref idref="f0003">3</figref>.
0027<figref idref="f0002">Figure 2</figref> is taken from the ITU-T G.711 Appendix specification, where it is used for illustrating lost packet concealment, while here it is used for illustrating the first approach , in which the ITU-T bad frame handler is called between adjacent packets for compensating for an increase of the jitter buffer size.
0028<figref idref="f0002">Figure 2</figref> shows three diagrams which depict the amplitude of signals over the sample number of the signals. In the first diagram the signals input to the jitter buffer are shown, while a second and third diagram show synthesized speech at two different points in time. The diagrams illustrate how the jitter buffer size is increased according to the first approach corresponding to a bad frame handling presented in the above cited ITU-T G.711 codec. As mentioned above, the cited standard describes a packet loss concealment method for the ITU-T G.711 codec based on pitch waveform replication.
0029The packet size employed in this approach is 20ms, which corresponds to 160 samples. The BFH was modified to be able to use 20 ms packets.
0030The arrived packets as well as the synthesised packets are saved in a history buffer of a length of 390 samples.
0031After an increase of the size of the jitter buffer by the length of two packets, there is an empty space in the jitter buffer corresponding to two lost packets, indicated in the first diagram of <figref idref="f0002">figure 2</figref> by a horizontal line connecting the received signals. At the start of each empty space, the contents of the history buffer are copied to a pitch buffer that is used throughout the empty space to find a synthetic waveform that can conceal the empty space. In the situation in the first diagram, the samples that are to the left of the two empty packets i.e. the samples that have arrived before the increase of size, form the current content of the pitch buffer.
0032A cross-correlation method is now used to calculate a pitch period estimate from the pitch buffer. As illustrated in the second diagram of <figref idref="f0002">figure 2</figref>, the first empty packet is then replaced by replicating the waveform that starts one pitch period length back from the end of the history buffer, indicated with a vertical line referred to by 21, in the required number. To ensure a smooth transition between the real and the synthesized speech, as well as between repeated pitch period length waveforms, the last 30 samples in the history buffer, in the region limited by a vertical and an inclined line referred to by 22 in the first diagram, are overlap added with the 30 samples preceding the synthetic waveform in the region limited by the vertical line 21 and a connected inclined line. The overlapped signal replaces the last 30 samples 22 in the pitch buffer. This overlap add procedure causes an algorithmic delay of 3.75 ms, or 30 samples. In the same way, a smooth transition between repeated pitch period length waveforms is ensured.
0033The synthetic waveform is moreover extended beyond the duration of the empty packets to ensure a smooth transition between the synthetic waveform and the subsequently received signal. The length of the extension 23 is 4 ms. In the end of the empty space, the extension is raised by 4 ms per additional added empty packet. The maximum extension length is 10 ms. In the end of the empty space this extension is overlapped with the signal of the first packet after the empty space, the overlap region being indicated in the figure with the inclined line 25. The second diagram of <figref idref="f0002">figure 2</figref> illustrates the state of the synthesized signal after 10 ms, when samples of one packet length have been replicated.
0034In case there is a second added empty packet, as in the first diagram of <figref idref="f0002">figure 2</figref>, another pitch period is added to the pitch buffer. Now the waveform to be replicated is two pitch periods long and starts from the vertical line referred to by 24. Next, the 30 samples 24 before the pitch buffer are overlap added with the last 30 samples 22 in the pitch buffer. Again, the overlapped signal replaces the last 30 samples in region 22 in the pitch buffer. A smooth transition between one and two pitch period length signals is ensured by performing an overlap add between the regions indicated by 23 and 26. Region 26 is placed by subtracting pitch periods until the pitch pointer is in the first wavelength of the currently used portion of the pitch buffer. The result of the overlap adding replaces the samples in region 23. The third diagram of <figref idref="f0002">figure 2</figref> shows the synthesized signal in which an empty space of the length of two packets added for an increase in the size of the jitter buffer was concealed.
0035If the size of the jitter buffer is further increased, another pitch period would be added to the pitch buffer. However, if the increase in jitter buffer size is large it is more likely that the replacement signal falsifies the original signal. Attenuation is used to diminish this problem. The first replacement packet is not attenuated. The second packet is attenuated with a linear ramp. The end of the packet is attenuated by 50 % compared to the start with the used packet size of 20 ms. This attenuation is also used for the following packets. This means that after 3 packets (60 ms) signal amplitude is zero.
0036Similarly, parametric speech coders' bad frame handling methods can be employed for compensating for an increase of the jitter buffer size.
0037<figref idref="f0003">Figure 3</figref> illustrates how the jitter buffer size is decreased according to the first approach by overlapping two adjacent packets. To this end, the figure shows three diagrams depicting the amplitude of signals over the sample number of the signals.
0038The first diagram of <figref idref="f0003">figure 3</figref> shows the signals of four packets 31-34 presently stored in a jitter buffer before a decrease in size, each packet containing 160 samples. Now, the size of the jitter buffer is to be decreased by one packet. To this end, two adjacent packets 32, 33 are multiplied with a downramp 36 and an upramp 37 function respectively, as indicated in the first diagram. Then, the multiplied packets 32, 33 of the signals are overlapped, which is shown in the second diagram of <figref idref="f0003">figure 3</figref>. Finally, the overlapped part of the signal 32/33 is added as shown in the third diagram of <figref idref="f0003">figure 3</figref>, the fourth packet now being formed by the packet 35 following the original fourth packet 34. The result of the overlap adding is a signal comprising one packet less than the original signal, and this removed packet enables a decrease of the size of the jitter buffer.
0039When the jitter buffer size is to be decreased by more than one packet at a time, not adjacent but spaced apart packets are overlap added, and the packets in between are discarded. For example, if the jitter buffer size is to be changed from three packets to one, the first packet in the jitter buffer is overlap added with the third packet in the jitter buffer as described for packets 32 and 33 with reference to <figref idref="f0003">figure 3</figref>, and the second packet is discard.
0040In a second approach, an immediate increase and decrease of a jitter buffer size is enabled by a time domain time scaling method, and more particularly by a waveform similarity overlap add (WSOLA) method described in the above mentioned document "An overlap-add technique based on waveform similarity (WSOLA) for high quality time-scale modification of speech".
0041The WSOLA method is illustrated for an exemplary time scaling resulting in a reduction of samples of a signal in <figref idref="f0004">figure 4</figref>, which comprises in the upper part an original waveform x(n) and in the lower part a synthetic waveform y(n) constructed with suitable values of the original waveform x(n). n indicates the respective sample of the signals. The WSOLA method is based on constructing a synthetic waveform that maintains maximal local similarity to the original signal. The synthetic waveform y(n) and original waveform x(n) have maximal similarity around time instances specified by a time warping function τ<sup>-1</sup>(n).
0042In <figref idref="f0004">figure 4</figref>, the input segment 41 of original waveform x(n) was the last segment excised from the original waveform x(n). This segment 41 is the last segment that was added as a synthesis segment A to the synthesized waveform y(n). Segment A was overlap-added to the output signal y(n) at time S<sub>k-1</sub>=(k-1)S, S being the interval between segments in the synthesized signal y(n).
0043The next synthesis segment B is to be excised from the input signal x(n) around time instant τ<sup>-1</sup>(S<sub>k</sub>), and overlap added to the output signal y(n) at time S<sub>k</sub>=kS. As can be seen in the figure, segment 41' of the input signal x(n) would overlap perfectly with segment 41 of the input signal x(n). Segment 41' is therefore used as a template when choosing a segment 42 around time instant τ<sup>-1</sup>(S<sub>k</sub>) of the input signal x(n) which is to be used as next synthesis segment B. A similarity measure between segment 41' and segment 42 is computed to find the optimal shifting value Δ that maximizes the similarity between the segments. The next synthesis segment B is thus selected by finding the best match 42 for the template 41' around time instant τ<sup>-1</sup>(S<sub>k</sub>). The best match must be within the tolerance interval of Δ, which tolerance interval lies between a predetermined minimum Δ<sub>min</sub> and a predetermined maximum Δ<sub>max</sub> value. After overlap-adding the synthesis segment 42 to the output signal as segment B, segment 42' of the input signal x(n) is used as the next template.
0044The WSOLA method uses regularly spaced synthesis instants S<sub>k</sub>=kS. The analysis and synthesis window length is constant. If the analysis/synthesis window is chosen in such a way that, <maths id="math0001" num="(1)"><math display="block"><mstyle displaystyle="false"><mstyle displaystyle="true"><munder><mo>∑</mo><mi>k</mi></munder></mstyle></mstyle><mi>v</mi><mo></mo><mfenced><mi>n</mi><mo>-</mo><mi mathvariant="italic">kS</mi></mfenced><mo>=</mo><mn>1</mn></math><img file="EP1536582B1_D0001.tif" /></maths> and if the analysis/synthesis window is symmetrical, the synthesis equation for the WSOLA method is <maths id="math0002" num="(2)"><math display="block"><mi>y</mi><mfenced><mi>n</mi></mfenced><mo>=</mo><mstyle displaystyle="false"><mstyle displaystyle="true"><munder><mo>∑</mo><mi>k</mi></munder></mstyle></mstyle><mi>v</mi><mo></mo><mfenced><mi>n</mi><mo>-</mo><mi mathvariant="italic">kS</mi></mfenced><mo></mo><mi>x</mi><mo></mo><mfenced><mi>n</mi><mo>+</mo><msup><mi>τ</mi><mrow><mo>-</mo><mn>1</mn></mrow></msup><mfenced><mi mathvariant="italic">kS</mi></mfenced><mo>-</mo><mi mathvariant="italic">kS</mi><mo>+</mo><msub><mi mathvariant="normal">Δ</mi><mi>k</mi></msub></mfenced><mn>.</mn></math><img file="EP1536582B1_D0002.tif" /></maths>
0045By selecting a different time warping function, the same method can be employed not only for reducing the samples of a signal but also for increasing the amount of samples of a signal.
0046It is important that the transition from the original signal to the time-scaled signal is smooth. In addition, the pitch period should not change during the jumps from the signal used as received to the time scaled signal. As was explained previously, WSOLA time scaling preserves the pitch period. However, when time scaling is performed for a part in the middle of the speech signal, some discontinuity on either the beginning, or the end of the time scaled signal can not be avoided sometimes.
0047In order to decrease the effect of such a phase mismatch, it is proposed for the second approach to slightly modify the method described with reference to <figref idref="f0004">figure 4</figref>. The modified WSOLA (MWSOLA) method uses history information and extra extension to decrease the effect of this problem.
0048A MWSOLA algorithm using an extension of the time scale for an extra half of a packet length will now be described with reference to the flow chart of <figref idref="f0005">figure 5</figref> and to the five diagrams of <figref idref="f0006">figure 6</figref>. The used packet size is 20 ms, or 160 samples, the sampling rate being 8 kHz. The analysis/ synthesis window used has the same length as the packets.
0049<figref idref="f0005">Figure 5</figref> illustrates the basic process of updating the jitter buffer size using the proposed MWSOLA algorithm. As shown on the left hand side of the flow chart of <figref idref="f0005">figure 5</figref>, first the packets to be time scaled are chosen. In addition, 1/2 packet length of the previously arrived signal, i.e. 80 samples, are selected as history samples. The selected samples are also indicated in the first diagram of <figref idref="f0006">figure 6</figref>. After being selected, they are forwarded to the MWSOLA algorithm.
0050The MWSOLA algorithm, which is shown in more detail on the right hand side of <figref idref="f0005">figure 5</figref>, is then used to provide the desired time scaling on the selected signals as described with reference to <figref idref="f0004">figure 4</figref>.
0051The analysis/synthesis window is created by modifying a Hanning window so that the condition of equation (1) is fulfilled. The time warping function τ<sup>-1</sup>(n) is constructed differently for time scale expansion and compression, i.e. for an increase and for a decrease of the jitter buffer size. The time warping function and the limits of the search region Δ (Δ= [Δ<sub>min</sub>...Δ<sub>max</sub>]) are chosen in such a way that a good signal variation is obtained. By setting the limits of the search region and the time warping function correctly, it can be avoided that adjacent analysis frames are chosen repeatedly. Finally, the first frame from the input signal is copied to an output signal which is to substitute the original signal. This ensures that the change from the preceding original signal to the time scaled signal is smooth.
0052After the initial parameters like the time warping function and the limits for the search region are set and an output signal is initialized, a loop is used to find new frames for the time scaled output signal as long as needed. A best match between the last L samples of the previous frame and the first L samples of the new frame is used as an indicator in finding the next frame. The used length L of the correlation is 1/2 * window length = 80 samples. The search region Δ (Δ= [Δ<sub>min</sub>...Δ<sub>max</sub>]) should be longer than the maximum pitch period in samples, so that a correct synchronization between consecutive frames is possible.
0053The second diagram of <figref idref="f0006">figure 6</figref> shows how the analysis windows 61-67 defining different segments are placed in the MWSOLA input signal when time scaling two packets to three packets.
0054The third diagram of <figref idref="f0006">figure 6</figref> shows how overlapping the synthesis segments succeeded. As can be seen, the different windows 61-67 overlap, in this case, quite nicely.
0055Overlap adding of all the analysis/synthesis frames results in the time scaled signal shown in the fourth diagram of <figref idref="f0006">figure 6</figref>, which constitutes the output signal of the MWSOLA algorithm. The MWSOLA algorithm returns the new time scaled packets and an extension to be overlap added with the first 1/2 packet length of the next arriving packet.
0056As shown again on the left hand side of the flow chart of <figref idref="f0005">figure 5</figref>, the jitter buffer is then updated with the time scaled signals and the extension is overlap added with the next arriving packet. The resulting signal can be seen in the fifth diagram of <figref idref="f0006">figure 6</figref>.
0057This procedure decreases the effect of the phase and amplitude mismatches between the time-scaled signal and the valid signal.
0058A phase vocoder based jitter buffer scaling method will now be described with reference to <figref idref="f0007 f0008 f0009">figures 7 to 9</figref> as third approach. This method constitutes a frequency domain time scaling method.
0059The phase vocoder time scale modification method is based on taking short-time Fourier transforms (STFT) of the speech signal in the jitter buffer as described in the above mentioned document "Applications of Digital Signal Processing to Audio and Acoustics". <figref idref="f0007">Figure 7</figref> illustrates this technique. The phase vocoder based time scale modification comprises an analyzing stage, indicated in the upper part of <figref idref="f0007">figure 7</figref>, a phase modification stage indicated in the middle of <figref idref="f0007">figure 7</figref>, and a synthesis stage indicated in the lower part of <figref idref="f0007">figure 7</figref>.
0060In the analyzing stage, short-time Fourier transforms are taken from overlapping windowed parts 71-74 of a received signal. In particular, discrete time Fourier transforms (DFT) as described by <nplcit id="ncit0003" npl-type="s"><text>J. Laroche and M. Dolson in "Improved Phase Vocoder Time-Scale Modification of Audio", IEEE Transactions on Speech and Audio Processing, Vol. 7, No. 3, May 1999. pp. 323-332</text></nplcit>, can be employed in the phase vocoder analysis stage. This means that both, the frequency scale and the time scale representation of the signal, are discrete. The analysis time instants <maths id="math0003"><math display="inline"><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup></math><img file="EP1536582B1_D0003.tif" /></maths> are regularly spaced by R<sub>a</sub> samples, <maths id="math0004"><math display="inline"><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup><mo>=</mo><mi mathvariant="normal">u</mi><mo mathvariant="normal">*</mo><msub><mi mathvariant="normal">R</mi><mi mathvariant="normal">a</mi></msub><mn>.</mn></math><img file="EP1536582B1_D0004.tif" /></maths> R<sub>a</sub> is called the analysis hop factor. The short time Fourier transform is then <maths id="math0005" num="(3)"><math display="block"><mi>X</mi><mrow><mo>(</mo><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup><mo>,</mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub><mo>)</mo><mo>=</mo></mrow><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>n</mi><mo>=</mo><mo>-</mo><mi>∞</mi></mrow><mi>∞</mi></munderover></mstyle><mi>h</mi><mfenced><mi>n</mi></mfenced><mo></mo><mi>x</mi><mo></mo><mfenced><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup><mo>+</mo><mi>n</mi></mfenced><mo></mo><msup><mi>e</mi><mrow><mo>-</mo><mi>j</mi><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub><mo></mo><mi>n</mi></mrow></msup><mo>,</mo></math><img file="EP1536582B1_D0005.tif" /></maths> where x is the original signal, h(n) the analysis window and Ω<sub>k</sub>=2pi*k/N the center frequency of the k<sup>th</sup> vocoder channel. The vocoder channels can also be called bins. N is the size of the DFT, where N must be longer than the length of the analysis window. In practical solutions, the DFT is usually obtained with the Fast Fourier Transform (FFT). The analysis window's cutoff frequency for the standard (Hanning, Hamming) windows requires the analysis windows to overlap by at least 75%. After the analysis FFT, the signal is represented by horizontal vocoder channels and vertical analysis time instants.
0061In the phase modification stage, the time scale of the speech signal is modified by setting the analysis hop factor R<sub>a</sub> different from a to be used synthesis hop factor R<sub>s</sub>, as described in the mentioned document "Improved Phase Vocoder Time-Scale Modification of Audio". The new time-evolution of the sine waves is achieved by setting <maths id="math0006"><math display="inline"><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">k</mi></msub><mo mathvariant="normal">)</mo><mo mathvariant="normal">|</mo><mo mathvariant="normal">=</mo><mfenced open="|" close="|"><mi mathvariant="normal">X</mi><mfenced><msup><msub><mi mathvariant="normal">t</mi><mi mathvariant="normal">a</mi></msub><mi mathvariant="normal">u</mi></msup><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">k</mi></msub></mfenced></mfenced></math><img file="EP1536582B1_D0006.tif" /></maths> and by calculating new phase values for <maths id="math0007"><math display="inline"><mi mathvariant="normal">Y</mi><mfenced><msubsup><mi>t</mi><mi>s</mi><mi>u</mi></msubsup><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">k</mi></msub></mfenced><mn>.</mn></math><img file="EP1536582B1_D0007.tif" /></maths>
0062The new phase values for <maths id="math0008"><math display="inline"><mi mathvariant="normal">Y</mi><mfenced><msubsup><mi>t</mi><mi>s</mi><mi>u</mi></msubsup><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">k</mi></msub></mfenced></math><img file="EP1536582B1_D0008.tif" /></maths> are calculated as follows. A process called phase unwrapping is used, where the phase increment between two consecutive frames is used to estimate the instantaneous frequency of a nearby sinusoid in each channel k. First the heterodyned phase increment is calculated by <maths id="math0009" num="(4)"><math display="block"><msubsup><mi mathvariant="normal">ΔΦ</mi><mi>k</mi><mi>u</mi></msubsup><mo>=</mo><mo>∠</mo><mi>X</mi><mrow><mo>(</mo><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup><mo>,</mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub><mo>)</mo><mo>-</mo></mrow><mo>∠</mo><mi>X</mi><mo></mo><mfenced><msubsup><mi>t</mi><mi>a</mi><mrow><mi>u</mi><mo>-</mo><mn>1</mn></mrow></msubsup><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub></mfenced><mo>-</mo><msub><mi>R</mi><mi>a</mi></msub><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub><mn>.</mn></math><img file="EP1536582B1_D0009.tif" /></maths>
0063Then by adding or subtracting multiples of 2π so that the result of (7) lies between ±π, the principal determination <maths id="math0010"><math display="inline"><mfenced><msub><mi mathvariant="normal">Δ</mi><mi>p</mi></msub><mo></mo><msubsup><mi mathvariant="normal">Φ</mi><mi>k</mi><mi>u</mi></msubsup></mfenced></math><img file="EP1536582B1_D0010.tif" /></maths> of the heterodyned phase increment is obtained.
0064The instantaneous frequency is then calculated using <maths id="math0011" num="(5)"><math display="block"><msub><mi>ω</mi><mi>k</mi></msub><mrow><mo>(</mo><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup><mo>)</mo><mo>=</mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub><mo>+</mo><mfrac><mn>1</mn><msub><mi>R</mi><mi>a</mi></msub></mfrac><mo></mo><msub><mi mathvariant="normal">Δ</mi><mi>p</mi></msub><mo></mo><msubsup><mi mathvariant="normal">Φ</mi><mi>k</mi><mi>u</mi></msubsup></mrow><mn>.</mn></math><img file="EP1536582B1_D0011.tif" /></maths>
0065The instantaneous frequency is determined because the FFT is calculated only for discrete frequencies Ω<sub>k</sub>. Thus the FFT does not necessarily represent the windowed signal exactly.
0066The time scaled phases of the STFT at a time <i>t<sup>u</sup><sub>s</sub></i> are calculated from <maths id="math0012" num="(6)"><math display="block"><mo>∠</mo><mi>Y</mi><mrow><mo>(</mo><msubsup><mi>t</mi><mi>s</mi><mi mathvariant="normal">u</mi></msubsup><mo>,</mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub><mo>)</mo><mo>=</mo></mrow><mo>∠</mo><mi>Y</mi><mo></mo><mfenced><msubsup><mi>t</mi><mi>s</mi><mrow><mi>u</mi><mo>-</mo><mn>1</mn></mrow></msubsup><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub></mfenced><mo>+</mo><msub><mi>R</mi><mi>s</mi></msub><mo></mo><msub><mi>ω</mi><mi>k</mi></msub><mfenced><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup></mfenced><mn>.</mn></math><img file="EP1536582B1_D0012.tif" /></maths>
0067The choice of initial synthesis phases <maths id="math0013"><math display="inline"><mo>∠</mo><mi>Y</mi><mfenced><msubsup><mi>t</mi><mi>s</mi><mn>0</mn></msubsup><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub></mfenced></math><img file="EP1536582B1_D0013.tif" /></maths> is important for good speech quality. In the above mentioned document "Improved Phase Vocoder Time-Scale Modification of Audio", a standard initialization setting of <maths id="math0014" num="(7)"><math display="block"><mo>∠</mo><mi>Y</mi><mrow><mo>(</mo><msubsup><mi>t</mi><mi>s</mi><mn>0</mn></msubsup><mo>,</mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub><mo>)</mo><mo>=</mo></mrow><mo>∠</mo><mi>X</mi><mfenced><msubsup><mi>t</mi><mi>s</mi><mn>0</mn></msubsup><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub></mfenced></math><img file="EP1536582B1_D0014.tif" /></maths> is recommended, which makes a switch from a non time scaled signal to a time scaled signal possible without phase discontinuity. This is an important attribute for jitter buffer time scaling.
0068After the phases values for <maths id="math0015"><math display="inline"><mi mathvariant="normal">Y</mi><mfenced><msubsup><mi>t</mi><mi>s</mi><mi>u</mi></msubsup><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi mathvariant="normal">k</mi></msub></mfenced></math><img file="EP1536582B1_D0015.tif" /></maths> are obtained, the signal can be reconstructed in a synthesis stage.
0069In the synthesis stage, the modified short time Fourier transforms Y(t<sub>s</sub><sup>u</sup>, Ω<sub>k</sub> ) are first inverse Fourier transformed with the equation <maths id="math0016" num="(8)"><math display="block"><msub><mi>y</mi><mi>n</mi></msub><mfenced><mi>n</mi></mfenced><mo>=</mo><mfrac><mn>1</mn><mi>N</mi></mfrac><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>k</mi><mo>=</mo><mn>0</mn></mrow><mrow><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover></mstyle><mi>Y</mi><mfenced><msubsup><mi>t</mi><mi>s</mi><mi>u</mi></msubsup><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub></mfenced><mo></mo><msup><mi>e</mi><mrow><mi>j</mi><mo></mo><msub><mi mathvariant="normal">Ω</mi><mi>k</mi></msub><mo></mo><mi>n</mi></mrow></msup><mn>.</mn></math><img file="EP1536582B1_D0016.tif" /></maths>
0070The synthesis time instants are set <maths id="math0017"><math display="inline"><msubsup><mi>t</mi><mi>s</mi><mi>u</mi></msubsup><mo>=</mo><mi mathvariant="normal">u</mi><mo mathvariant="normal">*</mo><msub><mi mathvariant="normal">R</mi><mi mathvariant="normal">s</mi></msub><mn>.</mn></math><img file="EP1536582B1_D0017.tif" /></maths> Finally the short-time signals are multiplied by a synthesis window w(n) and are summed, together giving the output signal y(n): <maths id="math0018" num="(9)"><math display="block"><mi>y</mi><mfenced><mi>n</mi></mfenced><mo>=</mo><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>u</mi><mo>=</mo><mo>-</mo><mi>∞</mi></mrow><mi>∞</mi></munderover></mstyle><mi>w</mi><mo></mo><mfenced><mi>n</mi><mo>-</mo><msubsup><mi>t</mi><mi>s</mi><mi>u</mi></msubsup></mfenced><mo></mo><msub><mi>y</mi><mi>u</mi></msub><mo></mo><mfenced><mi>n</mi><mo>-</mo><msubsup><mi>t</mi><mi>s</mi><mi>u</mi></msubsup></mfenced><mn>.</mn></math><img file="EP1536582B1_D0018.tif" /></maths>
0071The distance between the analysis windows is different from the distance between the synthesis windows due to the time scale modification, therefore a time extension or compression of the received jitter buffer data is achieved. Synchronisation between overlapping synthesis windows was achieved by modifying the phases in the STFT.
0072The use of the phase vocoder based time scaling for increasing or decreasing the size of a jitter buffer is illustrated in the flow chart of <figref idref="f0008">figure 8</figref>.
0073First, the input signal is received and a time scaling factor is set.
0074The algorithm is then initialized by setting analysis and synthesis hop sizes, and by setting the analysis and synthesis time instants. When doing this, a few constraints have to be taken into account, which have been listed e.g. in the above mentioned document "Applications of Digital Signal Processing to Audio and Acoustics". The cutoff frequency of the analysis window must satisfy w<sub>h</sub> < min<sub>i</sub>Δw<sub>i</sub>, i.e. the cutoff frequency must be less than the spacing between two sinusoids. Further, the length of the analysis window must be small enough so that the amplitudes and instantaneous frequencies of the sinusoids can be considered constants inside the analysis window. Finally, to enable phase unwrapping, the cutoff frequency and the analysis rate must satisfy w<sub>h</sub>Ra < π. The cutoff frequency for standard analysis windows (Hamming, Hanning) is W<sub>h</sub> ≈ 4π/Nw, where Nw is the length of the analysis window.
0075As further initial parameter, the number of frames to process is calculated. This number is used to determine how many times the following loop in <figref idref="f0008">figure 8</figref> must be processed. Finally, initial synthesis phases are set, according to equation (7).
0076After initialization, a vocoder processing loop follows for the actual time scaling. Inside the phase vocoder processing loop, the routine is a straightforward realization of the method presented above. First, the respective next analysis frame is obtained by multiplying the signal with the analysis window at time instant <maths id="math0019"><math display="inline"><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup></math><img file="EP1536582B1_D0019.tif" /></maths> Then the FFT of the frame is calculated. The heterodyned phase increment is calculated by setting R<sub>a</sub> in equation (4) to <maths id="math0020"><math display="inline"><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup><mo>-</mo><msubsup><mi>t</mi><mi>a</mi><mrow><mi>u</mi><mo>-</mo><mn>1</mn></mrow></msubsup><mn>.</mn></math><img file="EP1536582B1_D0020.tif" /></maths> Instantaneous frequencies are also obtained by setting R<sub>a</sub> in equation (5) to <maths id="math0021"><math display="inline"><msubsup><mi>t</mi><mi>a</mi><mi>u</mi></msubsup><mo>-</mo><msubsup><mi>t</mi><mi>a</mi><mrow><mi>u</mi><mo>-</mo><mn>1</mn></mrow></msubsup><mn>.</mn></math><img file="EP1536582B1_D0021.tif" /></maths> The time scaled phases are obtained from equation (6). Next, the IFFT of the modified FFT of the current frame is calculated according to equation (8). The result of equation (8) is then multiplied by the synthesis window and added to the output signal. Before going through the loop again, the previous analysis and synthesis phases to be used in equations (4) and (6) are updated.
0077Finally, before outputting the time scaled signal, transitions between the time scaled and the non time scaled signal are smoothed. After this, the jitter buffer size modification can be completed. <figref idref="f0009">Figure 9</figref> shows the resulting signal when time scaling two packets into three with the phase vocoder based time scaling. In a first diagram of <figref idref="f0009">figure 9</figref>, the amplitude of the signal over the samples before time scaling is depicted. In a second diagram of <figref idref="f0009">figure 9</figref>, the amplitude of the signal over the samples after time scaling is depicted. The two packets with samples 161 to 481 in the first diagram were expanded to three packets with samples 161 to 641.
0078Before the jitter buffer size is increased, an error concealment should be performed. Moreover, a predetermined number of packets should be received before the jitter buffer size is increased.
0079<figref idref="f0010">Figure 10</figref> is a flow chart illustrating a fourth approach, which can be used for changing a jitter buffer size in the parametric domain. The parametric coded speech frames are only decoded by a decoder after buffering in the jitter buffer.
0080In a first step, it is determined whether the jitter buffer size has to be changed. In case it does not have to be changed, the contents of the jitter buffer are directly forwarded to the decoder.
0081In case it is determined that the jitter buffer size has to be increased, the jitter buffer is increased and additional frames are generated by interpolating an additional frame from two adjacent frames in the parametric domain. The additional frames are used for filling the empty buffer space resulting from an increase in size. Only then the buffered frames are forwarded to the decoder.
0082In case it is determined that the jitter buffer size has to be decreased, the jitter buffer is decreased and two adjacent or spaced apart frames are interpolated in the parametric domain into one frame. The distance of the two frames used for interpolation to each other depends on the amount of the required decrease of the jitter buffer size. Only then the buffered frames are forwarded to the decoder.
0083<figref idref="f0011 f0012 f0013">Figures 11 to 13</figref> shows parts of three different voice over IP communications systems.
0084In the communications system of <figref idref="f0011">Figure 11</figref>, an encoder 111 and packetization means 112 belong to a transmitting end of the system. The transmitting end is connected to a receiving end via a voice over IP network 113. The receiving end comprises a frame memory 114, which is connected via a decoder 115 to an adaptive jitter buffer 116. The adaptive jitter buffer 116 further has a control input connected to control means and an output to some processing means of the receiving end which are not depicted.
0085At the transmitting end, speech that is to be transmitted is encoded in the encoder 111 and packetized by the packetization means 112. Each packet is provided with information about its correct position in a packet stream and about the correct distance in time to the other packets. The resulting packets are sent over the voice over IP network 113 to the receiving end.
0086At the receiving end, the received packets are first reordered in the frame memory 114 in order to bring them again into the original order in which they were transmitted by the transmitting end. The reordered packets are then decoded by the decoder 115 into linear PCM speech. The decoder 115 also performs a bad frame handling on the decoded data. After this, the linear PCM speech packets are forwarded by the decoder 115 to the adaptive jitter buffer 116. In the adaptive jitter buffer, a linear time scaling method can then be employed to increase or decrease the size of the jitter buffer and thereby get more time or less time for the packets to arrive to the frame memory.
0087The control input of the adaptive jitter buffer 116 is used for indicating to the adaptive jitter buffer 116 whether the size of the jitter buffer 116 should be changed. The decision on that is taken by control means based on the evaluation of the current overall delay and the current variation of delay between the different packets. The control means indicate more specifically to the adaptive jitter buffer 116 whether the size of the jitter buffer 116 is to be increased or decreased and by which amount and which packets are to be selected for time scaling.
0088In case the control means indicate to the adaptive jitter buffer 116 that its size is to be changed, the adaptive jitter buffer 116 time scales at least part of the presently buffered packets according to the received information, e.g. in a way described in the second or third approach. The jitter buffer 116 is therefore extended by time scale expansion of the currently buffered speech data and reduced by time scale compression of the currently buffered speech data. Alternatively, a method based on a bad frame handling method for increasing the buffer size could be employed for changing the jitter buffer size. This alternative method could for example be the method of the first approach, in which moreover data is overlapped for decreasing the buffer size.
0089The linear time scaling of <figref idref="f0011">Figure 11</figref> can be employed in particular for a low bit rate codec.
0090<figref idref="f0012">Figure 12</figref> shows a part of a communications system which is based on a linear PCM speech time scaling method. In this system, a transmitting end which corresponds to the one in <figref idref="f0011">figure 11</figref> and which is not depicted in <figref idref="f0012">figure 12</figref>, is connected again via a voice over IP network 123 to the receiving end. The receiving end, however, is designed somewhat differently from the receiving end in the system of <figref idref="f0011">figure 11</figref>. The receiving end comprises now means for A-law to linear conversion 125 connected to an adaptive jitter buffer 126. The adaptive jitter buffer 126 has again additionally a control input connected to control means and an output to some processing means of the receiving end which are not depicted.
0091Packets containing speech data which were transmitted by the transmitting end and received by the receiving end via the voice over IP network 123 are first input to the means for A-law to liner conversion 125 of the receiving end, where they are converted to linear PCM data. Subsequently, the packets are reorganized in the adaptive jitter buffer 126. Moreover, the adaptive jitter buffer 126 takes care of a bad frame handling, before forwarding the packets with a correct delay to the processing means.
0092Control means are used again for deciding when and how to change the jitter buffer size. Whenever necessary, some time scaling method for linear speech, e.g. one of the presented methods, is then used in the adaptive jitter buffer 126 to change its size according to the information received by the control means. Alternatively, a method based on a bad frame handling method could be employed again for changing the jitter buffer, e.g. the method of the first approach. This alternative method could also make use of the bad frame handling method implemented in the jitter buffer anyhow for bad frame handling.
0093<figref idref="f0013">Figure 13</figref>, finally, shows a part of a communications system in which a low bit rate codec and a parametric domain time scaling is employed.
0094Again, a transmitting end corresponding to the one in <figref idref="f0011">figure 11</figref> and not being depicted in <figref idref="f0013">figure 13</figref>, is connected via a voice over IP network 133 to a receiving end. The receiving end comprises a packet memory and organizer unit 134, which is connected via an adaptive jitter buffer 136 to a decoder 135. The adaptive jitter buffer 136 further has a control input connected to control means, and the output of the decoder 135 is connected to some processing means of the receiving end, both, control means and processing means not being depicted.
0095Packets containing speech data which were transmitted by the transmitting end and received by the receiving end via the voice over IP network 133 are first reordered in the packet memory and organizer unit 134.
0096The reordered packets are then forwarded directly to the adaptive jitter buffer 136. The jitter buffer 136 applies a bad frame handling on the received packets in the parametric domain. The speech contained in the packets is decoded only after leaving the adaptive jitter buffer 136 in the decoder 135.
0097As in the other two presented systems, the control means are used for deciding when and how to change the jitter buffer size. Whenever necessary, some time scaling method for parametric speech is then used in the adaptive jitter buffer 136 to change its size according to the information received by the control means. Alternatively, also a bad frame handling method designed for bad frame handling of packets in the parametric domain could be employed for increasing the jitter buffer size. As further alternative, additional frames could be interpolated from two adjacent frames as proposed with reference to <figref idref="f0010">figure 10</figref>. Decreasing the jitter buffer size could be achieved by discarding a packet or by interpolating two packets into one in the parametric domain as proposed with reference to <figref idref="f0010">figure 10</figref>. In particular, if a decrease by more than one packet is desired, the packets around the desired amount of packets could be interpolated into one packet.
0098An embodiment of the invention relating to time alignment will now be presented with reference to <figref idref="f0014">figure 14</figref>, which shows a GSM or 3G radio communications system.
0099The radio communications system comprises a mobile station 140, of which an antenna 141 and a decoder 142 are depicted. On the other hand, it comprises a radio access network, of which a base station and a radio network controller 143 is depicted as a single block with access to an antenna 144. Base station and radio network controller 143 are further connected to a network transcoder 145 comprising an encoder 146 and time alignment means 147 connected to each other. Base station or radio network controller 143 have moreover a controlling access to the time alignment means 147.
0100In the radio communications system, speech frames are transmitted in the downlink direction from the radio access network to the mobile station 140 and in the downlink direction from the mobile station 140 to the radio access network. Speech frames that are to be transmitted in the downlink direction are first encoded by the encoder 146 of the transcoder 145, transmitted via the radio network controller, the base station 143 and the antenna 144 of the radio access network, received by the antenna 142 of the mobile station 140 and decoded by the decoder 141 of the mobile station 140.
0101At the start of call or after a handover, the initial phase shift between uplink and downlink framing in the transcoder may be different from the phase shift of the radio interface, which prevents the required strict synchronous transmissions in uplink and downlink. In GSM, the base station 143 therefore guarantees that the phase shift is equal by buffering all encoded speech frames received from the transcoder 145 for a downlink transmission as long as required. Even though the base station 143 determines the required buffering delay by comparing uplink speech data received from the mobile station 140 with downlink speech data received from the transcoder 145, this means also a compensation of a phase shift of the downlink framing in the transcoder and at radio interface of the system accessed via the antenna 144. In a 3G network, this function is provided by the radio network controller 143. This buffering leads to an additional delay of up to one speech frame in the base station in downlink direction. In order to minimize the buffering delay required for synchronization, in GSM the base station and in 3G the radio network controller 143 requests from the time alignment means 147 of the transcoder 145 to apply a time alignment to the encoded speech frames.
0102In a time alignment, the transmission time instant of an encoded speech frame and the following frames is advanced or delayed for a specified amount of samples according to the information received the base station or the radio network controller 143 respectively, thus reducing the necessary buffering delay in the base station or the radio network controller 143.
0103According to the invention, the time alignment is now carried out by the time alignment means 147 by applying a time scaling on the speech frames encoded by the encoder 146, before forwarding them to the radio network controller or the base station 143. In particular, any of the time domain or frequency domain time scaling methods proposed for changing a jitter buffer size can be employed.
0104As a result, the buffering delay in the base station 143 is reduced as in a known time alignment, but the speech quality is affected less.
Contents5
35 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US11438025B2 | Cited by | United States of America | Applicant |
| WO2020060966A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US10931909B2 | Cited by | United States of America | Applicant |
| US10992336B2 | Cited by | United States of America | Applicant |
| US10958301B2 | Cited by | United States of America | Applicant |
| US12155408B2 | Cited by | United States of America | Applicant |
| US11177851B2 | Cited by | United States of America | Applicant |
| US11558579B2 | Cited by | United States of America | Applicant |
| US11671139B2 | Cited by | United States of America | Applicant |
| EP0279451A | Cites | European Patent Office (EPO) | – |
| EP0681398A | Cites | European Patent Office (EPO) | – |
| US4076958A | Cites | United States of America | – |
| US5664044A | Cites | United States of America | – |
| US5862232A | Cites | United States of America | – |
| US5920840A | Cites | United States of America | – |
18 members in 7 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 01931645 | European Patent Office (EPO) | A | |
| 0104608 | European Patent Office (EPO) | W |
Members18
| Document | Office | Kind | |
|---|---|---|---|
| WO02087137A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU2001258364A1 | Australia | A1 | |
| WO02087137A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1382143A2 | European Patent Office (EPO) | A2 | |
| US2004120309A1 | United States of America | A1 | |
| EP1536582A2 | European Patent Office (EPO) | A2 | |
| EP1536582A3 | European Patent Office (EPO) | A3 | |
| EP1382143B1 | European Patent Office (EPO) | B1 | |
| AT353503T | Austria | T | |
| ATE353503T1 | Austria | T1 | |
| DE60126513D1 | Germany | D1 | |
| ES2280370T3 | Spain | T3 | |
| DE60126513T2 | Germany | T2 | |
| EP1536582B1This record | European Patent Office (EPO) | B1 | |
| AT422744T | Austria | T | |
| ATE422744T1 | Austria | T1 | |
| DE60137656D1 | Germany | D1 | |
| ES2319433T3 | Spain | T3 |
60 legal events, as 7 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Announcement of lapse in spainLapsedFD2A | FD2A | ES | |
| Patent expired after termination of 20 yearsExpiredPE20 | PE20 | GB | |
| Expiry of rightR071 | R071 | DE | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Fee paymentPLFP | PLFP | FR | |
| Fee paymentPLFP | PLFP | FR | |
| Transmission of propertyTP | TP | FR | |
| Fee paymentPLFP | PLFP | FR | |
| Transfer of patentPC2A | PC2A | ES | |
| Change of applicant/patenteeR081 | R081 | DE | |
| Change of applicant/patenteeR081 | R081 | DE | |
| Amendments to the register in respect of changes of name or changes affecting rights (sect. 32/1977)REGISTERED BETWEEN 20150910 AND 20150916732E | 732E | GB | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Patent lapsedLapsedMM4A | MM4A | IE | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Patent ceasedCeasedPL | PL | CH | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Nl: lapsed or annulled due to failure to fulfill the requirements of art. 29p and 29m of the patents actLapsedNLV1 | NLV1 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Definitive protectionFG2A | FG2A | ES | |
| Corresponds to:REF | REF | EP | |
| European patents granted designating irelandGrantedFG4D | FG4D | IE | |
| European patent takes effect as a national patent in ch/liEP | EP | CH | |
| Divisional application: reference to earlier applicationAC | AC | EP | |
| Designated contracting statesAK | AK | EP | |
| European patent grantedGrantedFG4D | FG4D | GB | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Grant fee paidORIGINAL CODE: EPIDOSNIGR3GRAS | GRAS | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOSNIGR1GRAP | GRAP | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Designation fees paidAKX | AKX | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Information provided on ipc code assigned before grantRIC1 | RIC1 | EP | |
| Information provided on ipc code assigned before grantRIC1 | RIC1 | EP | |
| Information provided on ipc code assigned before grantRIC1 | RIC1 | EP | |
| Divisional application: reference to earlier applicationAC | AC | EP | |
| Designated contracting statesAK | AK | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 1536582
- Application
- 50048651
Titles3
- German
- Verfahren zum ändern der Grösse eines Zitterpuffers und zur Zeitausrichtung, Kommunikationssystem, Empfängerseite und Transcoder
- English
- Methods for changing the size of a jitter buffer and for time alignment, communications system, receiving end, and transcoder
- French
- Procédés de changement de la taille d'un tampon de gigue et pour l'alignement temporel, système de communications, extrémité de réception et transcodeur
Classification
- CPC, 4
- H04W88/181
- G10L21/04
- H04J3/0632
- G10L19/00
- IPC, 5
- H04J3 06
- H04W88 08
- G10L21 04
- G10L19 00
- H04W88 18
Designated states20
- Contracting states, 20
- Austria
- Belgium
- Switzerland
- Cyprus
- Germany
- Denmark
- Spain
- Finland
- France
- United Kingdom
- Greece
- Ireland
- Italy
- Liechtenstein
- Luxembourg
- Monaco
- Netherlands (Kingdom of the)
- Portugal
- Sweden
- Türkiye
