Frame loss correction by weighted noise injection
13 claims: 6 independent, 7 dependent
- 1デジタル信号の復号中に実施される、復号中に消失した一連のサンプルを置き換えるための、デジタル信号を処理するための方法であって、 消失した前記一連のサンプルに置き換えるための信号構造を生成するステップ(S6)であって、前記信号構造が、復号中(S1)にかつ消失した前記一連のサンプルよりも前に受信された有効サンプルから決定されたスペクトル成分を含む、ステップと、 復号器で利用可能な、受信された有効サンプルを含むデジタル信号(S602)と、前記スペクトル成分から生成された信号(S601)との間で残差を生成するステップ(S603)と、 前記残差からブロックを抽出するステップ(S605)とを含み、 前記ブロックは、重み付け窓による重複加算(ADD)手法を用いて前記信号構造に注入され(S608)、 連続する2つの 注入された前記ブロックは、時間の点で少なくとも部分的に重複する、方法。
- 2前記ブロックが、抽出ブロック開始時刻(i k )と、ブロック持続時間(L k )とによって規定されるので、前記抽出ブロック開始時刻と、前記ブロック持続時間とのうちの少なくとも1つのパラメータが、少なくとも2つの抽出ブロックの間で変動可能である、請求項1に記載の方法。
- 3前記ブロックが、抽出ブロック開始時刻(i k )と、ブロック持続時間(L k )とによって規定され、 前記抽出ブロック開始時刻と、前記ブロック持続時間とのうちの少なくとも1つのパラメータが、少なくとも1つの抽出ブロックについて擬似ランダムに決定される、請求項1または2に記載の方法。
- 4前記ブロックが、少なくとも2つの注入ブロック間で変動可能な少なくとも1つのパラメータを用いて注入され、 前記変動可能なパラメータは、 前記注入ブロックの書込み開始時刻(j k )と、 連続する2つの注入ブロック間の重複率とのうちの一方である、請求項1から3のいずれか一項に記載の方法。
- 5前記パラメータが、少なくとも1つの注入ブロックについて擬似ランダムに変動する、請求項4に記載の方法。
- 6連続する2つの注入ブロックに適用される前記重み付け窓の合計が、前記2つのブロック間で重複するセグメント(l k )の合計に等しい、請求項1から5のいずれか一項に記載の方法。
- 7連続する2つの注入ブロックに適用される前記重み付け窓の2乗の合計が、前記2つのブロック間で重複するセグメント(l k )の2乗の合計に等しい、請求項1から5のいずれか一項に記載の方法。
- 8少なくとも1つの注入ブロックの符号が変更される、請求項1から7のいずれか一項に記載の方法。
- 9少なくとも1つの注入ブロックが、時間反転される、請求項1から8のいずれか一項に記載の方法。
- 10前記ブロックが、まず、中間ノイズ信号に注入され、 前記中間ノイズ信号が、その後、前記信号構造に注入される、請求項1から9のいずれか一項に記載の方法。
- 11前記ブロックが、前記信号構造に実時間で注入される、請求項1から9のいずれか一項に記載の方法。
- 12プロセッサによって実行されると請求項1から11のいずれか一項に記載の方法を実施する命令を含むコンピュータプログラム。
- 13少なくとも1つの消失した信号フレームを置き換えるための手段(MEM、PROC)を備え、連続したフレームに分割された一連のサンプルを含む信号を復号化するためのデバイスであって、 消失した前記一連のサンプルを置き換えるための信号構造を生成する(S6)ための手段であって、前記信号構造が、復号中(S1)にかつ消失した前記一連のサンプルよりも前に受信された有効サンプルから決定されたスペクトル成分を含む、手段と、 復号器で利用可能な、受信された有効サンプルを含むデジタル信号(S602)と、前記スペクトル成分から生成された信号(S601)との間で残差を生成する(S603)ための手段と、 前記残差からブロックを抽出する(S605)ための手段と、 前記ブロックを前記信号構造に注入するための手段とを備え、 前記注入するための手段は、窓重み付けブロックを、重複加算手法を用いて利用し、 連続する2つの 注入された前記ブロックは、時間の点で少なくとも部分的に重複する、デバイス。
Independent claims13
113 paragraphs, as filed
The present invention relates to signal correction, and more particularly to signal correction in the decoder when the signal received by the decoder has frame loss.
A signal is a form in which a series of samples is divided into consecutive frames, where the term frame means a signal segment composed of at least one sample (having a frame containing a single sample). In that case, it simply corresponds to a signal in the form of a series of samples).
The present invention belongs to the field of digital signal processing, and is not limited thereto, and particularly belongs to the field of coding / decoding of audio signals. Frame loss is interrupted by channel conditions (due to radio problems, network congestion, etc.) where communication using the encoder and decoder (whether transmitted in real time or after storage) Occurs when you do.
In this case, the decoder uses a packet loss correction mechanism (or "masking") and uses information available within the decoder, such as a signal that has already been decoded or a parameter received within a previous frame. , Attempts to replace the missing signal with a rebuild signal. This technology makes it possible to maintain good service quality even when channel performance deteriorates.
Frame loss correction techniques are often highly dependent on the type of coding used.
When encoding a voice signal based on CELP (abbreviation for "Code Excited Linear Prediction") technology, frame loss correction applies the CELP model. For example, when encoding according to Recommendation G722.2, a solution to replace lost frames (or "packets") prolongs its use by attenuating long-term prediction (LTP) gain. In addition, the use of each ISF is prolonged by bringing each ISF (abbreviation of "Immittance Spectral Frequency") parameter closer to the average of each. The pitch period of the spoken signal (specified "LTP lag") is also repeated. In addition, the decoder is supplied with random values for the parameters that characterize "innovation" (excitation in CELP coding).
Applying this type of method to transform coding or PCM ("Pulse Code Modulation") coding requires CELP coding within the decoder, which adds complexity. Please note that
In ITU-T Recommendation G.711 for Waveform Encoders, the process for frame loss correction (exemplified in Appendix I of that Recommendation) finds the pitch period of the decoded voice signal and ends. The pitch period is repeated, and at that time, duplicate addition is used between the decoded signal and the repeated signal. This action "erases" the voice artifacts, but requires additional time for the decoder (the time corresponding to the duration of the duplication).
The most commonly used technique for correcting frame loss during transform coding consists of repeating the decoded spectrum in the last frame received. For example, for coding according to Recommendation G.722.1, MLT (modulated lapped transform), or modified discrete cosine (MDCT) using a sinusoidal window with 50% overlap. For a transform equal to transform), the transition (between the last vanishing frame and the repeating frame) is guaranteed to be slow enough to eliminate the artifacts caused by the simple repetition of the frame.
Advantageously, the technique takes advantage of the temporary aliasing of the MLT transform to perform duplicate addition with the reconstructed signal, so no additional time is required. This technology is very inexpensive in terms of resources.
However, this technique has a defect associated with a temporary inconsistency between the signal just before the frame disappears and the repeat signal. This inconsistency causes discontinuities in the audible phase, which can result in significant audio artifacts when the overlap between the two frames is small (as with using a "low latency" MDCT window). This short-duplication situation is illustrated in Figure 1B for a low-latency MLT transformation, and the usual situation in Figure 1A where a long sign window is used according to Recommendation G.722.1 (in this case, long duplication of very loose modulation). Compare with period ZRA provided). It can be seen that in the modulation by the low delay window, the audible phase shift occurs because the overlapping region ZRB is short as shown in FIG. 1B.
In this case, a solution that combines pitch detection (when encoding according to Appendix I of Recommendation G.711) with duplicate addition performed by the M DCT transform window was still associated with phase shift. Not enough to eliminate voice artifacts.
Another frame loss correction technique produces a composite signal from the signal structure extracted from the pitch period. It should be understood that the pitch period means the fundamental period, especially in the case of a voiced voice signal (the reciprocal of the fundamental frequency of the signal). However, the signal may also be derived, for example, from a musical signal whose entire tone is associated with a fundamental frequency and a fundamental period that may correspond to the repetition period.
However, the physical properties of the composite signal are inconsistent with the physical properties of the original signal (some frames are lost), causing unpleasant hearing loss. This causes more anomalies than the original signal. Moreover, the energy of the correctly received signal can be significantly different from the energy of the signal reconstructed from the signal structure described above. These differences can result in a "noise jump" audible sensation in which the noise level fluctuates sporadically. For example, if the noise signal is equivalent to the background noise, the listener will hear a jump in the background noise.
More generally, current technology introduces periodicity by generating a composite signal that replaces the vanishing frame and fills the frame, and for complex signals such as music, this periodicity is the range of all signal components to replace. The present inventors have noted that they are not compatible with.
For example, seeing Figure 1C, the signal S<sub>0</sub>But window F<sub>1</sub>From F<sub>7</sub>It is repeated 7 times. Window time characteristics (window start time v<sub>1</sub>From v<sub>7</sub>, And window duration L<sub>0</sub>From L<sub>7</sub>) Are the same, so periodicity is introduced.
This regular and improper periodicity results in a "metallic" and artificial (and therefore unpleasant to the listener) sound at each frame disappearance. Therefore, it is necessary to improve the existing replication method, including, but not limited to, an example of decoding using duplicate addition.
<p> The present invention remedies this situation.</p>
<p> To this end, the present invention proposes a method for processing a digital signal, which is performed during decoding of the digital signal to replace a series of samples lost during decoding. This method-is the step of generating a signal structure to replace the lost set of samples, which is determined from the valid samples received during decoding and prior to the lost set of samples. A step that includes a spectral component, and a step that creates an error between the digital signal that contains the received valid sample that is available in the decoder and the signal that is generated from the spectral component. --Includes the step of extracting blocks from the residuals. In particular, window weighted blocks are injected into the signal structure using the duplication addition technique, and these injected blocks overlap at least partially in terms of time.</p><p> Therefore, block injection allows the lost frame to be filled without perceptible loss of signal energy. The injection of this block smoothes the signal energy and artificially restores the spectral density to a certain level. The set of injected blocks corresponds, for example, to the noise signal injected into the replacement signal. In particular, the overlapping addition makes it possible to smooth the energy transition of the noise signal in the transition region.</p><p> In addition, the present invention proposes to re-inject the various extracted blocks without noticeable periodicity, thus avoiding the audible "metallic" effects associated with simple repetition of residuals. .. In particular, the partial overlap of blocks smoothes the transition of the noise signal between two consecutive blocks, thus reducing the periodic effect. Such overlap makes the transition from one cycle to another more indistinguishable, thereby limiting the periodic action.</p><p> It should be understood that the term "replacement signal structure" refers to a set of characteristics specific to a replacement signal, such as the spectral component of this signal, the amplitude associated with the spectral component, and the phase associated with the spectral component.</p><p> The overlap of blocks is at least partial, so blocks can be completely overlapped in a complementary manner, for example by two adjacent blocks. In another example, the first block completely overlaps with the beginning of the second block.</p><p> In one particular embodiment, the replacement signal structure can include spectral components determined from valid samples received during decoding and prior to the lost series of samples. Therefore, the replacement signal can be easily reproduced even in a period different from the period in which the spectral components are determined.</p><p> Further, this residual can be generated from the residual between a portion of the digital signal containing the received valid sample and the signal generated from the spectral components described above . Therefore, the block extracted from this residual is suitable for the signal to be reconstructed in that the lost energy component is injected into the replacement signal. In fact, the spectral components of the injected block correspond exactly to the missing spectral components in the signal generated from the replacement signal structure described above. The spectral density of the signal into which the block is injected corresponds to the spectral density of the signal to which the frame was correctly received. Therefore, the signal energy is advantageously harmonized (between the correctly received signal portion and the reconstructed portion).</p><p> In another embodiment, the block is defined by an extraction block start time and a block duration, so that at least one parameter of this extraction block start time and this block duration is at least two extractions. It may vary between blocks.</p><p> Alternatively, the block is injected with at least one parameter that is variable between at least two injection blocks, and this variable parameter is-the write start time of the injection block and-between two consecutive injection blocks. It is one of the duplication rates.</p><p> For example, inconsistency is introduced in the signal that replaces the lost sample. The volatility of the parameters described above eliminates the periodicity of the signal. When these parameters fluctuate, the signal no longer repeats in the same way after a certain time interval. Therefore, the impression of metallic sound caused by the repetition of noise signals is eliminated. Such variability can occur, for example, in these parameters by quasi-randomly or quasi-randomly determining at least one condition according to a predetermined rule.</p><p> In another alternative form, at least one of the above parameters can be varied pseudorandomly for at least one injection block.</p><p> It should be understood that the term "pseudo-random" means a series of numbers that are statistically close to perfect randomness. Depending on the algorithmic process used to generate this series of numbers and the source used, this series of numbers cannot be considered completely random. The conditions associated with the pseudo-random determination of at least one parameter can also be considered. For example, the average of all determined parameters can be fixed. In this situation, for example, parameters that are quasi-randomly derived and have the effect of establishing a predetermined interval average can be identified. The choice of parameter variability (pseudo-random, conditional pseudo-random, preset rules, etc.) includes the number of samples lost during decoding, the quality level of the signal required by the user, the resources available for reconstruction calculations, etc. It can meet the conditions by itself.</p><p> The above-mentioned parameters generated as described above introduce inconsistency into the noise signal, and this inconsistency gives the injected noise an imperceptible artificial property. The introduction of pseudo-randomly generated parameters means that any phenomenon, such as a repeating array of noise signals being heard, is extremely unlikely to occur. There is no logic between the different weighted windows. Therefore, the listener does not suffer from the impression that the noise signal (for example, background noise) is repeated.</p><p> In another embodiment, the above parameters for block extraction and / or block injection are pre-fixed. Therefore, the use of pre-defined blocks simplifies calculations, reduces processing time, and at the same time reduces the load on the processor used for these calculations.</p><p> In one embodiment, the sum of the weighted windows applied to two consecutive injection blocks is equal to the sum of the overlapping segments between these two blocks. Therefore, the amplitude of the replacement signal is constant and the signal is not interrupted by transition artifacts between the two blocks.</p><p> In another embodiment, the sum of the squares of the weighted windows applied to two consecutive injection blocks is equal to the sum of the squares of the overlapping segments between these two blocks. Therefore, the energy of the replacement signal is constant, and the energy of the signal becomes constant over time.</p><p> In one embodiment, the sign of at least one injection block can be changed. The blocks to be flipped are, for example, pseudo-randomly, pseudo-randomly with at least one condition (eg, modifying the maximum number of windows), or a given rule (every other window, all windows of a certain length). Etc.). Therefore, additional inconsistency is added to the noise signal. Also, this addition of inconsistency is done without complicating the steps for generating the replacement signal. Inversion of noise signals does not require large computational resources, reduces processing time, and reduces the load on the processor used for these calculations.</p><p> In one variant, at least one injection block is time inverted.</p><p> The term "time inverted" means applying the equation b (t) = b (FF + DF-t) to the time t-dependent block b in the weighted window [DF; FF]. Please understand that. Therefore, a new inconsistency is introduced in the replacement signal.</p><p> In another embodiment, the blocks are first injected into the intermediate noise signal, all blocks are injected into the intermediate noise signal, and then the intermediate noise signal itself is injected into the signal structure. Therefore, the noise signal to be injected into the replacement signal is injected after it is completely generated. By doing so, it becomes possible to establish a verification mechanism for verifying the intermediate audio signal before injecting the intermediate audio signal into the replacement signal.</p><p> Alternatively, the blocks are injected in real time without waiting for the intermediate noise signal to be fully generated. In that case, it should be understood that "real-time" injection means that the block is injected at a rate adapted to the temporal progression of the signal. In this situation, the time lag between the signal received by the decoder and the signal delivered to the listener's ear is as small as possible. For example, a replacement signal structure is generated at the beginning of a series of samples that are lost during decoding, and then as time passes and the signal progresses, blocks are injected without completely producing an intermediate noise signal. It is then injected into the replacement signal.</p><p> The present invention also provides a computer program containing instructions for carrying out the above method. For example, one or more of FIGS. 5-8 may be general algorithms for such computer programs.</p><p> The present invention comprises means for replacing at least one lost signal frame and can be implemented by a device for decoding a signal containing a series of samples divided into successive frames. This device is a means for generating a signal structure to replace the lost series of samples, from valid samples received during decoding and prior to the lost series of samples. Means containing the determined spectral component and-means for generating a residual between the digital signal containing the received valid sample available in the decoder and the signal generated from the spectral component. , --Means for extracting blocks from the residuals and --Means for injecting blocks into the signal structure are provided, and the means for injecting use window weighted blocks using the duplicate addition method. The injected blocks overlap at least partially in terms of time.</p><p> Such devices can take the form of physical forms, such as processors, and perhaps working memory, typically within a communication terminal.</p><p> Other features and advantages of the invention will become apparent by reading the following detailed description of some embodiments of the invention and referring to the drawings below.</p>
<figref num="1A">It is a figure which shows the overlap of the conventional window in MLT conversion.</figref><figref num="1B">It is a figure shown to compare the overlap of a low delay window with the overlap shown in FIG. 1A.</figref><figref num="1C">It is a figure which shows the periodic duplication of a noise signal.</figref><figref num="2">It is a figure which shows an example of the technical framework which can carry out this invention.</figref><figref num="3">FIG. 5 is a schematic representation of a device comprising means for carrying out the method according to the invention.</figref><figref num="4">It is a figure which shows an example of the general process of this invention.</figref><figref num="5">It is the schematic of the step of the method of this invention in one Embodiment.</figref><figref num="6">It is the schematic of the step of the method of this invention in another embodiment.</figref><figref num="7">It is the schematic of the step of the method of this invention in another embodiment.</figref><figref num="8">It is the schematic of the step of the method of this invention in another embodiment.</figref><figref num="9A">It is a figure which shows the continuous weighting window of this invention at a certain overlap rate determined according to one Embodiment.</figref><figref num="9B">It is a figure which shows the continuous weighting window of this invention at a certain overlap rate determined according to one Embodiment.</figref><figref num="9C">It is a figure which shows the continuous weighting window of this invention at a certain overlap rate determined according to one Embodiment.</figref><figref num="10">It is a figure which shows the continuous weighting window of this invention with a pseudo-random overlap rate determined according to one Embodiment.</figref><figref num="11">It is a figure which shows the continuous weighting window determined according to one Embodiment of this invention.</figref>
Next, an advantageous but optional example for carrying out the present invention will be described with reference to FIG. This figure relates to the processing performed in the decoder for the received signal. The decoder may be of any type, and the process as a whole is generally independent of the type of encoding / decoding. In the example described, this process applies to the received audio signal. However, this process can more generally be applied to any kind of signal analyzed by time-windowing and transformation, in which one or more at the time of synthesis. Harmonization is performed using the duplicate addition method for the replacement frame of.
It should be understood that the term "frame" means at least one block of samples. For most codecs, these frames consist of several samples. However, in some codecs, for example PCM (Pulse Code Modulation) according to Recommendation G.711, the signal simply consists of a series of samples (in this case, a "frame" in the sense of the present invention contains only one sample. Absent). The present invention can also be applied to this type of codec.
For example, the active signal can consist of the last active frame received before the frame disappeared. It is also possible to use one or several subsequent valid frames received after the lost frame (although in such an embodiment there will be a delay in decoding). The sample used from the active signal can be the frame sample itself, perhaps the sample corresponding to the conversion memory, or, in the case of conversion decoding using MDCT or MLT duplication, typically a sample containing aliasing.
In the first step S1 of the process of FIG. 2, N audio samples are sequentially stored in a buffer (FIFO buffer, etc.). These samples correspond to the already decoded samples and are therefore accessible when dealing with frame loss. If the first sample to be synthesized is a sample with a time index of N (among one or more consecutive vanishing frames), the audio buffer b (n) will be N with a time index of 0 to N-1. Corresponds to one preceding sample.
In the filtering step S2, the audio buffer b (n) is then separated into two frequency bands, a low frequency band BB and a high frequency band BH, at a separation frequency described below as Fc, for example, Fc = 4kHz.
Step S3 is applied to the low frequency band, which consists of searching for a turnaround point and a segment of length P corresponding to the basic period of buffer b (n) resampled at frequency Fc. In the case of a voiced voice signal, the fundamental period corresponds to, for example, the pitch period (the reciprocal of the fundamental frequency of the signal). However, the signal may also be derived from, for example, a music signal having a fundamental frequency and an overall tone associated with a fundamental period that may correspond to the repetition period.
Although the following assumes that only one basic period of length P is used for signal synthesis, it should be noted that this principle of processing applies equally well to segments over several basic periods. For some basic periods, even better results are obtained in terms of FFT accuracy and the richness of the resulting spectral components.
The next step S4 consists of decomposing the segment p (n) into the sum of the signs.
In step S5 of FIG. 2, the sinusoidal component is selected so that only the most important components are retained.
The next step S6 is the synthesis of the sine wave. In one exemplary embodiment, this step S6 comprises generating segments s (n) that are at least as long as the dimensions of the vanishing frame (T). In one particular embodiment, between the composite signal (frame loss corrected) and the signal decoded at such a next valid frame if the next valid frame is correctly received again, ( A length equal to two frames (eg 40ms) is generated to allow crossfade audio mixing (as a transition).
In anticipation of frame resampling (sample length indicated by LF), the number of samples to be combined can be increased by half the size of the resampling filter (LF). The combined signal s (n) is calculated as the sum of the selected sinusoidal components as follows.
<maths num="1"><img file="JP6469079B2_D0001.tif" /></maths>
In the equation, k is the index of the K components selected in step S5. There are several conventional methods possible to perform this sinusoidal synthesis.
Step S7 in FIG. 2 consists of injecting noise in the low frequency band to compensate for the energy loss caused by the omission of certain frequency components.
A simplified embodiment of the present invention can be described above with reference to FIG. In this embodiment, in step P5, the residual r (n) = between the signal block p (n) corresponding to the pitch extracted in step P1 and the composite signal s (n) generated in step P3 = It consists of calculating p (n) -s (n) as n [0; P-1] from the sinusoidal analysis performed in step S4.
This residual is dimensioned in step P6<maths num="2"><img file="JP6469079B2_D0002.tif" /></maths>Is converted to, and the signal b (n) is obtained in step P7.
Then, in step P8, the signal b (n) is injected into the signal s (n) generated in step P2 for a duration N corresponding to the duration in which the signal should be replaced.
The replacement signal f (n) is then mixed with the active signal in step P9. This mixing can include, for example, a duplicate addition RECOV performed over the overlap interval RO.
In one embodiment, the residual signal is replicated once or multiple times (depending on the time portion to be filled) with duplicate addition between the replicated signals.
In another embodiment, in each replication, various transformations can be applied to the blocks of the residual signal in a pseudo-random manner, thus inverting the sign of the signal and / or performing time inverting. Is.
Next, a method for generating a noise signal to be injected into the replacement signal structure according to the embodiment of the present invention will be described with reference to FIG.
In step S601, the signal s (n) is generated from the sinusoidal synthesis of step S6 (see also FIG. 2) over a time period corresponding to the time period of block p (n) extracted in step S602.
The residual r (n) is obtained by SUB subtracting the signal s (n) from the signal p (n). As a result, in step S603, r (n) such that r (n) = p (n) -s (n) is obtained.
In step S604, the counter variable k is initialized to 0, and the signal b (n, k) is initialized so that b (n, 0) = 0.
In step S605, block r (n, k) is extracted from signal r (n). In one embodiment, the temporal characteristics of this extraction (block i)<sub>k</sub>Start time, and block L<sub>k</sub>(Duration of) is calculated in a pseudo-random manner. In another embodiment, conditions can be imposed on this extraction. For example, the sum of the block start time value and the duration value must be less than the value corresponding to the duration value of block p (n) extracted in step S602.
In step S606, the duration L of the extracted block r (n, k)<sub>k</sub>Is sent to window configuration step S608.
In step S607, a set of weighted windows becomes available so that the weighted window can be configured in step S608. For example, the weighted window stored in memory is extracted and transferred to working memory.
In step S608, a weighting window is selected and configured so that block r (n, k) can be multiplied in step MULT. For window parameters, the duration L suitable for block r (n, k)<sub>k</sub>Is included.
Then block w<sub>k</sub>.r (n, k) is added in duplicate with the signal b (n, k-1) corresponding to the added (k-1) block, so b (n, k) = w<sub>k</sub>It becomes .r (n, k) + b (n, k-1). In one embodiment, the duplication addition is performed with a fixed duplication rate of 50%.
Test T609 verifies that the length of the generated signal b (n, k) is not greater than the value N corresponding to the duration of the signal to be replaced.
If greater than the value N, the signal b (n, k) is truncated in step S612 so that the time length of b (n, k) is equal to the value N corresponding to the duration of the signal to be replaced. , This truncated value is indicated by TQ. In step S613, the noise signal Y to be injected into the replacement signal instead of the vanishing frame is set in TQ and injected in step S7 (see also Figure 2).
If it is not greater than the value N, the value of b (n, k) is stored in the working memory MEM (see Figure 3) for later addition to the next block r (n, k + 1). In step S611, the counter variable k is incremented and the procedure returns to step S605.
Next, with reference to FIG. 6, a method for generating a noise signal to be injected into the replacement signal structure according to another embodiment of the present invention will be described.
In this embodiment, the residual signal is an overlay addition signal block r'obtained from the residual r (n).<sub>k</sub>It is continuously and repeatedly (k times) injected in (n).
At iteration k, the block reading is the block start index i<sub>k</sub>And block length L<sub>k</sub>How to inject this residual part into the target time slot is an optional transformation T<sub>k</sub>, Write index j<sub>k</sub>(Start copying block in time slot to fill), and duplicate addition window w<sub>k</sub>Specified by determining (n).
Complementary signals with dimensions of N samples to be generated from the residuals are shown as b (n). The procedure for generating a noise signal will be described below.
Initialization: b (n) = 0, 0 n <N k = 0 j<sub>0</sub>=0
j<sub>k</sub>+ L<sub>k</sub>Repeat as follows until = N. 1) i<sub>k</sub>And L<sub>k</sub>And, i<sub>k</sub>+ L<sub>k</sub> P and j<sub>k</sub>+ L<sub>k</sub>Select so that N, and extract block P (k). 2) r'<sub>k</sub>(n) = T<sub>k</sub>(r<sub>k</sub>(i<sub>k</sub>Transform T so that S (k) corresponding to + n)) is obtained<sub>k</sub>Select. This conversion will be described later. 3) j<sub>k</sub>+ L<sub>k</sub>If <N, j to be prepared for duplication with the next iteration<sub>k + 1</sub>j<sub>k</sub>+ L<sub>k</sub>(Preferably, to limit the number of duplicates at the same time to a maximum of 2 blocks j<sub>k + 1</sub> j<sub>k-1</sub>+ L<sub>k-1</sub>, For example S (k) and S (k + 1)) to extract block P (k + 1). 4) Weighted window w based on any overlap of adjacent blocks<sub>k</sub>Determine (n). 5) Window w<sub>k</sub>R'weighted by (n)<sub>k</sub>Paste (n) as follows. b (j<sub>k</sub>+ n) = b (j<sub>k</sub>+ n) + r'<sub>k</sub>(n) .w<sub>k</sub>(n), but 0 n <L<sub>k</sub> 6) Increment by k = k + 1.
In this embodiment, the write index j is according to the procedure described.<sub>k</sub>To increase. Any other progression (decrease, non-monotonic, etc.) can be selected.
In another embodiment, L<sub>k</sub>Is selected to be relatively large compared to the available reserve P so that it can proceed significantly in copying and avoids distortion of relatively low frequency components. For example, see Figure 11 where L<sub>0</sub>Is chosen to be relatively large so that the duplicate addition is applied only once.
In another embodiment, the dimensions of the overlapping region j to limit the number of addition and multiplication operations required.<sub>k</sub>+ L<sub>k</sub>-j<sub>k + 1</sub>To reduce. (Overlapping area dimensions j<sub>k</sub>+ L<sub>k</sub>-j<sub>k + 1</sub>The overlap rate adjustment (corresponding to) can also be configured so that the ratio between quality (elimination of artifacts) and processing cost fits the planned use of the decoder.
In a preferred embodiment, referring to FIG. 7, the weighted window ensures that the transitions in the pasted sections are smooth and that the resulting signal is continuous in terms of signal energy. Is stipulated to do. Typically, up to two blocks are planned to overlap at any point. Consider the overlap between blocks S (k) and S (k + 1). The frame ZP in FIG. 7 represents an enlarged view of the area ZM surrounded by the frame.
In the overlapping region, n [0; l<sub>k</sub>[, Here l<sub>k</sub>= j<sub>k</sub>+ L<sub>k</sub>-j<sub>k + 1</sub>Then, the resulting signal is as follows. b (j<sub>k + 1</sub>+ n) = r'<sub>k</sub>(j<sub>k + 1</sub>-j<sub>k</sub>+ n) .w<sub>k</sub>(j<sub>k + 1</sub>-j<sub>k</sub>+ n) + r'<sub>k + 1</sub>(n) .w<sub>k + 1</sub>(n)
In one embodiment, w<sub>k</sub>At the end of and w<sub>(k + 1)</sub>The beginnings of are combined as follows according to a criterion called "preservation of amplitude" below. w<sub>k</sub>(j<sub>k + 1</sub>-j<sub>k</sub>+ n) + w<sub>k + 1</sub>(n) = 1
Therefore, the crossfade function f, which is typically augmented and bounded by 0s and 1s.<sub>lk</sub>It is sufficient to select (n), and from there n [0; l<sub>k</sub>For [, it is sufficient to deduce the following. W<sub>k</sub>(j<sub>k + 1</sub>-j<sub>k</sub>+ n) = f<sub>out out</sub>(n) = 1-f<sub>lk</sub>(n) and w<sub>k + 1</sub>(n) = f<sub>in</sub>(n) = f<sub>lk</sub>(n)
For example, the crossfade function can be improved and specified by:
<maths num="3"><img file="JP6469079B2_D0003.tif" /></maths>
Function f in Figure 7<sub>in</sub>In another example represented by (n), the crossfade function may be a sine function and can be specified by:
<maths num="4"><img file="JP6469079B2_D0004.tif" /></maths>
In another embodiment, a criterion called "energy conservation" is selected, where the pasted signals can be combined without phase coherence and can be specified by: (w<sub>k</sub>(j<sub>k + 1</sub>-j<sub>k</sub>+ n))<sup>2</sup>+ (w<sub>k + 1</sub>(n))<sup>2</sup>=1
The crossfade function f proposed above<sub>k</sub>From (n), n [0; l<sub>k</sub>For [, the following can be deduced.
<maths num="5"><img file="JP6469079B2_D0005.tif" /></maths>
Each weighted window typically consists of three parts, from left to right: --Increased part (complementing the decreasing part of the preceding window), --Constant and conservative part (gain 1), and --Decreasing part
In one embodiment, for at least one weighted window, at least one of these parts is zero in length. For example, if the first block completely overlaps the beginning of the next injection block, the weighting window applied to this first injection block consists of only the reduced portion.
In another embodiment, the crossfade effect on the two blocks is managed simultaneously across their overlapping areas. This is simply a split of the steps described above and reassembled differently.
In that case, each iteration is-the stage of pasting without duplication and therefore without windowing (w)<sub>k</sub>(n) Elimination of multiplication by = 1), and / or-the above crossfade function f<sub>out out</sub>(n), and f<sub>in</sub>It consists of the steps of crossfading and pasting the end of the previous block and the beginning of the new block using (n).
The above will be described in more detail with the following procedure called "simultaneous crossfade".
Initialization: b (n) = 0, 0 n <N k = 0 j<sub>0</sub>= 0 l<sub>-1</sub>= 0 i<sub>0</sub>And L<sub>0</sub>And, i<sub>0</sub>+ L<sub>0</sub> P and j<sub>0</sub>+ L<sub>0</sub>Select so that N. From there, deduce the overlapping dimensions (l<sub>0</sub>= j<sub>0</sub>+ L<sub>0</sub>-j<sub>1</sub>), J<sub>1</sub> j<sub>0</sub>Select (however, j<sub>1</sub>j<sub>0</sub>+ L<sub>0</sub>). T<sub>0</sub>And T<sub>1</sub>Select the conversion of. R'<sub>0</sub>= T<sub>0</sub>(r<sub>0</sub>(i<sub>0</sub>+ n)) is calculated.
j<sub>k</sub>+ L<sub>k</sub>Repeat as follows until = N. 1) j<sub>k + 1</sub>> j<sub>k</sub>+ l<sub>k-1</sub>If, paste without duplication or windowing is performed as follows. b (j<sub>k</sub>+ n) = r'<sub>k</sub>(n), l<sub>k-1</sub>n <L<sub>k</sub>-l<sub>k</sub> 2) Crossfade paste in overlapping areas b (j)<sub>k + 1</sub>+ n) = r'<sub>k</sub>(L<sub>k</sub>-l<sub>k</sub>+ n) .f<sub>out out</sub>(n) + r'<sub>k + 1</sub>(n) .f<sub>in</sub>(n), 0 n <l<sub>k</sub> 3) When another iteration is required (especially j<sub>k</sub>+ L<sub>k</sub><N), a) j<sub>k + 1</sub>j<sub>k</sub>+ L<sub>k</sub>Select, but (to limit simultaneous duplication to a maximum of 2 blocks) j<sub>k + 1</sub> j<sub>k-1</sub>+ L<sub>k-1</sub>And. b) i<sub>k + 1</sub>And L<sub>k + 1</sub>And, i<sub>k + 1</sub>+ L<sub>k + 1</sub> P and j<sub>k + 1</sub>+ L<sub>k + 1</sub>Select so that N. c) Conversion T<sub>k + 1</sub>Select r'<sub>k + 1</sub>(n) = T<sub>k + 1</sub>(r<sub>k + 1</sub>(i<sub>k + 1</sub>+ n)) (see below for details). 4) Increment by k = k + 1.
In the variant, the principle of crossfade is applied between the newly pasted block and the signal generated at the overlap, b (j).<sub>k + 1</sub>+ n) = b (j<sub>k + 1</sub>+ n) f<sub>out out</sub>(n) + r'<sub>k + 1</sub>(n) .f<sub>in</sub>Apply so that it becomes (n). The present embodiment has an advantage that three or more blocks can be managed so as not to overlap at the same time without complicating the calculation.
Therefore, the parameter i<sub>k</sub>, J<sub>k</sub>, L<sub>k</sub>, And T<sub>k</sub>At least one of them fluctuates between one iteration and another to avoid periodic effects and the associated auditory artifacts (metallic artificial sounds).
In the filled time slot, the index i of one pasted block against another pasted block<sub>k</sub>, I<sub>k + 1</sub>, J<sub>k</sub>, And j<sub>k + 1</sub>Delay information d<sub>k, k + 1</sub>, D<sub>k, k + 1</sub>= (j<sub>k + 1</sub>-i<sub>k + 1</sub>)-(j<sub>k</sub>-i<sub>k</sub>) Can be deduced.
In a preferred but non-limiting form, d<sub>k, k + 1</sub>Is set differently for one iteration k and the next iteration k + 1.
In one embodiment, simple or complex transformations during iteration (T in the above) to improve the elimination of artifacts.<sub>k</sub>(Shown in) can be introduced in a variable manner, which has the advantage of introducing an uncorrelated form between the injected signal portions.
T, one of the simple transformations that can be performed<sub>k</sub>Consists of changing the sign of the signal as follows. r'<sub>k</sub>(n) = T<sub>k</sub>(r<sub>k</sub>(i<sub>k</sub>+ n))) = σ<sub>k</sub>r<sub>k</sub>(i<sub>k</sub>+ n), however, depending on the iteration σ<sub>k</sub>=±1
One of the feasible transformations that can be combined with the previous transformations and can be applied quasi-randomly consists of time inversion, which means that the reading or writing of the residuals is done in reverse order as follows: .. r'<sub>k</sub>(n) = T<sub>k</sub>(r<sub>k</sub>(i<sub>k</sub>+ n))) = σ<sub>k</sub>r<sub>k</sub>(i<sub>k</sub>+ L<sub>k</sub>-1-n), 0 n <L<sub>k</sub>
Other transformations with more complex computational costs, such as phase shift filters, are also possible. Phase-shift filters, also called all-pass filters, show the same gain over the entire frequency range used, but the relative phases of the frequencies that make up the signal fluctuate with frequency.
In the present specification, the intermediate variable r'is used for ease of explanation.<sub>k</sub>(n) is introduced, but the conversion T<sub>k</sub>Can be performed as a particular form for reading a digital sample without necessarily requiring intermediate storage in a buffer between reading from r (n) and writing to b (n).
In another embodiment, the kth injected signal portion is the generated complementary signal b (n), 0 n <j.<sub>k-1</sub>+ L<sub>k-1</sub>It can be obtained from, not just from the residual r (n).
Next, a modified embodiment including the above-mentioned "simultaneous crossfade" procedure incorporated in the digital speech decoder is shown as an example with reference to FIG.
Initialization j<sub>1</sub>= j<sub>0</sub>= 0: A crossfade of two blocks is applied when embedding begins. I<sub>0</sub>= P / 2 L<sub>0</sub>= P / 2
At each iteration-read index i<sub>k</sub>(If k> 0) is the start i of the calculated residual segment r (n)<sub>k</sub>Indicates = 0. --The crossfade function is the following sine function. F<sub>out out</sub>(n) = 1-f<sub>lk</sub>(n) f<sub>in</sub>(n) = f<sub>lk</sub>(n) However<maths num="6"><img file="JP6469079B2_D0006.tif" /></maths>
--If two blocks overlap at the same time, so k> 0, then j<sub>k + 1</sub>= j<sub>k</sub>+ l<sub>k-1</sub>= j<sub>k-1</sub>+ L<sub>k-1</sub>Will be. --The total size of each pasted block is the total L that connects the two overlapping areas.<sub>k</sub>= l<sub>k-1</sub>+ l<sub>k</sub>Corresponding to, this is the overlap region dimension lk obtained at each iteration, from which Lk and j<sub>k + 1</sub>Is deduced. This parameter l<sub>k</sub>Is calculated as follows in proportion to the dimension P / 2, which is half of the available residuals.
<maths num="7"><img file="JP6469079B2_D0007.tif" /></maths>
However, k'= mod (k + cnt_bfi), where cnt<sub>bfi</sub>Is the reciprocal of the number of lost frames, α = [1 0.8 0.6 0.9]. --Conversion T<sub>k</sub>Consists of changing the sign from time to time (no time inversion) and the coefficients<maths num="8"><img file="JP6469079B2_D0008.tif" /></maths>Indicated by.
The first step of the above method is shown in the table below with reference to FIG. Step INIT corresponds to the initialization of the method, and steps ST (0), ST (1), and ST (2) correspond to the first increment of the method.
<tables num="1"><img file="JP6469079B2_D0009.tif" /></tables>
When the complementary signal b (n) is generated for the desired time portion, this signal is added to the signal s (n) (where n> 0) generated by sinusoidal synthesis.
In a preferred embodiment, at least one parameter of the block is determined quasi-randomly to introduce inconsistency in the replacement signal, thus limiting the periodic phenomenon that causes hearing discomfort. The parameters of the weighted window are, for example, the extraction block start time and the block duration (parameter L described above).<sub>k</sub>(Similar to), the overlap rate of two consecutive blocks.
Referring to FIG. 9A, in one exemplary embodiment, the noise signal indicates that the entire block is injected once and then injected into the replacement signal, with the start time of writing the injection block at a constant overlap rate. Pseudo-randomly determined. In FIGS. 9A-11, the arrows indicate parameters determined in a pseudo-random manner. Once the first two parameters (block start time and duplication rate) are fixed, the block duration is deduced from these first two parameters. Other conditions can also work. For example, the total length of each block can be fixed so that the blocks do not exceed the duration N corresponding to the duration of the signal to be replaced. This condition can be expressed differently by considering that the sum of the start index of the last block and the length of the last block can be set to be shorter than the duration N. In practice, in a method for generating noise by repeating it continuously, these conditions can be checked with each overlapping addition.
For example, if the lost data to be replaced is 10 frames, the noise signal can be weighted by 20 weighted windows.
As mentioned above, the term pseudorandom is used in mathematics and computer science to refer to a series of numbers that are statistically close to perfect randomness. Depending on the algorithmic process used to generate this sequence of digits and the source used, this sequence of digits cannot be considered completely random. Of course, the parameters can be generated pseudo-randomly, but they still meet certain conditions, such as the length of the signal to be replaced.
Referring to FIG. 9B, in another embodiment, the block (L)<sub>0</sub>~ L<sub>5</sub>) Is pseudo-randomly determined with a constant overlap rate. Once the first two parameters are fixed, the starting index for writing the block is derived from these first two parameters. In this example, none of the parameters of the last block are determined pseudo-randomly, so the duration of the signal resulting from the overlap of all blocks is greater than the duration N corresponding to the duration of the signal to be replaced. It doesn't become.
Referring to FIG. 9C, in another embodiment, for an even window index, the duration of the block and the value of the start index for writing the injection block are determined quasi-randomly with a constant overlap rate. Therefore, j<sub>0</sub>, L<sub>0</sub>, J<sub>2</sub>, L<sub>2</sub>, J<sub>4</sub>, And L<sub>4</sub>Is determined pseudo-randomly, j<sub>1</sub>, L<sub>1</sub>, J<sub>3</sub>, L<sub>3</sub>, J<sub>5</sub>, And L<sub>5</sub>Is deduced from the parameters determined in a pseudo-random manner and the duplication rate. Conditions can be given to these parameters, so the duration of the signal resulting from the overlap of all s blocks does not exceed the duration N corresponding to the duration of the signal to be replaced.
With reference to FIG. 10, in another embodiment all parameters are determined quasi-randomly. However, these parameters may be conditioned so that the duration of the signal obtained from the overlapping injection blocks does not exceed the duration N corresponding to the duration of the signal to be replaced. In this configuration, the sum of two consecutive weighted windows is not equal to one, especially for the overlay segment between these two windows, and for the overlay segment between these two windows, two consecutive weights. The sum of the squares of the windows is also not equal to 1.
Next, returning to step S8 in FIG. 2, the high-frequency band that was not involved in steps S3 to S7 is processed by simply repeating the signal in this high-frequency band, so that the construction of the replacement signal can be continued arbitrarily. Can be done.
In step S9, the low frequency band is resampled at the original frequency Fc in step S70, and the resampled result is added to the signal obtained from the repetition of step S8 in the high frequency band to synthesize a signal.
In step S10, duplicate addition is performed so that the signal before the frame disappearance and the composite signal and between the composite signal and the signal after the frame disappearance are surely continuous.
Of course, the present invention is not limited to the above-described embodiment, but extends to other modifications.
For example, the separation into the high frequency band and the low frequency band in step S2 is optional. In an alternative embodiment, the signal from the buffer (step S1) is not separated into two subbands, and steps S3 through S10 remain identical to those described above. However, processing of spectral components at low frequencies can advantageously limit complexity.
The present invention can be practiced in a conversational decoder in the case of frame loss. Physically, the present invention can typically be implemented in a circuit for decoding within a telephone terminal. To this end, such a circuit CIR comprises or can be connected to a processor PROC, as shown in FIG. 3, and computer program instructions according to the invention for performing the above method are programmed. It can be equipped with a working memory MEM. For example, the present invention can be implemented in real-time conversion within a decoder.
More specifically, an embodiment based on a method for generating noise from the residuals between a known signal and a composite signal has been described above. Of course, it is also possible to calculate the residuals in the frequency domain (remove selected spectral components from the original spectrum) and obtain background noise by inverse transformation.
An embodiment based on a signal structure containing spectral components determined from valid samples received during decoding and prior to the lost series of samples has been described above. Of course, these spectral components may also be obtained from the samples received after this lost series of samples. These spectral components can also be determined from the samples received prior to this lost series of samples and the samples received after this lost series of samples. These spectral components may also be constant.
24 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP06149296A | Cites | Japan |
| JP2008261904A | Cites | Japan |
| JP2006053394A | Cites | Japan |
| JP2008058667A | Cites | Japan |
| US20080294428A1 | Cites | United States of America |
| JP2006526177A | Cites | Japan |
21 members in 12 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 1353551 | France | A | |
| 1353551 | France | A | |
| 1353551 | France | – | |
| 2014050945 | France | W | |
| 2014050945 | France | W | |
| 1353551 | – | – | – |
| FR20130053551 | – | – | – |
| FR2014050945 | – | – | – |
| WO2014FR50945 | – | – | – |
Members21
| Document | Office | Kind | |
|---|---|---|---|
| CA2909401A1 | Canada | A1 | |
| WO2014170617A1 | World Intellectual Property Organization (WIPO) | A1 | |
| FR3004876A1 | France | A1 | |
| KR20160002920A | Republic of Korea | A | |
| EP2987165A1 | European Patent Office (EPO) | A1 | |
| US2016055852A1 | United States of America | A1 | |
| CN105453172A | China | A | |
| JP2016515725A | Japan | A | |
| MX2015014650A | Mexico | A | |
| RU2015149384A | Russian Federation | A | |
| BR112015026153A2 | Brazil | A2 | |
| US9761230B2 | United States of America | B2 | |
| MX350721B | Mexico | B | |
| RU2647634C2 | Russian Federation | C2 | |
| EP2987165B1 | European Patent Office (EPO) | B1 | |
| JP6469079B2This record | Japan | B2 | |
| ES2704901T3 | Spain | T3 | |
| CN105453172B | China | B | |
| KR102184654B1 | Republic of Korea | B1 | |
| BR112015026153B1 | Brazil | B1 | |
| CA2909401C | Canada | C |
12 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6469079
- Publication, DOCDB
- 6469079
- Publication, EPODOC
- JP6469079B
- Application
- 2016508223
- Application, DOCDB
- 2016508223
- Application, EPODOC
- JP20160508223
Titles2
- Japanese
- 重み付けされたノイズの注入によるフレーム消失補正
- English
- Frame loss correction by weighted noise injection
Classification
- CPC, 4
- G10L19/005
- H03M13/00
- G10L21/0208
- G10L19/12
- IPC, 2
- G10L19 005
- G10L19 022
