Voice decoding method
Abstract
(57) A summary and subject CS*ACELP is applied to PHS and also let quality degradation to a transmission-line error be small. Solution means Every two frames, or every four-frame 8 kbits/s of coded voices of CS*ACELP of PHS are put, Error detection is carried out by the CRC check for every frame by a receiving side (S1), When usual will carry out decoding processing if there is no error (S3), and there was an error, by the error number in the past 100 ms Error rate P, Transmission-line error rate P is computed by the error number in the past 300 ms, and if the conditions which carried out error detection also of P>th1, th2<P<th3, and the front frame are fulfilled (S5), frame loss processing is performed (S6), and if the above-mentioned conditions are not fulfilled even if an error is detected, the usual decoding processing will be carried out.
Term
Term ended
Projected expiry passed 27 August 2016, 10.1 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
7 claims: 3 independent, 4 dependent
- 1[Claims] 1. In a method of decoding received encoded voice information on a frame-by-frame basis. The first process of detecting transmission line errors from received signals and The second process of calculating the average error rate from before a predetermined period to the present, When the transmission line error is detected, the third process of determining whether the frame loss condition is satisfied by using the average error rate, and When it is determined that the above frame loss condition is satisfied, the fourth process of performing frame loss processing for that frame and A voice decoding method including a fifth process of performing a decoding process without a frame loss process when the above frame loss condition is not satisfied and when a transmission line error is not detected. 【特許請求の範囲】 【請求項1】 受信符号化音声情報をフレーム単位で復号化する方法において、 受信信号から伝送路誤りを検出する第1過程と、 所定期間前から現在までの平均誤り率を算出する第2過程と、 上記伝送路誤りが検出されると、フレームロスの条件を満すかを上記平均誤り率を用いて判定する第3過程と、 上記フレームロスの条件を満すと判定されると、そのフレームについてフレームロス処理を行う第4過程と、 上記フレームロスの条件を満さない場合と伝送路誤りが検出されない場合はフレームロス処理を伴なわない復号化処理を行う第5過程とを有する音声復号化方法。
- 5In a method of decoding received encoded voice information on a frame-by-frame basis. The first process of detecting a predetermined bit error in the frame of the received signal, and The second process, which performs frame loss processing when a bit error is detected in the first process, and A voice decoding method including a third process of performing a decoding process without frame loss processing when a bit error is not detected in the first process. 【請求項5】 受信符号化音声情報をフレーム単位で復号化する方法において、 受信信号のフレーム中の予め決められたビットの誤りを検出する第1過程と、 その第1過程でビット誤りが検出されるとフレームロス処理を行う第2過程と、 上記第1過程でビット誤りが検出されないと、フレームロス処理を伴なわない復号化処理を行う第3過程とを有する音声復号化方法。
- 6In a method of decoding received encoded voice information on a frame-by-frame basis. The first process of detecting a predetermined bit error in the frame of the received signal, and The second process of calculating the average error rate from before a predetermined period to the present, When an error is detected in the first process, the third process of determining whether the frame loss condition is satisfied by using the average error rate, and When it is determined that the above frame loss condition is satisfied, the fourth process of performing frame loss processing for that frame and A voice decoding method including a case where the above-mentioned frame loss condition is not satisfied and a fifth process in which a decoding process is performed without frame loss processing when an error is not detected in the first process. 【請求項6】 受信符号化音声情報をフレーム単位で復号化する方法において、 受信信号のフレーム中の予め決められたビットの誤りを検出する第1過程と、 所定期間前から現在までの平均誤り率を算出する第2過程と、 上記第1過程で誤りが検出されると、フレームロスの条件を満すかを上記平均誤り率を用いて判定する第3過程と、 上記フレームロスの条件を満すと判定されると、そのフレームについてフレームロス処理を行う第4過程と、 上記フレームロスの条件を満さない場合と、上記第1過程で誤りが検出されない場合はフレームロス処理を伴なわない復号化処理を行う第5過程とを有する音声復号化方法。
Independent claims3
75 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
The present invention is applied to, for example, a PHS (Personal Hand-phone System), and relates to a decoding method for decoding voice information on a frame-by-frame basis, particularly a decoding method for compensating for deterioration of radio wave conditions.
【0002】
[Conventional technology]
In fields such as digital mobile communication, various high-efficiency coding methods are used in order to make effective use of radio waves. CELP (code-excited linear prediction), VSELP (vector addition-driven linear prediction), CS-ACELP, etc. are known as methods for coding with an amount of information of about 8 kbit / s. For each technology, MRSchroeder and BSAtal: Code-Excited Linear Prediction (CELP): High-quality Speech at Very LowRates, Proc. ICASSP '85, 25.1.1, pp.937-940,1985 (Reference 1) And IAGerson and MAJasiuk: Vector Sum Excited Linear Prediction (VSELP) Speech Coding at 8kps, proc. ICASSP'90, S9.3, pp.461-464, 1990 (Reference 2) and A. Kataoka et al: ITU-T 8-kbit / s Standard SpeechCodec for Personal Communication Services, Int. Conf. On Universal Personal It is described in Communications, pp.818-822, 1995 (Reference 3).
【0003】
In these methods, as shown in FIG. 3, a linear prediction coefficient is calculated by the filter coefficient determination unit 12 from a plurality of samples of a sample series of digitized audio signals from the input terminal, and the prediction coefficient is the filter coefficient quantization unit 13. Quantized with, and its quantization filter coefficient is set in the linear prediction synthesis filter 14. The transfer function of the composite filter 14 is A (z) Is. Pitch period components (residual signal vectors) extracted from multiple pitch period components (excitation candidates) of pitch excitation source (adaptive coding book excitation source) 15 and multiple noise waveform vectors of code book excitation source 16 (for example, random numbers). After an appropriate gain is added by the gain unit 17 to each of the candidates extracted from the vector and the excitation candidate), they are added by the adding means 18 and supplied to the synthesis filter 14 as a drive signal, and the sound is synthesized. The gain prediction unit 19 predicts an approximate gain from the past noise waveform vector and is set in the prediction gain unit 21. The difference between the synthetic voice from the synthetic filter 14 and the input voice from the input terminal 11 is taken by the subtraction means 22 and input to the distortion power calculation unit 23, and both excitations are performed so that the distortion of the synthetic voice with respect to the input voice is small. Each excitation candidate in the sources 15 and 16 is selected, and each gain of the gain unit 21 is set. The code output unit 24 outputs the prediction coefficient, the input voice power, the code number and the gain selected for each of the pitch period component candidate and the codebook candidate as codes.
【0004】
The decoder (Fig. 4) synthesizes the decoded voice based on the sent code. The received encoded voice information input to the input terminal 31 is separated into a pitch period component candidate, a codebook candidate, an input voice power, and a quantization filter coefficient, and the pitch period component candidate is taken out from the pitch excitation source 32 and has a code. The book candidate is taken out from the codebook excitation source 33, the input voice power is set in the gain unit 34, the gain prediction by the gain prediction unit 35, and the gain setting for the prediction gain unit 36 by the prediction are performed in the same manner as the encoder. , The pitch excitation candidate and the codebook excitation candidate to which the gain is given are added by the adding means 37 and supplied to the synthesis filter 38 as an excitation signal, and the separated quantization filter coefficient is decoded by the filter coefficient decoding unit 39. It is set as a filter coefficient in the composite filter 38, and the decoded voice is obtained from the composite filter 38.
【0005】
By the way, at present, PHS has a transmission capacity of 32 kbit / s, and 32 kbit / s, ADPCM is used as a voice coding method. PHS is described in "Second Generation Cordless Telephone System Standard (RCRSTD-28), Radio System Development Center" (Reference 4). However, by using ADPCM, all the transmission capacity of PHS is used for voice transmission, and other data cannot be transmitted. If the ITU-T international standard 8kbit / s audio coding method CS-ACELP (G.729) is used for this PHS, audio quality equivalent to ADPCM can be achieved at 8kbit / s, so the remaining 24kbit / s can be achieved. Can be distributed to other media such as images and data, and PHS can be extended to multimedia-compatible methods. The CS-ACELP audio coding method achieves the same quality as 32kbit / s and ADPCM at 8kbit / s. In addition, because it is designed for use in the next-generation publicly announced land mobile communication system (FPLMTS), it is resistant to transmission line errors, and interpolation processing at the time of frame disappearance is added as standard. For the interpolation process when this frame disappears, refer to ITU-T Recommendation. It is described in "G.729 Nov.1995" (Reference 5).
【0006】
The existing PHS system transmits 160-bit audio data and 80-bit control information including CRC every 5 ms frame, for a total of 240 bits. When CS-ACELP is applied to PHS, the frame of CS-ACELP is 10ms, so to minimize the delay for the audio data transmission method of 8kbit / s, 1 for 2 frames of PHS as shown in Fig. 5A. A method of transmitting data for 10 ms times and a method of transmitting data for 20 ms once every 4 frames can be considered as shown in Fig. 5B, which increases the transmission delay but simplifies the configuration.
【0007】
[Problems to be Solved by the Invention]
PHS can make a base station smaller, but if it goes out of the applicable area of the base station, the radio wave condition deteriorates rapidly. That is, the information received by the decoder will contain many errors. CS-ACELP (G.729) is more resistant to transmission line errors than ADPCM, but if the received information contains many errors, problems such as deterioration of decoded audio and abnormal noise are expected.
【0008】
An object of the present invention is to provide a decoding method for adaptively operating a decoder when a transmission line error occurs in a method for decoding on a frame-by-frame basis such as PHS to reduce quality deterioration. is there.
【0009】
[Means for solving problems]
According to the invention of claim 1, a transmission line error is detected from the received signal in the first process, an average error rate from a predetermined period before to the present is calculated in the second process, the transmission line error is detected, and a frame loss occurs. Whether or not the above condition is satisfied is determined in the third process using the above average error rate, and if it is determined that the frame loss condition is satisfied, frame loss processing is performed for that frame in the fourth process, and the frame loss condition is satisfied. If the above is not satisfied and if a transmission line error is not detected, a decoding process that does not involve frame loss processing is performed in the fifth process.
【0010】
In the invention of claim 2, the predetermined period of the second process is a period during which it can be estimated whether or not a burst error has occurred, and the condition of the frame loss is that the average error rate is the first threshold value th2 or more. is there. The invention of claim 3 has a sixth process of obtaining the error rate of the transmission line, and the condition of the frame loss includes an error of the transmission line from the second threshold value th1 of the first threshold value th2 or less. Including that the rate is high.
【0011】
In the invention of claim 4, in the invention of claim 2 or 3, the condition of the frame loss includes that the average error rate is smaller than the third threshold value th3, which is larger than the first threshold value th2. In the invention of claim 5, a predetermined bit error in the frame of the received signal is detected, and if a bit error is detected in the first process in the first process, a frame loss process is performed in the second process. If no bit error is detected in the first process, the decoding process without frame loss processing is performed in the third process.
【0012】
In the invention of claim 6, an error of a predetermined bit in the frame of the received signal is detected in the first process, the average error rate from before a predetermined period to the present is calculated in the second process, and in the first process described above. When an error is detected, it is determined in the third process whether or not the frame loss condition is satisfied using the above average error rate, and when it is determined that the above frame loss condition is satisfied, frame loss processing is performed for that frame. If the frame loss condition is not satisfied in the fourth process and if no error is detected in the first process, the decoding process without the frame loss process is performed in the fifth process.
【0013】
In the invention of claim 7, in claim 6, the condition of the frame loss is that the average error rate is the first threshold value th3 or less.
【0014】
BEST MODE FOR CARRYING OUT THE INVENTION
An example of the functional configuration of the invention of claim 1 is shown in FIG. 1A. The error detection unit 42 detects whether the received detection output data from the input terminal 41 contains an information error that occurred during transmission, and if an error is detected, the frame loss determination unit 44 processes the frame loss. A decision is made as to whether or not to do so. The decoder 45 actually decodes the audio, and includes a normal decoding processing unit 46 and a frame loss processing unit 47.
【0015】
The decoding processing method will be described below with reference to the flow chart of FIG. 1A. In the received coded voice, the presence or absence of a transmission line error is first detected by the error detection unit 42 (S1). In PHS, a 16-bit CRC (Cyclic Redundancy Check) is added to the frame every 5ms, and this error detection function can be used. If no error is detected (S2), the received coded voice is subjected to normal decoding processing (S3), and if an error is detected, the error rate is calculated using this error detection result in this example (S4). ). That is, a predetermined number of consecutive frames F<sub>C </sub>Number of frames in which an error was detected F<sub>E </sub>From error rate P<sub>E </sub>= F<sub>E </sub>/ F<sub>C </sub>Find (%).
【0016】
Next, check whether the frame loss condition is satisfied (S5). In other words, since the error in the wireless transmission line is bursty, it is better to treat the frame as not received (frame loss) rather than decoding the audio using information containing many errors, or the error is. Even if the amount is small, the quality will deteriorate if it is treated as a frame loss. Therefore, as a frame loss condition, the error rate P<sub>E </sub>If is equal to or higher than the first threshold value th2, it is judged as a burst error and processed as a frame loss (S6), and the error rate P<sub>E</sub>If is equal to or less than the first threshold value th2, it is not treated as a frame loss and is treated as a normal decoding process (S3).
【0017】
The frame loss processing is performed by interpolation processing at the time of frame loss. For example, the method described in Document 5 described above is used for this interpolation processing, but the frame-decoded audio waveform immediately before the frame loss may be used again for the frame. Alternatively, the decoding process may be performed using the coded voice before decoding immediately before the frame loss. If the frame loss process continues for a long time, the same voice waveform continues and the voice quality deteriorates. Therefore, the error rate P<sub>E </sub>When is particularly high, it is preferable not to perform the frame loss processing by presuming that the frame loss is continuous for a long time. That is, the error rate P<sub>E </sub>If is larger than the first threshold value th2 and larger than the second threshold value th3, it is also a condition that frame loss processing is not performed.
【0018】
Further, it is preferable to obtain the transmission line error rate in consideration of the state of the transmission line and add that this is larger than the third threshold value th1 to the condition of the frame loss processing. Let th1 <th2. Further, it is advisable to presume that the error is detected in the previous frame and the error is also detected in the current frame as a condition of the frame loss processing, that is, the error is burst-like.
【0019】
Error rate P above<sub>E </sub>It is conceivable that the calculation was performed for the past 50 to 200 ms, for example, 100 ms, and the transmission line error P<sub>L </sub>The error rate obtained for the past 300 ms or more can be considered. In this case, the average over a long period may be sufficient, but the longer the average, the more the past state needs to be stored, and the larger the capacity of the storage buffer becomes. The thresholds th1, th2, and th3 can be 0.1 to 0.3, 0.1 to 0.35, and 0.3 to 0.5, respectively, but always th1 <th2 <th3, for example, th1 = 0.1, th2 = 0.1, th3 = 0.3. Will be done.
【0020】
The first condition for frame loss processing P<sub>E </sub>> th2 And then P<sub>E </sub><th3 Further, one of the following is a condition of both.
【0021】
P<sub>L </sub>> th1, CRC<sub>m-1 </sub>= 1 and CRC<sub>m </sub>= 1 (CRC<sub>m </sub>= 1 means that an error was detected in frame m). Next, another embodiment of the present invention will be described. As described above, in the CS-ACELP voice coding method, various coding indexes are transmitted, and among them, if an error occurs, the voice quality is greatly affected, and the influence is relatively small. For example, the quantized index of the filter coefficient, the pitch excitation vector index, the gain index, etc. have a large effect on the voice quality when an error occurs. Further, among these indexes, for example, in the case of a quantized index of a filter coefficient, a bit error related to the reduction component has a particularly large effect on voice quality.
【0022】
Based on this relationship, an error detection bit such as a few bits of pacty is added to a specific bit that has a large effect if an error is made, and as shown in Fig. 2A, it is detected whether the specific bit has an error (S1). ), If there is no error (S2), normal decoding processing is performed (S3), and if an error is detected, frame loss processing is performed (S4). Alternatively, as shown in Fig. 2B, error detection of a specific bit is performed (S1), if there is no error (S2), normal decoding processing is performed (S3), and if there is an error, the error rate P<sub>E</sub>Is calculated (S4). The calculation of the error rate can be obtained, for example, by the number of erroneous frames that have occurred in the past or within a predetermined frame period for the determination bit. Next, it is checked whether the frame loss condition is satisfied (S5), and if it is not satisfied, the normal decoding process is performed (S3), and if it is satisfied, the frame loss process is performed (S6).
【0023】
As this frame loss condition, the error rate P<sub>E </sub>Is greater than the threshold t1, and if necessary, as in the previous embodiment, P<sub>E </sub>Can be conditioned on a threshold greater than t1 and less than t2. Similarly, the transmission line error rate P<sub>L </sub>And this is greater than or equal to t1 and greater than t0 or CRC<sub>m-1 </sub>= 1 and CRC<sub>m </sub>You may add = 1 to the condition.
【0024】
Also, in the above, the error rate P<sub>E </sub>, P<sub>L </sub>In the calculation of, the received electric field strength for each bit may be measured, the CNR (carrier power / noise power) may be obtained from the measurement result, and the error rate may be calculated from this CNR. In the example of FIG. 2B, the error rate P<sub>E </sub>, P<sub>L</sub>As described in FIG. 1B, the calculation based on the CRC inspection result for each frame may be used. Further, the coded voice is not limited to CS-ACELP, and various types described in the section of the prior art may be used.
【0025】
Next, the experimental results for the examples shown in FIG. 1 are shown. First, Fig. 6 shows the results of evaluating the quality of transmission line errors by MOS (Opinion Test). For the voice, 10 Japanese sentences (5 men and women) were used, and for the error pattern, 10 patterns with different fading frequency of 15 Hz were used. The subjects are 24 ordinary people. ADPCM and CS-ACELP are set to have the same error pattern and the same error rate. For the sake of accuracy, this error rate is expressed as the ratio of the number of erroneous bits to the total number of bits. If the measurement period is long, the error rate P<sub>E </sub>Is the same as this value. CS-ACELP is clearly better quality than ADPCM. Moreover, in CS-ACELP, when the error rate is small (0.1% or less), there is almost no quality deterioration even if the error rate increases. From this characteristic, it is understood that using CS-ACELP for PHS makes it resistant to transmission line errors.
【0026】
Fig. 7 shows the result of error processing based on the example of Fig. 1B. Since the quality deterioration is small as shown in Fig. 6 when the error rate is small, an experiment was conducted with an error rate of 0.3% or more. The experiment is a normal decryption of CS-ACELP that does not deal with errors. Improvement is seen at an error rate of 0.3 to 0.5%. On the contrary, it deteriorates as the error rate increases. As mentioned above, even if th3 = 0.3% is set, if the transmission line error rate is 0.5% or more, the condition of the transmission line is very bad, but P<sub>E </sub>Frequently occurs for a short time of 0.3% or less, and therefore, when 0.5% or more, the frequency of frame loss becomes too high, and the continuity of the same voice waveform becomes too long.
【0027】
[Effect of the invention]
As described above, according to the present invention, when decoding the coded voice in frame units, even if an error occurs depending on the state of the transmission line error, the normal decoding process is performed without performing the frame loss process. By doing so, even if the state of the transmission line is poor, the deterioration of the call quality can be reduced.
[Simple explanation of drawings]
[Figure 1]
A is a functional configuration diagram of a decoding apparatus to which the embodiment of the present invention is applied, and B is a flow chart showing a processing procedure of the embodiment of the present invention.
[Figure 2]
A is a flow chart showing a processing procedure of another embodiment of the present invention, and B is a flow chart showing a processing procedure of still another embodiment of the present invention.
[Fig. 3]
The figure which shows the functional structure of the conventional predictive encoder.
[Fig. 4]
The figure which shows the functional structure of the encoder of FIG. 3 and the corresponding decoder.
[Fig. 5]
The figure which shows the transmission method which applied 8kbit / s compressed audio code to PHS.
[Fig. 6]
The figure which shows each quality evaluation result with respect to the case of applying CS-ACELP to PHS, and the transmission line error rate with PHS using conventional ADPCM.
[Fig. 7]
The figure which shows each quality evaluation result with respect to the transmission line error rate in the case of applying this CS-ACELP to PHS, and the case of applying this invention to this.
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2002542518A | Cited by | Japan | Search report |
| JP2014206761A | Cited by | Japan | Search report |
| JP2013238894A | Cited by | Japan | Search report |
| JP2002542519A | Cited by | Japan | Search report |
| US9336783B2 | Cited by | United States of America | Applicant |
| JP2015180972A | Cited by | Japan | Search report |
| JP2015180972A | Cited by | Japan | Search report |
| US8731908B2 | Cited by | United States of America | Applicant |
| JP4966452B2 | Cited by | Japan | Search report |
| US8185386B2 | Cited by | United States of America | Applicant |
| JP2002542520A | Cited by | Japan | Search report |
| JP2012230419A | Cited by | Japan | Examiner |
| JP4966453B2 | Cited by | Japan | Search report |
| US8423358B2 | Cited by | United States of America | Applicant |
| JP4975213B2 | Cited by | Japan | Search report |
| JP2012230419A | Cited by | Japan | Search report |
| US8612241B2 | Cited by | United States of America | Applicant |
| JP2002542521A | Cited by | Japan | Search report |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 22522196 | Japan | A | |
| JP19960225221 | – | – | – |
Numbers
- Publication
- 10-69298
- Publication, DOCDB
- H1069298
- Publication, EPODOC
- JPH1069298
- Application
- 8225221
- Application, DOCDB
- 22522196
- Application, EPODOC
- JP19960225221
Titles2
- Japanese
- 【発明の名称】音声復号化方法
- English
- PROBLEM TO BE SOLVED: To decode a voice.
Classification
- IPC, 6
- G10L13 00
- G10L13 04
- G10L19 005
- G10L19 04
- G10L19 16
- H04B14 04