Digital speech signal coding/decoding method for transfer of data
Abstract
(57) A summary and the purpose When a digital sound signal is transmitted by a predetermined transmission line, for example, a telephone line, In the possible range of a real-time operation, compression encoding of the audio signal is carried out, it is sent out, and the method of the compression encoding and decryption for decrypting based on the coded signal compressed by the receiving side, and reproducing the original audio signal is offered. Composition It lets the digital sound input signal 4 introduced with the predetermined sampling rate pass to the narrow band filter 10, In quest of signal-component SF (F, T) and the time-axis maximum TMAX of this signal component (F), and the frequency-axis maximum FMAX (T), SF (F, T) is normalized to the frequency band F divided into M pieces, and the N processing time T to continue. When the audio signal in all the time value is below audible sound voice within the same frequency band as compared with the audible sound voice threshold TH (F) beforehand kept in the normalized audio signal NS (F, T), all audio signals, such as this, are not transmitted. The data volume of the audio signal which should be transmitted from this data volume that is not transmitted is reduced sharply, and it can code. Decryption performs the contrary of the above-mentioned coding and can reproduce an audio signal.
Term
No projected expiry on record.
- Priority and filed
- Published
- Today
10 claims: 3 independent, 7 dependent
- 1[Claims] 1. The following processing process in the transmission of a digital audio signal. a) A digital audio signal of a predetermined sample frequency is input to a narrow band multiplex digital filter, separated into multiple frequency bands, and the signal components of each frequency band (F) (S (F,,) at regular time intervals (T). T)), and b) Find the maximum time axis (TMAX (F)), which is the maximum absolute value of the signal component (S (F, T)) in each frequency band (F). c) Find the maximum frequency axis (FMAX (T)), which is the maximum value of the signal component (S (F, T)) within each time (T). d) Divide the signal component (S (F, T)) by the smaller of the maximum time axis value (TMAX (F)) and the maximum frequency axis value (FMAX (T)), and divide the normal signal component (NS (F, T)). )) Ask e) Compare the time axis maximum value (TMAX (F)) with the audible voice threshold value (TH (F)) in each frequency band (F), and use the former as the converted time axis maximum value (CTMAX (F)). When it is smaller than or equal to the latter, the value of 0 is used, and when the former is larger than the latter, the value of the maximum time axis (TMAX (F)) is used. f) When the maximum value of the converted time axis (CTMAX (F)) is 0, the subsequent transmission of the signal of all time (T) in the frequency band (F) is prohibited. g) The normal signal component (NS (F, T)) is converted into a quantized signal (QS (F, T)) with at least one code having a smaller number of bits. h) The frame information code is obtained by sequentially arranging the frame synchronization signal (SYNC), the maximum value on the conversion time axis (CTMAX (F)), the maximum value on the frequency axis (FMAX (T)), and the error code (CRC) as the data to be transmitted. It is formed, and all the quantization signals (QS (F, T)) on the same time axis of each frequency band F are further added as unique audio signal data (DT (F)). A method for compressing and encoding a digital audio signal, which comprises. 【特許請求の範囲】 【請求項1】 デジタル音声信号の伝送にあって下記処理過程、 a)所定標本周波数のデジタル音声信号を狭帯域多重デジタルフィルタに入力し、多重周波数帯域に分離し、一定時間間隔の順次時間(T)で各周波数帯域(F)の信号成分(S(F,T))を求め、 b)各周波数帯域(F)内で信号成分(S(F,T))の絶対値の最大値である時間軸最大値(TMAX(F))を求め、 c)各時間(T)内で信号成分(S(F,T))の最大値である周波数軸最大値(FMAX(T))を求め、 d)信号成分(S(F,T))を時間軸最大値(TMAX(F))と周波数軸最大値(FMAX(T))の小さい方で割り算して正規信号成分(NS(F,T))を求め、 e)各周波数帯域(F)で時間軸最大値(TMAX(F))を可聴音声しきい値(TH(F))と比較し、換算時間軸最大値(CTMAX(F))として、前者が後者より小さい時、あるいは等しい時、0の値を、また前者が後者より大きい時、時間軸最大値(TMAX(F))の値を使用し、 f)前記換算時間軸最大値(CTMAX(F))が0の時、その周波数帯域(F)内の全ての時間(T)の信号の以後の伝送を禁止し、 g)前記正規信号成分(NS(F,T))をビット数のより少ない少なくとも1種の符号で量子化信号(QS(F,T)に変換し、 h)送出するデータとしてフレーム同期信号(SYNC),換算時間軸最大値(CTMAX(F)),周波数軸最大値(FMAX(T))および誤り符号(CRC)を順次配列してフレーム情報符号を形成し、各周波数帯域Fの同一時間軸上の全ての量子化信号(QS(F,T))を固有な音声信号データ(DT(F))として更に付加する、 から成ることを特徴とするデジタル音声信号の圧縮符号化方法。
- 4Any of claims 1 to 3, wherein the maximum value on the conversion time axis (CTMAX (F)) and the maximum value on the frequency axis (FMAX (T)) are transmitted as logarithmically quantized values. The compression coding method described in item 1. 【請求項4】 換算時間軸最大値(CTMAX(F))と周波数軸最大値(FMAX(T))は対数量子化された値で伝送されることを特徴とする請求項1~3のいずれか1項に記載の圧縮符号化方法。
- 7The amount of data and the transmission speed of the inverse normalized voice signal (TNS (F, T)) that are not used for data transmission because the conversion time axis large value (CTMAX (F)) is 0. The claim is characterized in that a code having a large number of bits is allocated to a frequency band (F) having a large time axis maximum value (CTMAX (F)) in consideration of the data transmission amount specified by the above. Decoding method described in 6. 【請求項7】 前記逆正規化音声信号(TNS(F,T))の形成は、換算時間軸大値(CTMAX(F))が0であるため、データ伝送に供されないデータ量と伝送速度によって規定されるデータ伝送量と勘案して、時間軸最大値(CTMAX(F))の値が大きい周波数帯域(F)にビット数の大きい符号を配分して行われることを特徴とする請求項6に記載の復号化方法。
Independent claims3
113 paragraphs, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Industrial application field]
The present invention relates to a method for encoding and decoding a digital audio signal in data transmission, and more particularly to a method for compressing and encoding digital audio signal data in real time in ISDN and a method for decoding the encoded signal.
【0002】
[Conventional technology]
Coding of digital audio signals and digital recording / transmission of related audio signals have already been put into practical use. For example, regarding recording with a digital compact cassette, Kenfumi Fujimoto, "Key Points of Philips DCC System: Features and Details of Psycho-Acoustic PASC Code," AI Publishing Co., Ltd., Radio Technology Magazine, December 1991, No. 156-161 Please refer to the page. Here, Precision Adaptive Subband Coding (PASC) is used.
【0003】
In this coding scheme, the audio signal is first introduced into a bandpass filter and the signal is divided into, for example, 32 evenly spaced bands. DCC systems typically employ a bandwidth of 750 Hz because the sampling frequency is 48 kHz. Then, using 512 input data, these data are obtained in 32 bands, and the voice signal is quantized and encoded in consideration of the frequency dependence on the human audible voice signal level and voice sensitivity. There is.
【0004】
As is well known, there is a significant frequency dependence on the detection of audio signals. In other words, acoustic signals (sound pressure) with frequencies around 0 Hz and above about 15 kHz cannot be detected by the human ear. And, especially at 2 to 5kHz, the detection sensitivity of the acoustic signal is high, and paying attention to this point, PASC makes the coding of the audio signal more efficient and records the high quality audio signal without degrading the audio reception quality. It is possible.
【0005】
On the other hand, in the advanced information communication system (INS: Information Network System) as an ISDN plan already proposed by Nippon Telegraph and Telephone Public Corporation (NTT) in 1986, a data transmission speed of 64 kbps is used. Such a transmission speed naturally imposes a greater limitation on the transmission amount of the digital audio signal than the above DCC system, that is, a limitation on the amount of data transmitted per unit time.
【0006】
When a coding method such as PASC is used for a system having a limited amount of data transmission in this way, it is possible that a small signal often becomes zero due to coarse quantization. Therefore, when the sound is interrupted or a small signal is continuously flowing, when an additional pulse signal is input, the small real voice signal disappears.
【0007】
[Problems to be Solved by the Invention]
The object of the present invention is a compressed coded audio signal so that a high quality audio signal can be maintained even if the amount of data transmission such as ISDN is considerably limited, and nevertheless can be reproduced in real time. It is an object of the present invention to provide a method capable of efficiently compressing and coding data and at the same time decoding the compressed and coded audio signal.
【0008】
[Means for solving problems]
According to the present invention, the above-mentioned problems are described in the following processing process in the transmission of a digital audio signal, a) the digital audio signal of a predetermined sample frequency is input to a narrow band multiplex digital filter, separated into multiple frequency bands, and at regular time intervals. The signal component S (F, T) of each frequency band F is obtained from the sequential time T of, and b) the maximum value TMAX on the time axis, which is the maximum value of the absolute value of the signal component S (F, T) within each frequency band F. Find (F), c) Find the maximum frequency axis FMAX (T), which is the maximum value of the signal component S (F, T) within each time T, and d) Find the signal component S (F, T) on the time axis. Divide the maximum value TMAX (F) and the maximum value FMAX (T) on the frequency axis by the smaller one to obtain the normal signal component NS (F, T), and e) determine the maximum value TMAX (F) on the time axis in each frequency band F. Compared with the audible voice threshold TH (F), the converted time axis maximum value CTMAX (F) is 0 when the former is smaller or equal to the latter, and when the former is larger than the latter, the time axis. The value of the maximum value TMAX (F) is used, and f) When the conversion time axis maximum value CTMAX (F) is 0, the subsequent transmission of the signal of all time T in the frequency band F is prohibited, and g ) Convert the normal signal component NS (F, T) into a quantization signal QS (F, T) with at least one code with a smaller number of bits, and h) Frame synchronization signal SYNC as the data to be transmitted, conversion time axis. A frame information code is formed by sequentially arranging the maximum value CTMAX (F), the maximum value FMAX (T) of the frequency axis, and the error code CRC, and all the quantization signals QS (F, T) on the same time axis of each frequency band F. ) Is further added as unique audio signal data, which is solved by a compression coding method of a digital audio signal.
【0009】
Further, according to the present invention, the above-mentioned problem is the following processing process for receiving and decoding the compressed coded data signal obtained by the above processing process in the transmission of the digital audio signal, i) the received data. From the frame synchronization signal SYNC in the block, the conversion time axis maximum value CTMAX (F) and frequency axis maximum value FMAX (T) are decoded by the following data, and j) the error code in the data block is the data error from the CRC. Examine and extract the error-free data signal, k) extract the unique data attached to the audio signal from the data block, form the denormalized audio signal TNS (F, T) by decoding, l) maximum conversion time axis Multiply the inverse normalized audio signal TNS (F, T) by the smaller value of the value CTMAX (F) and the maximum frequency axis FMAX (T) to obtain the inverse audio signal (F, T), m) It is solved by a method of decoding a compressed audio signal that performs the determination of a digital output audio signal of a desired sampling frequency from the inverse audio signal TS (F, T) by a narrowband digital inverse filter.
【0010】
Other advantageous configurations according to the invention are described in the dependent claims of the claims.
【0011】
[Action]
According to the present invention, the ultra-narrow band multiplex digital filter already proposed by the present inventor is used (patent pending for this). Since this filter is capable of high-speed digital calculation, it is particularly suitable for real-time digital audio signal (PCM) transmission processing. In the case of one embodiment of the present invention, the bandwidth of each filter is 29 Hz because the filter provides 512 frequency bands between 0 and 15 kHz. In particular, in the filter used in the present invention, since the action of each band dividing filter extends only to adjacent bands, an extremely sharp-cut narrow band filter is realized. Therefore, the audio signal can be decomposed with very high resolution. Further, according to the present invention, a plurality of audio signal components having different times in the same frequency band are processed. In reality, since the number of frequency bands is large, even if the quantization of the PCM audio signal is considerably coarse, the distortion occurs only in a narrow band, so that the quality of the reproduced audio signal is not deteriorated.
【0012】
As used in the PASC described above, the signal processing of the inaudible audio signal can be omitted by taking into account the signal level curve of the audible audio threshold and the frequency dependence of the detected sound pressure. This makes it possible to effectively secure the amount of information to be transmitted in the transmission of audio signals on ISDN, which has a limited amount of transmission. In this way, the amount of signal data that characterizes the audio signal can be considerably compressed without degrading the quality of the audible audio signal. In the present invention, the compressed coded audio signal is decoded and reproduced, but the procedure is the reverse of the coding method, and an arithmetic routine similar to that at the time of coding can be used. ..
【0013】
[Example]
Hereinafter, the contents of the present invention will be described in more detail based on the preferred examples shown in the drawings.
【0014】
As shown in FIG. 1, a digital audio input signal (PCM signal) having a predetermined sampling frequency indicated by symbol 4 is introduced into the ultra-narrow band multiplexing filter 10 used in the present invention. With this filter, the signal component in a narrow band obtained by dividing the audible frequency band into M equal parts can be extracted. This frequency division processing is executed N times, and after all, MxN signal components S (F, T); 0 <F M, 0 <T N Is stored in the buffer 12. The stored signal components S (F, T) can be represented by a matrix arrangement specified by the index F of the frequency band and the index T of the time axis as shown in the figure. In this embodiment, the number of divided bandwidths M and the processing time N used are M = 512 N = 10 Is. The number of bits used in signal processing is 16 bits or more.
【0015】
Next, in order to normalize these frequency-divided signal components S (F, T) in the processing process 20, first, the maximum value of the absolute value of the signal component S (F, T) with respect to the frequency band axis and the time axis. Find FMAX (T) and TMAX (F) for each processing time T and each frequency band F. In other words FMAX (T) = MAX {S (F, T) ; F = 1 ~ M}, T = 1,2 N, TMAX (F) = MAX {S (F, T) ; T = 1 ~ N}, F = 1,2 M. [0016]
Next, with respect to the signal component S (F, T) specified by the frequency band F and the time axis T, the maximum value FMAX (T) of the signal component in the frequency band F and the maximum value of the signal component in the time axis T. The signal component S (F, T) divided by the smaller TMAX (F) is defined as the normalized signal component NS (F, T). In other words<img file="JPH07106978A_D0001.tif" />【0017】
The signal components NS (F, T) normalized in this way are obtained for the entire range of the frequency band F and the time axis T, and these are stored in the buffer 22.
【0018】
Here, the audible frequency characteristics of the audio signal are taken into consideration to prepare for data compression. To do this, first, the bit allocation determination unit 55 is stored in advance at a predetermined location 50 of the processing device, and the discrete value of the frequency characteristic of the audible voice threshold value introduced via the transfer path 74. TH (F); F = 1,2 M And the maximum value TMAX (F) in each frequency band F introduced by the transfer path represented by the symbol 70 are compared over all frequency bands F. If TMAX (F) is smaller than the discrete value TH (F) of the frequency characteristic of the audible voice threshold, it is assumed that the voice signal in this band F cannot be heard, and will be explained in more detail later. , In order to specify the processing that does not transmit the signal information thereafter over the entire processing time N of the frequency band F, this value is set to 0 as the maximum value CTMAX (F) for the new time axis. If not, use the raw value TMAX (F). In other words,<img file="JPH07106978A_D0002.tif" />Is converted to. This is shown in the first half (step S10) of the coded bit allocation determination processing routine of FIG.
【0019】
Further, the latter half (steps S11 to S15) of the coding bit allocation determination processing routine of FIG. 2 shows the preparatory processing before compressing the data to be transmitted. Here, as described above, since all the signals NS (F, T) in the frequency band F at CTMAX (F) = 0 are not transmitted, there is a margin in the amount of data actually transmitted. In addition, in the embodiment according to the present invention, the initial audio signal S (F, T) or NS (F, T) having 16 or more bits is usually compressed to 1.6 bits, but the value of CTMAX (F) is For larger ones, the number of bits of all data in the frequency band F is further increased, that is, the signal is compressed with higher resolution. In this example, CTMAX (F) with a value other than 0 is further classified into three groups. That is, the values are classified into large, medium, and small, and 4, 2.4 and 1.6 bits are assigned to the corresponding data during data compression, respectively. As an index showing this bit allocation<img file="JPH07106978A_D0003.tif" />To specify.
【0020】
Further, according to the embodiment of the present invention, the number of 4-bit compressed data and 2.4-bit compressed data is distributed 1: 2. When such a rule is applied, the data output speed to be finally transmitted, for example, 64 kbps in the case of this embodiment, and the amount of data that does not need to be transmitted due to the characteristics of the input audio signal (CTMAX described above). Number of data corresponding to ALOC (F) = O, ALOC (F) = 1 and ALOC (F) = 2 depending on the total number of bits of the signal NS (F, T) in the frequency band F at F) = 0) Can be calculated.
【0021】
In step S11, as the number of bits to spare within one processing time of buffer 22, which is determined by the transmission data output speed, as SBIT. SBIT = MxN Total number of bits given to buffer-Nx16 Is given. Here, the last term on the right side is calculated assuming that N (= 10) data on the time axis cannot be combined and all are 1.6 bits. The number obtained by adding 16 × P (P is the number of bands of CTMAX (F) = 0 obtained in step S10) of the non-transmitted data due to CTMAX (F) = 0 to SBIT is all of CTMAX (F) 0. When processing data with 1.6 bits, it is the total number of remaining bits still available. In step S11, the total number of remaining bits is divided by 40 to obtain the integer quotient Q and the remainder R (40 is the 1: 2 allocation of the 4-bit data and 2.4-bit data described above and the number of data in the band = 10). Determined from). This results in k as the number of bands F to allocate 2.4 bits in step S12.<sub>24</sub>And k as the number of bands F to allocate 4 bits<sub>40</sub>Is sought. Then, steps S13, S14 and S15 finally specify all the exponents ALOC (F) that indicate the bit allocation.
【0022】
After the coding bit allocation determination process in FIG. 2, the signal NS (F, T) normalized by the first quantization processing unit 30 (FIG. 1) is quantized to the coarse number of bits described above. .. This is done in the procedure shown in Figure 3. The index ALOC (F) that indicates the bit distribution of each frequency band F introduced from the bit allocation determination processing unit 55 via the transfer path 76 is determined in step S20, and a coefficient is determined according to the value of the index ALOC (F). Specify the PPX value and requantize the signal NS (F, T) in buffer 22 in step S21. Here, Int (X) means the maximum integer value that does not exceed X, >> means that it shifts to the right by 1 bit, divides by 2, and truncates the remainder. The final process is done to make the positive and negative bit distribution of the signal symmetric with respect to the 0 signal level (see Figure 4). The signal QS (F, T) quantized in this way is stored in the buffer 32.
【0023】
In the second quantization processing unit 60 (Fig. 1), the maximum value CTMAX (F) in each frequency band and the maximum value FMAX (T) in each time axis are introduced via transfer paths 78 and 72, respectively. Here it is 6-bit logarithmic quantization (2 dB step). Let these logarithmically quantized values be QTMAX (F) and QFMAX (T), respectively.
【0024】
Finally, in order to make the output signal 8 that sends the quantized signals QS (F, T), QTMAX (F) and QFMAX (T) to the output transmission line, this is done by the data coding processing unit 40 (Fig. 1). Etc. are encoded into a frame signal (frame information signal and compressed audio signal) that is blocked for each processing time.
【0025】
As shown in FIG. 5, first, the synchronization signal SYNC (8 bits) is added to the beginning of the frame by step S30. Next, QTMAX (F) is sent out in step S31, and QFMAX (T) is sent out in step S32. Further, a cyclic code CRC (7 bits) is added to complete the code of the frame information.
【0026】
Next, the compressed audio signal is postfixed to the free format section following the frame information by the data coding process shown in FIG. In this case, in addition to the quantized audio signal QS (F, T), the data coding processing unit 40 also introduces the bit distribution index ALOC (F) via the transfer path 77. This index ALOC (F) is first determined in step S40. In the case of 1.6 bits (at the time of ALOC (0)), as shown in the figure in step S41, the data from 1 to 5 on the time axis is set as one ternary display value SD0, and the data from 6 to 10 on the time axis is already used. As one ternary display value SD1, each is set to 8 bits in step S44 and sent out. In the case of 2.4 bits (at the time of ALOC (1)), as shown in the figure in step S42, the data from 1 to 5 on the time axis is set as one quintic display value SD0, and from 6 to 10 on the time axis. The data of is set to another quintic display value SD0, and is sent out as 12 bits in step S45. Alternatively, in the case of 4 bits (at the time of ALOC (2)), 7 is added to the data at each time in the corresponding frequency band F in step S43 to make 4 bits and send out.
【0027】
In the signal coding process described above, operations were performed over a large number of bands, and in particular, the maximum value FMAX (T) of the frequency axis was investigated in the range of M = 512. However, since voice has audible characteristics that differ greatly depending on the frequency, if this maximum value is also obtained within a frequency range further divided into several, the quality of the transmitted voice can be expressed more faithfully. When M = 512 was used in this example, the quality of the experimentally reproduced audio was the best when divided into eight. It is necessary to execute such a division process for all the processes of the normalization processing unit 20, the quantization processing unit 30, and the data coding processing unit 40.
【0028】
Figure 7a shows the compressed data when M = 512, N = 10 and the number of frequency band divisions = 8 created in this way. Note that the CTMAX (F) is 0 at F = 3, so the DT (3) of the frequency band data DT (F) is not sent. In addition, since the frequency band is divided into 8, the maximum value of the frequency axis is expressed as F (FB, T). In this case, the division index FB = 1,2 ... 8. Further, each data DT (F) in the free format section is formed as a set of all data on the same time axis. The data arrangement of 1.6, 2, 4 and 4 bits in this data DT (F) is shown in Fig. 7b.
【0029】
Next, the digital audio signal having the frame signal arrangement shown in FIG. 7, which is introduced via a certain transmission circuit or detected by some digital signal reader, is decoded and converted into the original audio signal. The processing method will be described.
【0030】
In FIG. 8, the digital audio signal 9 obtained by the compression coding method described with reference to FIGS. 1 to 7 above is introduced into the decoding unit 41 for the frame information on the receiving device side. This receiver may be an ISDN-compatible modem, or may be a waveform-shaped detection signal obtained from the read head of a digital compact cassette and shaped into a digital audio signal having the frame signal arrangement shown in FIG. Good. This received signal is returned to TS (F, T) to a value close to the original audio signal, further subjected to an inverse filter by the narrowband multiplexing filter 11, and finally the digital audio signal 5 is obtained. The details will be described below.
【0031】
The decoding unit 41 first receives the time axis maximum value QTMAX (F) and the frequency axis maximum value QFMAX (T) (step S60 in FIG. 9), and detects the cyclic code CRC as shown in step S61 in FIG. Then, a CRC code is created from the received frame synchronization signals, TMAX (F), and FMAX (T), and it is determined as CRCR whether or not the transmitted CRC code matches the CRCR. If they match, the process proceeds to the next process as it is, and if they do not match, the calculation using the data of that frame is practically not performed. In this way, the logarithmically quantized reception time axis maximum value QTMAX (F) and frequency axis maximum value QFMAX (T) are transferred to the inverse quantization processing unit 61 via transfer paths 86 and 85, respectively, where both values are transferred. Is converted to the inverse logarithm (exponential function), and the maximum value CTMAX (F) on the time axis and the maximum value FMAX (T) on the frequency axis are determined.
【0032】
Next, the bit allocation determination processing unit 56 sets the frequency band F (band when TMAX (F) = 0) of the data section that does not require transmission of audio data and the total number P of such frequency bands in the step of FIG. Calculate with S50. Then, from this total number P, the remaining number of bits used for transmission by the bit allocation determination processing unit 56 is calculated based on exactly the same calculation method as at the time of coding described in FIG. 2, that is, steps S51 and S52 in FIG. Obtained by S53, S54 and S55. In this case, as explained in the section of compression coding, the data in the frequency band with the same bit distribution, that is, the normal signal strength is 1.6 bits, and the data in the frequency band F with stronger signal strength and the data in the frequency band F with stronger signal strength are 1 Allocate to 2 bits with a ratio of: 2 and 2.4 bits. Determine the index ALOC (F), which indicates the bit allocation for their corresponding groups, and the number of frequency bands for each bit.
【0033】
Next, in the data decoding processing unit 43, as shown in FIG. 11, the data DT (F) on the time axis of the data free format section is based on the bit distribution instruction index ALOC (F) sent from the bit distribution determination processing unit 56. ) To obtain the quantized compressed coded audio signal QS (FT). This process corresponds to the inverse conversion of the process of FIG. in this case, HDATA (0, J) = 3<sup>4-J</sup>HDATA (1, J) = 5<sup>4-J</sup>Is.
【0034】
Further, as shown in FIG. 12, from the quantized compressed coded audio signal component QS (F, T) in the frame, the audio signal TNS (normalized) is processed in the reverse order of step S12 in FIG. Find F, T). Put this in buffer 23.
【0035】
Next, in the normalization processing unit 21, the time axis maximum value CTMAX (F) and the frequency axis maximum value FMAX (T) obtained by the inverse quantization processing unit 61 and introduced via the transfer paths 89 and 90 are used. The normalized audio signal TNS (F, T) is converted into the initial audio signal TS (F, T) and stored in the buffer 13. The transformation in this case is the reverse of the normalization in Figure 1. That is,<img file="JPH07106978A_D0004.tif" />Is. The normalized audio signal TNS (F, T) and the initial audio signal TS (F, T) referred to here are the normalized audio signal NS (F, T) and the initial audio signal S (F, T) shown in FIG. Substantially different from T). This is because, in the quantization process 30 of FIG. 1, the intensity of the audio signal component is converted with a resolution of a considerably coarse bit. Since the roughness of this signal level is determined in consideration of human audibility, when it is actually converted into voice, it can still be heard as high quality voice.
【0036】
Finally, the block of the matrix-shaped audio signal TS (F, T) of the buffer 13 can be taken out as the digital audio signal (PCM) represented by the symbol 5 by passing through the narrow-band multi-band inverse filter 11 according to the present invention. This voice signal can be stored in a storage medium by a predetermined storage device, or can be listened to as voice with the help of a predetermined voice conversion device (reproduction device).
【0037】
When the stereo left and right audio signals are sent to the 2B channels by ISDN, the skew of the 2B channel is not compensated, so the skew shift is absorbed by shifting the transmission time by about half the cycle of the frame in advance.
【0038】
The method for decoding the compressed digital audio signal according to the present invention has been described mainly in the case of ISDN. However, the present invention can be used not only for ISDN but also for reproduction on digital compact cassettes, magnetic tapes, and the like. In these cases, since the amount of data per unit time has more margin than in the case of ISDN, there is also a signal coding compression method that can maintain high sound quality by further increasing the bit allocation and maintaining high sound quality by fine steps, and a decoding method for it. It is possible.
【0039】
Furthermore, the present invention is not limited to the values of the parameters M = 512, N = 10, FB = 8 used in the above examples, and the type of ALOC (F) = 3, and is not limited to 2.4 bits and 4 bits. The ratio is not limited to 2: 1. These parameters can be appropriately selected, changed and used as necessary. Two of the three bits are evaluated by the absolute value of the CTMAX (F) value, rather than being determined in the order (that is, relative) of the CTMAX (F) magnitude of each frequency band as described above. You can also do it. As is well known, the characteristic curve TH (FB) of the audible speech threshold at the division frequency FB shows considerably different characteristics depending on the frequency range. In that case, the absolute value can be determined by appropriately weighting the characteristic curve TH (FB) of the audible voice threshold value at the division frequency FB. Such operations should be determined experimentally.
【0040】
[Effect of the invention]
As described above, according to the present invention, the method of compressing and encoding digital audio data and decoding after receiving it via a transmission circuit has a sufficiently high audio signal with a small amount of data transmission due to the following items. Can be maintained in quality. That is, (1) By dividing the signal into a large number of bands with a narrowband digital filter, even if the signal in each band is quantized very coarsely, the quantization distortion occurs only in the narrow band, so that it is audible. There is little deterioration in sound quality. (2) Since the signal before coarse quantization is normalized by the smaller maximum value of the time axis and frequency axis, it is normalized as a whole rather than normalized by only one of the maximum values. The later signal level rises, the effect of coarse quantization is reduced, and the signal level after normalization does not become extremely small even for sudden changes in the time axis and frequency axis, resulting in coarse quantization. The sound is no longer lost even if you do. (3) Since the indicated value of the coded bit allocation can be reproduced by calculation from the maximum value on the time axis on the receiving side, there is no need to transmit it. This measure makes it possible to significantly increase the amount of transmitted data.
[Simple explanation of drawings]
[Figure 1]
It is a processing figure which shows typically each process of the data compression coding processing by this invention and the signal attached to it.
[Figure 2]
It is a flowchart of a subroutine for determining a coding bit allocation.
[Fig. 3]
It is a flowchart of a subroutine that performs quantization.
[Fig. 4]
It is a figure which shows the bit arrangement of the signal which coarsened the resolution by quantization.
[Fig. 5]
It is a flowchart of a subroutine that performs data coding in order to process a data frame and give data information.
[Fig. 6]
It is a subroutine that encodes compressed voice data.
[Fig. 7]
The arrangement of the transmitted data in the frame (a) and the arrangement of the compressed audio signal in one data block (b) are shown.
[Fig. 8]
It is a processing diagram which shows typically each process of decoding the compressed coded voice data by this invention, and the signal attached |
[Fig. 9]
It is a schematic flowchart of a code error inspection at the time of reception and a reception routine of the maximum value of a frequency axis.
[Fig. 10]
It is a flowchart of a bit allocation determination routine.
[Fig. 11]
It is a flowchart of a data decoding routine.
[Fig. 12]
It is a flowchart of the routine which calculates the normalized voice signal component.
[Explanation of symbols]
4 Input audio signal (sampled PCM signal) 5 Output audio signal 8 Compressed output voice-coded signal 9 Compressed received audio signal 10 Narrowband multiplex filter 11 Narrowband demultiplexing filter 20 Normalization processing unit 21 Inverse normalization processing unit 30 Quantization processing unit I 40 Data coding processing unit 41 Maximum value decoding unit 43 Data Decryption Unit 50 A storage unit for numerical values of audible voice threshold characteristic curves 55-bit allocation decision unit 56-bit allocation decision unit 60 Quantization processing unit II 61 Inverse quantization processing unit
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8036162B2 | Cited by | United States of America | Applicant |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 27587393 | Japan | A | |
| JP19930275873 | – | – | – |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesR250 | R250 |
Numbers
- Publication
- 7-106978
- Publication, DOCDB
- H07106978
- Publication, EPODOC
- JPH07106978
- Application
- 5275873
- Application, DOCDB
- 27587393
- Application, EPODOC
- JP19930275873
Titles3
- English
- DIGITAL SPEECH SIGNAL CODING/DECODING METHOD FOR TRANSFER OF DATA
- Japanese
- 【発明の名称】データ伝送におけるデジタル音声信号の符号化と復号化方法
- English
- PROBLEM TO BE SOLVED: To encode and decode a digital audio signal in data transmission.
Classification
- IPC, 2
- H03M7 30
- H04B1 66