Speech coding/decoding method.
3 claims: 2 independent, 1 dependent
- 1Sprachkodierungsverfahren mit folgenden Schritten:Gewinnung eines eine Spektrumeinhüllende repräsentierenden Spektrumparameters und eines eine Tonhöhe repräsentierenden Tonhöhenparameters aus einem diskreten Eingabesprachsignal;Unterteilung eines Rahmenintervalls in Subintervalle in Übereinstimmung mit dem Tonhöhenparameter;Gewinnung eines Lautquellesignals in einem der Subintervalle;Gewinnung und Ausgabe von Korrekturinformation zum Korrigieren mindestens einer Amplitude und einer Phase des Lautquellesignals in anderen Subintervallen im Rahmen;gekennzeichnet dadurch, daß der Schritt zur Gewinnung des Lautquellesignals aufweist: (a) Gewinnung eines Differenzsignals durch Durchführung einer Tonhöhenprädiktion auf der Grundlage eines vorherigen Lautquellesignals;(b) Gewinnung eines Mehrfachpulses in bezug auf das Differenzsignal;und (c) Addition des Mehrfachpulses zum Tonhöhenprädiktionssignal.
- 2Sprachkodierungsverfahren mit folgenden Schritten:Gewinnung eines eine Spektrumeinhüllende repräsentierenden Spektrumparameters und eines eine Tonhöhe repräsentierenden Tonhöhenparameters aus einem diskreten Eingabesprachsignal;Unterteilung eines Rahmenintervalls in Subintervalle in Übereinstimmung mit dem Tonhöhenparameter;Gewinnung eines Lautquellesignals in einem der Subintervalle;Gewinnung und Ausgabe von Korrekturinformation zum Korrigieren mindestens einer Amplitude und einer Phase des Lautquellesignals in anderen Subintervallen im Rahmen;gekennzeichnet dadurch, daß der Schritt zur Gewinnung des Lautquellesignals aufweist: (a) Gewinnung eines Differenzsignals durch Durchführung einer Tonhöhenprädiktion auf der Grundlage eines vorherigen Lautquellesignals;(b) Auswahl eines Vektors des Lautguellesignals in bezug auf das Differenzsignal aus einem Kodeverzeichnis, in dem Lautquellesignalvektoren gespeichert sind;und (c) Addieren des ausgewählten Vektors zum Tonhöhenprädiktionssignal.
- 3Vorrichtung zum Ausführen eines Sprachkodierungssystems nach Anspruch 1 oder 2.
Independent claims3
109 paragraphs, as filed
The present invention relates to a method for speech coding and decoding for coding a speech signal with high quality at a low bit rate, in particular at 4.8 kb / s or less, by a relatively small amount of processing.
As a method for encoding a speech signal at a low bit rate of about 4.8 kb / s or less, speech encoding methods are known which are described, for example, in JP-A-58100/90 (reference 1) and in M. Schroeder and B. Atal, " Code-excited linear prediction: High quality speech at very low bit rates, "ICASSP, pp. 937-940, 1985 (reference 2).
According to the method in reference 1, a spectrum parameter, which represents the spectrum characteristic of a speech signal, and a pitch parameter, which represents its pitch, are extracted from a speech signal of each frame on the transmitter side. Speech signals are classified into several signal types (eg vowel, explosive, friction sound) using acoustic features. A one-frame sound source signal in a vowel volume interval is represented by improved pitch interpolation in the following manner. A signal component in a pitch interval (representative interval) of a plurality of pitch intervals obtained by dividing a frame is represented by a multiple pulse. In other pitch intervals in the same frame, amplitude and phase correction coefficients for correcting the amplitude and phase of the multiple pulse in the representative interval are obtained in units of the pitch interval. The amplitude and position of the multiple pulse in the representative interval, the amplitude and phase correction coefficients in other pitch intervals and the spectrum and pitch parameters are then transmitted. A multiple pulse is obtained in the entire frame in an explosive sound signal. In a friction noise interval, a type of noise signal is selected from a code list consisting of predetermined types of noise signals so as to minimize the difference in strength between a signal obtained by synthesizing a noise signal and the input speech signal, and an optimal gain is calculated. As a result, an index representing the type of the noise signal and the gain are transmitted. A description in connection with the reception side is omitted.
In the conventional method disclosed in Reference 1, with respect to a female speaker with a short pitch period, since there are a large number of pitch intervals in one frame, improved pitch interpolation can be effectively carried out, and accordingly, a sufficient number of pulses can be used for the whole frame can be obtained. For example, if the frame length is 20 ms, the pitch period is 4 ms, and the number of pulses in a typical interval is 4, 20 pulses can be obtained for the entire frame.
However, since a sufficient number of pulses for the entire frame cannot be obtained appropriately with respect to a male speaker with a long pitch period, an improved pitch interpolation is not satisfactory. Therefore there will be a problem with the sound quality. For example, if the pitch period is 10 ms and the number of pulses per pitch is 4, the number of pulses in the entire frame is 8, which is very small compared to the female speaker. To increase the number of pulses in the entire frame, the number of pulses per pitch must be increased. However, if this number is increased, the bit rate is increased. For this reason, it is difficult to increase the number of pulses.
In addition, if the bit rate is reduced from 4.8 kb / s to 3 kb / s or 2.4 kb / s, the number of pulses per pitch must be reduced to 2 to 3. Therefore, a more difficult problem will arise than that described above. At such a low bit rate, the performance of the improved pitch interpolation is insufficient, even for a female speaker.
In the CELP method disclosed in Reference 2, if the bit rate is reduced below 4.8 kb / s, the number of bits in a code dictionary must be reduced, resulting in an abrupt deterioration in sound quality. For example, at 4.8 kb / s, a 10-bit code directory is generally used for a subframe of 5 ms. However, at 2.4 kb / s, the number of bits in the code dictionary must be reduced to 5, provided that the period of the subframe is maintained at 5 ms. Since 5 bits as the number of bits are too few to detect different types of sound source signals, the sound quality deteriorates abruptly at a bit rate of less than 4.8 kb / s.
In addition to the method according to reference 1, the IEEE / IEICE GLOBAL TELECOMMUNICATION CONFERENCE, Tokyo, 15-18. Nov. 1987, Vol. 2, pages 752-756, IEEE, New York, US; S.Ono et al.:"2.4 kBPs pitch interpolation multi-pulse speech coding "a pitch interpolation method.
It is an object of the present invention to provide a speech coding and decoding method which performs high quality speech coding and decoding at 4.8 kb / s or less with a relatively small amount of processing. This object is achieved with the features of the claims.
A speech coding method as described has the steps:
Obtaining a spectrum parameter representing a spectrum envelope and a pitch parameter representing a pitch from a discrete input speech signal, dividing a frame interval into subintervals in accordance with the pitch parameter, obtaining a sound source signal in one of the subintervals by obtaining a multiple pulse with respect to a difference signal . which is obtained by making a prediction on the basis of a previous sound source signal, and obtaining and outputting correction data for correcting at least one amplitude and a phase of the sound source signal at other pitch intervals in the frame.
A sequence of processing steps based on the speech encoding and decoding method of the present invention will be described below.
In a voiced interval with periodic properties for each pitch, a pitch parameter representing a pitch period is obtained in advance from a speech signal in the frame. For example, the frame interval of a speech wave shown in Fig. 3 (a) is divided into a plurality of pitch intervals (subframes) in units of pitch periods as shown in Fig. 3 (b). A multiple pulse with a predetermined number of pulses is obtained with respect to a difference signal obtained by making a prediction in a pitch interval (representative interval) of the pitch intervals using a previous sound source signal. Gain and phase correction coefficients for correcting the gain and phase of the multiple pulse in the representative interval are then obtained for other subframes in the same frame.
A method of performing pitch prediction is described below. Assume that a drive sound source signal reproduced in the previous frame is described by v (n), and a prediction coefficient and a period are described by b and M, respectively. In addition, assume that an interval in Fig. 3 (c) is a representative interval of the current frame, and a voice signal in this interval by x & sub1; (n) is described. The coefficient b and period M are calculated to minimize the difference in the following equation:
E = [{x 1 (n) - bv (nM) * h (n)} * w (n)) 2 ... (1)
where w (n) is the impulse response of a perceptual weight filter (for the detailed description thereof, refer to Japanese Patent Application No. 57-231605, disclosed as Patent Application Laid-Open No. 59-116794 (reference 3) and the like), h (n) the impulse response of a synthesizing filter, formed from a spectrum parameter which is derived from the language of the current frame by known linear prediction analysis (LPC) (for a detailed description of which refer to reference 3 and the like) is obtained, and * is the folding operation.
In order to minimize equation (1), equation (1) is partially differentiated according to b and set to 0 in order to obtain the following equation:
Substituting equation (2) in equation (1) gives:
Since the first term of equation (4) is constant, equation (1) can be minimized by maximizing the second term of equation (4). The second term of equation (4) is calculated for different values of M, and the value of M that maximizes the second term is obtained. The value of b is then calculated from equation (2).
Pitch prediction is performed with respect to the interval using the obtained values b and M according to the following equation, so as to obtain a difference signal e (n):
e (n) = x 1 (n) - v (nM) * h (n) ... (5)
3 (c) shows an example of e (n).
A multiple pulse with a predetermined number of pulses is then obtained with respect to the difference signal e (n). As a practical method for obtaining a multiple pulse, a method using a cross correlation function Φxh and an autocorrelation function Rhh is known. Since this method is used, for example, in Reference 3 and in Araseki, Ozawa, Ono and Ochiai, "Multi-pulse Excited Speech Coder Based on Maximum Cross-Correlation Search A logarithm", GLOBECOM 83, IEEE Global Tele-communications Conference, contract number 23.3.1983 ( Reference 4) is disclosed, a description of this method is omitted. Fig. 3 (d) shows the multiple pulse obtained in the interval as an example in which two pulses are obtained.
As a result, a sound source signal d (n) is obtained in the interval according to the following equation:
d (n) = bv (nM) + gi.δ (n-mi) ... (6)
for δ (n-mj) =
where gi and mi are the amplitude and position of an i-th pulse of the multiple pulse.
In pitch intervals other than the representative interval, the gain and phase correction coefficients for correcting the gain and phase of the sound source signal in the representative interval are calculated in units of pitch intervals. If a gain correction coefficient or a phase correction coefficient in the last pitch interval is referred to as cj or dj, these values can be calculated to minimize the following equation:
E = [{xj (n) -cj d (n-T'-d 3) * h (n)} * w (n)] 2 ... (7)
Since the solution of the above equation is described in detail in Reference 3 and the like, its description is omitted. A sound source signal of the frame is obtained by obtaining gain and phase correction coefficients in different pitch intervals than the representative pitch interval according to equation (7).
Fig. 3 (e) shows, by way of example, the drive sound source signal of the current frame, which is reproduced by extracting the gain and phase correction coefficients at pitch intervals other than the interval.
In this case, a representative interval is fixed to the pitch interval. However, a pitch interval in which the volume difference between the input speech of a frame and the synthesized speech is minimized can be selected as a representative interval by checking a plurality of pitch intervals in the frame. Reference is made to reference 1 and the like for a detailed description of this method.
Information to be transmitted as sound source information for each frame includes the location of a representative pitch interval in a frame (not required if a representative interval is set); the prediction coefficient b, the period M, the amplitude and position of the multiple pulse in the representative interval; and the gain and phase correction coefficients in other pitch intervals in the same frame.
According to the second aspect of the present invention, instead of obtaining a multiple pulse with respect to a difference signal e (n) obtained by making a prediction in a representative interval, vector quantization is performed using a code dictionary. This method is described in detail below. It is assumed that 2B (B is the number of bits of a sound source) types of sound source signal vectors (code vectors) are stored in the code dictionary. If a sound source signal vector in the code dictionary is described by c (n), the sound source signal vector is selected from the code dictionary so that the following equation is minimized:
E = [{e (n) -gc (n) * h (n)} * w (n)] ² ... (8)
where is the amplification of the sound source signal. In order to minimize equation (8), equation (8) is partially differentiated and set to 0 in order to obtain the following equation:
in which
ew (n) = e (n) * h (n) ... (10)
w (n) = c (n) * h (n) * w (n) ... (11)
Substituting equation (9) into equation (8) gives:
Since the first term of equation (12) is constant, the second term is calculated for all values of the sound source vector c (n) and a value that maximizes the second term is selected. In this case the gain is obtained according to equation (9).
The code directory can be formed by learning on the basis of exercise signals, or can be formed, for example, from Gaussian random signals. The former method is described, for example, in Makhoul et al., "Vector Quantization in Speech Coding," Proc. IEEE, vol. 73, 11, 1551-1588, 1985 (reference 5). The latter method is described in Reference 2.
1 is a block diagram showing a system based on a speech coding and decoding method according to the first embodiment of the present invention;
2 is a block diagram showing a system based on a speech coding and decoding method according to the second embodiment of the present invention; and
3 (a) to 3 (e) are graphs for explaining a sequence of processing steps based on the method of the present invention.
1 shows a system for executing a method for speech coding and decoding according to the first embodiment of the present invention.
With reference to FIG. 1, a transmitter side receives a speech signal via an input connection 100 and stores a one-frame speech signal (eg 20 ms) in a buffer memory 110.
An LPC and pitch calculator 130 performs a known LPC analysis of the one-frame speech signal to calculate a K parameter corresponding to a predetermined degree P as a parameter representing the spectrum characteristics of the one-frame speech signal. With regard to a detailed description of this method for calculating the K parameter, reference is made to K parameter computers in the references 1 and 3 described above. It should be noted that a K parameter is identical to a PARCOR coefficient. A code 1k, which is obtained by quantizing the K parameter with a predetermined number of quantization bits, is output to a multiplexer 260 and is decoded into a linear prediction coefficient ai '(i = 1 to P). The coefficient ai 'is then output to a weighting circuit 200, an impulse response calculator 170 and a synthesizing filter 281. With regard to methods for coding the K parameter and for converting the K parameter into the linear prediction coefficients, reference is made to references 1 and 3 described above. An average pitch period T is calculated from the one-frame speech signal. A method based on an autocorrelation is known for this method. For a detailed description of this method, reference is made to a pitch extraction circuit in reference 1. In addition, other known methods (eg the cepstrum method, the SIFT method and the partial correlation method) can be used. A code obtained by quantizing the averaged pitch period T with a predetermined number of bits is output to the multiplexer 260. In addition, a decoded pitch period obtained by decoding this code is output to a subframe divider 195, a circuit 283 for reproducing the drive sound source, and a gain / phase correction calculator 270.
The impulse response calculator 170 calculates an impulse response hw (n) of the synthesizing filter that performs the perceptual weighting using the linear prediction coefficient ai ', and outputs it to an auto-correlation calculator 180 and a cross-correlation calculator 210.
The autocorrelation computer 180 calculates an autocorrelation function Rhh (n) of the impulse response and outputs it with a predetermined time delay. With regard to the operations of the impulse response computer 170 and the autocorrelation computer 180, reference is made to references 1 and 3.
A subtractor 190 subtracts a one-frame component of an output signal from the synthesizing filter 281 from a one-frame speech signal x (n) and outputs the subtraction result to the weighting circuit 200.
The weighting circuit 200 obtains and outputs a weighted signal xw (n) by filtering the subtraction result by a perceptual weight filter whose impulse response is described by w (n). With regard to the weighting method, reference is made to references 1 and 3 and the like.
Subframe divider 195 divides the weighted signal of the frame at pitch intervals of T '.
A prediction coefficient calculator 206 obtains a prediction coefficient b and a period M according to equations (1) to (4) using a previously reproduced drive sound source signal v (n), the impulse response hw (n) and one of the signals weighted at the pitch intervals of T 'in one predetermined representative interval (e.g. an interval in Fig. 3 (c)). The values obtained are then quantized with a predetermined number of bits in order to obtain values b 'and M'. The prediction coefficient calculator 206 further calculates a prediction sound source signal v '(n) according to the following equation and outputs it to a prediction circuit 205:
v '(n) = b' v (n-M ') ... (13)
The prediction circuit 205 predicts using the signal v '(n) according to the following equation to obtain a difference signal in the representative interval (the interval in Fig. 3 (c)):
ev (n) = xw (n) -v '(n) * hw (n) ... (14)
The cross-correlation function calculator 210 receives the values ev (n) and hv (n), calculates a cross-correlation function Φxh with a delay time and outputs the calculation result. With regard to this calculation method, reference is made to references 1 and 3 and the like.
A multiple pulse calculator 220 calculates a position mi and an amplitude gi of a multiple pulse with respect to the difference signal in the representative interval obtained from equation (14) using the cross-correlation function and the autocorrelation function.
A pulse encoder 225 encodes the amplitude gi and the position mi of the multiple pulse in the representative interval with a predetermined number of bits and outputs them to the multiplexer 260. At the same time, the pulse encoder 225 decodes the encoded multiple pulse and outputs it to the adder 235.
The adder 235 adds the decoded multiple pulse to the prediction sound source signal v '(n) output from the prediction coefficient calculator 206, so as to obtain a sound source signal d (n) in the representative interval.
As described in the summary, the gain / phase correction calculator 270 calculates and outputs a gain correction coefficient ck and a phase correction coefficient dk of the sound source d (n) in the representative interval, so as to reproduce a sound source signal in another pitch interval k in the same frame. Reference is made to reference 1 for a detailed description of this method.
An encoder 230 encodes the gain correction coefficient ck and the phase correction coefficient dk with a predetermined number of bits and outputs them to the multiplexer 260. In addition, encoder 230 decodes them and outputs the decoded values to circuit 283 for reproducing the drive sound source.
The drive sound source reproducing circuit 283 divides the frames according to averaged pitch periods T 'in the same manner as the subframe divider 195 and generates the sound source signal d (n) at a representative interval. Using the sound source signal and the decoded gain and phase correction coefficients in the representative interval, circuit 283 reproduces a drive sound source signal v (n) of the entire frame at pitch intervals other than the representative interval according to the following equation:
v (n) = Ck d (n-T'-dk) + d (n) ... (15)
The synthesizing filter 281 receives the reproduced drive sound source signal v (n) and the linear prediction coefficient ai 'and obtains a composite one-frame speech signal. The filter 281 also calculates a one frame influence signal that affects the next frame and outputs it to the subtractor 190. With regard to the method for calculating the influence signal, reference is made to reference 3.
The multiplexer 260 couples and provides the codes for the prediction coefficient, for the period, for the amplitude and for the position of the multipulse in the representative interval, the codes for the gain and phase correction coefficients and for the average pitch period and the code for the K- Parameters.
The above description is related to the transmitter side according to the first embodiment of the present invention.
On the decoding side, a demultiplexer 290 receives the coupled codes via a connection 285 and separates the code for the multiple pulse, the codes for the gain and phase correction coefficients, the codes for the prediction and for the period, the code for the average pitch period and the code for the K parameter from each other and outputs them.
A K-parameter / pitch decoder 330 decodes the codes for the K-parameter and the pitch period and outputs the decoded pitch period T 'to a circuit 340 for reproducing the drive sound source.
A pulse decoder 300 decodes the code for the multiple pulse, generates a multiple pulse at a representative interval, and outputs it to an adder 335.
The adder 335 adds the multiple pulse from the pulse decoder 300 to a prediction sound source signal v '(n) from a prediction circuit 345 so as to obtain a sound source signal d (n).
A gain / phase correction coefficient decoder receives the codes for the gain and phase correction coefficients, decodes them and outputs the obtained values.
A coefficient decoder 325 decodes and outputs the codes for the prediction coefficient and for the period to obtain a coefficient b 'and a period M'.
The prediction circuit 345 calculates a prediction sound source signal v '(n) from the drive sound source signal v (n) of the previous frame using the values b' and M 'in accordance with equation (13) and outputs it to the adder 335.
The drive sound source reproduction circuit 340 receives the output from the adder 335, the decoded pitch period T ', the decoded gain correction coefficient, and the decoded phase correction coefficient. Subsequently, the circuit 340 reproduces and outputs the one-frame drive sound source signal v (n) by the same operation as that performed by the transmission sound source reproducing circuit 283 on the transmitter side.
A synthesizing filter 350 receives the reproduced one-frame drive sound source signal and the linear prediction coefficient ai ', calculates a synthesized one-frame speech x (n) and outputs it via a connector 360.
The above description is related to the receiving side according to the first embodiment of the present invention.
Fig. 2 shows the second embodiment of the present invention. The same reference numerals in FIG. 2 denote the same parts as in FIG. 1, and the description thereof is omitted.
In this embodiment, an optimal code vector is selected from a code dictionary 520 with respect to a prediction difference signal calculated according to equations (1) to (4) and (14), and a gain g of the code vector is calculated. In this case, a code vector c (n) is selected and the gain g is calculated with respect to a value ew (n) obtained from equation (14) so as to minimize equation (8). It is assumed that the number of dimensions of a code vector of the code dictionary is given by L and the type of the code vector is 2B. In addition, it is assumed that the code directory consists of Gaussian random signals as in reference 2.
A cross correlation calculator 505 calculates a cross correlation function Φ and an autocorrelation function R according to the following equations:
Φ = ew (n) w (n) ... (16)
R = w (n) w (n) ... (17)
where ew (n) and w (n) are calculated according to equations (10) and (11). Equation (16) or (17) also corresponds to the numerator or denominator of equation (9). Calculations based on equations (16) and (17) are performed for all code vectors, and the values of Φ and R of each code vector are output to a code dictionary selector 500.
The code dictionary selector 500 selects a code vector that maximizes the second term of equation (12). The second term of equation (12) can be rewritten as follows:
D = Φ2 / R ... (18)
Therefore, a code vector is chosen that maximizes equation (18). The gain g of the selected code vector can be calculated using the following equation:
g = Φ / R ... (19)
The code dictionary selector 500 outputs the data for the index of the selected code dictionary to a multiplexer and outputs the calculated gain g to a gain encoder 510.
The gain encoder 510 quantizes the gain by a predetermined number of bits and outputs the code to the multiplexer 260. Using a decoded value g ', the encoder 510 simultaneously obtains a sound source signal z (n) based on the selected code directory and outputs it to an adder 525 according to the following equation:
z (n) = g 'c (n) ... (20)
The adder 525 adds a prediction sound source signal v '(n) obtained from the equation (13) to the value z (n) according to the following equation to obtain a sound source signal d (n) in the representative interval, and is supplied to a drive sound source decoder 283 and a gain / phase correction calculator 270 from:
d (n) = v '(n) + z (n) ... (21)
The above description is related to the transmitter side according to the second embodiment of the present invention.
The reception side of the system according to the second embodiment will be described below. A gain decoder 530 decodes the gain code and outputs a decoded gain g '. A generator 540 receives the code for the index of the selected code directory and, in accordance with the index, selects a code vector c (n) from a code directory 520. The generator 540 then generates a sound source signal z (n) using the decoded gain g ′ according to equation (20) and outputs it to an adder 550.
The adder 550 performs the same operation as the adder on the transmitter side so as to produce a sound source signal d (n) in the representative interval by adding the value z (n) to a prediction sound source signal v '(n) output from a prediction circuit 345. to win, and outputs it to a circuit 340 for reproducing the drive sound source.
The above description is related to the receiving side according to the second embodiment of the present invention.
The above-described embodiments are only examples of the present invention, and various modifications can be made.
In the first embodiment, the amplitude and position of the multiple pulse obtained with respect to a prediction difference signal in the representative interval are quantized (SQed). However, to reduce the amount of information, these values can be vector quantized (VQed). For example, only the position can be vectorized while the amplitude is scalar quantized, or the amplitude is scalar quantized while the position is vector quantized. Alternatively, both the amplitude and the position can be quantized vectorially. With regard to a detailed description of the method for vectorial quantization of the position, reference is made, for example, to R. Zinser et al., "4800 and 7200 bit / sec Hybrid Codebook Multipulse Coding," (ICASSP, pp. 747-750, 1989) (reference 6) ,
Further, in the first embodiment, the gain correction coefficient ck and the phase correction coefficient dk are obtained and transmitted at pitch intervals other than the representative interval. However, the decoded average pitch period T 'can be interpolated using the adjacent pitch period for each pitch interval, so that the transmission of a phase correction coefficient can be omitted. In addition to transmitting a gain correction coefficient every pitch interval, a gain correction coefficient obtained every pitch interval can be approximated by a least squares curve or a least squares line, and a transmission can be performed by coding the coefficient of the curve or the line. These methods can be used in any combination. With these arrangements, the amount of information for transmitting the correction information can be reduced.
Instead of obtaining a phase correction coefficient in each pitch interval, a linear phase term τ can be obtained from an end portion of a frame so as to be assigned to each pitch interval, as described, for example, in Ono and Ozawa et al., "2.4 kbps pitch prediction multi-pulse Speech Coding ", Proc. ICASSP 54.9, 1988) (reference 7). According to another method, a phase correction coefficient obtained every pitch interval is approximated by a least squares line or a least squares curve, and transmission is performed by coding the coefficient of the line or the curve.
Furthermore, in the first embodiment of the present invention, various sound source signals in accordance with the features of a one-frame speech signal as in Reference 1 can be used. For example, voice signals are classified into vowel, nasal, rubbing and explosive sounds, and the arrangement of the first embodiment can be used in a vowel volume interval.
In the first and second embodiments, a K parameter is encoded as a spectrum parameter, and an LPC analysis is used as an analysis method. However, other known parameters such as LSP, LPC-cepstrum, cepstrum, improved cepstrum, general cepstrum and melcepstrum can be used as spectrum parameters. An optimal analysis method can be used for each parameter.
Furthermore, in the first and second embodiments, when prediction is to be performed, a representative interval is set to a predetermined pitch interval in one frame. However, prediction can be performed in each pitch interval in one frame to calculate a sound source signal with respect to a given difference signal, and gain and phase correction coefficients in other pitch intervals are calculated. Further, a weighted volume difference between a voice signal reproduced by the above operation and an input signal is calculated, and a pitch interval that minimizes the volume difference is selected as a representative interval. Reference is made to reference 1 for a detailed description of this method. Although the processing effort is increased and the information about the location of the representative interval must also be transmitted, the properties of the system are further improved with this arrangement.
In the subframe divider 195, a frame is divided into pitch intervals, each of which has the same length as a pitch period. However, a frame can be divided into pitch intervals, each of a given length (e.g. 5 ms). With this arrangement, although the pitch period does not have to be extracted and the processing cost is reduced, the sound quality is slightly deteriorated.
In order to reduce the processing effort, the calculation of an influence signal can also be omitted on the transmitter side. With this waiver, the circuit 283 for reproducing the drive sound source, the synthesizing filter 281 and the subtractor 190 can be omitted on the transmitter side, but the sound quality is deteriorated.
To improve the sound quality by shaping the quantization noise, an adaptive post-filter, which responds to at least one pitch or spectrum envelope, can be connected to the output connection of the synthesizing filter on the decoding side. With regard to the arrangement of the adaptive postfilter, reference is made, for example, to Kroon et al., "A Class of Analysis-by-synthesis Predictive Coders for High Quality Speech Coding at Rates between 4.8 and 16 kb / s," (IEEE JSAC, Vol. 6.2, 353-363, 1988) (reference 8).
As is known in the field of digital signal processing, the autocorrelation function or the cross-correlation function corresponds to a power density spectrum or a cross power density spectrum on the frequency axis, and can therefore be calculated on the basis of these spectra. With regard to the method for calculating these functions, reference is made to Oppenheim et al., "Digital Signal Processing" (Prentice-Hall, 1975) (reference 9).
As described above, according to the present invention, a sound source signal in a representative interval can be very effectively divided by dividing a frame into units of pitch periods, the prediction for a pitch interval (representative interval) being made based on a previous sound source signal, and be represented by a suitable representation of a prediction error by a multiple pulse or a sound source signal vector (code vector). In addition, in other pitch intervals of the same frame, the gain and phase of the sound source signal in the representative interval are corrected to obtain the sound source signal of the frame, so that the sound source signal of the speech of the frame can be appropriately represented by a small amount of sound source information. Therefore, according to the present invention, decoded / reproduced speech can be obtained in excellent sound quality compared to the conventional method.
3 sheets
Sheet 1 Sheet 2 Sheet 3
8 members in 4 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 18908489 | Japan | A | |
| 18908489 | Japan | A | |
| 18908489 | Japan | – | |
| 18908489 | – | – | – |
| JP19890189084 | – | – | – |
Members8
| Document | Office | Kind | |
|---|---|---|---|
| EP0409239A2 | European Patent Office (EPO) | A2 | |
| JPH0353300A | Japan | A | |
| EP0409239A3 | European Patent Office (EPO) | A3 | |
| US5142584A | United States of America | A | |
| EP0409239B1 | European Patent Office (EPO) | B1 | |
| DE69023402D1 | Germany | D1 | |
| DE69023402T2This record | Germany | T2 | |
| JP2940005B2 | Japan | B2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Ceased/non-payment of the annual feeCeased8339 | 8339 | |
| No opposition during term of oppositionOpposition8364 | 8364 |
Numbers
- Publication
- 69023402
- Publication, DOCDB
- 69023402
- Publication, EPODOC
- DE69023402T
- Application
- 69023402
- Application, DOCDB
- 69023402
- Application, EPODOC
- DE19906023402T
Titles2
- German
- Verfahren zur Sprachkodierung und -dekodierung.
- English
- Speech coding and decoding methods.
Classification
- CPC, 2
- G10L19/10
- G10L25/90
- IPC, 6
- G10L19 06
- G10L19 10
- G10L19 125
- G10L25 51
- G10L25 90
- H03M7 30
