Detecting speech frames belonging to a low energy sequence
Claim Score by NHIP
Abstract
In order to enable a detection of speech frames belonging to a low energy sequence of a speech signal, it is proposed that a speech energy in a current speech frame is determined. Further, a speech energy level is estimated based on a speech energy in a plurality of speech frames. The determined speech energy in the current speech frame is scaled, if the estimated speech energy level deviates at least by a predetermined amount from a predetermined nominal speech energy level. Then, it can be decided that the current speech frame belongs to a low energy sequence, if the, potentially scaled, frame energy is lower than a predetermined low energy threshold value.
Term
Term ended
Projected expiry passed 4 April 2025, 1.5 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
25 claims: 5 independent, 20 dependent
- 1Broadest claimClaim Score 57, broad(NHIP)A method for detecting speech frames belonging to a low energy sequence of a speech signal, said method comprising:determining a speech energy in a current speech frame;estimating a speech energy level based on a speech energy in a plurality of speech frames;if said estimated speech energy level deviates at least by a predetermined amount from a predetermined nominal speech energy level, scaling said determined speech energy in said current speech frame;and deciding that said current speech frame belongs to a low energy sequence, if said, potentially scaled, frame energy is lower than a predetermined low energy threshold value.
- 14An encoding module comprising:a frame energy detector adapted to determine a speech energy in a current speech frame of a speech signal;a speech level estimator adapted to estimate a speech energy level based on a speech energy in a plurality of speech frames;a scaling portion adapted to scale a speech energy in a current speech frame determined by said frame energy detector, if a speech energy level estimated by said speech level estimator deviates at least by a predetermined amount from a predetermined nominal speech energy level;and a selector adapted to decide that said current speech frame belongs to a low energy sequence, if a, potentially scaled, frame energy provided by said scaling portion is lower than a predetermined low energy threshold value.
- 21An encoding module comprising:means for determining a speech energy in a current speech frame of a speech signal;means for estimating a speech energy level based on a speech energy in a plurality of speech frames;means for scaling a determined speech energy in a current speech frame, if an estimated speech energy level deviates at least by a predetermined amount from a predetermined nominal speech energy level;and means for deciding that said current speech frame belongs to a low energy sequence, if a, potentially scaled, frame energy is lower than a predetermined low energy threshold value.
- 22An electronic device comprising:a frame energy detector adapted to determine a speech energy in a current speech frame of a speech signal;a speech level estimator adapted to estimate a speech energy level based on a speech energy in a plurality of speech frames;a scaling portion adapted to scale a speech energy in a current speech frame determined by said frame energy detector, if a speech energy level estimated by said speech level estimator deviates at least by a predetermined amount from a predetermined nominal speech energy level;and a selector adapted to decide that said current speech frame belongs to a low energy sequence, if a, potentially scaled, frame energy provided by said scaling portion is lower than a predetermined low energy threshold value.
- 25A software program product in which a software code for detecting speech frames belonging to a low energy sequence of a speech signal is stored, said software code realizing the following steps when running in a processing unit of an electronic device:determining a speech energy in a current speech frame;estimating a speech energy level based on a speech energy in a plurality of speech frames;if said estimated speech energy level deviates at least by a predetermined amount from a predetermined nominal speech energy level, scaling said determined speech energy in said current speech frame;and deciding that said current speech frame belongs to a low energy sequence, if said, potentially scaled, frame energy is lower than a predetermined low energy threshold value.
Independent claims5
115 paragraphs in 5 sections, as filed
FIELD OF THE INVENTION
0001The invention relates to a method for detecting speech frames belonging to a low energy sequence of a speech signal. The invention relates equally to encoding modules, to an electronic device and to a software program product.
BACKGROUND OF THE INVENTION
0002Speech signals can be encoded for enabling an efficient transmission or storage of these speech signals. The encoding can be based on a single coding mode. Alternatively, it can be based on different coding modes, resulting in different bit rates of the encoded speech. In this case, the respectively appropriate coding mode can be selected based on current conditions. The encoded signal may be decoded again taking account of the coding mode employed for the encoding.
0003Well-known speech codecs, which are employed for packet based transmissions of speech, are the Adaptive Multi-Rate (AMR) speech codec and the Adaptive Multi-Rate Wideband (AMR-WB) speech codec.
0004The AMR speech codec was developed for Global System for Mobile communications (GSM) channels and Enhanced Data rates for GSM Evolution (EDGE) channels, whereas the AMR-WB speech codec was developed for Wideband Code Division Multiple Access (WCDMA) channels. In addition, the codecs can be utilized in packet switched networks. Aspects of the AMR speech codec are defined in the standards 3GPP TS 26.071 V6.0.0 (2004-12): “AMR Speech Codec; General description” and 3GPP TS 26.090 V6.0.0 (2004-12): “AMR Speech Codec; Transcoding Functions”, which are incorporated by reference herein. Aspects of the AMR WB speech codec are defined in the standards 3GPP TS 26.171 V6.0.0 (2004-12): “AMR Wideband Speech Codec; General description” and 3GPP TS 26.190 V6.0.0 (2004-12): “AMR Wideband Speech Codec; Transcoding Functions”, which are equally incorporated by reference herein.
0005The AMR codec samples incoming speech with a sampling frequency of 8 kHz. The AMR-WB speech codec samples incoming speech with a sampling frequency of 16 kHz. The sampled speech is then subjected to an encoding process.
0006Both codecs are based on the conventional Algebraic Code Excitation Linear Prediction (ACELP) technology. Both codecs are further multi-rate codecs, which are able to employ a plurality of independent coding modes. However, the codecs may also be operated using a variable rate scheme, in which the output bit rate is not fixed to a number of predetermined values but can be selected freely.
0007A voice activity detection (VAD) is used to lower the bit rate during silence periods by employing a discontinuous transmission (DTX) functionality in case the speech signal does not comprise an active voice signal. In addition, the AMR codec comprises eight active speech coding modes with bit-rates of 12.2, 10.2, 7.95, 7.40, 6.70, 5.90, 5.15 and 4.75 kbit/s, while the AMR WB codec comprises nine active speech coding modes with bit-rates of 23.85 23.05, 19.85, 18.25, 15.85, 14.25, 12.65, 8.85 and 6.60 kbit/s. Typically the AMR and AMR WB codecs select the codec mode based only on the network capacity and radio channel conditions. GSM radio networks utilize the codec mode selection to handle the channel fading and error bursts, whereas WCDMA networks rely on a fast power control and make use of the codec mode selection for controlling the capacity in the network.
0008The codec mode can be selected independently for each analysis speech frame, having a length of 20 ms, depending on the supported mode set.
0009For illustration, an AMR based mobile communication system is depicted in <figref idref="DRAWINGS">FIG. 1</figref>.
0010The system comprises a mobile station (MS) <b>10</b>, a base transceiver station (BTS) <b>11</b> and a transcoder (TC) <b>12</b>.
0011The mobile station <b>10</b> comprises a multi-rate speech encoder <b>101</b> and a multi-rate channel encoder <b>102</b>. The multi-rate speech encoder <b>101</b> receives input speech. The signals output by the multi-rate channel encoder <b>102</b> are transmitted via an uplink radio channel <b>141</b> to the BTS <b>11</b>. The mobile station <b>10</b> further comprises a multi-rate channel decoder <b>103</b> and a multi-rate speech decoder <b>104</b>. The multi-rate channel decoder <b>103</b> receives signals from the BTS <b>11</b> via a downlink radio channel <b>142</b>. The multi-rate speech decoder <b>104</b> provides a speech output. The mobile station <b>10</b> further comprises a link adaptation unit <b>105</b> with a downlink quality measurement component <b>106</b> and a mode request generator <b>107</b>.
0012The BTS <b>11</b> comprises a multi-rate channel decoder <b>111</b>, which receives signals from a mobile station <b>10</b> via the uplink radio channel <b>141</b>. The signals output by the multi-rate channel decoder <b>111</b> are transferred via an A<sub>bis/ter </sub>interface to the TC <b>12</b>. The BTS <b>11</b> further comprises a multi-rate channel encoder <b>112</b>, which receives an input from the TC <b>12</b> via the A<sub>bis/ter </sub>interface. The signals output by the multi-rate channel encoder <b>112</b> are transmitted via the downlink radio channel <b>142</b> to the mobile station <b>10</b>. The BTS <b>11</b> further comprises a link adaptation unit <b>113</b>, including an uplink quality measurement component <b>114</b>, an uplink mode control component <b>115</b> and a downlink mode control <b>116</b>.
0013The TC <b>12</b> comprises a multi-rate speech decoder <b>121</b>, which receives signals from the BTS <b>11</b> via the A<sub>bis/ter </sub>interface, and which provides a speech output. The TC <b>12</b> further comprises a multi-rate signal encoder <b>122</b>, which receives a speech input. The signals output by the multi-rate signal encoder <b>122</b> are transferred via the A<sub>bis/ter </sub>interface to the BTS <b>11</b>.
0014The uplink quality measurement component <b>114</b> of the BTS <b>11</b> performs quality measurements on received uplink signals and provides a corresponding quality indicator QI<sub>u </sub>to the uplink mode control component <b>115</b>. The uplink mode control component <b>115</b> receives in addition information about network constraints and determines a codec mode command MC<sub>u</sub>. This command MC<sub>u </sub>indicates a codec mode that should be employed by the mobile station <b>10</b> for encoding speech signals for uplink transmissions in view of the current uplink radio channel conditions and the current network capacity. The command MC<sub>u </sub>is transmitted after a channel encoding by the multi-rate channel encoder <b>112</b> as an inband signaling to the mobile station <b>10</b>, together with speech data S and a downlink codec mode indicator MI<sub>d</sub>.
0015In the mobile station <b>10</b>, the codec mode command MC<sub>u </sub>is available at the output of the multi-rate channel decoder <b>103</b> and provided to the multi-rate speech encoder <b>101</b>. The multi-rate speech encoder <b>101</b> uses thereupon the codec mode indicated by the command MC<sub>u </sub>for encoding input speech signals on a frame basis. For the encoding, it detects whether any voice activity is present. If no voice activity is present, it uses a discontinued transmission. Otherwise, it computes linear prediction coding (LPC) coefficients and long term prediction (LTP) parameters, and performs a fixed codebook excitation search for obtaining a parametrical representation of the speech. The encoded speech S and an indication of the employed codec mode MI<sub>u </sub>are then provided to the multi-rate channel encoder <b>102</b> for a channel encoding. Moreover, the downlink quality measurement component <b>106</b> performs quality measurements on received downlink signals and provides a corresponding quality indicator QI<sub>d </sub>to the mode request generator <b>107</b>. The mode request generator <b>107</b> determines based on the quality indicator QI<sub>d </sub>a downlink codec mode request MR<sub>d</sub>. This request MR<sub>d </sub>is transmitted after a channel encoding by the multi-rate channel encoder <b>102</b> in an inband signaling to the BTS <b>11</b>, together with speech data S and an uplink codec mode indicator MI<sub>u</sub>.
0016In the BTS <b>11</b>, the codec mode request MR<sub>d </sub>is available again at the output of the multi-rate channel decoder <b>111</b> and provided to the downlink mode control <b>116</b>, which receives in addition network constraint information as well. Based on the received information, the downlink mode control <b>116</b> determines a downlink codec mode command MC<sub>d</sub>, which is then transferred via the A<sub>bis/ter </sub>interface to the multi-rate speech encoder <b>122</b> of the TC <b>12</b>. This command MC<sub>d </sub>indicates a codec mode that should be employed by the TC <b>12</b> for encoding speech signals for downlink transmissions in view of the current downlink radio channel conditions and the current network capacity. The multi-rate speech encoder <b>122</b> thus uses the codec mode indicated by the command MC<sub>d </sub>for encoding input speech signals, as outlined above for the multi-rate speech encoder <b>101</b>, and outputs in addition to the encoded speech S a corresponding downlink codec mode indicator MI<sub>d</sub>.
0017As becomes apparent, the rate adaptation is based exclusively on channel and network conditions. The speech signal itself is only evaluated for deciding on a discontinuous transmission. While a VAD driven discontinuous transmission is the most common approach for optimizing the network capacity based on the source signal, other perceptual criteria could be utilized to select the optimal codec mode during active speech. Current AMR and AMR-WB codec implementations in GSM and UMTS do not support such a source based rate adaptation, though, nor do they support an average bit rate control.
0018In a source based rate adaptation, the encoded speech sequence is classified into classes based on speech characteristics. Used speech classes can be, for example, low energy sequence, transient sequence, unvoiced sequence and voiced sequences. The codec mode which is to be employed may then be selected based as well on the detected speech class. For example, low energy sequences can be encoded with lower bit rates without any degradation in speech quality than other sequences. On the other hand, for example, during transient sequences the speech quality degrades very rapidly, if codec modes with a lower bit rate are used. The appropriate codec modes for voiced and unvoiced speech sequences depend on the frequency content of the sequences. For example, low frequency voiced sequence can be coded with a lower bit rate without speech quality degradation than high frequency voiced sequence. Usually noise-like unvoiced sequences requires a high bit rate representation.
0019Speech frames can be classified for a codec mode selection based on information that is already available in the speech encoder. Such information may comprise for instance already calculated values, which are obtained from VAD, LPC and LTP routines, like spectral content, LTP and fixed codebook gains of previous speech frames, etc. Therefore, a source adaptation algorithm may be rather simple and not increase the complexity of the encoding process substantially.
0020A source based rate adaptation is currently used for example in IS-95 CDMA networks, in which the Enhanced Variable Rate Codec (EVRC) is used as a source controlled variable rate codec. The EVRC selects the bit-rate of the encoded parameters before encoding the signal. In an exemplary EVRC implementation, the source based rate adaptation is performed based on the input signal and on LPC filter parameters, before the quantization of the filter parameters as well as before the search for LTP filter parameters and for an excitation signal. In the source based rate adaptation of the EVRC, the frame energy is calculated in two frequency bands and compared to thresholds for a mode selection. The thresholds are updated using background noise estimates and an autocorrelation function from the LPC analysis. Hence, the rate is selected using the reflection coefficients of the LPC analysis and input signal before the encoding functions. The EVRC has been described for instance in the document IS-127: “Enhanced Variable Bit Rate Codec, Speech Service option 3 for Wideband spread spectrum digital systems”.
0021In general, a source controlled variable rate operation aims at reducing the average source bit rate without any perceptual degradation in the decoded speech quality. The advantage of a lower average bit rate is a lower transmission power and hence a higher capacity in the networks. A reduced bit rate also results in a smaller storage size in a voice recording application.
0022<figref idref="DRAWINGS">FIG. 2</figref> is a diagram depicting the energy of a speech sequence over time and in addition a possible source adaptation exploiting codec modes with 6.60 kbit/s, 12.65 kbit/s and 23.05 kbit/s in addition to a discontinuous transmission mode. It can be seen from <figref idref="DRAWINGS">FIG. 2</figref> that a considerable bit rate reduction can be achieved by coding low energy sequences with 6.60 kbit/s. The usage of discontinuous transmission (DTX) is not possible during such low energy sequences, because a discontinuous transmission may cause audible speech clipping.
0023However, the absolute speech quality will degrade as a function of the bit-rate in a multi-rate speech codec. This is especially true, when strong environmental noise, for instance in a car, on the street or in a cafeteria, is present during a call. It is thus a problem, if the low energy threshold has been set by too high value and the low bit rate mode is used more frequently than appropriate.
0024J. Mäkinen and J. Vainio propose in the conference paper: “Source signal based rate adaptation for GSM AMR speech codec”, Proc ITCC 2004, Las Vegas, USA, 2004, to scale the low energy threshold based on an estimate of a long-term energy level of the speech.
SUMMARY OF THE INVENTION
0025It is an object of the invention to enable a reliable determination whether a speech frame belongs to a low energy sequence of a speech signal.
0026It is also an object of the invention to provide an alternative to existing approaches for determination of whether a speech frame belongs to a low energy sequence of a speech signal.
0027A method for detecting speech frames belonging to a low energy sequence of a speech signal is proposed. The method comprises determining a speech energy in a current speech frame. The method further comprises estimating a speech energy level based on a speech energy in a plurality of speech frames. The method further comprises scaling the determined speech energy in the current speech frame, if the estimated speech energy level deviates at least by a predetermined amount from a predetermined nominal speech energy level. The method further comprises deciding that the current speech frame belongs to a low energy sequence, if the, potentially scaled, frame energy is lower than a predetermined low energy threshold value.
0028Moreover, an encoding module is proposed, which comprises a frame energy detector adapted to determine a speech energy in a current speech frame of a speech signal. The electronic device further comprises a speech level estimator adapted to estimate a speech energy level based on a speech energy in a plurality of speech frames. The electronic device further comprises a scaling portion adapted to scale a speech energy in a current speech frame determined by the frame energy detector, if a speech energy level estimated by the speech level estimator deviates at least by a predetermined amount from a predetermined nominal speech energy level. The electronic device further comprises a selector adapted to decide that the current speech frame belongs to a low energy sequence, if a, potentially scaled, frame energy provided by the scaling portion is lower than a predetermined low energy threshold value.
0029Moreover, an encoding module is proposed, which comprises in general means for determining a speech energy in a current speech frame of a speech signal, means for estimating a speech energy level based on a speech energy in a plurality of speech frames, means for scaling a speech energy in a current speech frame determined by the frame energy detector, if a speech energy level estimated by the speech level estimator deviates at least by a predetermined amount from a predetermined nominal speech energy level, and means for deciding that the current speech frame belongs to a low energy sequence, if a, potentially scaled, frame energy provided by the scaling portion is lower than a predetermined low energy threshold value.
0030Moreover, an electronic device is proposed, which comprises at least the same features as one of the proposed encoding modules.
0031Finally, a software program product is proposed, in which a software code for detecting speech frames belonging to a low energy sequence of a speech signal is stored. When running in a processing unit of an electronic device, the software code realizes the steps of the proposed method.
0032The invention proceeds from the consideration that the optimal low energy threshold for determining whether a speech frame belongs to a low energy sequence or not depends on the speech energy level. The term speech energy level denotes the average energy level of active speech over a longer time period, for instance during one to five seconds. The optimal low energy threshold can be found easily, if the energy level of speech remains constant. However, that is an ideal case. In practice, the speech energy level varies for instance from one conversation to another and also during a single conversation. Even during a single, long sentence, the speech energy level may vary considerably.
0033In order to take account of variations in the speech energy level, the low energy threshold could be adapted based on an estimate of a long-term energy level of the speech, as proposed in the above cited document by J. Makinen and J. Vainio.
0034The invention proposes instead that the speech energy in the current speech frame, which will also be referred to as frame energy, is scaled depending on deviations of a speech energy level estimate from a predetermined nominal speech energy level. Thereby, the scaled frame energy can be kept independent of the current general speech level and, consequently, it can always be compared to a fixed low energy threshold.
0035It is an advantage of the invention that it enables a reliable association of speech frames to low energy sequences. The invention may thus be used for optimizing the speech codec mode selection in a source based rate adaptation. By an improved basis for a speech codec mode selection, a more efficient source based rate adaptation can be achieved, and thus a better and more constant compromise between the contrasting requirements of a low bit rate and a high speech quality.
0036In one embodiment of the invention, the speech energy in the current speech frame is determined by averaging speech energies in a plurality of frequency sub bands in the current speech frame.
0037In a further embodiment of the invention, only those speech frames are considered as a basis for determining the speech energy level, in which a voice activity has been detected, since only the active speech may contribute to the speech energy level.
0038In a further embodiment of the invention, the speech energy level is updated for a respective current speech frame by combining a speech energy level available for a speech frame preceding the current speech frame with the speech energy determined for the current speech frame with variable coefficients, wherein a coefficient for the speech energy determined for the current speech frame may be equal to zero.
0039The combining can be carried out in various ways and taking account of various criteria.
0040For example, in case the determined speech energy in the current speech frame is higher than a speech energy in the preceding speech frame, the available speech energy level may be weighted with a first coefficient. In case the determined speech energy in the current speech frame is lower than the speech energy in the preceding speech frame, the available speech energy level may be weighted with a second coefficient. The first coefficient is then advantageously higher than the second coefficient. The coefficient for the speech energy determined for the current speech frame may be adapted for example in an opposite manner.
0041In addition, different coefficients may be selected for speech frames at the beginning of a respective speech signal than for speech frames at a later stage. The difference between the first coefficient for the available speech energy level and the second coefficient for the available speech energy level may in particular be larger for a predetermined number of speech frames at a beginning of a respective speech sequence. Thereby, a more aggressive adaptation can be achieved at the beginning of a speech sequence, at which only a few speech frames are available for determining the speech energy level.
0042For the very first speech frame, for example simply the energy of this first speech frame could be used as the speech energy level. Alternatively, the nominal speech energy level could be considered as the speech energy level which is available for a theoretical preceding speech frame.
0043In a further embodiment of the invention, the speech energy level is only updated, in case the updated speech energy level exceeds a background noise estimate for the current speech frame by a predetermined factor.
0044Advantageously, the frame energy is scaled to the nominal speech level. That is, in case the estimated speech energy level exceeds the nominal speech level at least by a predetermined amount, the determined speech energy for the current frame is scaled to a lower value, while in case the estimated speech energy level falls short of the nominal speech level at least by a predetermined amount, the determined speech energy for the current frame is scaled to a higher value.
0045In one embodiment of the invention, the scaling is performed based on one of a plurality of correction functions. Each correction function is valid for another range of speech energy levels. Using a plurality of predetermined correction functions makes the implementation of the scaling easier than the use of a correction function that is adapted exactly for each occurring speech energy level. It is to be understood that the plurality of predetermined correction functions may be based on a single correction function including different coefficients. Such coefficients can be stored for instance in the form of a matrix, from which they are retrieved depending on the respective speech energy level.
0046In one embodiment of the invention, the scaling comprises at least a multiplication of the speech energy in the current speech frame by a selected value and an adding of a selected value. The values may constitute the coefficients of a correction function and be selected depending on the speech energy level.
0047The invention can be employed for instance in the scope of a speech coding. The current speech frame may be encoded with a dedicated low bit rate coding mode in case it is detected to belong to a low energy sequence. Encoding low energy sequences with a lower rate codec mode than higher energy sequences may contribute significantly to a decrease of the average bit rate. If the low energy sequences are further detected in a flexible way, as proposed, a speech quality degradation due to a too high low energy threshold can be avoided.
0048The low energy sequence detection can be performed before an encoding of the current speech frame or during an encoding of the current speech frame.
0049In AMR-WB modes from 12.2 kbit/s to 23.85 kbit/s, for example, the only difference in the encoding process is the fixed codebook excitation search. Lower modes, like the 12.2 kbit/s mode, use less pulses for excitation than higher modes, like the 23.85 kbit/s mode. The 23.85 kbit/s has also a high frequency extension gain parameter, which is calculated after the fixed codebook excitation search. For these modes, a source base rate adaptation could be performed during the encoding process before the fixed codebook excitation search. The lowest AMR-WB modes, 6.60 kbit/s and 8.85 kbit/s, in contrast, have a different LPC and LTP parameter representation. If these modes are considered as an option in a source based rate adaptation, the source based rate adaptation has to be performed at the beginning of encoding process before the calculation and quantization of the LPC and the LTP parameters. The mode selection can be performed more accurately, if it is carried out during the encoding process. In this case, the mode selection can exploit as well information about the current frame, that is, basically LPC and LTP information of the current encoded frame, in addition to information history.
0050The invention can be employed for example, though not exclusively, for a source based rate adaptation in AMR or AMR-WB speech codecs used in GSM and WCDMA communication systems. It can further be used for variable and multi-rate speech coding. It can be employed for example for the transmission of encoded speech over erroneous and capacity limited transmission channels, in both, circuit switched and packet switched domains.
0051Each of the proposed electronic devices can be for example a device comprising a speech encoder, in particular any device adapted to transmit or to store speech data. It can be for instance a mobile communication device or a network element of a mobile communication network, etc.
0052Other objects and features of the present invention will become apparent from the following detailed description considered in conjunction with the accompanying drawings. It is to be understood, however, that the drawings are designed solely for purposes of illustration and not as a definition of the limits of the invention, for which reference should be made to the appended claims. It should be further understood that the drawings are not drawn to scale and that they are merely intended to conceptually illustrate the structures and procedures described herein.
BRIEF DESCRIPTION OF THE FIGURES
0053<figref idref="DRAWINGS">FIG. 1</figref> is a schematic block diagram of a conventional communication system;
0054<figref idref="DRAWINGS">FIG. 2</figref> is a diagram illustrating data rates employed in an adaptive rate coding for an exemplary speech sequence;
0055<figref idref="DRAWINGS">FIG. 3</figref> is a schematic block diagram of a mobile station which can be implemented in accordance with an embodiment of the invention;
0056<figref idref="DRAWINGS">FIG. 4</figref> is a schematic block diagram of a processing unit of the mobile station of <figref idref="DRAWINGS">FIG. 3</figref> operating in accordance with an embodiment of the invention;
0057<figref idref="DRAWINGS">FIG. 5</figref> is a flow chart illustrating a source based coding mode selection according to an embodiment of the invention;
0058<figref idref="DRAWINGS">FIG. 6</figref> is a diagram presenting the same speech sequence at different speech levels;
0059<figref idref="DRAWINGS">FIG. 7</figref> is a diagram illustrating energy level estimates of the same speech sequence for different speech levels; and
0060<figref idref="DRAWINGS">FIG. 8</figref> is a diagram illustrating a scaling of the frame energy in accordance with an embodiment of the invention.
DETAILED DESCRIPTION OF THE INVENTION
0061<figref idref="DRAWINGS">FIG. 3</figref> is a block diagram of a mobile station <b>30</b> that has been implemented according to an embodiment of the invention, in order to enable an enhanced source based coding mode selection.
0062The mobile station <b>30</b> can be a part of a communication system like the one that has been described above with reference to <figref idref="DRAWINGS">FIG. 1</figref>. It may comprise the same structure of a receiving chain, a transmitting chain and a link adaptation portion. <figref idref="DRAWINGS">FIG. 3</figref> presents only selected parts of this transmitting chain.
0063A speech input (not shown) of the mobile station <b>30</b>, for instance a microphone, is connected via a pre-processing component <b>31</b>, an AMR encoder <b>32</b> and a post-processing component <b>38</b> to an output of the mobile station <b>30</b>, for instance a transmit antenna (not shown).
0064The pre-processing component <b>31</b> is designed to perform all processing preceding the AMR encoding of input speech, including for instance a sampling and quantization step (i.e., analog-to-digital conversion). The post-processing component <b>38</b> is designed to perform all processing following the AMR encoding of input speech that is required for a transmission via the radio interface, including for instance a channel encoding. That is, the post-processing component <b>38</b> may comprise at least the multi-rate channel encoder <b>102</b> of the mobile station <b>10</b> of <figref idref="DRAWINGS">FIG. 1</figref>.
0065The AMR encoder <b>32</b> can be a modification of the multi-rate speech encoder <b>101</b> of the mobile station <b>10</b> of <figref idref="DRAWINGS">FIG. 1</figref>.
0066The input of the AMR encoder <b>32</b> is connected to a voice activity detector (VAD) <b>33</b>. A first output of the VAD <b>33</b> is connected via a speech encoder <b>34</b> to the output of the AMR encoder <b>32</b>. A second output of the VAD <b>33</b> is connected via a discontinous transmission (DTX) component <b>35</b> to the output of the AMR encoder <b>32</b>.
0067Within the speech encoder <b>34</b>, an LPC (linear prediction coding) calculation component <b>341</b>, an LTP (long term prediction) calculation component <b>342</b> and a fixed codebook excitation component <b>343</b> are connected in sequence between the input and the output of the speech encoder <b>34</b>.
0068The AMR encoder <b>32</b> further comprises a source adaptation component <b>36</b> which is designed according to an embodiment of the invention. The source adaptation component <b>36</b> is arranged to access the speech encoder <b>34</b> either at its input or between the LTP calculation component <b>342</b> and the fixed codebook excitation component <b>343</b>. Optionally, the source adaptation component <b>36</b> may further interact with an additional rate determination algorithm (RDA) <b>37</b>.
0069It has to be noted that the mobile station <b>30</b> of <figref idref="DRAWINGS">FIG. 3</figref> may correspond entirely to a conventional mobile station, like a mobile phone, except for the source adaptation component <b>36</b>, which is modified in accordance with the invention.
0070<figref idref="DRAWINGS">FIG. 4</figref> is a schematic block diagram, which presents in particular in more detail parts of this modified source adaptation component <b>36</b>.
0071The source adaptation component <b>36</b> may be realized in software (SW), as shown, which can be executed by a processing unit <b>40</b> of the mobile station <b>30</b>. The processing unit <b>40</b> may be a dedicated processing unit or a processing unit that is adapted to execute various software codes implemented in the mobile station <b>30</b>, including for example any software code for other components of the AMR encoder <b>32</b>, like the additional rate determination algorithm <b>37</b>.
0072A speech frame input of the source adaptation component <b>36</b> is linked to a frame energy detector <b>361</b> and to a noise estimator <b>362</b>. The output of the frame energy detector <b>361</b> and the output of the noise estimator <b>362</b> are linked to an input of a speech level estimator <b>363</b>. The output of the frame energy detector <b>361</b> is linked in addition to an input of a scaling portion <b>364</b>. The output of the speech level estimator <b>363</b> is equally linked to an input of the scaling portion <b>364</b>. The output of the scaling portion <b>364</b> is linked to a mode selector <b>365</b>. The mode selector <b>365</b> may interact with the rate determination algorithm <b>37</b>. The output of the mode selector <b>365</b> forms the output of the source adaptation component <b>36</b>.
0073The operation of the mobile station <b>30</b> in accordance with an embodiment of the invention will now be described with reference to the flow chart of <figref idref="DRAWINGS">FIG. 5</figref>.
0074When a speech signal, which is received by the mobile station <b>30</b> for instance via a microphone, is to be transmitted to a mobile communication network, an encoding with source adaptation is started (step <b>501</b>).
0075The actual encoding of the speech signal is performed in a conventional manner. That is, the speech signal is first pre-processed by the pre-processing component <b>31</b>, including a sampling and quantization of the speech signal, etc. Resulting speech frames are provided to the AMR encoder <b>32</b>.
0076The VAD <b>33</b> of the AMR encoder <b>32</b> then detects for each speech frame whether there is any voice activity in the signal.
0077If no voice activity detected in the current speech frame, a discontinuous transmission is taken care of by the DTX component <b>35</b>. The signal is then provided directly to the post processing component <b>38</b>.
0078If a voice activity detected in the current speech frame, the frame is input to the speech encoder <b>34</b> and in addition to the source adaptation component <b>36</b>.
0079In the source adaptation component <b>36</b>, the frame energy detector <b>361</b> detects the frame energy total_band<sup>j </sup>of the current speech frame j, while the noise estimator <b>362</b> estimates the background noise NE<sup>j </sup>for the current speech frame j (step <b>502</b>).
0080The frame energy total_band<sup>j </sup>of the current speech frame is a total band energy corresponding to the average of the energy in the speech frame over various sub bands. It is calculated as follows: <maths id="MATH-US-00001" num="1"><math overflow="scroll"><mrow><mrow><msup><mi>total_band</mi><mi>j</mi></msup><mo>=</mo><mfrac><mrow><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>m</mi></munderover><mo></mo><mrow><mo>(</mo><mrow><mi>vad_filt</mi><mo></mo><msubsup><mi>_band</mi><mi>i</mi><mi>j</mi></msubsup></mrow><mo>)</mo></mrow></mrow><mi>m</mi></mfrac></mrow><mo>,</mo></mrow></math></maths><br /> where vad_filt_band<sub>i</sub><sup>j </sup>is the energy level of i<sup>th </sup>band of the available sub bands in j<sup>th </sup>speech frame, and where m is the number of available sub bands in the total frequency band.
0081The available sub bands can be taken for example from the following table, in which the total frequency band has been divided into m=12 sub bands, similarly as known from the VAD algorithm described in the standards 3GPP TS 26.093 V6.0.0 (2003-03): “AMR Speech Codec; Source Controlled Rate operation”, which is incorporated by Reference herein. <tables id="TABLE-US-00001" num="1"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="OFFSET" colwidth="35PT" align="left" /><colspec colname="1" colwidth="56PT" align="center" /><colspec colname="2" colwidth="126PT" align="center" /><thead><row><entry /><entry /></row><row><entry /><entry namest="OFFSET" nameend="2" align="center" rowsep="1" /></row><row><entry /><entry>Sub band number</entry><entry>Frequencies</entry></row><row><entry /><entry namest="OFFSET" nameend="2" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="4"><colspec colname="OFFSET" colwidth="35PT" align="left" /><colspec colname="1" colwidth="56PT" align="char" char="." /><colspec colname="2" colwidth="77PT" align="right" /><colspec colname="3" colwidth="49PT" align="left" /><tbody valign="top"><row><entry /><entry>1</entry><entry>0-200</entry><entry>Hz</entry></row><row><entry /><entry>2</entry><entry>200-400</entry><entry>Hz</entry></row><row><entry /><entry>3</entry><entry>400-600</entry><entry>Hz</entry></row><row><entry /><entry>4</entry><entry>600-800</entry><entry>Hz</entry></row><row><entry /><entry>5</entry><entry>800-1200</entry><entry>Hz</entry></row><row><entry /><entry>6</entry><entry>1200-1600</entry><entry>Hz</entry></row><row><entry /><entry>7</entry><entry>1600-2000</entry><entry>Hz</entry></row><row><entry /><entry>8</entry><entry>2000-2400</entry><entry>Hz</entry></row><row><entry /><entry>9</entry><entry>2400-3200</entry><entry>Hz</entry></row><row><entry /><entry>10</entry><entry>3200-4000</entry><entry>Hz</entry></row><row><entry /><entry>11</entry><entry>4000-4800</entry><entry>Hz</entry></row><row><entry /><entry>12</entry><entry>4800-6400</entry><entry>Hz</entry></row><row><entry /><entry namest="OFFSET" nameend="3" align="center" rowsep="1" /></row></tbody></tgroup></table></tables><br /> The noise estimator <b>362</b> estimates the background noise NE<sup>j </sup>in the j<sup>th </sup>speech frame as follows: <maths id="MATH-US-00002" num="2"><math overflow="scroll"><mrow><mrow><msup><mi>NE</mi><mi>j</mi></msup><mo>=</mo><mfrac><mrow><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>m</mi></munderover><mo></mo><mrow><mo>(</mo><msubsup><mi>bckr_est</mi><mi>i</mi><mi>j</mi></msubsup><mo>)</mo></mrow></mrow><mi>m</mi></mfrac></mrow><mo>,</mo></mrow></math></maths><br /> where the variable bckr_est<sub>i</sub><sup>j </sup>is the background noise estimate of i<sup>th </sup>sub band in the j<sup>th </sup>speech frame. The calculation of the background noise estimate as such is well known and can be performed for example as described in the above mentioned standard TS 26.093 and in the further standard 3GPP TS 26.193 V6.0.0 (2004-12): “AMR Wideband Speech Codec; Source Controlled Rate operation”, the latter being equally incorporated by reference herein.
0082The detected frame energy total_band<sup>j </sup>and the estimated background noise NE<sup>j </sup>for the current speech frame j are both forwarded to the speech level estimator <b>363</b>. The frame energy total_band<sup>j </sup>for the current speech frame j is forwarded in addition to the scaling portion <b>364</b>.
0083Next, the speech level estimator <b>363</b> calculates or updates a long term energy level E<sub>L</sub><sup>j </sup>of the speech, which is valid for the current speech frame j. The energy level estimate is determined only during speech segments for which the discontinuous transmission is off, in order to calculate the energy level estimate for the active speech level (step <b>503</b>).
0084Only for the first speech frame of the speech sequence, the energy level estimate E<sub>L</sub><sup>j </sup>is calculated. That is, for this first speech frame, the energy level estimate is simply set equal to the frame energy total_band<sup>1 </sup>of the first speech frame, which is provided by the frame energy detector <b>361</b>. For any subsequent speech frame j, the energy level estimate E<sub>L</sub><sup>j−1 </sup>determined for the respective preceding speech frame j−1 is updated to obtain the energy level estimate E<sub>L</sub><sup>j </sup>for the current speech frame j.
0085An update of the energy level estimate E<sub>L</sub><sup>j </sup>for the j<sup>th </sup>speech frame in the speech level estimator <b>363</b> according to an embodiment of the invention will now be described.
0086A flag BOS is used for indicating the beginning of a speech segment as follows: <maths id="MATH-US-00003" num="3"><math overflow="scroll"><mrow><mi>BOS</mi><mo>=</mo><mrow><mo>{</mo><mtable><mtr><mtd><mrow><mn>1</mn><mo>,</mo></mrow></mtd><mtd><mrow><msub><mi>C</mi><mi>activity</mi></msub><mo><</mo><mn>0</mn></mrow></mtd></mtr><mtr><mtd><mrow><mn>0</mn><mo>,</mo></mrow></mtd><mtd><mrow><msub><mi>C</mi><mi>activity</mi></msub><mo>≥</mo><mn>0</mn></mrow></mtd></mtr></mtable></mrow></mrow></math></maths><br /> where C<sub>activity </sub>is a counter having the initial value of 100. For the first 100 speech frames of the speech segment, the counter C<sub>activity </sub>is decremented by one with each new speech frame for which a frame energy value is received: <maths id="MATH-US-00004" num="4"><math overflow="scroll"><mrow><msub><mi>C</mi><mi>activity</mi></msub><mo>=</mo><mrow><mo>{</mo><mtable><mtr><mtd><mrow><mn>0</mn><mo>,</mo></mrow></mtd><mtd><mrow><msub><mi>C</mi><mi>activity</mi></msub><mo><</mo><mn>0</mn></mrow></mtd></mtr><mtr><mtd><mrow><mrow><msub><mi>C</mi><mi>activity</mi></msub><mo>-</mo><mn>1</mn></mrow><mo>,</mo></mrow></mtd><mtd><mrow><msub><mi>C</mi><mi>activity</mi></msub><mo>≥</mo><mn>0</mn></mrow></mtd></mtr></mtable></mrow></mrow></math></maths>
0087The counter C<sub>activity </sub>thus has a value of zero after the first 100 speech frames of a new speech segment have been received. This enables a differentiation between the beginning of a speech segment and a later stage of a speech segment when updating the energy level estimate.
0088In addition, the updating differentiates between the cases that the frame energy total_band<sup>j </sup>of the current speech frame j is higher or lower than the frame energy total_band<sup>j−1 </sup>of the preceding speech frame j−1.
0089The update is performed as follows: <maths id="MATH-US-00005" num="5"><math overflow="scroll"><mrow><msubsup><mi>E</mi><mi>L</mi><mi>j</mi></msubsup><mo>=</mo><mrow><mo>{</mo><mtable><mtr><mtd><mrow><msubsup><mi>E</mi><mi>L</mi><mrow><mi>j</mi><mo>-</mo><mn>1</mn></mrow></msubsup><mo>,</mo></mrow></mtd><mtd><mrow><mi>BOS</mi><mo>=</mo><mrow><mrow><mn>0</mn><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>and</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><msup><mi>total_band</mi><mi>j</mi></msup></mrow><mo>></mo><msup><mi>total_band</mi><mrow><mi>j</mi><mo>-</mo><mn>1</mn></mrow></msup></mrow></mrow></mtd></mtr><mtr><mtd><mrow><mrow><mrow><mn>0.95</mn><mo>*</mo><msubsup><mi>E</mi><mi>L</mi><mrow><mi>j</mi><mo>-</mo><mn>1</mn></mrow></msubsup></mrow><mo>+</mo><mrow><mn>0.05</mn><mo>*</mo><msup><mi>total_band</mi><mi>j</mi></msup></mrow></mrow><mo>,</mo></mrow></mtd><mtd><mrow><mi>BOS</mi><mo>=</mo><mrow><mrow><mn>0</mn><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>and</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><msup><mi>total_band</mi><mi>j</mi></msup></mrow><mo>≤</mo><msup><mi>total_band</mi><mrow><mi>j</mi><mo>-</mo><mn>1</mn></mrow></msup></mrow></mrow></mtd></mtr><mtr><mtd><mrow><mrow><mrow><mn>0.999</mn><mo>*</mo><msubsup><mi>E</mi><mi>L</mi><mrow><mi>j</mi><mo>-</mo><mn>1</mn></mrow></msubsup></mrow><mo>+</mo><mrow><mn>0.001</mn><mo>*</mo><msup><mi>total_band</mi><mi>j</mi></msup></mrow></mrow><mo>,</mo></mrow></mtd><mtd><mrow><mi>BOS</mi><mo>=</mo><mrow><mrow><mn>1</mn><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>and</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><msup><mi>total_band</mi><mi>j</mi></msup></mrow><mo>></mo><msup><mi>total_band</mi><mrow><mi>j</mi><mo>-</mo><mn>1</mn></mrow></msup></mrow></mrow></mtd></mtr><mtr><mtd><mrow><mrow><mrow><mn>0.992</mn><mo>*</mo><msubsup><mi>E</mi><mi>L</mi><mrow><mi>j</mi><mo>-</mo><mn>1</mn></mrow></msubsup></mrow><mo>+</mo><mrow><mn>0.008</mn><mo>*</mo><msup><mi>total_band</mi><mi>j</mi></msup></mrow></mrow><mo>,</mo></mrow></mtd><mtd><mrow><mi>BOS</mi><mo>=</mo><mrow><mrow><mn>1</mn><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>and</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><msup><mi>total_band</mi><mi>j</mi></msup></mrow><mo>≤</mo><msup><mi>total_band</mi><mrow><mi>j</mi><mo>-</mo><mn>1</mn></mrow></msup></mrow></mrow></mtd></mtr></mtable></mrow></mrow></math></maths><br /> where E<sub>L</sub><sup>j </sup>is the energy level estimate of current speech frame j and E<sub>L</sub><sup>j−1 </sup>is the energy level estimate of previous speech frame j−1.
0090The impact of the frame energy total_band<sup>j </sup>of the current speech frame j is thus lower, when the frame energy total_band<sup>j </sup>of the current speech frame j is higher than the frame energy total_band<sup>j−1 </sup>of the preceding speech frame j−1 as compared to when the frame energy total_band<sup>j </sup>of the current speech frame j is lower than or equal to the frame energy total_band<sup>j−1 </sup>of the preceding speech frame j−1. Further, the energy level estimate adapts more aggressively at the beginning of a speech sequence. Therefore the correct speech level estimation can be calculated faster, even if the speech sequence is short.
0091The updated energy level estimate E<sub>L</sub><sup>for j</sup><sup>th </sup>speech frame is only used, however, if the estimate E<sub>L</sub><sup>j </sup>lies clearly above the provided background noise estimate NE<sup>j</sup>, for instance when E<sub>L</sub><sup>j</sup>>2.5*NE<sup>j </sup>(step <b>503</b>). Otherwise, the energy level estimate E<sub>L</sub><sup>j−1 </sup>for the previous speech frame j−1 is used as well as the energy level E<sub>L</sub><sup>j </sup>for the current speech frame j.
0092The energy level estimate E<sub>L</sub><sup>j </sup>for the current speech frame j is then provided to the scaling portion <b>364</b> for performing a scaling of the frame energy total_band<sup>j </sup>provided by the frame energy detector <b>361</b>, if required.
0093<figref idref="DRAWINGS">FIG. 6</figref> is a diagram illustrating the frame energies for the same speech sequence with four different speech energy levels, namely −16 dBov, −22 dBov, −26 dBov and −32 dBov. The respective frame energies total_band<sup>j </sup>in [dBov] are depicted roughly over speech frames j=1030 to 1130, indicated on the x-axis. As can be seen from <figref idref="DRAWINGS">FIG. 6</figref>, the frame energy varies significantly from frame to frame.
0094<figref idref="DRAWINGS">FIG. 7</figref> is a diagram illustrating the energy levels E<sub>L</sub><sup>j </sup>in [dBov] estimated for the four speech sequences of <figref idref="DRAWINGS">FIG. 6</figref>, with the respective speech frame j=1 to 6000 indicated on the x-axis.
0095In order to eliminate the impact of a deviation of the speech energy level from a nominal speech level onto the frame energy, the scaling portion <b>364</b> scales the frame energy into a nominal speech level set to approximately −26 dBov.
0096The scaling portion <b>364</b> determines to this end first whether the energy level estimate E<sub>L</sub><sup>j </sup>for the current speech frame j is close to the nominal speech level (step <b>504</b>).
0097If the energy level E<sub>L</sub><sup>j </sup>is close to the nominal speech level, no frame energy scaling is performed. Instead, the frame energy is provided directly to the mode selector <b>365</b> (step <b>505</b>).
0098If the energy level E<sub>L</sub><sup>j </sup>is not close to the nominal speech level, in contrast, a correction function for scaling the frame energy is selected, depending on the difference between the nominal energy level and the energy level estimate E<sub>L</sub><sup>j </sup>for the current speech frame j. The correction function is selected such that if the energy level is lower than the nominal energy level, the frame energy is scaled upwards, and if the energy level is higher than the nominal energy level, the frame energy is scaled downwards (step <b>506</b>).
0099A new frame energy is then calculated by applying the determined correction function, and the scaled frame energy is provided to the mode selector <b>365</b> (step <b>507</b>).
0100The following algorithm is a simple example for realizing steps <b>504</b> to <b>507</b>:
0101As long as the flag BOS is equal to one (BOS=1), that is, at the beginning of a speech sequence, no scaling is performed. If the flag BOS is equal to zero (BOS=0), the following algorithm is performed: <tables id="TABLE-US-00002" num="2"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="14PT" align="left" /><colspec colname="1" colwidth="203PT" align="left" /><thead><row><entry /><entry /></row><row><entry /><entry namest="OFFSET" nameend="1" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /><entry>/* If energy level estimate < −28dBov */</entry></row><row><entry /><entry>if (E<sub>L </sub>< Energy_Level[1])</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>/* If −32dBov < energy level estimate < −28dBov*/</entry></row><row><entry /><entry>if E<sub>L </sub>> Energy_Level [0]</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="42PT" align="left" /><colspec colname="1" colwidth="175PT" align="left" /><tbody valign="top"><row><entry /><entry>i = 3</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>/* If energy level estimate < −32dBov */</entry></row><row><entry /><entry>else</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="42PT" align="left" /><colspec colname="1" colwidth="175PT" align="left" /><tbody valign="top"><row><entry /><entry>i = 4</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>/* Frame energy is scaled towards nominal level*/</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="70PT" align="left" /><colspec colname="2" colwidth="119PT" align="left" /><tbody valign="top"><row><entry /><entry>total_band<sub>Scaled </sub>=</entry><entry>M[i][2]*total_band*total_band +</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="98PT" align="left" /><colspec colname="1" colwidth="119PT" align="left" /><tbody valign="top"><row><entry /><entry>M[i][1]*total_band + M[i][0]</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="14PT" align="left" /><colspec colname="1" colwidth="203PT" align="left" /><tbody valign="top"><row><entry /><entry>/* If energy level estimate > −24dBov */</entry></row><row><entry /><entry>else if E<sub>L </sub>> Energy_Level[2]</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>/* If −18dBov > energy level estimate > −24dBov*/</entry></row><row><entry /><entry>if E<sub>L </sub>< Energy_Level[3]</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="42PT" align="left" /><colspec colname="1" colwidth="175PT" align="left" /><tbody valign="top"><row><entry /><entry>i = 2</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>/* If energy level estimate > −18dBov*/</entry></row><row><entry /><entry>else</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="42PT" align="left" /><colspec colname="1" colwidth="175PT" align="left" /><tbody valign="top"><row><entry /><entry>i = 1;</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>/* Frame energy is scaled towards nominal level*/</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="70PT" align="left" /><colspec colname="2" colwidth="119PT" align="left" /><tbody valign="top"><row><entry /><entry>total_band<sub>Scaled </sub>=</entry><entry>M[i][2]*total_band*total_band +</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="98PT" align="left" /><colspec colname="1" colwidth="119PT" align="left" /><tbody valign="top"><row><entry /><entry>M[i][1]*total_band + M[i][0].</entry></row><row><entry /><entry namest="OFFSET" nameend="1" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0102The frame energy total_band is thus scaled by using a second order equation (a*x<sup>2</sup>+b*x+c) as follows: <br />total_band<sub>scaled</sub>=(<i>M</i>(<i>i,</i>3))*total_band<sup>2</sup><i>+M</i>(<i>i,</i>2)*total_band+<i>M</i>(<i>i,</i>1), <br /> where M is a matrix including the coefficients for the scaling equation. The indices 1, 2 and 3 point to the first, second and third column, respectively, of the matrix. The index i points to a respective row of the matrix, the value of i depending on the energy level estimate E<sub>L</sub><sup>j </sup>for the current speech frame j.
0103An exemplary scaling matrix M is give by: <maths id="MATH-US-00006" num="6"><math overflow="scroll"><mrow><mi>M</mi><mo>=</mo><mrow><mo>[</mo><mtable><mtr><mtd><mrow><mrow><mstyle><mtext>-</mtext></mstyle><mo></mo><mn>2.3740</mn></mrow><mo>,</mo></mrow></mtd><mtd><mrow><mn>0.3161</mn><mo>,</mo></mrow></mtd><mtd><mrow><mstyle><mtext>-</mtext></mstyle><mo></mo><mn>0.2863</mn></mrow></mtd></mtr><mtr><mtd><mrow><mrow><mstyle><mtext>-</mtext></mstyle><mo></mo><mn>1.2955</mn></mrow><mo>,</mo></mrow></mtd><mtd><mrow><mn>0.6309</mn><mo>,</mo></mrow></mtd><mtd><mrow><mrow><mstyle><mtext>-</mtext></mstyle><mo></mo><mn>0.0004</mn></mrow><mo>,</mo></mrow></mtd></mtr><mtr><mtd><mrow><mn>1.4239</mn><mo>,</mo></mrow></mtd><mtd><mrow><mn>1.4527</mn><mo>,</mo></mrow></mtd><mtd><mrow><mn>0.3637</mn><mo>,</mo></mrow></mtd></mtr><mtr><mtd><mrow><mn>2.7193</mn><mo>,</mo></mrow></mtd><mtd><mrow><mn>1.7620</mn><mo>,</mo></mrow></mtd><mtd><mn>0.4832</mn></mtd></mtr></mtable><mo>]</mo></mrow></mrow></math></maths>
0104The first row of the matrix M, used for i=1, includes the coefficients for a speech level exceeding −18 dBov. The second row of the matrix M, used for i=2, includes the coefficients for a speech level between −18 dBov and −24 dBov. The third row of the matrix M, used for i=3, includes the coefficients for a speech level between −32 dBov and −28 dBov. The fourth row of the matrix M, used for i=4, includes the coefficients for a speech level falling short of −32 dBov. No scaling is performed, when the energy level estimate E<sub>L</sub><sup>j </sup>lies between −24 dBov and −28 dBov, that is, in case it lies around the nominal level of −26 dBov.
0105<figref idref="DRAWINGS">FIG. 8</figref> is a diagram illustrating the scaling, in which the y-axis describes an amount of energy and in which the x-axis indicates the number of frames. Five curves are depicted for a speech sequence having a speech energy level of −16 dBov, −28 dBov, −22 dBov, −26 dBov and −32 dBov, respectively. For each speech energy level, the speech frames have been sorted to have ascending frame energies. In addition, a curve representing the nominal speech energy level of approximately −26 dBov is indicated. According to the presented embodiment of the invention, the frame energies are scaled to a nominal speech level around −26 dBov by exploiting the correction functions designed for each speech level. The frame energies of the curve having a speech energy level of −26 dBov are not scaled, because they are very close to the frame energies of the nominal speech level. Once the correction function specified for each speech level has been applied to the frame energies of the other four curves, the frame energies of these curves are scaled towards the frame energies of the curve for the nominal speech energy level.
0106The mode selector <b>365</b> compares the, potentially scaled, frame energy with a low energy threshold. Since the frame energy is scaled depending on the energy level estimate for the speech frame, it is independent of the general speech energy level, and it can be compared to a fixed low energy threshold (step <b>508</b>).
0107More specifically, the mode selector <b>365</b> has an threshold vector Energy_Level: <maths id="MATH-US-00007" num="7"><math overflow="scroll"><mrow><mrow><mi>Energy_Level</mi><mo>=</mo><mrow><mo>[</mo><mtable><mtr><mtd><mi>A</mi></mtd></mtr><mtr><mtd><mi>B</mi></mtd></mtr><mtr><mtd><mi>C</mi></mtd></mtr><mtr><mtd><mi>D</mi></mtd></mtr></mtable><mo>]</mo></mrow></mrow><mo>,</mo></mrow></math></maths><br /> where A, B, C and D are various energy level thresholds. In the above presented algorithm, for example, these thresholds are A=energy level[0]=−32 dBov, B=energy level[1]=−28 dBov, C=energy level[2]=−24 dBov and D=energy level[3]=−18 dBov.
0108If the scaled frame energy total_band<sub>Scaled </sub>is smaller than the low energy threshold (step <b>509</b>), a low energy sequence is detected and a low rate coding mode is selected (step <b>510</b>).
0109If the scaled frame energy total_band<sub>Scaled </sub>is larger than the low energy threshold (step <b>509</b>), a low energy sequence is not detected, and a further adaptation algorithm is called for determining the appropriate coding mode (step <b>511</b>). The further adaptation algorithm can be for example the rate detection algorithm <b>37</b>. The rate detection algorithm <b>37</b> may receive as well uplink codec mode commands (MC<sub>u</sub>) from the mobile communication network embedded in received signals, as indicated above with reference to <figref idref="DRAWINGS">FIG. 1</figref>.
0110The mode selector <b>365</b> then provides the selected coding mode to the speech encoder <b>34</b>, which finishes the encoding of the current frame j making use of the selected coding mode. The encoding comprises a calculation of LP coefficients in the LPC calculation component <b>341</b>, a calculation of LTP parameters in the LTP calculation component <b>342</b> and a determination of fixed codebook parameters for an LPC excitation. As known from a conventional coding mode selection, it depends on the selected mode the processing of which portions <b>341</b>, <b>342</b>, <b>343</b> of the speech encoder are affected (step <b>512</b>).
0111For the respected next frame j+1, steps <b>502</b> to <b>512</b> are repeated in a loop, until the speech sequence has been encoded completely for transmission.
0112In practice, the low energy threshold A should be tuned to the point, where the information content of the signal starts to increase dramatically.
0113In <figref idref="DRAWINGS">FIG. 8</figref>, a ring <b>80</b> indicates the point where the information content is dramatically increased in the case of a nominal speech energy level, and which is thus selected as the low energy threshold. If the frame energies of the −16 dBov speech level, for example, are not scaled and the low energy threshold for the nominal speech energy level is employed nevertheless, a short low energy sequence is detected and hardly any bit rate reduction is achieved, because most of the low energy sequence is encoded by a higher bit rate. On the other hand, if the frame energies of the −32 dBov speech level, for example, are not scaled and the low energy threshold for the nominal speech energy level is employed nevertheless, a very long low energy sequence is detected and it is encoded by lower bit rates. Obviously a considerable bit rate reduction can be achieved, but it also degrades the speech quality dramatically. Due to the scaling to the nominal speech energy level, in contrast, the low energy threshold for the nominal speech energy level is suitable for all speech energy levels. A vertical dashed line <b>81</b> through ring <b>80</b> shows the effective low energy thresholds which is applied for the various original speech energy levels.
0114It is to be understood that the source based rate adaptation presented for a mobile station <b>30</b> could equally be implemented in other electronic devices, for instance, in the multi-rate speech encoder <b>122</b> of the transcoder <b>12</b> of <figref idref="DRAWINGS">FIG. 1</figref>.
0115While there have been shown and described and pointed out fundamental novel features of the invention as applied to preferred embodiments thereof, it will be understood that various omissions and substitutions and changes in the form and details of the devices and methods described may be made by those skilled in the art without departing from the spirit of the invention. For example, it is expressly intended that all combinations of those elements and/or method steps which perform substantially the same function in substantially the same way to achieve the same results are within the scope of the invention. Moreover, it should be recognized that structures and/or elements and/or method steps shown and/or described in connection with any disclosed form or embodiment of the invention may be incorporated in any other disclosed or described or suggested form or embodiment as a general matter of design choice. It is the intention, therefore, to be limited only as indicated by the scope of the claims appended hereto.
Contents5
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US2009144062A1 | Cited by | United States of America | Pre-grant |
| US2011066429A1 | Cited by | United States of America | Pre-grant |
| US8248935B2 | Cited by | United States of America | Search report |
| US2011112844A1 | Cited by | United States of America | Pre-grant |
| US2007165644A1 | Cited by | United States of America | Pre-grant |
| US2010057449A1 | Cited by | United States of America | Pre-grant |
| US9749741B1 | Cited by | United States of America | Search report |
| US2011264449A1 | Cited by | United States of America | Pre-grant |
| US9135925B2 | Cited by | United States of America | Search report |
| US9135926B2 | Cited by | United States of America | Search report |
| US9142222B2 | Cited by | United States of America | Search report |
| US2011035213A1 | Cited by | United States of America | Pre-grant |
| US2006069553A1 | Cited by | United States of America | Pre-grant |
| US7860509B2 | Cited by | United States of America | Search report |
| US8909522B2 | Cited by | United States of America | Search report |
| WO2009009522A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US8688441B2 | Cited by | United States of America | Applicant |
| US2010198587A1 | Cited by | United States of America | Pre-grant |
| US8463599B2 | Cited by | United States of America | Applicant |
| US2010049342A1 | Cited by | United States of America | Pre-grant |
| CN106796801A | Cited by | China | Search report |
| US8527283B2 | Cited by | United States of America | Applicant |
| US10418052B2 | Cited by | United States of America | Applicant |
| US9818433B2 | Cited by | United States of America | Applicant |
| US2009198498A1 | Cited by | United States of America | Pre-grant |
| US2009201983A1 | Cited by | United States of America | Pre-grant |
| US8990073B2 | Cited by | United States of America | Search report |
| US8433582B2 | Cited by | United States of America | Applicant |
| US10586557B2 | Cited by | United States of America | Applicant |
| US2013066627A1 | Cited by | United States of America | Pre-grant |
| US2013073282A1 | Cited by | United States of America | Pre-grant |
| US8463412B2 | Cited by | United States of America | Applicant |
| US9990938B2 | Cited by | United States of America | Applicant |
| US9773511B2 | Cited by | United States of America | Search report |
| GB2450886B | Cited by | United Kingdom | Search report |
| US2011112845A1 | Cited by | United States of America | Pre-grant |
| US2003009325A1 | Cites | United States of America | Pre-grant |
| US2004064309A1 | Cites | United States of America | Pre-grant |
| US5742734A | Cites | United States of America | Pre-grant |
| US5778338A | Cites | United States of America | Pre-grant |
| US5854845A | Cites | United States of America | Pre-grant |
| US6003004A | Cites | United States of America | Pre-grant |
| US6104993A | Cites | United States of America | Pre-grant |
| US6535846B1 | Cites | United States of America | Pre-grant |
| US6647366B2 | Cites | United States of America | Pre-grant |
| US7054809B1 | Cites | United States of America | Pre-grant |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 9940805 | United States of America | A | |
| US20050099408 | – | – | – |
45 transactions on the USPTO file
Abandoned after 2 non-final rejections, 1 final rejection and 1 RCE.
- Non-final rejections
- 2
- Final rejections
- 1
- RCEs
- 1
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Mail Abandonment for Failure to Respond to Office ActionAbandonedMABN2 | MABN2 | |
| Aband. for Failure to Respond to O. A.AbandonedABN2 | ABN2 | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| to Close the A/R Record and Reset the Status for Expired Suspensions.EOSP | EOSP | |
| Mail Letter Suspending Prosecution at Applicant's RequestMAISP | MAISP | |
| Suspension Letter- Applicant InitiatedAISP | AISP | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Letter Requesting Suspension of ProsecutionM856 | M856 | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Mail Advisory Action (PTOL - 303)MCTAV | MCTAV | |
| Advisory Action (PTOL-303)CTAV | CTAV | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| IFW TSS Processing by Tech Center CompleteTSSCOMP | TSSCOMP | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Application Is Now CompleteCOMP | COMP | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Payment of additional filing fee/PreexamFLFEE | FLFEE | |
| A statement by one or more inventors satisfying the requirement under 35 USC 115, Oath of the ApplicOATHDECL | OATHDECL | |
| Notice Mailed--Application Incomplete--Filing Date AssignedINCD | INCD | |
| Cleared by OIPE CSRL194 | L194 | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Initial Exam Team nnIEXX | IEXX |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Information on status: application discontinuationABANDONED -- FAILURE TO RESPOND TO AN OFFICE ACTIONSTCB | STCB | |
| AssignmentAS | AS |
Numbers
- Publication
- 20060224381
- Publication, DOCDB
- 2006224381
- Publication, EPODOC
- US2006224381
- Application
- 11099408
- Application, DOCDB
- 9940805
- Application, EPODOC
- US20050099408
Titles
- English
- Detecting speech frames belonging to a low energy sequence
Classification
- CPC, 2
- G10L19/20
- G10L25/78
- IPC, 1
- G10L19 12
- USPC, 3
- 704223000
- 704E11003
- 704E19042