Method and device for estimating the tonality of a sound signal
Abstract
A method for estimating a tone of a sound signal, in which the procedure comprises: calculating a current residual spectrum of the sound signal; detect the peaks in the current residual spectrum; calculate a correlation map between the current residual spectrum and a previous residual spectrum for each peak detected; and calculate a long-term correlation map based on the calculated correlation map, in which the long-term correlation map is indicative of a tone in the sound signal.
Term
1.7 yearsto projected expiry
Projected expiry 20 June 2028, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
27 claims: 4 independent, 23 dependent
- 1ES 2 533 358 T3 REIVINDICACIONES 1. Un procedimiento para estimar una tonalidad de una señal de sonido, en el que el procedimiento comprende:calcular un espectro residual actual de la señal de sonido;detectar los picos en el espectro residual actual;calcular un mapa de correlación entre el espectro residual actual y un espectro residual previo para cada pico detectado;y calcular un mapa de correlación a largo plazo basado en el mapa de correlación calculado, en el que el mapa de correlación a largo plazo es indicativo de una tonalidad en la señal de sonido.
- 2Procedimiento según la reivindicación 1, en el que el cálculo del espectro residual actual comprende:buscar los mínimos en el espectro de la señal de sonido en una trama actual;estimar un suelo espectral conectando los mínimos entre sí;y restar el suelo espectral estimado del espectro de la señal de sonido en la trama actual para producir el espectro residual actual.
- 3Procedimiento según la reivindicación 1 o 2, en el que la detección de los picos en el espectro residual actual comprende localizar un máximo entre cada par de dos mínimos consecutivos.
- 4Procedimiento según la reivindicación 1, 2 o 3, en el que el cálculo del mapa de correlación comprende:para cada pico detectado en el espectro residual actual, calcular un valor de correlación normalizado con el espectro residual anterior, sobre los contenedores de frecuencia entre dos mínimos consecutivos en el espectro residual actual que delimitan el pico;y asignar una puntuación a cada pico detectado, en el que la puntuación corresponde al valor de correlación normalizado;y para cada pico detectado, asignar el valor de correlación normalizado del pico sobre los contenedores de frecuencia entre los dos mínimos consecutivos que delimitan el pico para formar el mapa de correlación.
- 5Procedimiento según cualquiera de las reivindicaciones anteriores, en el que el cálculo del mapa de correlación a largo plazo comprende:filtrar el mapa de correlación a través de un filtro de un polo de contenedor de frecuencias en contenedor de frecuencias;y sumar el mapa de correlación filtrado sobre los contenedores de frecuencia para producir un mapa de correlación sumado a largo plazo.
- 6Procedimiento para detectar actividad sonora en una señal de sonido, en el que la señal de sonido es clasificada como una de entre una señal de sonido inactiva y una señal de sonido activa según la actividad sonora detectada en la señal de sonido, en el que el procedimiento comprende:estimar un parámetro relacionado con una tonalidad de la señal de sonido usada para distinguir una señal musical de una señal de ruido de fondo;en el que la estimación del parámetro relacionado con la tonalidad de la señal de sonido previene la actualización de las estimaciones de energía de sonido cuando se detecta una señal musical;en el que la estimación de tonalidad es realizada según una cualquiera de las reivindicaciones 1 a 5.
- 7Procedimiento según la reivindicación 6, que comprende además calcular un parámetro no estacionariedad complementaria y un parámetro carácter de ruido con el fin de distinguir una señal musical de una señal de ruido de fondo y evitar la actualización de las estimaciones de energía de ruido en la señal musical.
- 8Procedimiento según la reivindicación 7, en el que el cálculo del parámetro no estacionariedad complementaria comprende calcular un parámetro similar a un no estacionariedad convencional con restablecimiento de energía a largo plazo cuando se detecta un ataque espectral.
- 9Procedimiento según la reivindicación 8, en el que la detección del ataque espectral y el restablecimiento de la energía a largo plazo comprende calcular un parámetro diversidad espectral y en el que el cálculo del parámetro diversidad espectral comprende:ES 2 533 358 T3 calcular una relación entre una energía de la señal de sonido en una trama actual y una energía de la señal de sonido en un trama previa, para las bandas de frecuencia más altas que un número determinado;y calcular la diversidad espectral como una suma ponderada de la relación calculada sobre todas las bandas de frecuencia más altas que el número determinado.
- 10Procedimiento según la reivindicación 8 o 9, en el que el cálculo del parámetro carácter de ruido comprende:dividir una pluralidad de bandas de frecuencia en un primer grupo de un cierto número de primeras bandas de frecuencia y un segundo grupo de un resto de las bandas de frecuencia;calcular un primer valor de energía para el primer grupo de bandas de frecuencia y un segundo valor de energía del segundo grupo de bandas de frecuencias;calcular una relación entre los valores de energía primero y segundo para producir el parámetro carácter de ruido;y calcular un valor a largo plazo del parámetro carácter ruido en base al parámetro carácter de ruido calculado;en el que la actualización de las estimaciones de energía de ruido se evita si el parámetro carácter de ruido es menor que un umbral fijo determinado.
- 11Un procedimiento de clasificación de una señal de sonido con el fin de optimizar la codificación de la señal de sonido usando la clasificación de la señal de sonido, en el que el procedimiento comprende:detectar una actividad sonora en la señal de sonido;clasificar la señal de sonido como una de entre una señal de sonido inactivo y una señal de sonido activo según la actividad sonora detectada en la señal de sonido;y en respuesta a la clasificación de la señal de sonido como una señal de sonido activo, clasificar adicionalmente la señal de sonido activo como una de entre una señal de voz sorda y una señal de voz no sorda;en el que la clasificación de la señal de sonido activo como una señal de voz sorda comprende la estimación de una tonalidad de la señal de sonido con el fin de evitar la clasificación de las señales musicales como señales de voz sorda, en el que la estimación de tonalidad es realizada según una cualquiera de las reivindicaciones 1 a 5.
- 12Procedimiento según la reivindicación 11, que comprende además codificar la señal de sonido según la clasificación de la señal de sonido, en el que la codificación de la señal de sonido según la clasificación de la señal de sonido comprende codificar la señal de sonido inactivo usando generación de ruido de confort.
- 13Procedimiento según la reivindicación 11 o 12, en el que la clasificación de la señal de sonido activo como una señal de voz sorda comprende calcular una regla de decisión en base a al menos una de entre una medida de sonoridad, una medida de inclinación espectral media, un aumento máximo de energía de corto tiempo a bajo nivel, una estabilidad tonal y una energía relativa de trama.
- 14Un procedimiento para codificar una banda superior de una señal de sonido usando una clasificación de la señal de sonido, en el que el procedimiento comprende:clasificar la señal de sonido como una de entre una señal de sonido tonal y una señal de sonido no tonal;en el que la clasificación de la señal de sonido como una señal tonal comprende estimar una tonalidad de la señal de sonido según una cualquiera de las reivindicaciones 1 a 5.
- 15Procedimiento según la reivindicación 14, en el que la estimación de la tonalidad de la señal de sonido según una cualquiera de las reivindicaciones 1 a 5 comprende además el uso de un procedimiento alternativo para calcular un suelo espectral, en el que el uso del procedimiento alternativo para calcular el suelo espectral comprende filtrar un espectro logarítmico de energía de la señal de sonido en una trama actual usando un filtro de media móvil.
- 16Procedimiento según la reivindicación 14 o 15, en el que la estimación de la tonalidad de la señal de sonido según una cualquiera de las reivindicaciones 1 a 5 comprende además suavizar el espectro residual por medio de un filtro de media móvil de tiempo corto.
- 17Procedimiento según la reivindicación 14 o 16, que comprende además codificar la banda superior de la señal de sonido según la clasificación de dicha señal de sonido. ES 2 533 358 T3
- 18Procedimiento según cualquiera de las reivindicaciones 14 a 17, en el que la banda superior de la señal de sonido comprende un intervalo de frecuencias por encima de 7 kHz.
- 19Un dispositivo para estimar una tonalidad de una señal de sonido, en el que el dispositivo comprende:un calculador para calcular un espectro residual actual de la señal de sonido;un detector para detectar los picos en el espectro residual actual;un calculador para calcular un mapa de correlación entre el espectro residual actual y un espectro residual previo para cada pico detectado;y un calculador para calcular un mapa de correlación a largo plazo en base al mapa de correlación calculado, en el que el mapa de correlación a largo plazo es indicativo de una tonalidad en la señal de sonido.
- 20Un dispositivo según la reivindicación 19, en el que el calculador del espectro residual actual comprende:un localizador de mínimos en el espectro de la señal de sonido en una trama actual;un estimador de un suelo espectral que conecta los mínimos entre sí;y un restador del suelo espectral estimado del espectro para producir el espectro residual actual.
- 21Dispositivo según la reivindicación 19 o 20, en el que el calculador del mapa de correlación a largo plazo comprende:un filtro para filtrar el mapa de correlación de contenedor de frecuencias en contenedor de frecuencias;y un sumador para sumar el mapa de correlación filtrado sobre los contenedores de frecuencia con el fin de producir un mapa sumado de correlación a largo plazo.
- 22Un dispositivo para detectar la actividad sonora en una señal de sonido, en el que la señal de sonido es clasificada como una de entre una señal de sonido inactivo y una señal de sonido activo según la actividad sonora detectada en la señal de sonido, en el que el dispositivo comprende:un estimador de tonalidad para la señal de sonido, usado para distinguir una señal musical de una señal de ruido de fondo;en el que el estimador de tonalidad comprende un dispositivo según una cualquiera de las reivindicaciones 19 a 21.
- 23Un dispositivo para clasificar una señal de sonido con el fin de optimizar la codificación de la señal de sonido usando la clasificación de la señal de sonido, en el que el dispositivo comprende:un detector para detectar una actividad sonora en la señal de sonido;un primer clasificador de señal de sonido para clasificar la señal de sonido como una de entre una señal de sonido inactivo y una señal de sonido activo según la actividad sonora detectada en la señal de sonido;un segundo clasificador de señal de sonido en conexión con el primer clasificador de sonido para clasificar la señal de sonido activo como una de entre una señal de voz sorda y una señal de voz no sorda;en el que el detector de actividad sonora comprende un estimador de tonalidad para estimar una tonalidad de la señal de sonido con el fin de evitar la clasificación de las señales musicales como señales de voz sorda en el que el estimador de tonalidad comprende una dispositivo según una cualquiera de las reivindicaciones 19 a 21.
- 24Dispositivo según la reivindicación 23, que comprende además un codificador de sonido para codificar la señal de sonido según la clasificación de la señal de sonido, en el que el codificador de sonido es seleccionado de entre el grupo que consiste en:un codificador de ruido para codificar las señales de sonido inactivas, un codificador optimizado para voz sorda, un codificador optimizado para voz sonora para codificar señales sonoras estables, y un codificador de señal de sonido genérico para codificar señales sonoras de evolución rápida.
- 25Un dispositivo para codificar una banda superior de una señal de sonido usando una clasificación de la señal de sonido, en el que el dispositivo comprende:un clasificador de señal de sonido para clasificar la señal de sonido como una de entre una señal de sonido tonal y una señal de sonido no tonal;y un codificador de sonido para codificar la banda superior de la señal de sonido clasificada;en el que el clasificador de señal de sonido comprende un dispositivo para estimar una tonalidad de la señal de sonido según una cualquiera de las reivindicaciones 19 a 21. ES 2 533 358 T3
- 26Dispositivo según la reivindicación 25, que comprende además un filtro de media móvil para calcular un suelo espectral derivado de la señal de sonido, en el que el suelo espectral se usa en la estimación de la tonalidad de la señal de sonido.
- 27Dispositivo según la reivindicación 25 o 26, que comprende además un filtro de media móvil de tiempo corto para suavizar un espectro residual de la señal de sonido, en el que el espectro residual se usa en la estimación de la tonalidad de la señal de sonido.
Independent claims27
359 paragraphs in 15 sections, as filed
ES 2 533 358 T3
DESCRIPTION
Procedure and device for estimating the tonality of a sound signal
Field of Invention
The present invention relates to the detection of sound activity, the estimation of background noise and the classification of the sound signal, where it is understood that sound is a useful signal. The present invention also relates to the sound activity detector, the background noise estimator and the corresponding sound signal classifier.
In particular, but not exclusively:
- Audible activity detection is used to select the frames to be encoded using techniques optimized for idle frames.
- Sound signal classifier is used to discriminate between different classes of voice and music signals to enable more efficient encoding of sound signals, i.e. optimized encoding of unvoiced speech signals, optimized encoding of voiced speech signals stable, and generic encoding of other sound signals.
- An algorithm is provided and uses various relevant parameters and characteristics to allow a better choice of encoding mode and a more robust estimation of background noise.
- Key estimation is used to improve the performance of the detection of sound activity in the presence of musical signals, and to better discriminate between unvoiced sounds and music. For example, tonality estimation can be used in a super wideband codec to decide the codec model to encode the signal above 7 kHz.
Background of the Invention
The demand for efficient, narrowband and broadband digital voice coding techniques with a good balance between subjective quality and bit rate is increasing in various application areas such as teleconferencing, multimedia and wireless communications. Until recently, the telephone bandwidth limited to a range of 200-3,400 Hz has been used mainly in voice coding applications (signal sampled at 8 kHz). However, broadband voice applications provide greater intelligibility and naturalness in communication compared to conventional telephone bandwidth. In broadband services, the input signal is sampled at 16 kHz and the encoded bandwidth is in the range of 50 to 7,000 Hz. This bandwidth has been found to be sufficient to provide good quality giving an impression. of almost a face-to-face communication. A further improvement in quality is achieved with the so-called super wideband, where the signal is sampled at 32 kHz and the encoded bandwidth is in the range of 50 to 15,000 Hz. For voice signals, this provides a face-to-face quality, since almost all the energy in the human voice is less than 14,000 Hz. This bandwidth also provides a considerable quality improvement over general audio signals including music (wide band is equivalent to AM radio and super wide band is equivalent to FM radio). Higher bandwidth has been used for general audio signals with the full band 20-20,000 Hz (CD quality sampled at 44.1 kHz or 48 kHz).
A sound encoder converts a sound signal (voice or audio) into a digital bit stream that is transmitted through a communication channel or stored on a storage medium. The sound signal is digitized, that is, it is sampled and quantized generally with 16 bits per sample. The sound encoder plays the role of representing these digital samples with a smaller number of bits while maintaining good subjective quality. The sound decoder operates the transmitted or stored bit stream on the stream and converts it back to a sound signal.
Coding based on Code-Excited Linear Prediction (CELP) is one of the best prior art techniques for achieving a good compromise between subjective quality and bit rate. This encoding technique is a foundation for various voice encoding standards, in both wireless and landline applications. In CELP encoding, the sampled speech signal is processed in successive blocks of L samples, generally called frames, where L is a predetermined number typically corresponding to 10-30 ms. A linear prediction filter (LP) is calculated and transmitted every frame. The frame of sample L is divided into smaller blocks called subframes. In each subframe, a drive signal is typically derived from two components, the past drive and the groundbreaking, fixed codebook drive. The component formed from past excitation is often referred to as an adaptive codebook or pitch excitation. The parameters that characterize the excitation signal are encoded and transmitted to the decoder, where the reconstructed excitation signal is used.
ES 2 533 358 T3 as the input to the LP filter.
The use of source controlled Variable Bit Rate (VBR) speech coding greatly improves system capacity. In source-controlled VBR encoding, the codec uses a signal classification module and an optimized encoding model is used to encode each speech frame based on the nature of the speech frame (e.g. voiced, unvoiced, transient , background noise). Also, different bit rates can be used for each class. The simplest form of source controlled VBR encoding is to use Voice Activity Detection (VAD) and encode the idle speech frames (background noise) at a very low bit rate. In addition, Discontinuous transmission (DTX) can be used where no data is transmitted in the case of stable background noise. The decoder uses Comfort Noise Generation (CNG) to generate the background noise characteristics. VAD / DTX / CNG results in a considerable reduction in the average bit rate and in packet switched applications considerably reduces the number of packets routed. VAD algorithms work well with voice signals, but can cause serious problems with musical signals. Music signal segments can be classified as unvoiced signals and therefore can be encoded with a model optimized for unvoiced signals that seriously affects the quality of the music. Also, some segments of stable music signals can be classified as stable background noise and this can cause background noise updating in the VAD algorithm, resulting in degradation of algorithm performance. Therefore, it would be advantageous to extend the VAD algorithm to better discriminate musical signals. In the present description, this algorithm will be referred to as the Sound Activity Detection (SAD) algorithm in which the sound could be speech or music or any useful signal. The present description also describes a key detection method used to improve the performance of the SAD algorithm in the case of musical signals.
Another aspect in speech and audio coding is the concept of built-in coding, also known as layered coding. In embedded encoding, the signal is encoded in a first layer to produce a first bit stream, and then the error between the original signal and the first layer encoded signal is further encoded to produce a second bit stream. This can be repeated for more layers by encoding the error between the original signal and the encoded signal of all the previous layers. Bit streams from all layers are concatenated for transmission. The advantage of layered encoding is that parts of the bit stream (corresponding to the upper layers) can be discarded in the network (for example, in case of congestion) while still being able to decode the signal at the receiver depending on the number of layers received. Layer encoding is also useful in multicast applications where the encoder produces the bit stream for all layers and the network decides to send different bit rates to different endpoints based on the bit rate available on each link. .
Embedded or layered coding can also be useful to improve the quality of existing, widely used codecs, while still maintaining interoperability with these codecs. Adding more layers to the standard codec core layer can improve the quality and even increase the bandwidth of the encoded audio signal. Examples are the recently standardized ITU-T Recommendation G.729.1, where the core layer is interoperable with the widely used narrowband G.729 standard at 8 kbit / s and higher layers produce bit rates up to 32 kbit / s. s (with a broadband signal from 16 kbit / s). The goal of current standardization work is to add more layers to produce a super wideband codec (14 kHz bandwidth) and stereo extensions. Another example is ITU-T Recommendation G.718 for encoding super-wideband signals at 8, 12, 16, 24 and 32 kbit / s. The codec is also being expanded to encode stereo and super-wideband signals with higher bit rates.
The requirements for built-in codecs typically demand good quality for both audio signals and voice signals. Because speech can be encoded at relatively low bit rates using a model-based approach, the first layer (or the first two layers) is encoded (or encoded) using a technique specific to speech and signal. error for the upper layers is encoded using a more generic audio encoding technique. This provides good speech quality at low bit rates and good audio quality as the bit rate is increased. In G.718 and G.729.1, the first two layers are based on the ACELP (Algebraic Code-Excited Linear Prediction) technique which is suitable for encoding speech signals. In the upper layers, appropriate transform-based encoding is used for audio signals to encode the error signal (the difference between the original signal and the output of the first two layers). The well-known MDCT transform (Modified Discrete Cosine Transform) is used, in which the error signal is transformed in the frequency domain. In the super-wideband layers, the signal above 7 kHz is encoded using a generic coding model or a tonal coding model. The tonality detection indicated above can also be used to select the appropriate coding pattern to be used.
ES 2 533 358 T3
An example of a method known apparatus for determining the tonality of an input audio signal is described in patent document US 2004/181393 A1.
Compendium of the Invention
According to a first aspect of the present invention, a method is provided for estimating a tonality of a sound signal. The method comprises: calculating a current residual spectrum of the sound signal; detect peaks in the current residual spectrum; calculating a correlation map between the current residual spectrum and a previous residual spectrum for each detected peak; and calculating a long-term correlation map based on the calculated correlation map, wherein the long-term correlation map is indicative of a tonality in the sound signal.
According to a second aspect of the present invention, a device is provided for estimating a tonality of a sound signal. The device comprises: a calculator of a current residual spectrum of the sound signal; a detector for detecting peaks in the current residual spectrum; a calculator for calculating a correlation map between the current residual spectrum and a previous residual spectrum for each detected peak; and a calculator for calculating a long-term correlation map based on the calculated correlation map, wherein the long-term correlation map is indicative of a tonality in the sound signal.
The foregoing and other objects, advantages and features of the present invention will become more apparent upon reading the following non-restrictive description of an illustrative embodiment thereof, provided by way of example only with reference to the accompanying drawings.
Brief description of the drawings
In the attached drawings:
Figure 1 is a schematic block diagram of a portion of an exemplary sound communication system including sound activity detection, background noise estimate update, and sound signal classification;
Figure 2 is a non-limiting illustration of the use of windows in spectral analysis;
Figure 3 is a non-restrictive graphic illustration of the spectral soil calculation principle and residual spectrum;
Figure 4 is a non-limiting illustration of spectral correlation map calculation in a current frame; Figure 5 is an example functional block diagram of a signal classification algorithm; and Figure 6 is an example decision tree for unvoiced speech discrimination.
Detailed description
In the illustrative, non-limiting embodiment of the present invention, sound activity detection (SAD) is performed within a sound communication system to classify short-time frames of the signals as sound or background noise / silence. Sound activity detection is based on a frequency dependent signal-to-noise ratio (SNR) and uses an estimated background noise energy for each critical band. A decision about updating the background noise estimator is based on various parameters, including parameters that discriminate between background noise / silence and music, thus preventing the updating of the background noise estimator in music signals.
SAD corresponds to a first stage of signal classification. This first stage is used to discriminate idle frames for optimized idle signal encoding. In a second stage, the unvoiced speech frames are discriminated for an optimized encoding of an unvoiced signal. In this second stage, music detection is added in order to prevent the music from being classified as a deaf signal. Finally, in a third stage, the sound signals are discriminated by means of a further examination of the frame parameters.
The techniques described herein can be implemented with narrow band sound signals (Narrow Band, NB) sampled at 8,000 samples / s or wideband sound signals (Wide Band, WB) sampled at 16,000 samples / s, or any other sampling frequency. The encoder used in the non-limiting illustrative embodiment of the present invention is based on the AMR-WB codecs [AMR Wideband Speech Codec: Transcoding Functions, 3GPP Technical Specification TS 26.190 (http://www.3gpp.org)] and VMR-WB [Source-Controlled Variable-Rate Multimode Wideband Speech Codec (VMR-WB), Service Options 62 and 63 for Spread Spectrum Systems , 3GPP2 Technical Specification C.S0052-A v1.0, April 2005 (http://www.3gpp2.org)] which use an internal sample conversion to convert the signal sample rate to 12,800 samples / s (working in 6.4 kHz bandwidth). Thus, the sound activity detection technique in the illustrative non-restrictive embodiment operates on narrowband or broadband signals after
ES 2 533 358 T3 of a sampling conversion at 12.8 kHz.
Figure 1 is a block diagram of a sound communication system 100 in accordance with the illustrative, non-limiting embodiment of the invention, including sound activity detection.
The sound communication system 100 of Figure 1 comprises a pre-processor 101. The pre-processing performed by the module 101 can be performed as described in the following example (high-pass filtering, resampling, and pre-emphasis).
Before frequency conversion, the input sound signal is filtered with a high pass filter. In this illustrative, non-restrictive embodiment, the cutoff frequency of the high pass filter is 25 Hz for WM and 100 Hz for NB. The high pass filter serves as a precaution against unwanted low frequency components. For example, the following transfer function can be used:
<img file="ES2533358T3_D0001.tif" />
where, for WB, bo = 0.9930820, bi = -1.98616407, b2 = 0.9930820, ai = -1.9861162, a2 = 0.9862119292 and, for NB, bo = 0.945976856, bi = -1.891953712, b2 = 0.945976856, ai = -1.889033079, a2 = 0.894874345. Obviously, high-pass filtering can alternatively be carried out after resampling at 12.8 kHz.
In the case of WB, the input sound signal is decimated from 16 kHz to 12.8 kHz. Decimation is performed by an oversampler that oversamples the sound signal by a factor of 4. The resulting output is then filtered through a FIR (Finite Impulse Response) low pass filter with a cutoff frequency. 6.4 kHz. The signal filtered with the low pass filter is then subsampled by a factor of 5 by an appropriate subsampler. The filtering delay is 15 samples at a sampling frequency of 16 kHz.
In the case of NB, the sound signal is oversampled from 8 kHz to 12.8 kHz. For that purpose, an oversampler oversamples by a factor of 8 on the sound. The resulting output is then filtered through a low pass FIR filter with a cutoff frequency of 6.4 kHz. A subsampler then subsamples the low-pass filtered signal by a factor of 5. The filtering delay is 16 samples at a sample rate of 8 kHz.
After sample conversion, a pre-emphasis is applied to the sound signal before the encoding procedure. In pre-emphasis, a first-order high-pass filter is used to emphasize the higher frequencies. This first-order high-pass filter forms a pre-emphasizer and uses, for example, the following transfer function:
= 1-0.68Z- '
Pre-emphasis is used to improve codec performance at high frequencies and to improve perceptual weighting in the error minimization procedure used in the encoder.
As described above, the input sound signal is converted to a sampling frequency of 12.8 kHz and is pre-processed, for example, as described above. However, the techniques described can be applied equally to signals at other sampling frequencies, such as 8 kHz or 16 kHz, with a different preprocessing or without preprocessing.
In the non-limiting illustrative embodiment of the present invention, the encoder 109 (Figure 1) using sound activity detection operates on 20 ms frames containing 256 samples at the 12.8 kHz sample rate. In addition, the encoder 109 uses a 10 ms look-ahead of the future frame to perform its analysis (Figure 2). Sound activity detection follows the same frame structure.
Referring to Figure 1, spectral analysis is performed on spectral analyzer 102. Two are made
ES 2 533 358 T3 analysis on each frame using 20 ms windows with 50% overlap. The principle of windows is illustrated in Figure 2. The signal energy is calculated for the frequency bins and critical bands [JD Johnston, Transform coding of audio signal using perceptual noise criteria, IEEE J. Select. Areas Commun., Vol. 6, pp. 314-323, February 1988].
Sound activity detection (first stage of signal classification) is performed in sound activity detector 103 using noise energy estimates calculated in the previous frame. The output of the sound activity detector 103 is a binary variable that is further used by the encoder 109 and that determines whether the current frame is encoded as active or inactive.
The noise estimator 104 updates a noise estimate downwards (first level of estimation and noise update), that is, if in a critical band the energy of the frame is lower than an estimated energy of the background noise, the energy of the noise estimate is updated in that critical band.
Noise reduction is optionally applied by an optional noise reducer 105 to the speech signal using, for example, a spectral subtraction procedure. An example of such a noise reduction scheme is described in [M. Jelinek and R. Salami, Noise Reduction Method for Wideband Speech Coding, in Proc. EUSIPCO, Vienna, Austria, September 2004].
An LP analyzer and pitch tracker 106 perform linear prediction (LP) analysis and open loop pitch analysis (typically as part of the speech coding algorithm). In this non-restrictive illustrative embodiment, the resulting parameters from the LP analyzer and tone tracker 106 are used in the decision to update the noise estimates in the critical bands as performed in module 107. Alternatively, the noise activity detector 103 can also be used to make the noise update decision. According to a further alternative, the functions implemented by the LP analyzer and the tone tracker 106 may be an integral part of the sound coding algorithm.
Before updating the noise energy estimates in module 107, a music detection is performed to prevent a false update on the active music signals. Music detection uses spectral parameters calculated by spectral analyzer 102.
Finally, the noise energy estimates are updated in module 107 (second level of noise estimation and update). This module 107 uses all available parameters previously calculated in modules 102 to 106 to decide on updating the energies of the noise estimate.
In signal classifier 108, the sound signal is further classified as unvoiced, voiced stable, or generic. Several parameters are calculated to support this decision. In this signal classifier, the encoding mode of the sound signal of the current frame is chosen so that it best represents the kind of signal being encoded.
The sound encoder 109 encodes the sound signal based on the encoding mode selected in the sound signal classifier 108. In other applications, the sound signal classifier 108 may be an automatic speech recognition system.
Spectral analysis
Spectral analysis is performed by the spectral analyzer 102 of Figure 1.
The Fourier transform is used to perform spectral analysis and spectral energy estimation. Spectral analysis is performed twice for each frame using a 256-point Fast Fourier Transform (FFT) with 50 percent overlap (as illustrated in Figure 2). The analysis windows are placed in such a way as to take advantage of all the anticipation. The beginning of the first window is at the beginning of the current encoder frame. The second window is placed 128 samples beyond. A square root Harming window (which is equivalent to a sinusoidal window) has been used to weight the input sound signal for spectral analysis. This window is particularly suitable for overlap-sum procedures (thus, this particular spectral analysis is used in spectral subtraction-based noise suppression and overlap-sum analysis / synthesis). The square root Harming window is given by:
ES 2 533 358 T3
<img file="ES2533358T3_D0002.tif" />
where Lfft = 256 is the size of the FTT analysis. Here, only half of the window is calculated and stored since this window is symmetric (from 0 to Lfft / 2).
The window signals for both spectral analyzes (first and second spectral analyzes) are obtained using the following two relationships:
^ FFT * “θ ·» »^ FFΓ 1
Where s' (0) is the first sample in the current frame. In the non-limiting illustrative embodiment of the present invention, the beginning of the first window is placed at the beginning of the current frame. The second window is placed 128 samples further.
The FFT is performed on both window signals to obtain the following two sets of spectral parameters for each frame:
(*) - Σ (nk '<sup>Jlr</sup>\ A - 0 L<sub>FFT</sub> -1
Jf<sup>m</sup>(A) = * = 0 .....- 1
Jico in which N = Lfft.
The FFT provides the real and imaginary parts of the spectrum denoted by XR (k), k = 0 to 128, and Xf {k), k = 1 to 127. Xr (0) corresponds to the spectrum at 0 Hz (DC) and Xr (128) corresponds to the spectrum at 6,400 Hz. The spectrum at these points only has real values.
After FFT analysis, the resulting spectrum is divided into critical bands using the intervals having the following upper limits [M. Jelinek and R. Salami, Noise Reduction Method for Wideband Speech Coding, in Proc. Eusipco, Vienna, Austria, September 2004] (20 bands in the frequency range 0-6,400 Hz):
Critical bands = {100.0, 200.0, 300.0, 400.0, 510.0, 630.0, 770.0, 920.0, 1,080.0, 1,270.0, 1,480.0, 1,720, 0, 2,000.0, 2,320.0, 2,700.0, 3,150.0, 3,700.0, 4,400.0, 5,300.0, 6,350.0} Hz.
The 256-point FFT results in a frequency resolution of 50 Hz (6,400 / 128). Thus, after ignoring the DC component of the spectrum, the number of frequency bins for each critical band is Mcb = {2, 2, 2, 2, 2, 2, 3, 3, 3, 4, 4, 5 , 6, 6, 8, 9, 11, 14, 18, 21}, respectively.
The average energy in a critical band is calculated using the following relationship:
ES 2 533 358 T3
<img file="ES2533358T3_D0003.tif" />
(2) where KR (k) and Xf (k) are, respectively, the real and imaginary parts of the k-th frequency container and j is the index of the first container in the ith critical band provided by ji = { 1, 3, 5, 7, 9, 1 1, 13, 16, 19, 22, 26, 30, 35, 41, 47, 55, 64, 75, 89, 107}.
The spectral analyzer 102 also calculates the normalized energy for each frequency container, EcoNT (k), in the range of 0 to 6400 Hz, using the following relationship:
^ contW - 7F - (^ (A) + Y<sup>2</sup> (*))> A -1 ..... 1Ξ7 (3)
Furthermore, the energy spectra for each frequency container in both analyzes are combined with each other to obtain the logarithmic spectrum of mean energy (in decibels), that is,
<img file="ES2533358T3_D0004.tif" />
where the superscripts (1) and (2) are used to denote the first and second spectral analyzes, respectively.
Finally, the spectral analyzer 102 calculates the mean total energy for both the first and second spectral analyzes in a 20 ms frame by summing the mean energies of the critical bands Ece. That is, the spectral energy for a given spectral analysis is calculated using the following relationship:
ιϊ j = 0 and the total frame energy is calculated as the average of the spectral energies of both the first and second spectral analyzes in a frame. Namely
E, = 101óg (0<sub>r</sub>5 (£ ^ (0) + ^ (!)), DB. (6)
The output parameters of the spectral analyzer 102, that is, the average energy for each critical band, the energy for each frequency container, and the total energy, are used in the sound activity detector 103, and in the selection of the rate. The logarithmic spectrum of mean energy is used in music detection.
In narrowband input signals sampled at 8,000 samples / s, after a sample conversion at 12,800 samples / s, there is no content at both ends of the spectrum, so the first lower frequency critical band as well as all three Last high-frequency bands are not considered in the calculation of the relevant parameters (only the bands from i = 1 to 16 are considered). However, equations (3) and (4) are not affected.
Sound Activity Detection (SAD)
Sound activity detection is performed by sound activity detector 103 based on the SNR of Figure 1.
Analyzer 102 performs the spectral analysis described above twice for each frame. Consider that E<sup>(1></sup>cb (í) and E<sup>(2)</sup>cb (t), as calculated in equation (2), denote the average energy for each critical band information in the first and second spectral analyzes, respectively. The average energy for each critical band for the entire frame and part of the previous frame is calculated using the following relationship:
ES 2 533 358 T3
<img file="ES2533358T3_D0005.tif" />
in which E<sup>(0)</sup>cb (¡) denotes the energy for each critical band information from the second spectral analysis of the previous frame. Next, the signal-to-noise ratio (SNR) for each critical band is calculated using the following relationship:
SNR<sub>eB</sub>(i) = E ^ (i} / N<sub>Cti</sub>(i) with (8) where Ncb (¡) is the estimated noise energy for each critical band, as will be explained later. The mean SNR for each frame is then calculated as
<img file="ES2533358T3_D0006.tif" />
where bmin = 0 and bmax = 19 in the case of wideband signals, and bmin = 1 and bmax = 16 in the case of narrowband signals.
Sound activity is detected by comparing the mean SNR for each frame with a certain threshold that is a function of the long-term SNR. The long-term SNR is given by the following relationship:
<img file="ES2533358T3_D0007.tif" />
where Ef and Nf are calculated using equations (13) and (14), respectively, which will be described later. The initial value of Ef is 45 dB.
The threshold is a linear function at intervals of the long-term SNR. Two functions are used, one optimized for clean speech and one optimized for loud speech.
For wideband signals, if SNRlt <35 (loud speech), then the threshold is equal to:
^^^ = 0.41287 SNR<sub>LJ</sub> +13.259625 if not (clean voice):
threshold ^ = 1.0333 SNR<sub>LT</sub> -18
For narrowband signals, if SNRlt <20 (noisy speech), then the threshold is equal to:
threshold ^ = 0.1071 SNR<sub>tr</sub> +16.5 if not (clean voice):
threshold<sub>S / iD</sub> = 0.4773 SNR<sub>LT</sub> - 6. J 364
ES 2 533 358 T3
In addition, a hysteresis is added in the SAD decision to avoid frequent switching at the end of an active sound period. The hysteresis strategy is different for wideband and narrowband signals and takes effect only if the signal is noisy.
For broadband signals, the hysteresis strategy is applied in the case where the frame is in a holding period whose length varies according to the long-term SNR as follows:
U-0 u, = l Si 15SSV »<sub>tr</sub><35.
= 2 <sup>yes</sup> SNR<sub>LT</sub> <15
The maintenance period begins on the first inactive sound frame after three (3) consecutive active sound frames. Its function is to force each idle frame during the maintenance period as an active frame. The SAD decision will be explained later.
For narrowband signals, the hysteresis strategy is to lower the SAD decision threshold as follows:
threshold <sub>YesfD</sub> = threshold ^ -5.2 yes
SNR<sub>Lr</sub><] 9 = threshold- 2 yes
WíSNR<sub>lt</sub><35 threshold = threshold <sub>SAf}</sub> yes <SNR<sub>Lr</sub>
In this way, for noisy signals with low SNR, the threshold is made smaller to give preference to the active signal decision. There is no maintenance for narrowband signals.
Finally, the sound activity detector 103 has two outputs - a SAD flag and a local SAD flag. Both flags are set to one if an active signal is detected and set to zero otherwise. Also, the SAD flag is set to one in the maintenance period. The SAD decision is made by comparing the average SNR for each frame with the SAD decision threshold (through a comparator, for example), that is:
If SNRav> thresholdSAD
SADlocal = 1
SAD = 1 else
SADlocal = 1
If in maintenance period
SAD = 1 else
SAD = 0 end end.
ES 2 533 358 T3
First level of noise estimation and update
A noise estimator 104 as illustrated in Figure 1 calculates the total noise energy, the relative energy of the frame, updates the long-term average noise energy and the long-term frame average energy, the average energy for each critical band, and a noise correction factor. In addition, the noise estimator 104 performs a noise energy initialization and performs a downgrade.
The total noise energy for each frame is calculated using the following relationship:
ΣμΙ (in i = OJ
<img file="ES2533358T3_D0008.tif" />
where Ncb (i) is the estimated noise energy for each critical band.
The relative energy of the frame is given by the difference between the frame energy in dB and the long-term average energy. The relative energy of the frame is calculated using the following relationship:
<img file="ES2533358T3_D0009.tif" />
where Et is given by equation (6).
The long-term average noise energy or long-term frame average energy is updated in each frame. In case of active signal frames (indicator SAD = 1), the long-term average frame energy is updated using the relation:
<img file="ES2533358T3_D0010.tif" />
with initial value Ef = - 45 dB.
In case of idle speech frames (SAD indicator = 0), the long-term average noise energy is updated as follows:
^=0,99^+0,0¾ (14)
The initial value of Nf is set equal to Ntot for the first 4 frames. Also, in the first four (4) frames, the value of Ef is limited by<sup>AND</sup>f> N<sub>tot</sub> + 10.
The frame energy for each critical band for the entire frame is calculated by averaging the energies of both the first and second spectral analyzes in the frame using the following relationship:
AND<sub>cs</sub> (i) = O.SEg * (i) + 0, SE ™ (í) (15)
The noise energy for each critical band Ncb (i) is initialized to 0.03.
At this stage, only a downward update of the noise energy is performed for the critical bands so that the energy is less than the background noise energy. First, the temporal updated noise energy is calculated using the following relationship:
<img file="ES2533358T3_D0011.tif" />
in which E<sup>0)</sup>cb () denotes the energy for each critical band corresponding to the second spectral analysis from 11
ES 2 533 358 T3 the previous frame.
Then, for i = 0 to 19, if Ntmp (i) <Ncb (í), then Ncb (í) = Ntmp (i).
Later, a second level of noise estimation and updating is performed by setting Ncb (i) = Ntmp (i) if the frame is declared as an inactive frame.
Second level of noise estimation and update
The parametric noise activity detection and noise estimate update module 107 updates the noise energy estimates for each critical band to be used in the noise activity detector 103 in the next frame. The update is performed during idle signal periods. However, the SAD decision made above, based on the SNR for each critical band, is not used to determine whether or not the noise energy estimates are updated. Another decision is made based on other independent parameters instead of the SNR for each critical band. The parameters used for updating the noise energy estimates are: tone stability, signal non-stationarity, loudness and the ratio between the 2nd order and 16th order LP residual error energies and generally have a low sensitivity to noise level variations. The decision for updating the noise energy estimates is optimized for voice signals. To improve the detection of active music signals, the following different parameters are used: spectral diversity, complementary non-stationarity, noise character, and tonal stability. Music detection will be explained in detail in the following description.
The reason for not using the SAD decision for updating noise energy estimates is to make the noise estimate robust at rapidly changing noise levels. If the SAD decision was used for updating the noise energy estimates, a sudden increase in the noise level would cause an increase in the SNR even for idle signal frames, preventing updating the noise energy estimates, thereby which in turn would keep the SNR high in subsequent frames, and so on. Consequently, the update would be blocked and some other logic would be needed to resume the noise adaptation.
In the illustrative, non-restrictive embodiment of the present invention, an open-loop tone analysis is performed on an LP analyzer and tone tracker module 106 in Figure 1) to calculate three open-loop tone estimates per frame: do, di and d2 corresponding to the first half of the frame, second half of the frame and the anticipation, respectively. This procedure is well known to those of ordinary skill in the art and will not be described further in the present description (for example, VMR-WB [Source Controlled Variable-Rate Multimode Wideband Speech Codec (VMR-WB)], Service Options 62 and 63 for Spread Spectrum Systems, 3GPP2 Technical Specification C.S0052-A v1.0, April 2005 (http7 / www 3gpp2 org)]) The Pitch Tracker and LP Analyzer Module 106 calculates a Tone Stability Counter using the following relationship:
pc = ¡d<sub>Q</sub>-d_<sub>}</sub> | + | ¿, -D<sub>Q</sub> I + IíÍj - </, | (19) where d-1 is the delay of the second half of the frame of the previous frame. For pitch delays greater than 122, the LP analyzer and pitch tracker module 105 sets d2 = dr. In this way, for said delays, the value of pc in equation (19) is multiplied by 3/2 to compensate for the third missing term in the equation. Tone stability is true if the PC value is less than 14. Also, for low loud frames, pc is set to 14 to indicate pitch instability. More specifically:
Yes «!)) 3+ r<sub>and</sub> <th<sub>Cpc</sub> then pc = 14, (20) where Cnorm (d) is the normalized raw correlation and re is an optional correction added to the normalized correlation to compensate for the reduction of the normalized correlation in the presence of background noise. The loudness threshold cpc = 0.52 for WB and cpc threshold = 0.65 for NB. The correction factor can be calculated using the following relationship:
r, = 0.00024492 -0.022
ES 2 533 358 T3 in which Ntot is the total noise energy for each frame calculated according to equation (11).
The normalized gross correlation can be calculated based on the decimated weighted acoustic swd (n) signal using the following equation:
<img file="ES2533358T3_D0012.tif" />
in which the limit of the sums depends on the delay itself. The weighted swd (n) signal is that used in open-loop tone analysis and provided by filtering the pre-processed input sound signal from pre-processor 101 through a weighting filter of the form A (z / Y) / (1-pz '<sup>1</sup>). The weighted swd (n) signal is decimated by a factor of 2 and the limits of the summations are given as a function of:
Lsec = 40 for d = 10, ..., 16
Lsec d = 10, ..., 16
Lsec d = 32, ..., 61
Lsec <sup>=</sup> = 62,..., 115
These lengths ensure that the length of the correlated vector comprises at least one pitch period which helps to obtain robust open-loop pitch detection. The initial moments are related to the beginning of the current plot and are given by:
start = 0 for the first half of the frame start = 138 for the second half of the frame start = 256 for the look-ahead at a sampling frequency of 12.8 kHz.
The parametric module 107 for detecting sound activity and updating the noise estimate performs an estimation of the non-stationarity of the signal based on the product of the relationships between the energy for each critical band and the long-term average energy for each critical band. .
The long-term average energy for each critical band is updated using the following relationship:
“^ E ^ CB.LT + 0 ¢ 0 i P<sup>ara</sup> tO (21) in which bmin = 0 and bmax = 19 in the case of wideband signals, and bmin = 1 and bmax = 16 in the case of narrowband signals, and E<sub>CB</sub> (i) is the frame energy for each critical band defined in equation (15). The update factor ae is a linear function of the total frame energy, defined in Equation (6), and is calculated as follows:
For broadband signals: αe = 0.024Et - 0.235 with 0.5 <ae <0.99.
For narrow band signals: ae = 0.00091 Et + 0.3185 with 0.5 <ae <0.999.
Et is calculated using equation (6).
The non-stationarity of the frame is given by the product of the relations between the frame energy and the long-term average energy for each critical band. More specifically:
ES 2 533 358 T3
<img file="ES2533358T3_D0013.tif" />
The parametric noise activity detection and noise estimation update module 107 further produces a loudness factor for the noise update using the following relationship:
loudness - (C ^<sub>H</sub> (d<sub>0</sub> ) + C<sub>aofrH</sub> (^))/2 + ^ (23)
Finally, the parametric noise activity detection and noise estimation update module 107 calculates a relationship between LP residual energy after 2nd order and 16th order LP analysis using the relationship:
ratioresid = E (2) / £ (16) (24) where E (2) and E (16) are the residual energies LP after 2nd order and 16th order LP analysis calculated on module 106 analyzer LP and pitch tracker using Levinson-Durbin recursion, which is a procedure well known to those of ordinary skill in the art. This relationship reflects the fact that to represent a spectral envelope of the signal, a higher order of LP is generally needed for a speech signal than for noise. In other words, the difference between E (2) and E (16) is assumed to be less for noise than for active speech.
The update decision made by the parametric module 107 for detecting sound activity and updating the noise estimate is determined based on a variable update_noise that is initially set to 6 and is decremented by 1 if an inactive frame is detected and increases by 2 if an active frame is detected. Also, the update_noise variable is limited between 0 and 6. Noise energy estimates are only updated when update_noise = 0.
The value of the update_noise variable is updated in each frame as follows:
If (nostat> thresholdstac) OR (pc <14) OR (loudness> thresholdCnorm) OR (ratio_resid> thresholdresid) update_noise = update_noise + 2 else update_noise = update_noise - 2 where for broadband signals, thresholdstac = thresholdCnorm = 0.85 and thresholdresid = 1.6, and for narrowband signals, thresholdresid = 500,000, thresholdCnorm = 0.7 and thresholdresid = 10.4.
In other words, frames are declared inactive for noise update when (noestac <thresholdstac) AND (pc> 14) AND (loudness <thresholdcnorm) AND (residence_relation <thresholdresid) and a 6-frame hold is used before performing the noise update.
In this way, if update_noise = 0 then for i = 0 to 19 Ncb (í) = Ntmp (i) where Ntmp (i) is the temporal updated noise energy already calculated in Equation (18).
Improved noise detection from music signals
The noise estimation described above has its limitations for certain musical signals, such as piano or rock concerts and instrumental pop, as it was developed and optimized primarily for speech detection. To improve the detection of musical signals in general, the parametric module 107 for noise activity detection and noise estimation update uses other parameters or techniques in conjunction with existing ones. These other parameters or techniques comprise, as previously described herein, spectral diversity, complementary non-stationarity, noise character and tonal stability, calculated by a spectral diversity calculator, a complementary non-stationarity calculator , a noise character calculator and a tonality estimator, respectively. They will be described in detail herein below.
ES 2 533 358 T3
Spectral diversity
Spectral diversity provides information about significant changes in the signal in the frequency domain. The changes are tracked in the critical bands by comparing the energies in the first spectral analysis of the current frame and the second spectral analysis two more frames ago . The energy in a critical i band of the first spectral analysis in the current frame is denoted as E<sup>(1></sup>cb (í). Denote the energy in the same critical band calculated in the second spectral analysis two frames ago as E '~<sup>2)</sup>cb (í). These two energies are initialized to 0.0001. Then, for all critical bands above 9, the maximum and minimum of the two energies are calculated as follows:
<img file="ES2533358T3_D0014.tif" />
Subsequently, a relationship between the maximum and minimum energy in a specific critical band is calculated as
<img file="ES2533358T3_D0015.tif" />
Finally, the parametric noise activity detection and noise estimation update module 107 calculates a spectral diversity parameter as a normalized weighted sum of the relationships in which self-weight is the maximum energy Emax (i). This spectral diversity parameter is determined by the following relationship:
<img file="ES2533358T3_D0016.tif" />
The spec_div parameter is used in the final decision about music activity and noise energy update. The spec_div parameter is also used as an auxiliary parameter for the computation of a complementary non-stationarity parameter that is described later.
Complementary non-stationarity
The inclusion of a complementary non-stationarity parameter is motivated by the fact that the non-stationarity parameter, defined in equation (22), fails when a sharp attack of energy in a musical signal is followed by a slow decrease in energy. In this case, the long-term average energy per critical band, Ecb, lt (í), defined in Equation (21), increases slowly during the attack, while the frame energy per critical band, defined in the Equation (15), decreases slowly. In a certain frame after the attack, these two energy values meet and the parameter noestac results in a small value that indicates an absence of active signal. This leads to a false noise update and subsequently a false SAD decision.
To overcome this problem, an alternative long-term average energy is calculated for each critical band using the following relationship:
<img file="ES2533358T3_D0017.tif" />
The variable E2cb, lt (í) is initialized to 0.03 for all i. Equation (26) is very similar to Equation (21), the only difference being the update factor pe, which is calculated as follows:
if (divespec> umbraldiv_espec) pe = 0 else pe = cte eni.
ES 2 533 358 T3 in which umbraldiv_espec = 5. In this way, when an energy attack is detected (div_espec> 5) the alternative long-term average energy is immediately set to the average frame energy, that is, E2cb, lt (í) = E<sub>cb</sub> (i). Otherwise, this long-term alternative average energy is updated in the same way as conventional non-stationarity, that is, using the exponential filter with the factor a<sub>and</sub> update. The complementary non-stationarity parameter is calculated in the same way as noestac, but using E2cb, lt (í), that is, ^<sub>h</sub>max (£<sub>ffl</sub>Ci)<sub>t</sub>E2 „<sub>£ r</sub>(0) no_estac2 = | I ------ = —------------- (¿7} íriínfí ^ (t), E2<sub>C3 item</sub> (; ))
The complementary nonstationary parameter, noestac2, may fail a few frames right after an energy attack, but should not fail during steps characterized by slowly decreasing energy. Because the noestac parameter works well in power attacks and a few frames later, therefore, a logical disjunction of noestac and noestac2 solves the problem of inactive signal detection in certain music signals. However, the disjunction applies only to passages that are likely to be active. The probability is calculated as follows:
If (noestac> thresholdstac) OR (tonal_stability = 1)) pred_act_LT = ka pred_act_LT + (1 - ka). 1 else pred_act_LT = ka pred_act_LT + (1 - ka). 0 end.
The coefficient ka is set to 0.99. The pred_act_LT parameter that is in the range <0: 1> can be interpreted as a predictor of activity. When it is close to 1, the signal is likely to be active, and when it is close to 0, it is likely to be inactive. The pred_act_LT parameter is initialized to one. In the above condition, tonal_stability is a binary parameter used to detect a stable tonal signal. This tonal_stability parameter will be described in the following description.
The parameter noestac2 is taken into consideration (in disjunction with noestac) in updating the noise energy only if pred_act_LT is greater than a certain threshold, which has been set to 0.8. The logic of the noise energy update is explained in detail at the end of this section.
Noise character
Noise Character is another parameter used in detecting certain noise-like musical signals, such as cymbals or low-frequency drums. This parameter is calculated using the following relationship:
<img file="ES2533358T3_D0018.tif" />
The noise_character parameter is calculated only for frames whose spectral content has at least a minimum energy, which is fulfilled when the numerator and denominator of Equation (28) are greater than 100. The noise_character parameter has an upper limit of 10 and its long-term value is updated using the following relationship:
«JractrrndelT '—a carnct_noise_LT + (1 - a) character_raErfo (29)
The initial value of LT_noise_character is 0 and On is set to a value of 0.9. This LT_noise_charact parameter is used in the noise energy update decision explained at the end of this section.
ES 2 533 358 T3
Tonal stability
Tonal stability is the last parameter used to prevent a false update of the noise energy estimates. Tonal stability is also used to avoid declaring some music segments as unvoiced frames. Tonal stability is further used in a built-in super wideband codec to decide which encoding model to use to encode the sound signal above 7 kHz. Tonal stability detection makes use of the tonal nature of musical signals. In a typical musical signal there are tones that are stable for several consecutive frames. To make use of this feature, it is necessary to keep track of the positions and shapes of the strong spectral peaks, as these can correspond to the tones. Detection of tonal stability is based on a correlation analysis between spectral peaks in the current frame and those in the past frame. The input is the logarithmic mean energy spectrum defined in Equation (4). The number of spectral containers is denoted as Nspec (container 0 is the CC component and Nspec = Lfft / 2). In the following description, the term spectrum will refer to the logarithmic spectrum of mean energy, defined by Equation (4).
Tonal stability detection is done in three stages. In addition, the tonal stability detection uses a calculator of a current residual spectrum, a detector of peaks in the current residual spectrum, and a calculator of a correlation map and a long-term correlation map, which will be described later here. memory.
In the first stage, the indices of the local minimums of the spectrum are searched (by a spectrum minimum locator for example), in a loop described by the following formula and stored in an imin buffer that can be expressed as go on:
ta = (Vr (£<sub>from</sub>(i-l)> E<sub>gave</sub>(0) A (£<sub>-B</sub>CO <£<sub>to</sub>(<sub>í</sub> +])) /=1,,..,^^-2 (30) in which the symbol Λ means logical AND.
In Equation (30), EdB (i) denotes the logarithmic spectrum of mean energy calculated by Equation (4). The first index in imin is 0, if EdB (O) <EdB (1). Therefore, the last index in imin is Nespec-1, if EdB (NESPEc-1) <EdB (NESPEc-2). Denote the number of minima found as Nmin.
The second stage consists of calculating a spectral floor (using a spectral floor estimator, for example) and subtracting it from the spectrum (using a suitable subtractor, for example). The spectral floor is a piecewise linear function that extends through the detected local minima. Each linear section between two consecutive minima imin (x) and imin (x + 1) can be described as:
= * (/ - ta (x)) + 9 J = ta (χ) ..... ta (x + where k is the slope of the line and q = EdB (im¡n (x)). k can be calculated using the following relationship:
¿<Ta (X + 0) ~ (ta <- * »ta (* <sup>+ 1</sup>> ta (x)
In this way, the spectral ground is a logical connection of all the spans:
soil_espect (j) = E ^ J) soil_espect (y) = fl (suetoesped (J) = /-o.....Μ°)J ~ ta (0) '··> ta <sup>—</sup>'
J 'taítain “0 .....
(31)
ES 2 533 358 T3
The main containers up to imin (0) and the terminating containers from imin (Nmin -1) of the spectral ground are set to the spectrum itself. Finally, the spectral ground is subtracted from the spectrum using the following relationship:
AND<sub>dB</sub>(j) -<sup>dreams</sup>P<sup>ect</sup>(J) jt<sup>32</sup>) and the result is called the residual spectrum. The spectral ground calculation is illustrated in Figure 3.
In the third stage, a correlation map and a long-term correlation map are calculated from the residual spectrum of the current frame and the previous frame. Again, this is a piecemeal operation. In this way, the correlation map is calculated peak to peak since the minima delimit the peaks. In the following description, the term peak will be used to denote a stretch between two minima in the residual spectrum Edb, res.
Denote the residual spectrum of the previous plot as E '~<sup>1)</sup>dB, res (j). For each peak in the current residual spectrum a normalized correlation is calculated in which the shape in the previous residual spectrum corresponds to the position of this peak. If the signal was stable, the peaks should not move considerably from frame to frame and their positions and shapes should be roughly the same. In this way, the correlation operation takes into account all the indices (bins) of a specific peak, which is bounded by two consecutive lows. More specifically, the normalized correlation is calculated using the following relationship:
(33)
<img file="ES2533358T3_D0019.tif" />
The main map_cor wrappers up to imin (0) and the map_cor termination wrappers from imin (Nmin 1) are set to zero. The correlation map is shown in Figure 4.
The correlation map of the current frame is used to update its long-term value which is described by:
mtip_cor_LT (k ^ -a ^ map_cor_LT (k) + O -,
1 * (34) in which αmap = 0.9. The map_cor_LT cor is initialized to zero for all k.
Finally, all the values of map_cor_LT are added together (using an adder, for example) as follows:
<img file="ES2533358T3_D0020.tif" />
If any value of map_cor_LT (j), j = 0, ... N-spec 1, exceeds a threshold of 0.95, a flag cor_strong (which can be considered as a detector) is set to one, otherwise, set to zero.
The decision about tonal stability is calculated by subjecting sum_map_cor to an adaptive threshold, tonal_threshold. This threshold is initialized to 56 and each frame is updated as follows:
ES 2 533 358 T3 if (sum_map_cor> 56) tonal_threshold = tonal_threshold - 0.2 else tonal_threshold = tonal_threshold + 0.2 end.
The adaptive tonal threshold threshold has an upper limit of 60 and a lower limit of 49. Thus, the adaptive tonal threshold threshold decreases when the correlation is relatively good, indicating an active signal segment, and if it does not increase. When the threshold is lower, it is more likely that more frames will be classified as active, especially at the end of active periods. Therefore, the adaptive threshold can be considered as a maintenance.
The tonal_stability parameter is set to one each time sum_map_cor is greater than tonal_threshold or when the cor_strong flag is set to one. More specifically:
If ((sum_map_cor> tonal_threshold) OR (strong_cor = 1)) tonal_stability = 1 else tonal_stability = 0 end.
Using Music Detection Parameters in Noise Energy Update
All music detection parameters are incorporated into the final decision made in the parametric module 107 for noise activity detection and noise estimate update (Act) about updating the noise energy estimates. The noise energy estimates are updated whenever the update_noise value equals zero. Initially, it is set to 6 and each frame is updated as follows:
if (noestac> threshold_estac) OR (pc <14) OR (loudness> thresholdCnorm) OR (ratio_resid> threshold_resid) OR (tonal_stability = 1) OR (car_noise_LT> 0.3) OR ((pred_act_LT> 0.8) AND (noestac2 > threshold_stat)) update_noise = update_noise + 2 else update_noise = update_noise - 1 end.
If the combined condition has a positive result, the signal is active and the update_noise parameter is incremented. Otherwise, the signal is inactive and the parameter is decremented. When it reaches 0, the noise energy is updated with the current signal energy.
In addition to updating noise energy, the tonal_stability parameter is also used in the unvoiced sound signal classification algorithm. Specifically, the parameter is used to improve the robustness of the unvoiced signal classification over music, as will be described in the next section.
Sound signal classification (108 sound signal classifier)
The general philosophy underlying the sound signal classifier 108 (Figure 1) is depicted in Figure 5. The approach can be described as follows. Sound signal classification is performed in three stages in logic modules 501, 502, and 503, each of which discriminates a specific signal class. First, a signal activity detector (SAD) 501 discriminates between active and inactive signal frames. This signal activity detector 501 is the same as that of the signal activity detector 103 in Figure 1. The signal activity detector has already been described in the above description.
If the signal activity detector 501 detects an inactive frame (background noise signal), then the classification chain ends and, if there is support for Discontinuous Transmission (DTX), a coding module 541 that can be incorporated. in encoder 109 (Figure 1) encodes the frame with comfort noise generation (CNG). If there is no DTX support, the frame continues in the active signal classification, and more frequently it is classified as unvoiced speech frame.
If an active signal frame is detected by the voiced activity detector 501, the frame is subjected to a second classifier 502 dedicated to discriminating unvoiced speech frames. If classifier 502 classifies the frame
ES 2 533 358 T3 as unvoiced speech signal, the classification chain ends, an encoding module 542 that can be incorporated in encoder 109 (Figure 1) encodes the frame with an optimized encoding procedure for unvoiced speech signals.
Otherwise, the signal frame is processed through a stable voiced classifier 503. If the frame is classified as a stable voiced frame by classifier 503, then an encoding module 543 that can be incorporated into encoder 109 (Figure 1) encodes the frame using a coding procedure optimized for stable or quasi voiced signals. periodic.
Otherwise, the frame is likely to contain a non-stationary signal segment, such as a voiced start or rapidly evolving voiced speech or a musical signal. Typically, these frames require a general purpose encoding module 544 that can be incorporated into encoder 109 (Figure 1) to encode the frame at a high bit rate to maintain good subjective quality.
Next, the classification of voiced and unvoiced signal frames will be described. The SAD detector 501 (or 103 in Figure 1) used to discriminate idle frames has already been described in the previous description.
The unvoiced parts of the speech signal are characterized by lacking the periodic component and can be divided into unstable frames, in which the energy and spectrum change rapidly, and stable frames, in which these characteristics remain relatively stable. The illustrative non-restrictive embodiment of the present invention proposes a method for classifying unvoiced frames using the following parameters:
• loudness measure, calculated as a mean normalized correlation (r<sub>x</sub> );
• mean spectral tilt measurement (e);
• short-time maximum energy rise from low level (dEO) designed to efficiently detect speech plosives in a signal;
• the tonal stability to discriminate music from a voiceless signal (described in the description above); and • frame relative energy (Ere) to detect very low energy signals.
Loudness measurement
The normalized correlation, used to determine the loudness measurement, is calculated as part of the open-loop tone analysis performed in the LP analyzer and tone tracker module 106 of Figure 1. For example, 20 ms frames can be used. The LP analyzer and tone tracker module 106 typically outputs an open loop tone estimate every 10 ms (twice for each frame). Here, the LP analyzer and tone tracker module 106 is also used to produce and output the normalized correlation measurements. These normalized correlations are calculated on a weighted signal and a weighted signal passed in the open-loop pitch delay. The weighted speech signal sw (n) is calculated using a perceptual weighting filter. For example, a fixed denominator perceptual weighting filter, suitable for broadband signals, can be used. An example of a transfer function for the perceptual weighting filter is determined by the following relationship:
<img file="ES2533358T3_D0021.tif" />
in which 0 <γ2 <γ1 <1 in which A (z) is the transfer function of a linear prediction filter (LP) calculated in the module 106 LP analyzer and tone tracker, which is determined by the following relationship :
The details of LP analysis and open-loop tone analysis will not be described further herein, as they are believed to be well known to those of ordinary skill in the art.
The loudness measure is given by the mean correlation C<sub>norm</sub> which is defined as:
<img file="ES2533358T3_D0022.tif" />
(36)
ES 2 533 358 T3 in which Cnorm (do), C norm (di) and c norm (d2) are, respectively, the normalized correlation of the first half of the current frame, the normalized correlation of the second half of the current frame , and the normalized look-ahead correlation (the beginning of the next frame). The arguments to the correlations are the above-noted open-loop pitch delays calculated in the LP analyzer and pitch tracker module 106 of Figure 1. For example, a lead time of 10 ms can be used. A factor r is added<sub>and</sub> correction to the mean correlation in order to compensate for background noise (in the presence of background noise the correlation value decreases). The correction factor is calculated using the following relationship:
r, = 0.00024492 -0.022 (37) where Ntot is the total noise energy for each frame calculated according to Equation (11).
Spectral tilt
The spectral tilt parameter contains information about the frequency distribution of the energy. The spectral tilt can be estimated in the frequency domain as a ratio between the energy concentrated in the low frequencies and the energy concentrated in the high frequencies. However, it can also be estimated using other procedures, such as a ratio between the first two autocorrelation coefficients of the signal.
The spectral analyzer 102 in Figure 1 is used to perform two spectral analyzes for each frame, as described in the description above. The energy in the high frequencies and in the low frequencies is calculated following the perceptual critical bands [M. Jelinek and R. Salami, Noise Reduction Method for Wideband Speech Coding, in Proc. Eusipco, Vienna, Austria, September 2004], repeated here for convenience
Critical bands = {100.0, 200.0. 300,0,400,0, 510.0. 630.0, 770.0. 920.0.
1.080,0,1.270,0,1.480,0,1.720,0, 2.000,0, 2.320,0,2.700.0, 3.150,0,
3,700.0, 4,400.0, 5,300.0, 6,350.0} Hz.
The energy at high frequencies is calculated as the average of the energies of the last two critical bands using the following relationships:
<img file="ES2533358T3_D0023.tif" />
(39) in which the energies of the critical bands Ecb (i) are calculated according to Equation (2). The calculation is performed twice for both spectral analyzes.
The energy at low frequencies is calculated as the average of the energies in the first 10 critical bands (for NB signals, the first band is not included), using the following relationship:
<img file="ES2533358T3_D0024.tif" />
(40)
The intermediate critical bands have been excluded from the calculation to improve discrimination between frames with high energy concentration in the low frequencies (generally voiced) and with high energy concentration in the high frequencies (generally unvoiced). In the middle part, the energy content is not characteristic of any of the classes and increases the confusion of the decision.
However, energy at low frequencies is calculated differently for high-energy harmonic voiceless signals at low frequencies. This is due to the fact that for female voiced segments, the harmonic structure of the spectrum can be exploited to increase voiced-unvoiced discrimination. The affected signals are those whose pitch period is shorter than 128 or those that are not considered a priori as unvoiced. Sound signals considered a priori as deaf must meet the following condition:
ES 2 533 358 T3
<img file="ES2533358T3_D0025.tif" />
In this way, for the signals discriminated by the previous condition, the energy in the low frequencies is calculated per container, only the containers whose frequencies are sufficiently close to the harmonics are taken into account in the sum. More specifically, the following relationship is used:
<img file="ES2533358T3_D0026.tif" />
in which Kmin is the first container (Kmin = 1 for WB and Kmin = 3 for NB) and EcONT (k) are the container energies, as defined in Equation (3), in the first 25 frequency containers (the CC component is omitted). These 25 containers correspond to the first 10 critical bands. In the previous sum, only the terms close to the harmonics of the tone are considered; wh (i) is set to 1 if the distance between the closest harmonics is not greater than a certain frequency threshold (for example, 50 Hz) and is set to 0 otherwise; therefore, only bins are taken into account that are less than 50 Hz away from the closest harmonics. The counter cnt is equal to the number of nonzero terms in the sum. Therefore, if the structure is harmonic at low frequencies, only high energy terms will be included in the sum. On the other hand, if the structure is not harmonic, the selection of the terms will be random and the sum will be less. In this way, even high-energy unvoiced sound signals at low frequencies can be detected.
The spectral tilt is given by the following relationship:
(43) in which N<sub>h</sub> and N, are the mean noise energies in the last two (2) critical bands and the first 10 critical bands (or the first 9 critical bands for NB), respectively, calculated in the same way as <sup>AND</sup>hy <sup>AND</sup>i in Equations (39) and (40). The estimated noise energies have been included in the slope calculation to account for the presence of background noise. For the NB signals, the missing bands are compensated by multiplying et by 6. The calculation of the spectral inclination is performed twice for each frame to obtain et (0) and et (1) corresponding to both first and second spectral analyzes for each plot. The mean spectral tilt used in the classification of voiceless frames is given by
<img file="ES2533358T3_D0027.tif" />
where eantigua is the slant in the second half of the previous frame.
Low-level short-term maximum energy boost
The maximum short-term energy increase at the low level dE0 is evaluated on the sound signal s (n), where n = 0 corresponds to the beginning of the current frame. For example, 20 ms speech frames are used and each frame is divided into 4 sub-frames for speech coding purposes. The signal energy is evaluated twice for each subframe, that is, 8 times for each frame, based on short-duration segments of 32 samples length (at a sampling rate of 12.8 kHz). In addition, the short-term energies of the last 32 samples of the previous frame are also calculated. Short duration energies are calculated using the following relationship:
<img file="ES2533358T3_D0028.tif" />
where j = -1 and j = 0, ..., 7 correspond to the end of the previous frame and the current frame, respectively. Another set of 9 maximum energies is calculated by shifting the signal indices in Equation (45) in 16 samples. It is
ES 2 533 358 T3 say
<img file="ES2533358T3_D0029.tif" />
For those energies that are low enough, that is, that satisfy the condition 10log (Est (j)) <37, the following relationship is calculated:
<img file="ES2533358T3_D0030.tif" />
for the first set of indices and the same calculation is repeated for E<sup>(2)</sup>st (j) to get two sets of rat relationships<sup>(1)</sup>(j) and rat<sup>(2)</sup>(j). The only maximum in these two sets is searched as follows:
<img file="ES2533358T3_D0031.tif" />
which is the maximum short duration energy surge at the low level.
Measurement of the flatness of the noise spectrum
In this example, idle frames are normally encoded with a encoding mode designed for unvoiced speech in the absence of DTX operation. However, in the case of a quasi-periodic background noise, such as some car noises, a more faithful reproduction of noise is achieved if generic coding for WB is used instead.
To detect this type of background noise, a measure of the flatness of the background noise spectrum is calculated and averaged over time. First, the mean noise energy is calculated for the first and last four critical bands as follows:
you = 15
Next, the measure of flatness is calculated using the following relationship:
<img file="ES2533358T3_D0032.tif" />
and is averaged over time using the following relationship:
<img file="ES2533358T3_D0033.tif" />
/ '[- l] / * [0] where J<sub>rui</sub>d<sub>or</sub> piano is the averaged flatness measure of the past frame and J<sub>rui</sub>d<sub>or</sub> piano is the updated value of the averaged flatness measure of the current frame.
ES 2 533 358 T3
Unvoiced signal classification
The classification of voiceless signal frames is based on the parameters described above, specifically: measure C<sub>norm</sub> loudness, the mean spectral tilt and<sub>t</sub>, the maximum short duration energy increase a by the tonal stability parameter and the relative frame energy calculated during the noise energy update phase (modulo 107 in Figure 1). Relative frame energy is calculated using the following relationship:
<img file="ES2533358T3_D0034.tif" />
(50) where Et is the total frame energy (in dB) calculated in Equation (6) and Ef is the long-term frame average energy, updated in each active frame using the following relationship:
AND<sub>F</sub> = 0,99^-0,01^ .
The update takes place only when the SAD flag is set (SAD variable equal to 1).
The rules for classifying as deaf for WM signals are summarized below:
[((Cnorm <0.695) AND (<sup>and</sup><sub>t</sub> <4,0)) OR (Ee <-14)] AND
[Last frame INACTIVE or DEAF OR ((old <2.4) AND (C norm (do) + re <0.66))] AND
[dE0 <250] AND
[ef (1) <2,7] AND
[(local SAD indicator = 1) OR (_<sub>platw</sub> <1.45) OR (N<sub>F</sub> <20)] AND
NOT [(tonal_stability AND (((C<sub>norm</sub> > 0.52) AND (e<sub>t</sub> > 0.5)) OR (e<sub>t</sub> > 0.85)) AND (Ee> -14) AND SAD flag set to 1]
The first line of the condition is related to low-energy signals and low-correlation signals that focus their energy on high frequencies. The second line covers the voiced offsets, the third line covers the explosive segments of a signal, and the fourth line is for the voiced starts. The fifth line ensures a flat spectrum in case of noisy idle frames. The last line discriminates the musical signals that would otherwise be declared as deaf.
For NB signals, the unvoiced classification condition has the following form:
[Local SAD flag set to 0 OR (Erel <-25) OR ((C<sub>norm</sub> <0.61) AND e<sub>t</sub> <7.0) AND (last frame INACTIVE OR
SORDA OR ((old <7.0) AND (Cnorm (do) + re <0.52))))] AND
[dEO <250] AND
[e<sub>t</sub> <390] AND
NOT [(tonal_stability AND (((C<sub>norm</sub> > 0.52) AND (e<sub>t</sub> > 0.5)) OR (e<sub>t</sub> > 0.75)) AND (Erel> -10) AND SAD flag set to 1]
The decision trees for the WB case and the NB case are shown in Figure 6. If the combined conditions are met, the classification ends by selecting the voiceless encoding mode.
Sound signal classification
If a frame is not classified as an idle frame or a voiceless frame, then it is checked whether it is a stable voiced frame. The decision rule is based on the normalized correlation in each subframe (with resolution
ES 2 533 358 T3% subsample), mean spectral skew, and open-loop pitch estimates in all subframes (with% subsample resolution).
The open-loop tone estimation procedure res performed by the LP analyzer and tone tracker module 106 of Figure 1. In Equation (19), three open-loop tone estimates are used: do, di and d2, corresponding to the first half of the plot, the second half of the plot and anticipation. In order to obtain accurate pitch information on all four subframes, a fractional pitch enhancement with% sample resolution is calculated. This improvement is calculated on the weighted sound signal swd (n). In this exemplary embodiment, the weighted signal swd (n) is not decimated by the open-loop pitch estimation enhancement. At the beginning of each subframe, a short correlation analysis (64 samples at a sampling frequency of 12.8 kHz) is performed with a resolution of 1 sample in the range (-7, +7) using the following delays: for first and second subframes and di for the third and fourth subframes. The correlations are then interpolated around their maxima at the fractional positions dmax - 3/4, dmax - 1/2, dmax - 1/4, dmax, dmax + 1/4, dmax + 1/2, dmax +% . The value that produces the maximum correlation is selected as the enhanced pitch delay.
Denote the enhanced open-loop pitch delays in the four subframes as T (0), T (1), T (2), and T (3) and their corresponding normalized correlations as C (0), C (1), C (2) and C (3). Then, the classification condition of the sound signal is determined by:
[C (0)> 0.605] AND
[C (1)> 0.605] AND
[C (2)> 0.605] AND
[C (3)> 0.605] AND
[and<sub>t</sub> > 4] AND
[| T (1) - T (0) | <3] AND
[| T (2) - T (1) | <3] AND
[| T (3) - T (2) | <3]
The condition says that the normalized correlation is high enough in all the subframes, the pitch estimates do not diverge throughout the frame, and the energy is concentrated in the low frequencies. If this condition is met, the classification ends by selecting the audio signal encoding mode, otherwise, the signal is encoded using a generic signal encoding mode. The condition applies to both WB signals and NB signals.
Hue estimation in super wideband content
In super-wideband signal coding, a specific coding mode is used for sound signals with tonal structure. The frequency range of interest is primarily 7,000-14,000 Hz, but it can also be different. The goal is to detect frames that have strong tonal content in the range of interest so that the pitch-specific encoding mode can be used efficiently. This is done using the tonal stability analysis described earlier in the present description. However, there are some aberrations that are described in this section.
First, the spectral ground that is subtracted from the logarithmic energy spectrum is calculated as follows. The logarithmic energy spectrum is filtered using a Moving Average (MA) filter, or an FIR filter, whose length is Lma = 15 samples. The filtered spectrum is determined by:
] specsoil (f) = —------ £ E<sub>M</sub> (J + k), for - -J ^ spec-Lma- 1 <sup>+ 1</sup>
To avoid computational complexity, the filtering operation is performed only for j = Lma and for the other delays, it is calculated as:
specsoil (f) ~ spec_soil (j -1) for / = £ ^ + 1. .JVsspec-Lmt1
<img file="ES2533358T3_D0035.tif" />
ES 2 533 358 T3
For the lags 0, .., Lma-1 and Nspec-Lma, ..., Nspecs-1, the spectral floor is calculated by extrapolation. More specifically, the following relationship is used:
specsoil (j ') = 0.9spec_soil (j +1) + 0, (j), for J = Lma- L · Ά sue / specQ) = 0.9sue / spec (/ - 1) + 0, (j) , for / = N<sub>yes</sub>> ec- £ mj<sub>í</sub>... Xs? £ c-1
In the first equation above the update continues from Lma-1 down to 0.
Next, the spectral ground is subtracted from the logarithmic energy spectrum in the same manner as previously described in the present description.
Next, the residual spectrum, denoted as Eres, dB (j), is smoothed over 3 samples as follows using a short-time moving average filter:
The search for spectral minima and their indices, the calculation of the correlation map and the long-term correlation map are the same as in the procedure described above in the present description, using the smoothed spectrum E '<sub>re</sub>s, dB (¡)
The decision about the signal tonality in the super wideband content is also the same as that described earlier in the present description, that is, based on an adaptive threshold. However, in this case a different fixed threshold and stage are used. The threshold threshold_tonal is initialized to 130 and updated on each frame as follows:
if (sum_map_cor> 130) tonal_threshold = tonal_threshold - 1.0 else tonal_threshold = tonal_threshold + 1.0 end.
The adaptive tonal_threshold threshold has an upper limit of 140 and a lower limit of 120. The fixed threshold has been set with respect to the frequency range 7,000 to 14,000 Hz. For a different range, it will have to be adjusted. As a general rule, the following relationship tonal_threshold = Nspec / 2 can be applied.
The last difference to the method described above in the present description is that loud tone detection is not used in super wideband content. This is motivated by the fact that loud tones are not perceptually suitable for the purpose of encoding the tonal signal in super wideband content.
The present invention has been described in the foregoing description by means of an illustrative, non-restrictive embodiment thereof. The scope of the present invention is defined by the appended claims.
Contents15
14 members in 7 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 929336P | United States of America | – | |
| 92933607 | United States of America | P | |
| 2008001184 | Canada | W |
Members14
| Document | Office | Kind | |
|---|---|---|---|
| CA2690433A1 | Canada | A1 | |
| WO2009000073A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009000073A8 | World Intellectual Property Organization (WIPO) | A8 | |
| EP2162880A1 | European Patent Office (EPO) | A1 | |
| JP2010530989A | Japan | A | |
| US2011035213A1 | United States of America | A1 | |
| RU2010101881A | Russian Federation | A | |
| RU2441286C2 | Russian Federation | C2 | |
| EP2162880A4 | European Patent Office (EPO) | A4 | |
| JP5395066B2 | Japan | B2 | |
| EP2162880B1 | European Patent Office (EPO) | B1 | |
| US8990073B2 | United States of America | B2 | |
| ES2533358T3This record | Spain | T3 | |
| CA2690433C | Canada | C |
Numbers
- Publication
- 2533358
- Application
- 8783143
Titles2
- Spanish
- Procedimiento y dispositivo para estimar la tonalidad de una señal de sonido
- English
- Procedure and device to estimate the tone of a sound signal
Classification
- CPC, 2
- G10L25/78
- G10L19/22
- IPC, 3
- G10L25 78
- G10L19 22
- G10L25 93