Speech enhancement in entertainment audio
Abstract
improvement of speech in entertainment audio. the present invention relates to audio signal processing. more specifically, the invention refers to the improvement of entertainment audio, such as television audio, to improve the clarity and intelligibility of speech, such as dialogue and auditory narrative. the invention relates to methods, apparatus for performing such methods, and software stored on a computer-readable medium to cause a computer to perform such methods.
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
28 claims: 3 independent, 25 dependent
- 1CLAIMS REIVINDICAÇÕES 1. Method to improve speech in entertainment audio, comprising the steps of:1. Método para aperfeiçoar a fala em áudio de entretenimento, compreendendo as etapas de: processar, em resposta a um ou mais controles, o áudio de entretenimento para aperfeiçoar a clareza e inteligibilidade de partes da fala do áudio de entretenimento, o processamento incluindo variar o nível do áudio de entretenimento em cada uma das faixas de frequência múltipla de acordo com uma característica de ganho que relaciona o nível de sinal de faixa ao ganho, e gerar um controle para variar a característica de ganho em cada faixa de frequência, a geração incluindo caracterizar segmentos de tempo do áudio de entretenimento como (a) fala ou não-fala ou (b) como provável para ser fala ou não-fala, em que as caracterizações operam em uma única faixa de frequência banda larga, obter, em cada uma das faixas da frequência múltipla, uma estimativa da potência do sinal, caracterizado pelo fato de que o método compreende ainda: processing, in response to one or more controls, the entertainment audio to improve the clarity and intelligibility of speech parts of the entertainment audio, the processing including varying the level of the entertainment audio in each of the multiple frequency ranges according to a gain characteristic that relates the range signal level to the gain, and generate a control to vary the gain characteristic in each frequency range, the generation including characterizing time segments of the entertainment audio as (a) speech or non-speech or (b) as likely to be speech or non-speech, in which the characterizations operate in a single broadband frequency range, obtain, in each of the multiple frequency bands, an estimate of the signal strength, characterized by the fact that the method also comprises: monitor, in each of the multiple frequency ranges, the minimum audio level in the range, the monitoring response time responding to the signal strength estimate, transform the monitored minimum in each range into a corresponding adaptive threshold level, and deviate each corresponding level of adaptive limit with the result of the characterization to produce the control for each range. monitorar, em cada uma das faixas da frequência múltipla, o mínimo do nível de áudio na faixa, o tempo de resposta do monitoramento respondendo à estimativa da potência do sinal, transformar o mínimo monitorado em cada faixa em um correspondente nível de limite adaptativo, e desviar cada correspondente nível de limite adaptativo com o resultado da caracterização para produzir o controle para cada faixa.
- 27Computer-readable non-transient storage medium, characterized by the fact that it is encoded with a method to get a computer to perform the steps of the method as defined in claim 1. 27. Meio de armazenamento não-transitório legível por computador, caracterizado pelo fato de que é codificado com um método para fazer com que um computador execute as etapas do método conforme definido na reivindicação 1.
- 28Computer-readable non-transient storage medium, characterized by the fact that it is encoded with a method to get a computer to perform the steps of the method as defined in claim 14. 28. Meio de armazenamento não-transitório legível por computador, caracterizado pelo fato de que é codificado com um método para fazer com que um computador execute as etapas do método conforme definido na reivindicação 14.
Independent claims3
127 paragraphs, as filed
Descriptive Report of the Invention Patent for METHOD TO IMPROVE THE SPEAK IN AUDIO OF ENTERTAINMENT AND NON-TRANSITORY STORAGE MEDIA READABLE BY COMPUTER.
description
Technical field
[0001] The present invention relates to audio signal processing. More specifically, the invention relates to entertainment audio processing, such as television audio, to improve speech clarity and intelligibility, such as dialogue and audio narrative. The invention relates to methods, to apparatus for performing such methods, and to software stored on a computer-readable medium to cause a computer to perform such methods.
Background of the Technique
[0002] Audiovisual entertainment has evolved within a fast paced sequence of dialogue, narrative, music, and effects. The high realism achievable with modern audio entertainment technologies and output methods has encouraged the use of conversational styles of speaking on television that differ substantially from presentation as a clearly announced stage from the past. This situation poses a problem not only for the growing population of elderly viewers who, in the face of sensory impairment and language processing skills, must strive to follow the schedule, but also for people with normal hearing, for example, when listening at low acoustic levels.
[0003] How speech is understood depends on several factors.
Examples are the care of outgoing speech (clear or conversational speech), speech rate, and speech audibility. The spoken language is remarkably robust and can be understood under conditions
Petition 870200055174, of 05/04/2020, p. 4/38
2/24 smaller than ideal. For example, hearing impaired listeners can typically follow clear speech even when they cannot hear parts of speech due to impaired hearing. However, as the speech rate increases and the speech output becomes less accurate, listening and understanding requires increasing effort, particularly if parts of the speech spectrum are inaudible.
[0004] Due to the fact that television audiences cannot do anything to affect the clarity of broadcasting speech, hearing impaired listeners may try to compensate for inadequate audibility by increasing the listening volume. Aside from being objectionable to people of normal hearing in the same room or to neighbors, this approach is only partially effective. This is because most hearing losses are not uniform across frequencies; they affect high frequencies more than low and medium frequencies. For example, the typical ability of a 70 year old male to hear sounds at 6 kHz is about 50 dB worse than that of a young person, but at frequencies below 1 kHz the hearing impairment of the older person is less than 10 dB (ISO 7029, Acoustics - Statistical distribution of hearing thresolds as a function of age). Increasing the volume makes low and medium frequency sounds louder without significantly increasing their contribution to intelligibility because for those frequencies the audibility is already adequate. Increasing the volume also does little to overcome significant hearing loss at high frequencies. A more appropriate correction is a tone control, as provided by a graphic equalizer.
[0005] Although a better option than simply increasing the volume control, a tone control is still insufficient for most hearing loss. The large gain required from
Petition 870200055174, of 05/04/2020, p. 5/38
3/24 high frequency to make soft passages audible to the hearing impaired listener is likely to be uncomfortably loud during high level passages and may even overload the audio playback chain. A better solution is to amplify depending on the signal level, providing higher gains for low-level signal parts and smaller gains (or no gain at all) for high-level parts. Such systems, known as automatic gain controls (AGC) or dynamic range compressors (DRC) are used in hearing aid and their use to improve intelligibility for impaired hearing in telecommunication systems has been proposed (for example, US patent 5,388,185 , US Patent 5,539,806, and US Patent 6,061.43 1).
[0006] Because hearing loss usually develops gradually, most hearing impaired listeners have grown accustomed to hearing loss. As a result, they often object to the entertainment audio sound quality when it is processed to compensate for their hearing impairment. Hearing impaired audiences are more likely to accept compensated audio sound quality when it provides a tangible benefit to them, such as when it increases the intelligibility of dialogue and narrative or reduces the mental effort required for understanding. It is therefore advantageous to limit the application of hearing loss compensation to those parts of the audio program that are dominated by speech. In doing so, it optimizes the choice between potentially objectionable sound quality modifications of music and ambient sounds on the one hand and the desirable intelligibility benefits on the other.
Description of the Invention
[0007] According to one aspect of the invention, speech in entertainment audio can be improved by processing, in response to a
Petition 870200055174, of 05/04/2020, p. 6/38
4/24 or more controls, entertainment audio to enhance the clarity and intelligibility of speech parts of entertainment audio, generate control for processing, generation including characterizing time segments of entertainment audio as (a) speech or non-speech or (b) as likely to be speech or non-speech, and respond to changes in the level of entertainment audio to provide control for processing, in which such changes are responded to within a shorter period of time than the time segments, and a criterion for deciding the response is controlled by characterization. Each of the processing and response can operate in corresponding multiple frequency ranges, the response providing control for processing for each of the multiple frequency ranges.
[0008] Aspects of the invention can operate in a way of looking ahead such that when there is access to a time evolution of the entertainment audio before and after a processing point, and in which the generating control responds to at least some audio after the processing point.
[0009] Aspects of the invention may employ temporal and / or spatial separation such that processing steps, how to characterize and respond are performed at different times or in different places. For example, the characterization can be performed in a first time or place, the processing and response can be performed in a second time or place, and information about the characterization of time segments can be stored or transmitted to control the decision criteria of the answer.
[00010] Aspects of the invention may also encode entertainment audio according to a perceptual encoding scheme or a lossless encoding scheme, and decode entertainment audio according to the same scheme of entertainment.
Petition 870200055174, of 05/04/2020, p. 7/38
5/24 coding used by coding, in which processing steps, how to characterize and respond are performed together with coding or decoding. Characterization can be performed along with coding and processing and / or the response can be performed along with decoding.
[00011] According to the aforementioned aspects of the invention, processing can operate according to one or more processing parameters. The adjustment of one or more parameters can be in response to the entertainment audio in such a way that a speech intelligibility metric of the processed audio is either maximized or boosted above a desired threshold level. According to aspects of the invention, entertainment audio can comprise multiple audio channels in which one channel is mainly speech and the one or more other channels are mainly non-speech, where the metrics of speech intelligibility is based on the level of the speech channel and the level in one or more other channels. The speech intelligibility metric can also be based on the noise level in the listening environment in which the processed audio is played. Adjusting one or more parameters can be in response to one or more of the long-term description of entertainment audio. Examples of long-term descriptors include the average level of dialogue for entertainment audio and a processing estimate already applied to entertainment audio. The adjustment of one or more parameters can be according to a prescriptive formula, in which the prescriptive formula relates the hearing acuity of a listener or group of listeners to one or more parameters. Alternatively, or in addition, the adjustment of one or more parameters can be according to the preferences of one or more listeners.
[00012] According to the aforementioned aspects of the invention, processing may include multiple functions acting on
Petition 870200055174, of 05/04/2020, p. 8/38
6/24 parallel. Each of the multiple functions can operate in one of the multiple frequency ranges. Each of the multiple functions can provide, individually or collectively, dynamic range control, dynamic equalization, spectral narrowing, frequency transposition, speech extraction, noise reduction, or other speech enhancement action. For example, dynamic range control can be provided by multiple compression / expansion functions or devices, each of which processes a frequency region of the audio signal.
[00013] The processing part includes or does not include multiple functions acting in parallel, the process can provide dynamic range control, dynamic equalization, spectral narrowing, frequency transposition, speech extraction, noise reduction, or other speech enhancement action . For example, dynamic range control can be provided by a dynamic range compression / expansion feature or device.
[00014] One aspect of the invention is to control speech enhancement suitable for hearing loss compensation in such a way that, ideally, it operates only on the speech parts of an audio program and does not operate on the remaining (non-speech) parts of program, thus tending not to change the timbre (spectral distribution) or perceived sound of the remaining parts (non-speech) of the program.
[00015] According to another aspect of the invention, improving speech in entertainment audio includes analyzing entertainment audio to classify audio time segments as being speech or other audio, and applying dynamic band compression to one or multiple tracks of audio. frequency of entertainment audio during time segments classified as speech.
Description Of Drawings
[00016] Figure 1a is a schematic functional block diagram
Petition 870200055174, of 05/04/2020, p. 9/38
7/24 illustrating an exemplary implementation of aspects of the invention.
[00017] Figure 1b is a schematic functional block diagram showing an exemplary implementation of a modified version of Figure 1a in which devices and / or functions can be separated temporally and / or spatially.
[00018] Figure 2 is a schematic functional block diagram showing an exemplary implementation of a modified version of Figure 1a in which speech improvement control is derived from a way of looking ahead.
[00019] Figures 3a-c are examples of power gain transformations useful in understanding the example in Figure 4.
[00020] Figure 4 is a schematic functional block diagram showing how the gain in speech improvement in a frequency range can be derived from the estimation of the signal strength of that range according to aspects of the invention.
Best Mode for Carrying Out the Invention.
[00021] Techniques for classifying audio into speech and non-speech (such as music) are known in the art and are sometimes known as a speech-versus-other (SVO) discriminator. See, for example, US Patents 6,785,645 and 6,570,991 as well as Published Patent Applications US 20040044525, and the references contained therein. Speech-versus-other audio discriminators analyze time segments of an audio signal and extract one or more signal descriptors (characteristics) from every time segment. These characteristics are passed to a processor that either produces an estimate of the probability that the time segment is speech, or makes an arduous speech / non-speech decision. Most features reflect the evolution of a signal over time. Typical examples of characteristics are the rate at which the signal spectrum changes with
Petition 870200055174, of 05/04/2020, p. 10/38
8/24 over time or the slope of the rate distribution at which the polarity of the signal changes. To reliably reflect the distinct characteristics of speech, the time segments must be of sufficient length. Because many features are based on signal features that reflect the transitions between adjacent syllables, time segments typically cover at least the duration of two syllables (that is, more or less 250 ms) to capture such a transition. However, time segments are often longer (for example, by a factor of about 10) to obtain more reliable estimates. Although relatively slow in operation, SVOs are reasonably reliable and accurate in classifying audio into speech and non-speech. However, to selectively improve speech in an audio program according to aspects of the present invention, it is desirable to control speech improvement on a finer time scale than the duration of the time segments analyzed by a speech discriminator. versus-other.
[00022] Another class of techniques, sometimes known as voice activity detectors (VADs), indicates the presence or absence of speech in a relatively stable noise background. VADs are used extensively as part of noise reduction schemes in speech communication applications. Unlike speech-versus-other discriminators, VADs usually have a temporal resolution that is adequate for the control of speech improvement according to aspects of the present invention. VADs interpret a sudden increase in signal strength as the beginning of a speech sound and a sudden decrease in signal strength as the end of a speech sound. In doing so, they signal the demarcation between speech and background almost instantly (that is, within a time integration window to measure signal strength, for example, more or less 10 ms). However,
Petition 870200055174, of 05/04/2020, p. 11/38
9/24 Because VADs react to any sudden change in signal strength, they cannot differentiate between speech and other dominant signals, such as music. Therefore, if used alone, VADs are not suitable for controlling speech improvement to selectively improve speech in accordance with the present invention.
[00023] It is an aspect of the invention to combine the speech versus non-speech specificity of speech-versus-other (SVO) discriminators with the temporal acuity of voice activity detectors (VADs) to facilitate speech improvement that selectively responds to speech in an audio signal with a temporal resolution that is finer than that found in speech-versus-other discriminators in the prior art.
[00024] Although, in principle, aspects of the invention can be implemented in analog and / or digital domains, practical implementations are likely to be implemented in the digital domain in which each of the audio signals are represented by individual samples or samples within blocks of data.
[00025] Referring now to Figure 1a, a schematic function block diagram is shown illustrating aspects of the invention in which an audio input signal 1 is passed to a speech enhancement function or device (Speech Enhancement ') 102 which, when enabled by a control signal 103, produces an enhanced speech audio output signal 104. The control signal is generated by a Speech Enhancement Controller control function or device 105 that operates on time segments buffered from the audio input signal 101. Speech Enhancement Controller 105 includes a discriminating function or device speech-versus-other (SVO) 107 and a set of one or more functions or activity-detecting devices (VAD) 108. SVO 107 analyzes the signal over a
Petition 870200055174, of 05/04/2020, p. 12/38
10/24 duration of time that is longer than that analyzed by VAD. The fact that SVO 107 and VAD 108 operate over time with different lengths of time is illustrated by illustrating a parenthesis accessing a wide region (associated with SVO 107) and another parenthesis accessing a more (associated with VAD 108) of a function or storage device (Buffer) 106. The wide region and the narrowest region are schematic and not to scale. In the case of a digital implementation in which audio data is carried in blocks, each part of Buffer 106 can store a block of audio data. The region accessed by the VAD includes the most recent parts of the signal storage in Buffer 106. The probability that the current signal section is spoken, as determined by SVO 107, serves to control 109 the VAD 108. For example, it can control a VAD 108 decision criterion, thereby diverting VAD decisions.
[00026] Buffer 106 symbolizes memory inherent in processing and may or may not be implemented directly. For example, if processing is performed on an audio signal that is stored on a medium with random memory access, that medium can serve as a buffer. Similarly, the history of the audio input can be reflected in the internal state of the speech-versus-other 107 and the internal state of the voice activity detector, in which case no separate buffer is required.
[00027] Speech improvement 102 can consist of multiple devices or audio processing functions that work in parallel to improve speech. Each device or function can operate in a frequency region of the audio signal in which speech is to be improved. For example, devices or functions can provide, individually or as a whole, control
Petition 870200055174, of 05/04/2020, p. 13/38
11/24 dynamic range, dynamic equalization, spectral narrowing, frequency transposition, speech extraction, noise reduction, or other action to improve speech. In the detailed examples of aspects of the invention, dynamic range control provides compression and / or expansion in frequency ranges of the audio signal. In this way, for example, the improvement of Speech 102 can be a bank of range compressors / expanders or dynamic compression / expansion functions, in which each one processes a frequency region of the audio signal (a compressor / expander or function of expansion). multiple range compression / expansion). The frequency specificity arranged by multiple band compression / expansion is useful not only because it allows to sew the speech improvement pattern to the pattern of a given hearing loss, but also because it allows to respond to the fact that at any given moment it may be present speaks in one region of frequency but absent in another.
[00028] To take full advantage of the frequency specificity offered by multiple range compression, each compression / expansion range can be controlled by its own voice activity detector or voice detection function. In such a case, each voice activity detector or voice detection function can signal voice activity in the frequency region associated with the compression / expansion range it controls. While there are advantages to Speech Enhancement 102 being composed of several audio processing devices or functions that work in parallel, simple versions of aspects of the invention can employ Speech Enhancement 2 which is composed of only one audio processing device or function .
[00029] Even when there are many voice activity detectors, there can be only one speech-versus-other discriminator
Petition 870200055174, of 05/04/2020, p. 14/38
12/24
107 generating a single output 9 to control all voice activity detectors that are present. The choice to use only one speech-versus-other discriminator reflects two observations. One is that the rate at which the passband pattern of voice activity changes over time is typically much faster than the temporal resolution of the speech-versus-other discriminator. The other observation is that the characteristics used by the falaversus-other discriminator are typically derived from spectral characteristics that can best be observed in a broadband signal. Both observations make the use of specific-versus-other speech discriminators impractical.
[00030] A combination of SVO 107 and VAD 108 as illustrated in the Speech Enhancement Controller 105 can also be used for purposes other than to improve speech, for example to estimate speech loudness in an audio program, or to measure the speech rate.
[00031] The speech improvement scheme just described can be deployed in many ways. For example, the entire scheme can be implemented within a television or set-top box converter to operate on the audio signal received from a television broadcast. Alternatively, it can be integrated with a perceptual audio encoder (for example, AC-3 or AAC) or it can be integrated with a lossless audio encoder.
[00032] Speech Improvement according to aspects of the present invention can be performed at different times or in different places. Consider an example in which speech enhancement is integrated or associated with an audio encoder or encoding processing. In such a case, the speech discriminator-versus-other (SVO) part 107 of the
Petition 870200055174, of 05/04/2020, p. 15/38
13/24
Speech 105, which is often computationally expensive, can be integrated with or associated with the audio encoder or encoding processing. Output 109 of the SVO, for example a flag indicating the presence of speech, can be embedded in the encoded audio stream. Such information embedded in an encoded audio stream is often referred to as metadata. Speech Enhancement 102 and Speech Enhancement Controller VAD 108 can be integrated or associated with an audio decoder and operate on previously encoded audio. The set of one or more speech activity detectors (VAD) 108 also uses output 109 from speech-versus-other (SVO) 107, which it extracts from the encoded audio stream.
[00033] Figure 1b shows an exemplary implementation of such a modified version of Figure 1a. Devices or functions in Figure 1b that correspond to those in Figure 1 are given the same reference numbers. The audio input signal 101 is passed to an encoder or encoding function (encoder) 110 and a Buffer 106 that covers the length of time required by SVO 107. Encoder 110 can be part of a perceptual or encoding system without loss. The output from encoder 110 is passed to a multiplexer or multiplex function (Multiplexer) 112. The SVO output (109 in Figure 1a) is shown as being applied 109a to encoder 110 or, alternatively, applied 109b to Multiplexer 112 which also receives the output from encoder 110. The SVO output, such as a flag as in Figure 1a, is either carried on the bit stream output of encoder 110 (as metadata, for example) or is multiplexed with encoder output 110 to provide a packet and bit stream mounted 114 for storage or transmission to a demultiplexer or demultiplex function (Demultiplexer) 116 that unpacks bit stream 114 for
Petition 870200055174, of 05/04/2020, p. 16/38
14/24 go to a decoder or a decoding function 118. If the output from SVO 107 was passed 109b to Multiplexer 112, then it is received 109b 'from Demultiplexer 116 and passed to VAD 108. Alternatively, if the output from SVO 107 has been passed 109a to encoder 110, so it is received 109a 'from Decoder 118. As in the example in Figure 1a, VAD 108 can comprise multiple voice activity functions or devices. A function or signal buffer device (Buffer) 120 powered by Decoder 118 that covers the length of time required by VAD 108 provides another supply for VAD 108. The output of VAD 103 is passed to a Speech Enhancement 102 that proves the enhanced speech audio output as in Figure 1a. Although shown separately for clarity in the presentation, the SVO 107 and / or Buffer 106 can be integrated with the encoder 110. Similarly, although shown separately for clarity in the presentation, the VAD 108 and / or Buffer 120 can be integrated with Decoder 118 or Speech Enhancement 102.
[00034] If the audio signal to be processed has been pre-recorded, for example as when playing from a DVD in a consumer's home or when processing offline in a broadcasting environment, the speech-versus-other discriminator and / or the voice activity detector can operate on signal sections that include parts of signal that, during playback, happens after the current signal sample or signal block. This is illustrated in Figure 2, where the symbolic signal buffer 201 contains signal sections that, during playback, happen after the current signal sample or signal block (look ahead). Although the signal was not pre-recorded, looking ahead can still be used when the audio encoder has a significant inherent processing delay.
[00035] Processing Parameters for Improvement of
Petition 870200055174, of 05/04/2020, p. 17/38
15/24 speech 102 can be updated in response to the processed audio signal at a rate that is lower than the dynamic response rate of the compressor. There are several objectives that can be pursued when updating the processor parameters. For example, the speech enhancement processor's processing function gain parameter can be adjusted in response to the program's average speech level to ensure that the change in the long-term medium speech spectrum is independent of the speech level. To understand the effect of such an adjustment and the need for it, consider the following example. Speech enhancement is applied only to a high frequency part of a signal. At a given average speech level, the power estimate 301 of the high frequency signal portion is the mean P1, where P1 is greater than the compression power limit 304. The gain associated with this power estimate is G1, which is the average gain applied to the high frequency part of the signal. Because the low frequency part receives no gain, the average speech spectrum is shaped to be G1 dB higher at high frequencies than at low frequencies. Now what is considered is what happens when the average level of speech increases by a certain amount, AL. An increase in the average speech level per AL dB increases the average power estimate 301 of the high frequency signal part to P2 = P1 + AL. As can be seen from Figure 3a, the higher power estimate P2 gives rise to a gain, G2 which is less than G1. Consequently, the average speech spectrum of the processed signal shows the lower emphasis of high frequency when the average input level is high than when it is low. Because listeners compensate for differences in the average level of speech with their volume control, level dependence on the medium high frequency emphasis is undesirable. It can be eliminated by modifying the gain curve of Figures 3a-c in response
Petition 870200055174, of 05/04/2020, p. 18/38
16/24 at the average level of speech. Figures 3a-c are discussed below.
[00036] The Speech Enhancement processing parameters 102 can also be adjusted to ensure that a speech intelligibility metric is either maximized or boosted above a desired threshold level. The speech intelligibility metric can be computed from the relative levels of the audio signal and a sound competing in the listening environment (such as aircraft cabin noise). When the audio signal is a multichannel audio signal with speech on one channel and non-speech signals on the remaining channels, the speech intelligibility metric can be computed, for example, of the relative levels of all channels and the distribution spectral energy in them. Appropriate intelligibility metrics are well known [for example, ANSI S3.5-1997 Method for Calculation of the Speech Intelligibility Index, American National Standards Institute, 1997; or Musch and Buus, Using statistical decision theory to predict speech intelligibility J Model Structure, Journal of the Acoustical Society of America, (2001) 109, pp 2896 - 2909].
[00037] Aspects of the invention shown in the function block diagrams of Figure 1a and 1b and described here can be implemented as in the example of Figures 3a-c and 4. In this example, the compression amplification to conform frequency of speech components and release of the Processing for the non-speech components can be performed through a dynamic multiband band processor (not shown) that implements both compressive and expansive features. Such a processor can be characterized by a set of gain functions. Each gain function relates to the input power in a frequency range for a corresponding range gain, which can be applied to the signal components in that range. Such a relationship is illustrated in Figures 3a-c.
Petition 870200055174, of 05/04/2020, p. 19/38
17/24
[00038] Referring to Figure 3a, the estimate of the input power power of range 301 is related to a desired range gain 302 by a gain curve. That gain curve is taken as the minimum of two constituent curves. A constituent curve, shown by the solid line, has a compression characteristic with an appropriately chosen compression ratio (CR) 303 for power estimates 301 above a compression limit 304 and a constant gain for power estimates below the compression limit . The other constituent curve, shown by the dashed line, has an expansive characteristic with an appropriately chosen expansion ratio (ER) 305 for power estimates above the expansion limit 306 and a gain of zero for the power estimates below. The final gain curve is taken as the minimum of these two constituent curves.
[00039] The compression limit 304, the compression ratio 303, and the gain in the compression limit are fixed parameters. Your choice determines how the envelope and spectrum of the speech signal are processed in a particular range. Ideally, they are selected according to a prescriptive formula that determines appropriate gain and compression ratios in respective ranges for a group of listeners given their acuity of hearing. An example of such a prescriptive formula is NAL-NLI, which was developed by the National Acoustics Laboratory, Australia, and is described by H. Dillon in Prescribing hearing aid performance [H. Dillon (Ed.), Hearing Aids (pp. 249-261); Sydney; Boomerangue Press, 2001.] However, they can also be based simply on listener preference. The compression limit 304 and the compression ratio 303 in a particular track can additionally depend on specific parameters for a given audio program, such as the average level of dialogue in a soundtrack of
Petition 870200055174, of 05/04/2020, p. 20/38
18/24 film.
[00040] Considering that the compression limit can be fixed, the expansion limit 306 is preferably adaptable and varies in response to the input signal. The expansion limit can assume any value within the dynamic range of the system, including values greater than the compression limit. When the input signal is dominated by speech, a control signal described below triggers the expansion limit towards low levels so that the input level is higher than the range of power estimates for which the expansion is applied (see Figures 3a and 3b). In that condition, the gains applied to the signal are dominated by the compression characteristic of the processor. Figure 3b shows an example of a gain function representing such a condition.
[00041] When the input signal is dominated by audio other than speech, the control signal triggers the expansion limit to high levels so that the input level tends to be lower than the expansion limit. In that condition most signal components do not receive any gain. Figure 3c shows an example of a gain function representing such a condition.
[00042] The range power estimates from the preceding discussion can be derived by analyzing the outputs of a filter bank or the output of a time-frequency-domain transformation, such as DFT (discrete Fourier transform), MDCT (transform modified discrete cosine) or wavelet transformed. Power estimates can also be replaced by measures that are related to signal strength such as the average absolute value of the signal, Teager energy, or perceptual measures such as loudness. In addition, bandwidth estimates can be smoothed over time to control the rate at which gain changes.
Petition 870200055174, of 05/04/2020, p. 21/38
19/24
[00043] According to one aspect of the invention, the expansion limit is ideally placed such that when the signal is speech the signal level is above the expansive region of the gain function and when the signal is audio different from say the signal level is below the expansive region of the gain function. As explained below, this can be achieved by monitoring the non-speaking audio level and setting the expansion limit in relation to that level.
[00044] Certain prior art level monitors set a limit below which downward expansion (or squelch) is applied as part of a noise reduction system that seeks to discriminate between desirable and undesirable audio noise. See, for example, US Patents 3803357, 5263091,
[00045] 5774557, and 6005953. In contrast, aspects of the present invention require differentiating between speech on the one hand and all other audio signals, such as music and effects, on the other. The noise monitored in the prior art is characterized by temporal and spectral envelopes that fluctuate much less than those of desirable audios. In addition, noise often has distinctive spectral shapes that are known a priori. Such distinctive features are exploited by noise monitors in the prior art. In contrast, aspects of the present invention monitor the level of non-speech audio signals. In many cases, such non-speech audio signals exhibit variations in their envelope and spectral shape that are at least as large as those of speech audio signals. Consequently, a level monitor employed in the present invention requires analyzing signal characteristics suitable for the distinction between speech and non-speech audio rather than between speech and noise.
[00046] Figure 4 shows how the speech improvement gain in a frequency range can be derived from the signal strength estimate of that range. Referring now to Figure 4, a
Petition 870200055174, of 05/04/2020, p. 22/38
20/24 representation of a signal from a limited range 401 is passed to a power estimator or estimation device (Power Estimate) 402 that generates an estimate of signal strength 403 in that frequency range. That signal strength estimate is passed to a power transformation for gain or transformation function (Gain Curve) 404, which can be in the form of the example illustrated in Figures 3a-c. The power conversion to par or transform function 404 generates a band gain 405 that can be used to modify the signal strength in the band (not shown).
[00047] The signal strength estimate 403 is also passed to a device or function (Level Monitor) 406 that monitors the level of all signal components in the range that are non-speech. The 406 level monitor may include a 407 minimum leak (Maintain Minimum) leak circuit or function with an adaptive leak rate. This leak rate is controlled by a time constant 408 which tends to be low when the signal strength is dominated by speech and high when the signal strength is dominated by audio other than speech. The time constant 408 can be derived from information contained in the estimated signal strength 403 in the range. Specifically, the time constant can be monotonically related to the energy of the band signal envelope in the frequency range between 4 and 8 Hz. That feature can be extracted by an appropriately tuned bandpass filter or 409 filtering function (bandpass).
[00048] The output of the Bandpass 409 can be related to the time constant 408 by a transfer function (Power-to-Time Constant) 410. The level estimate of the non-speech components 411, which is generated by Level 406 Monitor, is the input to a transformation or transformation function
Petition 870200055174, of 05/04/2020, p. 23/38
21/24 (Power-to-Expansion Limit) 412 which relates the bottom level estimate to an expansion limit 414. The combination of the level monitor 406, transformation 412, and downward expansion (characterized by the expansion ratio 305) corresponds to VAD 108 of Figures 1a and 1b.
[00049] Transformation 412 can be a simple addition, that is, the expansion limit 306 can be a fixed number of decibels above the estimated level of non-speech audio 411. Alternatively, transformation 412 which relates the background level estimated 411 at the expansion limit 306 may depend on an independent estimate of the probability of the bandwidth signal being spoken 413. Thus, when estimate 413 indicates a high probability of the signal being spoken, the expansion limit 306 is lowered. Conversely, when estimate 413 indicates a low probability of the signal being spoken, the expansion limit 306 is increased. The speech probability estimate 413 can be derived from a single signal characteristic or a combination of signal characteristics that distinguish speech from other signals. It corresponds to exit 109 of SVO 107 in FIGS 1a and 1b.
[00050] Suitable signal characteristics and methods of processing them to derive an estimate of speech probability 413 are known to those skilled in the art. Examples are described in US Patents 6,785,645 and 6,570,991, as well as in patent application 20040044525, and in the references contained therein. Incorporation by Reference
[00051] The following patents, patent applications and publications are hereby incorporated by reference, each in its entirety.
[00052] United States patent 3,803,357; Sacks, April 9, 1974, Noise Filter.
Petition 870200055174, of 05/04/2020, p. 24/38
22/24
[00053] United States patent 5,263,091; Waller, Jr., November 16, 1993, Intelligent automatic threshold circuit.
[00054] United States patent 5,388,185; Terry, et al., February 7, 1995, System for adaptive processing of telephone voice signals.
[00055] United States patent 5,539,806; Allen, et al., July 23, 1996, Method for customer selection of telephone sound enhancement.
[00056] United States patent 5,774,557; Slater, June 30, 1998, Autotracking microphone squelch for aircraft intercom systems.
[00057] United States patent 6,005,953; Stuhlfelner, December 21, 1999, Circuit arrangement for improving the signal-tonoise ratio.
[00058] United States patent 6,061,431; Knappe, et al., 9 May 2000, Method for hearing loss compensation in telephony systems based on telephone number resolution.
[00059] United States patent 6,570,991; Scheirer, et al., May 27, 2003, Multi-feature speech / music discrimination system.
[00060] United States patent 6,785,645; Khalil, et al., August 31, 2004, Real-time speech and music classifier.
[00061] United States patent 6,914,988; Irwan, and others,
July 5, 2005, Audio reproducing device.
[00062] Published Patent Application US 2004/0044525; Vinton,
Mark Stuart, et al., March 4, 2004 Controlling loudness of speech in signals that contain speech and other types of audio material.
[00063] Dynamic Range Control via Metadata by Charles Q. Robinson and Kenneth Gundry, Convention Paper 5028, 107th
Petition 870200055174, of 05/04/2020, p. 25/38
23/24
Audio Engineering Society Convention, New York, September 24-27, 1999.
Implementation
[00064] The invention can be implemented in hardware or software, or a combination of both (for example, programmable logic sets). Unless otherwise specified, the algorithms included as part of the invention are not inherently related to any particular computer or other device. In particular, several general purpose machines can be used with programs written in accordance with its precepts, or it may be more convenient to build more specialized devices (for example, integrated circuits) to perform the steps required by the method. Thus, the invention can be implemented in one or more computer programs running on one or more programmable computer systems, each comprising at least one processor, at least one data storage system (including volatile and non-volatile and / or storage elements), at least one device or port, and at least one device or port. The program code is applied to the input data to perform the functions described here and generate output information. The output information is applied to one or more output devices, in a known way.
[00065] Each of these programs can be implemented in any desired computer language (including, machine, assembly, or high-level procedure, logic, or object-oriented programming languages) to communicate with a computer system. In any case, the language can be a compiled or interpreted language.
[00066] Each such computer program is preferably stored on a medium or device
Petition 870200055174, of 05/04/2020, p. 26/38
24/24 storage or loaded into it (for example, solid state memory, or magnetic or optical medium) readable by a general purpose or special programmable computer, to configure and operate the computer when the storage medium or device is read by the system computer to perform the procedures described here. The inventive system can also be considered to be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium thus configured causes a computer system to operate in a specific and predefined way to perform the functions described here.
[00067] A large number of versions of the invention have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of the invention. For example, some of the steps described here can be order independent, and thus can be performed in an order different from that described.
7 priority claims, no other members on record
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 60903392 | United States of America | – | |
| 90339207 | United States of America | P | |
| 2008002238 | United States of America | W | |
| 60903392 | – | – | – |
| PCTUS2008002238 | – | – | – |
| US20070903392P | – | – | – |
| WO2008US02238 | – | – | – |
Numbers
- Publication
- PI0807703
- Publication, DOCDB
- PI0807703
- Publication, EPODOC
- BRPI0807703
- Application
- 7703
- Application, DOCDB
- PI0807703
- Application, EPODOC
- BR2008PI07703
Titles2
- Portuguese
- método para aperfeiçoar a fala em áudio de entretenimento e meio de armazenamento não-transitório legível por computador
- English
- METHOD FOR IMPROVING SPEECH IN ENTERTAINMENT AUDIO AND COMPUTER-READABLE NON-TRANSITIONAL MEDIA
Classification
- CPC, 8
- G10L25/78
- G10L19/012
- G10L19/018
- G10L21/02
- G10L21/0364
- G10L25/93
- G10L2025/932
- G10L2025/937
- IPC, 6
- G10L25 78
- G10L19 012
- G10L19 018
- G10L21 02
- G10L21 0364
- G10L25 93