Method and system for scaling ducking of speech-relevant channels in multi-channel audio
Abstract
Procedure to filter a multichannel audio signal that has a voice channel and at least one channel without voice, to improve the intelligibility of the voice determined by the signal, in which said procedure includes the steps of: (a) determine at least one attenuation control value indicative of a measure of similarity between the voice related content determined by the voice channel and the voice related content determined by at least one voiceless channel of the signal of multichannel audio; and (b) attenuate at least one voiceless channel of the multichannel audio signal in response to at least one attenuation control value.
Term
4.4 yearsto projected expiry
Projected expiry 28 February 2031, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
17 claims: 4 independent, 13 dependent
- 1REIVINDICACIONES 1. Procedimiento para filtrar una senal de audio multicanal que tiene un canal de voz y al menos un canal sin voz, para mejorar la inteligibilidad de la voz determinada por la senal, en el que dicho procedimiento incluye las etapas de:(a) determinar al menos un valor de control de atenuacion indicativo de una medida de similitud entre el contenido relacionado con la voz determinado por el canal de voz y el contenido relacionado con la voz determinado por al menos un canal sin voz de la senal de audio multicanal;y (b) atenuar al menos un canal sin voz de la senal de audio multicanal en respuesta al por lo menos un valor de control de atenuacion.
- 2Procedimiento segun la reivindicacion 1, en el que cada valor de control de atenuacion determinado en la etapa (a) es indicativo de una medida de similitud entre el contenido relacionado con la voz determinado por el canal de voz y el contenido relacionado con la voz determinado por un canal sin voz de la senal de audio, y la etapa (b) incluye una etapa de atenuar dicho canal sin voz en respuesta a cada uno de dichos valores de control de atenuacion.
- 3Procedimiento segun la reivindicacion 1, en el que la etapa (a) incluye una etapa de derivar un canal sin voz derivado a partir del por lo menos un canal sin voz de la senal de audio, y el al menos un valor de control de atenuacion es indicativo de una medida de similitud entre el contenido relacionado con la voz determinado por el canal de voz y el contenido relacionado con la voz determinado por el canal sin voz derivado.
- 4Procedimiento segun la reivindicacion 3, en el que el canal sin voz derivado es derivado combinando un primer canal sin voz de la senal de audio multicanal y un segundo canal sin voz de la senal de audio multicanal.
- 5Procedimiento segun la reivindicacion 3, en el que la senal de audio multicanal tiene al menos dos canales sin voz, y la etapa (b) incluye la etapa de atenuar algunos, pero no todos, de los canales sin voz en respuesta al por lo menos un valor de control de atenuacion.
- 6Procedimiento segun la reivindicacion 3, en el que la senal de audio multicanal tiene al menos dos canales sin voz, y la etapa (b) incluye la etapa de atenuar todos los canales sin voz en respuesta al por lo menos un valor de control de atenuacion.
- 7Procedimiento segun la reivindicacion 1, en el que la etapa (b) comprende escalar una senal de control de atenuacion no procesada para el canal sin voz en respuesta al por lo menos un valor de control de atenuacion.
- 8Procedimiento segun la reivindicacion 1, en el que la etapa (a) incluye la etapa de generar una senal de control de atenuacion indicativa de una secuencia de valores de control de atenuacion, en el que cada uno de los valores de control de atenuacion es indicativo de una medida de similitud en un tiempo diferente entre el contenido relacionado con la voz determinado por el canal de voz y el contenido relacionado con la voz determinado por el al menos un canal sin voz de la senal de audio multicanal, y la etapa (b) incluye las etapas de:escalar una senal de control de ganancia de atenuacion en respuesta a la senal de control de atenuacion para generar una senal de control de ganancia escalada;y aplicar la senal de control de ganancia escalada para atenuar al menos un canal sin voz de la senal de audio multicanal.
- 9Procedimiento segun la reivindicacion 8, en el que la etapa (a) incluye una etapa de comparar una primera secuencia de caractensticas relacionadas con la voz, indicativas del contenido relacionado con la voz determinada por el canal de voz, con una segunda secuencia de caractensticas relacionadas con la voz indicativas del contenido relacionado con la voz determinada por el al menos un canal sin voz de la senal de audio multicanal para generar la senal de control de atenuacion, y cada uno de los valores de control de atenuacion indicados por la senal de control de atenuacion es indicativo de una medida de similitud en un tiempo diferente entre la primera secuencia de caractensticas relacionadas con la voz y la segunda secuencia de caractensticas relacionadas con la voz.
- 10Procedimiento segun la reivindicacion 1, en el que cada uno de dichos valores de control de atenuacion esta relacionado monotonicamente con la probabilidad de que el al menos un canal sin voz de la senal de audio multicanal sea indicativo de contenido mejorador de voz, en el que dicho contenido mejorador de voz comprende contenido que mejora la inteligibilidad u otra cualidad percibida del contenido de voz determinado por el canal de voz.
- 11Procedimiento para filtrar una senal de audio multicanal que tiene un canal de voz y al menos un canal sin voz, para mejorar la inteligibilidad de la voz determinada por la senal, en el que dicho procedimiento incluye las etapas de:(a) comparar una caractenstica del canal de voz y una caractenstica del canal sin voz para generar al menos un valor de atenuacion para controlar la atenuacion del canal sin voz con relacion al canal de voz;y (b) ajustar el al menos un valor de atenuacion en respuesta al por lo menos un valor de probabilidad de mejora de voz para generar al menos un valor de atenuacion ajustado para controlar la atenuacion del canal sin voz con relacion al canal de voz.
- 12Procedimiento segun la reivindicacion 11, en el que el al menos un valor de probabilidad de mejora de voz es una secuencia de valores de comparacion, y el procedimiento incluye una etapa de:determinar la secuencia de valores de comparacion comparando una primera secuencia de caractensticas relacionadas con la voz, indicativa del contenido relacionado con la voz determinada por el canal de voz, con una segunda secuencia de caractensticas relacionadas con la voz, indicativa del contenido relacionado con la voz determinada por el canal sin voz, en el que cada uno de los valores de comparacion es una medida de similitud en un tiempo diferente entre la primera secuencia de caractensticas relacionadas con la voz y la segunda secuencia de caractensticas relacionadas con la voz.
- 13Procedimiento segun la reivindicacion 11, en el que cada uno de dichos valores de atenuacion generado en la etapa (a) es un primer factor indicativo de una cantidad de atenuacion del canal sin voz necesario para limitar la relacion de la potencia de senal en el canal sin voz a la potencia de la senal en el canal de voz de manera que exceda un umbral predeterminado, escalado por un segundo factor relacionado monotonicamente con la probabilidad de que el canal de voz sea indicativo de voz.
- 14Procedimiento segun la reivindicacion 11, en el que cada uno de dichos valores de atenuacion generado en la etapa (a) es un primer factor indicativo de una cantidad de atenuacion del canal sin voz suficiente para causar que la inteligibilidad predicha de la voz determinada por el canal de voz en presencia del contenido determinado por el canal sin voz exceda un valor de umbral predeterminado, escalado por un segundo factor relacionado monotonicamente con la probabilidad de que el canal de voz sea indicativo de voz.
- 15Procedimiento segun la reivindicacion 11, en el que la generacion de cada uno de dichos valores de atenuacion en la etapa (a) incluye las etapas de:determinar un espectro de potencia indicativo de la potencia como una funcion de la frecuencia del canal de voz y un segundo espectro de potencia indicativo de la potencia como una funcion de la frecuencia del canal sin voz, y realizar una determinacion en el dominio de la frecuencia del valor de atenuacion en respuesta al espectro de potencia y al segundo espectro de potencia.
- 16Medio de almacenamiento legible por ordenador que comprende instrucciones, que cuando son ejecutadas con uno o mas procesadores, controlan el uno o mas procesadores para realizar el procedimiento descrito en cualquiera de las reivindicaciones 1-15.
- 17Sistema configurado para realizar un procedimiento segun cualquiera de las reivindicaciones 1-15.
Independent claims17
116 paragraphs in 1 section, as filed
DESCRIPTION
Procedure and scaling system for attenuation of relevant voice channels in multichannel audio
Cross reference to related applications
The present application claims priority to the provisional US patent application No. 61 / 311,437, filed on March 8, 2010.
Background of the invention
one. Field of the Invention
The invention relates to systems and methods for improving the intelligibility of the human voice (eg, a dialogue) determined by a multichannel audio signal. In some embodiments, the invention is a method and system for filtering an audio signal having a voice channel ("speech channel") and a voiceless channel ("non-speech channel") to improve speech intelligibility. determined by the signal, by determining at least one attenuation control value indicative of a measure of similarity between the voice related content determined by the voice channel and the voice related content determined by the voiceless channel, and the channel attenuation No voice in response to the attenuation control value.
2. Background of the invention
Throughout the present description, including the claims, the term "voice" is used in a broad sense to indicate the human voice. Therefore, the "voice" determined by an audio signal is an audio content of the signal that is perceived as a human voice (for example, dialogue, monologue, song or other human voice) during the reproduction of the signal by a speaker (or other sound emitting transducer). According to the typical embodiments of the invention, the audibility of the voice determined by an audio signal is improved in relation to other audio content (for example, instrumental music or sound effects without voice) determined by the signal, thus improving the intelligibility (for example, clarity or ease of understanding) of the voice.
Throughout the present description, including the claims, the expression "voice enhancer content" of a channel of a multichannel audio signal is a content (determined by the channel) that improves the intelligibility or other perceived quality of the voice content determined by another channel (for example, a voice channel) of the signal.
Typical embodiments of the invention assume that most of the voice determined by a multichannel audio input signal is determined by the central channel of the signal. This assumption is consistent with the convention in the production of surround sound according to which most of the voice is normally placed on a single channel (the central channel), and most of the music, the ambient sound and the effects of Sound mixes normally on all channels (for example, the Left, Right, Left Surround and Right Surround channels, as well as the center channel).
Thus, in the present specification, the central channel of a multichannel audio signal will sometimes be referred to as the "voice" channel and, in the present specification, sometimes the other channels will be referred to (for example, Left, Right, Left Surround and Right Surround) of the signal as "voiceless" channels. Similarly, in this report, reference will sometimes be made to a "central" channel generated by adding the left and right channels of a stereo signal whose voice is centrally paired as a "voice" channel, and, in the present memory, sometimes a reference will be made to a "lateral" channel generated by subtracting said central channel from the left (or right) channel of the stereo signal as a "voiceless" channel.
Throughout the present description, including the claims, the expression perform an operation "on" signals or data (for example, filter, scale or transform signals or data) is used in a broad sense to indicate the performance of the operation directly on signals or data, or on processed versions of signals or data (for example, on versions of the signals that have undergone a preliminary filtering before the operation on them).
Throughout the present description, including the claims, the expression "system" is used in a broad sense to indicate a device, system or subsystem. For example, a subsystem that implements a decoder can be called a decoder system, and a system that includes said subsystem (for example, a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other XM inputs are received from an external source) can also be called a decoder system.
Throughout the description, including the claims, the expression "relationship" of a first value ("A") to a second value ("B") is used in a broad sense to indicate A / B, or B / A , or relation of an escalated or displaced version of one from A to B to a scaled or shifted version of the other from A to B (for example, (A x) / (B y), where x and y are offset values).
Throughout the description, including the claims, the expression "reproduction" of signals by sound emitting transducers (eg, loudspeakers) refers to causing the transducers to produce sound in response to the signals, including the realization of any amplification. required and / or other signal processing.
When a voice is heard in the presence of competitive sounds (such as when a friend is heard about the noise of people in a restaurant), a part of the acoustic features that mark the phonological content of the voice (voice features or signals) ) are masked by competitive sounds and are no longer available to the listener to decode the message. As the level of the competitive sound increases in relation to the level of the voice, the number of voice features that are received correctly decreases and the perception of the voice becomes increasingly uncomfortable until, at some level of competitive sound, The voice perception process deteriorates. Although this relationship is valid for all listeners, the tolerable competitive sound level for any voice level is not the same for all listeners. Some listeners, for example, those with hearing loss due to age (presbiacusis) or those who listen to language they acquired after puberty, tolerate less competitive sounds than listeners with good hearing or those who listen to their native language.
The fact that listeners have different abilities to understand the voice in the presence of competitive sounds has implications for the level at which ambient sounds and background music in the news or in an entertainment audio are mixed with the voice. Listeners with hearing loss or those who listen to a foreign language often prefer a lower relative level of voiceless audio than that provided by the content creator.
To meet these special needs, it is known to apply a "ducking" to the voiceless channels of a multichannel audio signal, but less (or no) attenuation to the signal's voice channel, to improve the intelligibility of the voice determined by the signal.
For example, PCT international application publication number WO 2010/011377, which names Hannes Muesch as inventor and assigned to Dolby Laboratories Licensing Corporation (published on January 28, 2010), describes that voiceless channels (for example, Left and right channels) of a multichannel audio signal can mask the voice in the signal's voice channel (for example, the center channel) to the point that a desired level of speech intelligibility is satisfied. WO 2010/011377 describes how to determine an attenuation function to be applied by an attenuation circuit to voiceless channels in an attempt to unmask the voice in the voice channel while retaining as much of the creator's intention as possible. Of content. The technique described in WO 2010/011377 is based on the assumption that content in a voiceless channel never improves the intelligibility (or other perceived quality) of the voice content determined by the voice channel.
The present invention is based, in part, on the recognition that, although this assumption is correct for the large majority of multichannel audio content, it is not always valid. The present inventor has recognized that when at least one voiceless channel of a multichannel audio signal includes content that improves the intelligibility (or other perceived quality) of the voice content determined by the signal's voice channel, the signal filtering According to the procedure of WO 2010/011377, it can negatively affect the entertainment experience of a person listening to the reproduced filtered signal. According to the typical embodiments of the present invention, the application of the procedure described in WO 2010/011377 is suspended or modified during times when the content does not conform to the assumptions underlying the procedure of WO 2010/011377.
There is a need for a procedure and a system to filter a multichannel audio signal to improve speech intelligibility in the common case in which at least one voiceless channel of the audio signal includes content that improves content intelligibility. of voice in the voice channel of the audio signal.
Brief Description of the Invention
In a first class of embodiments, the invention is a method for filtering a multichannel audio signal having a voice channel and at least one voiceless channel, to improve the intelligibility of the voice determined by the signal. The procedure includes the steps of: (a) determining at least one attenuation control value indicative of a measure of similarity between the voice related content determined by the voice channel and the voice related content determined by at least one voiceless channel of the multichannel audio signal; and (b) attenuate at least one voiceless channel of the multichannel audio signal in response to at least one attenuation control value. Typically, the attenuation step comprises scaling an unprocessed attenuation control signal (eg, an attenuation gain control signal) for the voiceless channel in response to at least one attenuation control value. Preferably, the voiceless channel is attenuated to improve the intelligibility of the voice determined by the voice channel without undesirably attenuating the voice enhancing content determined by the voiceless channel. In some embodiments, each attenuation control value determined in step (a) is indicative of a measure of similarity between the voice related content determined by the voice channel and the voice related content determined by a voiceless channel of the audio signal, and step (b) includes the stage of attenuating this voiceless channel in response to said attenuation control value. In some other embodiments, step (a) includes a step of deriving a voiceless channel derived from at least one voiceless channel of the audio signal, and the at least one attenuation control value is indicative of a measurement. of similarity between the content related to the voice determined by the voice channel and the content related to the voice determined by the channel without derived voice. For example, the channel without derived voice can be generated by adding or mixing or combining at least two channels without voice of the audio signal. The determination of each attenuation control value from a single channel without derived voice can reduce the cost and complexity of implementing some embodiments of the invention, in relation to the cost and complexity of determining different subsets of a set of values of attenuation from different voiceless channels. In embodiments in which the input audio signal has at least two voiceless channels, step (b) may include the stage of attenuating a subset of the voiceless channels (eg, each voiceless channel from which a channel without voice has been derived) or all channels without voice, in response to at least one attenuation control value (for example, in response to a unique sequence of attenuation control values).
In some embodiments in the first class, step (a) includes a step of generating an attenuation control signal indicative of a sequence of attenuation control values, in which each of the attenuation control values is indicative of a measure of similarity between the voice-related content determined by the voice channel and the voice-related content determined by the at least one voiceless channel at a different time (for example, in a different time interval), and stage (b) includes the stages of: scale an attenuation gain control signal in response to the attenuation control signal to generate a gain gain control signal, and apply the gain gain control signal to attenuate the at least one voiceless channel (for example, activate, enable or issue (“assert”) the gain control signal scaled to the attenuation circuit to control the attenuation of at least one voiceless channel through the attenuation circuit). For example, in some of said embodiments, step (a) includes a step of comparing a first sequence of voice related features (indicative of the voice related content determined by the voice channel) with a second sequence of related features with the voice (indicative of the content related to the voice determined by the at least one channel without voice) to generate the attenuation control signal, and each of the attenuation control values indicated by the attenuation control signal is indicative of a measure of similarity between the first sequence of voice-related characteristics and the second sequence of voice-related characteristics at a different time ( for example, in a different time interval). In some embodiments, each attenuation control value is a gain control value.
In some embodiments in the first class, each attenuation control value is monotonically related to the probability that at least one voiceless channel of the audio signal is indicative of voice enhancing content that improves intelligibility (or other perceived quality) of the voice content determined by the voice channel.
In a second class of embodiments, the invention is a method for filtering a multichannel audio signal having a voice channel and at least one voiceless channel, to improve the intelligibility of the voice determined by the signal. The procedure includes the steps of: (a) comparing a voice channel characteristic and a voiceless channel characteristic to generate at least one attenuation value to control the attenuation of the voiceless channel relative to the voice channel; and (b) adjust the at least one attenuation value in response to at least one voice enhancement probability value to generate at least one adjusted attenuation value to control the attenuation of the voiceless channel relative to the voice channel. Typically, the adjustment step is (or includes) scaling each attenuation value in response to one of said voice enhancement probability values to generate one of said adjusted attenuation values. Typically, each voice enhancement probability value is indicative (for example, is monotonically related to) the probability that the voiceless channel (or a voiceless channel derived from the voiceless channel or from a set of channels without voice of the input audio signal) is indicative of voice enhancing content (content that improves intelligibility or other perceived quality of the voice content determined by the voice channel).
In some embodiments in the second class, the at least one voice improvement probability value is a sequence of comparison values (for example, difference values) determined by a procedure that includes a step of comparing a first sequence of voice related features indicative of the voice related content determined by the voice channel with a second sequence of voice related features indicative of the related content with the voice determined by the channel without voice, and each of the comparison values is a measure of similarity between the first sequence of voice-related features and the second sequence of voice-related features at a different time (for example, at a different time interval). In typical embodiments in the third class, the procedure also includes the step of attenuating the voiceless channel in response to at least one set attenuation value. Step (b) may comprise scaling the at least one attenuation value (which is typically, or is determined by, an attenuation gain control signal or other unprocessed attenuation control signal) in response to at least one voice enhancement probability value.
In some embodiments in the second class, each attenuation value generated in step (a) is a first factor indicative of an amount of attenuation of the voiceless channel necessary to limit the signal power ratio in the voiceless channel to the power of signal in the voice channel so as not to exceed a predetermined threshold, scaled by a second factor monotically related to the probability that the voice channel is indicative of voice.
In some embodiments in the second class, step (a) includes the steps of generating each of said attenuation values, including by determining a power spectrum (indicative of power as a function of frequency) of each of the voice channels and the voiceless channel, and make a determination in the frequency domain of the attenuation value in response to said power spectrum. Preferably, the attenuation values generated in this way determine the attenuation as a function of the frequency to be applied to the frequency components of the voiceless channel.
Aspects of the invention include a system configured (for example, programmed) to perform any realization of the method of the invention, and a computer-readable medium (for example, a disk) that stores code to implement any embodiment of the method of the invention. .
Brief description of the drawings
Fig. 1 is a block diagram of an embodiment of the system of the invention.
Fig. 1A is a block diagram of another embodiment of the system of the invention.
Fig. 2 is a block diagram of another embodiment of the system of the invention. Fig. 2A is a block diagram of another embodiment of the system of the invention. Fig. 3 is a block diagram of another embodiment of the system of the invention.
Fig. 4 is a block diagram of an audio digital signal processor (DSP) that is an embodiment of the system of the invention.
Fig. 5 is a block diagram of a computer system, which includes a computer readable storage medium 504 that stores computer code to program the system to perform an embodiment of the method of the invention.
Detailed description of preferred embodiments
Many embodiments of the present invention are technologically possible. From the present description, the implementation of the embodiments will be evident for people with ordinary knowledge in the field. Embodiments of the system, method and means of the invention will be described with reference to Figs. 1, 1A, 2, 2A and 3-5, and are defined in the appended claims.
The present inventor has observed that some multichannel audio content has different, but still related, voice content in the voice channel and at least in a voiceless channel. For example, multichannel audio recordings of some theatrical shows are mixed so that the "dry" voice (that is, the voice without noticeable reverberation) is placed in the voice channel (typically, the central channel, C, of the signal) and the same voice, but with a significant reverberation component ("wet" voice) is placed on the channels without the signal's voice. In a typical scenario, the dry voice is the signal from the microphone that the artist holds near his mouth and the wet voice is the signal from the microphones placed in the audience. The wet voice is related to the dry voice, since it is the interpretation as it is heard by the public on the site. However, it is different from the dry voice. Typically, the wet voice is delayed in relation to the dry voice, and has a different spectrum and different additive components (eg, public noise and reverberation).
Depending on the relative levels of the dry voice and the wet voice, it is possible that the wet voice component masks the dry voice component to such an extent that the attenuation of the voiceless channels in the attenuation circuit (for example, as in the procedure described in WO 2010/011377 indicated above) it undesirably attenuates the wet voice signal. Although the dry and wet voice components can be described as separate entities, a listener perceptually merges the two and listens to them as a single voice stream. The attenuation of the wet voice component (for example, in the attenuation circuitine) may have the effect of reducing the perceived intensity of the fused voice flow along with the collapse of its image width. The present inventor has recognized that for multichannel audio signals that have wet and dry voice components of the indicated type, often it will be significantly more pleasant, as well as more favorable for speech intelligibility, if the level of the wet voice components It will not be altered during signal enhancement processing.
The invention is based, in part, on the recognition that, when at least one voiceless channel of a multichannel audio signal includes content that improves the intelligibility (or other perceived quality) of the voice content determined by the voice channel of the signal, the filtering of the channels without voice of the signal using attenuation circuit (for example, according to the procedure of WO 2010/011377) it can negatively affect the entertainment experience of a person listening to the reproduced filtered signal. According to the typical embodiments of the invention, the attenuation (in attenuation circuit) of at least one voiceless channel of a multichannel audio signal is suspended or modified during the times when the voiceless channel includes voice enhancing content (content which improves the intelligibility or other perceived quality of the voice content determined by the signal's voice channel). In times when the voiceless channel does not include voice enhancer content (or does not include voice enhancement content that meets a predetermined criteria), the voiceless channel is normally attenuated (attenuation is not suspended or modified).
A typical multi-channel signal (which has a voice channel) for which a conventional filtering in attenuation circuit is inappropriate is one that includes at least one voiceless channel that conveys voice characteristics that are substantially identical to the voice characteristics in the voice channel According to the typical embodiments of the present invention, a sequence of voice-related features in the voice channel is compared to a sequence of voice-related features in the voiceless channel. A substantial similarity of the two feature sequences indicates that the voiceless channel (i.e., the signal in the voiceless channel) contributes to the useful information to understand the voice in the voice channel and that the attenuation of the channel without voice.
To appreciate the importance of examining the similarity between these sequences of voice-related features rather than the signals themselves, it is important to recognize that the "dry" and "wet" voice content (determined by voice and voiceless channels ) is not identical; the signals indicative of the two types of voice content are typically temporarily outdated, and have undergone different filtering processes and different foreign components have been added to them. Therefore, a direct comparison between the two signals will produce a low similarity, regardless of whether or not the voiceless channel contributes to voice characteristics that are the same as in the voice channel (as in the case of dry voice and humid), unrelated voice features (as in the case of two unrelated voices in the voice channel and without voice [for example, an objective conversation in the voice channel and background murmur in the voiceless channel]), or no voice signal at all (for example, the voiceless channel carries music and effects). By basing the comparison on the characteristics of the voice (as in the preferred embodiments of the present invention), a level of abstraction is achieved that decreases the impact of irrelevant aspects of signals, such as small amounts of delay, spectral and signal differences. strange added. In this way, the preferred implementations of the invention typically generate at least two streams of voice features: one that represents the signal in the voice channel; and at least one that represents the signal of a voiceless channel.
A first example (125) of a system that implements the claimed methods will be described with reference to Fig. 1. In response to a multichannel audio signal comprising a voice channel 101 (central C channel) and two channels 102 and 103 without voice (left and right L and R channels), the system of Fig. one filters the voiceless channels to generate a filtered multichannel output audio signal comprising voice channel 101 and channels 118 and 119 without filtered voice (filtered left and right L 'and R' channels). Alternatively, one or both of the channels 102 and 103 without voice may be another type of voiceless channel of a multichannel audio signal (for example, left rear and / or right rear channels of a 5.1 channel audio signal) or may to be a derived in-voice channel, derived from (for example, a combination of) any of the many different subsets of voiceless channels of a multichannel audio signal. Alternatively, the system can be implemented to filter only one channel without voice, or more than two channels without voice, of a multichannel audio signal.
With reference once more to Fig. 1, the channels 102 and 103 without voice are emitted or activated ("asserted") to the attenuation amplifiers 117 and 116, respectively. During operation, the attenuation amplifier 116 is driven by a control signal S3 (which is indicative of a sequence of control values and, thus, is also called a sequence S3 of control values) emitted from the element 114 of multiplication, and the attenuation amplifier 117 is driven by the control signal S4 (which is indicative of a sequence of control values, and in this way it is also called sequence S4 of control values) emitted from the multiplication element 115.
The power of each channel of the multichannel input signal is measured with a bank of power estimators (104, 105 and 106) and is expressed in a log scale [dB]. These power estimators can implement a smoothing mechanism, such as a leaking integrator, so that the measured power level reflects the average power level during the duration of a sentence or a complete passage. The power level of the signal in the voice channel is subtracted from the power level in each of the voiceless channels (subtracting elements 107 and 108) to give a measure of the power relationship between the two types of signal. The output of element 107 is a measure of the ratio of power in channel 103 without voice to power in voice channel 101. The output of element 108 is a measure of the ratio of power in channel 102 without voice to power in voice channel 101.
The comparison circuit 109 determines, for each voiceless channel, the number of decibels (dB) in which the voiceless channel must be attenuated so that its power level remains at least O dB below the signal power level in the voice channel (where the symbol O also known as "theta script", indicates a threshold value predetermined). In an implementation of circuit 109, the addition element 120 adds the threshold value 0 (stored in element 110, which may be a record) to the difference in power level (or "margin") between channel 103 without voice and the voice channel 101, and the addition element 121 adds the threshold value 0 to the difference in power level between the voiceless channel 102 and the voice channel 101. Elements 111-1 and 112-1 change the sign of the output of addition elements 120 and 121, respectively. This change sign operation converts the attenuation values into gain values. Elements 111 and 112 limit each result to be equal to or less than zero (the output of element 111-1 is output to limiter 111 and the output of element 112-1 is output to limiter 112). The current C1 value emitted from the limiter 111 determines the gain (attenuation denied) in dB that must be applied to channel 103 without voice to maintain its power level 0 dB below the power level of voice channel 101 (over time relevant or in the relevant time window, of the multichannel input signal). The current C2 value emitted from the limiter 112 determines the gain (attenuation denied) in dB that must be applied to channel 102 without voice to maintain its power level 0 dB below the power level of voice channel 101 (over time relevant, or in the relevant time window, of the multichannel input signal). A typical suitable value for 0 is 15 dB.
Because there is a unique relationship between a measure expressed on a logantmic scale (dB) and that same measure expressed on a linear scale, a circuit (or programmed or otherwise configured processor) can be constructed that is equivalent to elements 104, 105, 106, 107, 108 and 109 of Fig. 1, in which power, gain and threshold are expressed in a linear scale. In such implementation, all level differences are replaced by relationships of the linear measurements. Alternative implementations can replace the power measurement with measures that are related to signal strength, such as the absolute value of the signal.
The signal C1 emitted from the limiter 111 is an unprocessed attenuation control signal for the voiceless channel 103 (a gain control signal for the attenuation amplifier 116) that could be output directly to the amplifier 116 to control the attenuation of the Channel 103 without voice. The signal C2 emitted from the limiter 112 is an unprocessed attenuation control signal for the voiceless channel 102 (a gain control signal for the attenuation amplifier 117) that could be output directly to the amplifier 117 to control the attenuation of the Channel 102 without voice.
The unprocessed attenuation control signals C1 and C2 are scaled in the multiplication elements 114 and 115 to generate the gain control signals S3 and S4 to control the attenuation of the voiceless channels by the amplifiers 116 and 117. Signal C1 is scaled in response to a sequence of attenuation control S1 values, and signal C2 is scaled in response to a sequence of attenuation control S2 values. Each control value S1 is emitted from the output of the processing element 134 (which will be described later) to an input of the multiplication element 114, and the signal C1 (and, thus, each control value C1 of "unprocessed" gain determined in this way) is emitted from limiter 111 to the other input of element 114. Element 114 scales the current C1 value in response to the current S1 value by multiplying these values by sf to generate the current S3 value, which is emitted to the amplifier 116. Each control value S2 is emitted from the output of the processing element 135 (which will be described below) to an input of the multiplication element 115, and the signal C2 (and, thus, each gain control value C2 " unprocessed ”determined) is emitted from limiter 112 to the other input of element 115. Element 115 scales the current C2 value in response to the current S2 value by multiplying these values by sf to generate the current S4 value, which is emitted to amplifier 117.
The control values S1 and S2 are generated as follows.
In the voice probability processing elements 130, 131 and 132, a voice probability signal (each of the signals P, Q and T of Fig. 1) is generated for each channel of the multichannel input signal. Voice probability signal P is indicative of a sequence of voice probability values for channel 102 without voice; the voice probability signal Q is indicative of a sequence of voice probability values for the voice channel 101, and the voice probability signal T is indicative of a sequence of voice probability values for the voiceless channel 103 .
The voice probability signal Q is a value monotonically related to the probability that the signal in the voice channel is, in fact, indicative of voice. The voice probability signal P is a value monotonically related to the probability that the signal in the channel 102 without voice is a voice, and the voice probability signal T is a value monotonically related to the probability that the signal in Channel 103 without voice is a voice. Processors 130, 131 and 132 (which are typically identical to each other, but are not identical to each other in some embodiments) can implement any of several procedures to automatically determine the probability that the input signals emitted to them are indicative of voice. In one example, the probability processors 130, 131 and 132 are identical to each other, the processor 130 generates the signal P (from the information on the channel 102 without voice), so that the signal P is indicative of a sequence of voice probability values, each monotonically related to the probability that the signal on channel 102 at a different time (or time window) is a voice, processor 131 generates signal Q (from the information in channel 101), so that the signal Q is indicative of a sequence of voice probability values, each related monotonically with the probability that the signal in channel 101 at a different time (or time window) is a voice, the processor 132 generates the signal T (from the information in the channel 103 without voice) so that the signal T is indicative of a sequence of voice probability values, each monotonically related to the probability that the signal on channel 102 at a different time (or time window) is a voice, and each of the processors 130, 131 and 132 does so by implementing (in the relevant channel between channels 102, 101 and 103) the mechanism described by Robinson and Vinton in "Automated Speech / Other Discrimination for Loudness Monitoring" (Audio Engineering Society, pre-publication number 6437 of Convention 118, May 2005). Alternatively, the signal P can be created manually, for example, by the content creator, and can be transmitted along with the audio signal on channel 102 to the end user, and processor 130 can simply extract said created signal P previously from channel 102 (or processor 130 can be removed and previously created signal P can be output directly to processor 134). Similarly, signal Q can be created manually and can be transmitted along with the audio signal on channel 101, processor 131 can simply extract said signal Q previously created from channel 101 (or processor 131 can be removed and the previously created signal Q can be output directly to processors 134 and 135), signal T can be created manually and can be transmitted along with the audio signal on channel 103, and processor 132 may simply extract said previously created signal T from channel 103 (or processor 132 may be removed and previously created signal T may be issued directly to processor 135).
In a typical implementation of processor 134, the voice probability values determined by signals P and Q are compared in pairs to determine the difference between the current values of signals P and Q for each of a sequence of current values of the signal P. In a typical implementation of processor 135, the voice probability values determined by signals T and Q are compared in pairs to determine the difference between the current values of signals T and Q for each of a sequence of current values of signal Q. As a result, each of processors 134 and 135 generates a time sequence of difference values for a pair of voice probability signals.
Processors 134 and 135 are preferably implemented to smooth each of said difference value sequences by time average, and optionally to scale each resulting sequence of averaged difference values. Scaling of the sequences of averaged difference values may be necessary so that the scaled averaged values emitted from processors 134 and 135 are in a range such that the outputs of multiplication elements 114 and 115 are useful for driving amplifiers 116 and Attenuation 117
In a typical implementation, the signal S1 emitted from the processor 134 is a sequence of scaled averaged difference values (in which each of these scaled averaged difference values is a scaled average of the difference between current values of the difference values of signals P and Q in a different time window). The signal S1 is an attenuation gain control signal for the voiceless channel 102, and is used to scale the unprocessed attenuation gain control signal C1 generated independently for the voiceless channel 102. Similarly, in a typical implementation, the signal S2 emitted from processor 135 is a sequence of scaled averaged difference values (in which each of these scaled averaged difference values is a scaled average of the difference between the current values of signals T and Q in a different time window). The signal S2 is an attenuation gain control signal for the voiceless channel 103, and is used to scale the unprocessed attenuation gain control signal C2 generated independently for the voiceless channel 103.
The scaling of the unprocessed gain control signal C1 in response to the attenuation gain control signal S1 can be performed by multiplying (in item 114) each unprocessed gain control value of the signal C1 by a corresponding value between the averaged difference values scaled from signal S1, to generate signal S3. The scaling of the unprocessed gain control signal C2 in response to the attenuation gain control signal S2 can be performed by multiplying (in item 115) each unprocessed gain control value of the signal C2 by a corresponding value between the scaled averaged difference values of signal S2, to generate signal S4.
Another example (125 ') of the system will be described with reference to Fig. 1A. In response to a multichannel audio signal comprising a voice channel 101 (central C channel) and two voiceless channels 102 and 103 (left and right L and R channels), the system of Fig. 1A filters the voiceless channels to generate a filtered multichannel output audio signal comprising voice channel 101 and voiceless channels 118 and 119 (left and right L 'and R' channels, filtered).
In the system of Fig. 1A (as in the system of Fig. 1), channels 102 and 103 without voice are output to amplifiers 117 and 116, respectively. During operation, the attenuation amplifier 117 is driven by a control signal S4 (which is indicative of a sequence of control values and, thus, is also called a sequence S4 of control values) emitted from the control element 115 multiplication, and the attenuation amplifier 116 is driven by a control signal S3 (which is indicative of a sequence of control values and, of this way, It is also called sequence S3 of control values) emitted from the multiplication element 114. Elements 104, 105, 106, 107, 108, 109 (including elements 110, 120, 121, 111-1, 112-1, 111 and 112), 114, 115, 130, 131, 132, 134 and 135 of Fig. 1A are identical to (and function identically to) the elements of Fig. 1, and the above description thereof will not be repeated.
The system of Fig. 1A differs from that of Fig. one in which a control signal V1 (emitted at the output of the multiplier 214) is used to scale the control signal C1 (emitted at the output of the limiting element 111) instead of the control signal S1 (emitted at the output of the processor 134), and a control signal V2 (emitted at the output of multiplier 215) is used to scale the control signal C2 (emitted at the output of the limiting element 112) instead of the control signal S2 (emitted at the output processor 135). In Fig. 1A, the scaling of the attenuation gain control signal C1 unprocessed in response to the sequence of attenuation control values V1 according to the invention is performed by multiplying (in item 114) each unprocessed gain control value of the signal C1 by a corresponding value from among the attenuation control values V1, to generate the signal S3, and the scaling of the unprocessed gain control signal C2 in response to the sequence of attenuation control values V2 according to the invention is performed by multiplying (in item 115) each unprocessed gain control value of the signal C2 by a corresponding value from among the attenuation control values V2, to generate the signal S4.
To generate the sequence of the attenuation control values V1, the signal Q (emitted at the output of the processor 131) is output to an input of the multiplier 214, and the control signal S1 (issued at the output of the processor 134) is issued to the other input of the multiplier 214. The output of the multiplier 214 is the sequence of the attenuation control values V1. Each of the attenuation control values V1 is one of the voice probability values determined by the signal Q, scaled by a corresponding value from the attenuation control values S1.
Similarly, to generate the sequence of attenuation control values V2, the signal Q (emitted at the output of the processor 131) is output to an input of the multiplier 215, and the control signal S2 (emitted at the output of the processor 135) is emitted to the other input of the multiplier 215. The output of the multiplier 215 is the sequence of the attenuation control values V2. Each of the attenuation control values V2 is one of the voice probability values determined by the signal Q, scaled by a corresponding value from the attenuation control values S2.
The system of Fig. 1 (or that of Fig. 1A) can be implemented in a software by a processor (for example, processor 501 of Fig. 5) that has been programmed to implement the operations described in the system of Fig. 1 (or 1A). Alternatively, it can be implemented in hardware with connected circuit elements as shown in Fig. 1 (or 1A).
In the variations in the example of Fig. 1 (or in that of Fig. 1A), the scaling of the attenuation gain control signal C1 in response to the gain control signal S1 (or V1) of attenuation according to the invention (to generate an attenuation gain control signal to operate amplifier 116) can be performed in a non-linear manner. For example, said nonlinear scaling can generate an attenuation gain control signal (which replaces signal S3) that does not cause attenuation on the part of amplifier 116 (i.e., the application of a unit gain by amplifier 116 and, thus, a zero attenuation of channel 103) when the current value of the signal S1 (or V1) is below a threshold, and causes the current value of the attenuation gain control signal (which replaces signal S3) to be equal to the current value of signal C1 (so that signal S1 (or V1) does not change the current value of C1 ) when the current value of signal S1 exceeds the threshold. Alternatively, another linear or non-linear scaling of signal C1 (in response to the attenuation gain control signal of the invention) can be performed to generate an attenuation gain control signal to drive the amplifier 116 . For example, said scaling of signal C1 may generate an attenuation gain control signal (which replaces signal S3) that does not cause attenuation on the part of amplifier 116 (i.e., the application of a unit gain by the amplifier 116) when the current value of signal S1 (or V1) is below a threshold, and causes the current value of the attenuation gain control signal (which replaces signal S3) to be equal to the current value of signal C1 multiplied by the current value of signal S1 or V1 (or some other value determined to from this product) when the current value of signal S1 (or V1) exceeds the threshold.
Similarly, in variations of the example of Fig. 1 (or Fig. 1A), the scaling of the attenuation gain control signal C2 in response to the gain control signal S2 (or V2) of attenuation according to the invention (to generate an attenuation gain control signal to operate the amplifier 117) can be performed in a non-linear manner. For example, said non-linear scaling can generate an attenuation gain control signal (which replaces signal S4) that does not cause attenuation on the part of amplifier 117 (i.e., the application of a unit gain by amplifier 117 and, thus, a null attenuation of channel 102) when the current value of signal S2 (or V2) is below a threshold, and causes the current value of the attenuation gain control signal (which replaces signal S4) to be equal to the current value of signal C2 (so that signal S2 or V2 does not change the current value of C2) when The current value of signal S2 (or V2) exceeds the threshold. From alternatively, another linear or non-linear scaling of the C2 signal (in response to the S2 or V2 signal of attenuation gain control) can be performed to generate an attenuation gain control signal to drive the amplifier 117. For example, said scaling of the signal C2 can generate an attenuation gain control signal (which replaces the signal S4) that does not cause an attenuation on the part of the amplifier 117 (i.e., the application of a unit gain by the amplifier 117) when the current value of signal S2 (or V2) is below a threshold, and causes the current value of the attenuation gain control signal (which replaces signal S4) to be equal to the current value of signal C2 multiplied by the current value of signal S2 or V2 (or some other value determined at from this product) when the current value of signal S2 (or V2) exceeds the threshold.
Another example (225) of a system of the invention will be described with reference to Fig. 2. In response to a multichannel audio signal comprising a voice channel 101 (central C channel) and two voiceless channels 102 and 103 ( L and R channels left and right), the system of Fig. 2 filters the voiceless channels to generate a filtered multichannel output audio signal comprising voice channel 101 and voiceless channels 118 and 119 (L 'channels and R 'left and right filtered).
In the system of Fig. 2 (as in the system of Fig. 1), channels 102 and 103 without voice are output to the attenuation amplifiers 117 and 116, respectively. During operation, the attenuation amplifier 117 is driven by a control signal S6 (which is indicative of a sequence of control values and, thus, is also called a sequence S6 of control values) emitted from the control element 115 multiplication, and the attenuation amplifier 116 is actuated by the control signal S5 (which is indicative of a sequence of control values and, thus, is also called the sequence S5 of control values), issued from multiplication element 114. Elements 114, 115, 130, 131, 132, 134 and 135 of Fig. 2 are identical to (and function identically a) the elements of Fig. 1, and their description will not be repeated.
The system of Fig. 2 measures the power of the signals in each of the channels 101, 102 and 103 with a bank of estimators 201, 202 and 203 of power. Unlike their counterparts in Fig. 1, each of the power estimators 201, 201 and 203 measures the distribution of the signal power along the frequency (i.e., the power in each frequency band different from among a set of frequency bands of the relevant channel), resulting in a power spectrum instead of a single number for each channel. The spectral resolution of each power spectrum ideally matches the spectral resolution of the intelligibility prediction models implemented by elements 205 and 206 (described below).
The power spectra are supplied to the comparison circuit 204. The purpose of circuit 204 is to determine the attenuation to be applied to each voiceless channel to ensure that the signal in the voiceless channel does not reduce the intelligibility of the signal in the voice channel so that it is less than a predetermined criterion. This functionality is achieved using an intelligibility prediction circuit (205 and 206) that predicts the intelligibility of the voice from the power spectra of the signal (201) of the voice channel and signals (202 and 203) of the channels without voice. Intelligibility prediction circuits 205 and 206 can implement an adequate intelligibility prediction model according to design choices and commitments. Examples are the voice intelligibility index as specified in ANSI S3.5- 1997 ("Methods for Calculation of the Speech Intelligibility Index") and the voice recognition sensitivity model of Muesch and Buus ("Using statistical decision theory to predict speech intelligibility. I. Model structure "Journal of the Acoustical Society of America, 2001, Vol. 109, p. 2896-2909). It is clear that the output of the intelligibility prediction model has no meaning when the signal in the voice channel is somewhat different from a voice. Despite this, from now on, reference will be made to the output of the intelligibility prediction model as predicted speech intelligibility. The perceived error is accounted for in subsequent processing by scaling the gain values emitted from circuit 204 compared to parameters S1 and S2, each of which is related to the probability that the signal in the voice channel is indicative of a voice.
The intelligibility prediction models have in common that they predict an increased or unaltered voice intelligibility as a result of the reduction of the level of the signal without voice. Continuing in the process flow of Fig. 2, comparison circuits 207 and 208 compare the predicted intelligibility with a predetermined criterion value. If element 205 determines that the level of channel 103 without voice is so low that the predicted intelligibility exceeds the criterion, a gain parameter, which is initialized to 0 dB, is recovered from circuit 209 is provided to circuit 211 as the output C3 of comparison circuit 204. If element 206 determines that the level of channel 102 without voice is so low that the predicted intelligibility exceeds the criterion, a gain parameter, which is initialized to 0 dB, is recovered from circuit 210 and is provided to circuit 212 as the C4 output of comparison circuit 204. If element 205 or 206 determines that the criterion is not met, the gain parameter (in the relevant element between elements 209 and 210) is reduced by a fixed amount and the intelligibility prediction is repeated. An adequate step size to reduce the gain is 1 dB. The iteration just described continues until the predicted intelligibility meets or exceeds the value of the criterion.
Of course, it is possible that the signal in the voice channel is such that the intelligibility criterion cannot be reached even in the absence of a signal in the voiceless channel. An example of such a situation is a very low level voice signal or with a severely restricted bandwidth. In this case, a point will be reached where any further reduction of the gain applied to the voiceless channel will not affect the predicted speech intelligibility and the criteria will never be met. Under said condition, the loop formed by elements 205, 207 and 209 (or elements 206, 208 and 210) continues indefinitely, and additional logic (not shown) can be applied to interrupt the loop. A particularly simple example of such logic is to count the number of iterations and exit the loop once a predetermined number of iterations has been exceeded.
The scaling of the unprocessed gain control signal C3 in response to the attenuation gain control signal S1 according to the invention can be performed by multiplying (in item 114) each unprocessed gain control value of the signal C3 by a corresponding value from among the scaled averaged difference values of signal S1, to generate signal S5. The scaling of the unprocessed gain control signal C4 in response to the attenuation gain control signal S2 according to the invention can be performed by multiplying (in element 115) each unprocessed gain control value of the signal C4 by a corresponding value from among the scaled averaged difference values of the signal S2, to generate the signal S6.
The system of Fig. 2 can be implemented in software by a processor (for example, processor 501 of Fig. 5) that has been programmed to implement the operations described in the system of Fig. 2. Alternatively, It can be implemented in hardware with connected circuit elements as shown in Fig. 2.
In variations of the example of Fig. 2, the scaling of the unprocessed gain control signal C3 in response to the attenuation gain control signal S1 (to generate an attenuation gain control signal to drive the amplifier 116) can be performed in a non-linear manner. For example, said non-linear scaling can generate an attenuation gain control signal (which replaces signal S5) that does not cause attenuation on the part of amplifier 116 (i.e., the application of a unit gain by amplifier 116 and, thus, a null attenuation of channel 103) when the current value of the signal S1 is below a threshold, and causes the current value of the attenuation gain control signal (which replaces signal S5) to be equal to the current value of signal C3 (so that signal S1 does not change the current value of C3) when the value Current of signal S1 exceeds the threshold. Alternatively, another linear or non-linear scaling of the signal C3 (in response to the attenuation gain control signal of the invention) can be performed to generate an attenuation gain control signal to drive the amplifier 116. For example, said scaling of the signal C3 can generate an attenuation gain control signal (which replaces the signal S5) that does not cause an attenuation on the part of the amplifier 116 (i.e., the application of a unit gain by the amplifier 116) when the current value of signal S1 is below a threshold. and causes the current value of the attenuation gain control signal (which replaces signal S5) to be equal to the current value of signal C3 multiplied by the current value of signal S1 (or some other value determined from this product) when the current value of signal S1 exceeds the threshold.
Similarly, in variations of the example of Fig. 2, the scaling of the attenuation gain control signal C4 in response to the attenuation gain control signal S2 (to generate a gain control signal of Attenuation to operate the amplifier 117) can be performed in a non-linear manner. For example, said non-linear scaling can generate an attenuation gain control signal (which replaces signal S6) that does not cause an attenuation of amplifier 117 (i.e., the application of a unit gain by amplifier 117 and, in this way, a null attenuation of channel 102) when the current value of signal S2 is below a threshold and causes the current value of the attenuation gain control signal (which replaces signal S6) to be equal to the current value of signal C4 (so that signal S2 does not change the current value of C4) when the current value of signal S2 exceeds the threshold. Alternatively, another linear or non-linear scaling of the signal C4 (in response to the attenuation gain control signal S2 of the invention) can be performed to generate an attenuation gain control signal to drive the amplifier 117. For example, said scaling of the signal C4 can generate an attenuation gain control signal (which replaces the signal S6) that does not cause an attenuation of the amplifier 117 (i.e., the application of a unit gain by the amplifier 117) when the current value of signal S2 is below a threshold, and causes the current value of the attenuation gain control signal (which replaces signal S6) to be equal to the current value of signal C4 multiplied by the current value of signal S2 (or some other value determined from this product) when the current value of signal S2 exceeds the threshold.
Another example (225 ') of the system will be described with reference to Fig. 2A. In response to a multichannel audio signal comprising a voice channel 101 (central C channel) and two channels 102 and 103 without voice (left and right L and R channels), the system of Fig. 2A filters the voiceless channels to generate a filtered multichannel output audio signal comprising voice channel 101 and channels 118 and 119 without filtered voice (left and right L 'and R' channels filtered).
In the system of Fig. 2A (as in the system of Fig. 2), channels 102 and 103 without voice are output to attenuation amplifiers 117 and 116, respectively. During operation, the attenuation amplifier 117 is driven by a control signal S6 (which is indicative of a sequence of control values and, thus, is also called a sequence S6 of control values) emitted from the control element 115 multiplication, and the amplifier Attenuation 116 is driven by the control signal S5 (which is indicative of a sequence of control values and, thus, is also called the sequence S5 of control values), issued from multiplication element 114. Elements 201, 202, 203, 204, 114, 115, 130 and 134 of Fig. 2A are identical to (and function identically to) elements with identical numbering in Fig. 2, and the description will not be repeated thereof.
The system of Fig. 2A differs from that of Fig. 2 in two main aspects. First, the system is configured to generate (that is, to derive) a channel without voice "derived" (LR) from two individual channels without voice (102 and 103) of the input audio signal, and to determine the attenuation control values (V3) in response to this channel without derived voice. On the contrary, the system of Fig. 2 determines the attenuation control values S1 in response to a voiceless channel (channel 102) of the input audio signal and determines the attenuation control values S2 in response to another voiceless channel (channel 103) of the signal input audio During operation, the system of Fig. 2A attenuates each voiceless channel of the input audio signal (each of channels 102 and 103) in response to the same set of attenuation control values V3. During operation, the system of Fig. 2 attenuates the voiceless channel 102 of the input audio signal in response to the attenuation control values S2, and attenuates the channel 103 without voice of the input audio signal in response to a different set of attenuation control values (S1 values).
The system of Fig. 2A includes an addition element 129 whose inputs are coupled to receive channels 102 and 103 without voice from the input audio signal. The channel without derived voice (LR) is emitted at the output of element 129. The voice probability processing element 130 emits the voice probability signal P in response to the LR channel without voice derived from the element 129. In Fig. . 2A, the signal P is indicative of a sequence of voice probability values for the channel without derived voice. Typically, the voice probability signal P of Fig. 2A is a value monotonically related to the probability that the signal in the channel without derived voice is a voice. The voice probability signal Q (generated by the processor 131) of Fig. 2A is identical to the voice probability signal Q of Fig. 2 described above.
A second main aspect in which the system of Fig. 2A differs from that of Fig. 2 is as follows. In Fig. 2A, the control signal V3 (emitted at the output of the multiplier 214) (instead of the control signal S1 emitted at the output of the processor 134) is used to scale the control signal C3 gain control unprocessed (emitted at the output of the element 211), and the control signal V3 is also used (instead of the control signal S2 emitted at the output of the processor 135 of Fig. 2) to scale the unprocessed gain gain control signal C4 (emitted at the output of element 212). In Fig. 2A, the scaling of the attenuation gain control signal C3 not processed in response to the sequence of attenuation control values indicated by the signal V3 (referred to as attenuation control values V3) is performed by multiplying (in item 114) each unprocessed gain control value of signal C3 by a corresponding value from among the attenuation control values V3, to generate signal S5, and the scaling of the unprocessed gain control signal C4 in response to the sequence of attenuation control values V3 is performed by multiplying (in item 115) each unprocessed gain control value of the C4 signal by a corresponding value from among the attenuation control values V3, to generate the signal S6.
During operation, the system of Fig. 2A generates the sequence of attenuation control values V3 as follows. The voice probability signal Q (emitted at the output of the processor 131 of Fig. 2A) is output to an input of the multiplier 214, and the attenuation control signal S1 (emitted at the output of the processor 134) is output at the other input of multiplier 214. The output of multiplier 214 is the sequence of attenuation control values V3. Each of the attenuation control values V3 is one of the voice probability values determined by the signal Q, scaled by a corresponding value from the attenuation control values S1.
Another example (325) of a system will be described with reference to Fig. 3. In response to a multichannel audio signal comprising a voice channel 101 (central C channel) and two voiceless channels 102 and 103 (L and R left and right), the system of Fig. 3 filters the voiceless channels to generate a filtered multichannel output audio signal comprising voice channel 101 and voiceless channels 118 and 119 (channels L 'and R' left and right filtered).
In the system of Fig. 3, each of the signals in the three input channels is divided into its spectral components by the filter bank 301 (for channel 101), the filter bank 302 (for channel 102) and filter bank 303 (for channel 103). Spectral analysis can be achieved with banks of N-channel filters in the time domain. According to one example, each filter bank divides the frequency range into 1/3 octave bands or mimics the filtering that is supposed to occur in the human internal wave. The fact that the signal emitted from each filter bank consists of N sub-signals is illustrated by the use of thick lines.
In the system of Fig. 3, the frequency components of the signals on channels 102 and 103 without voice are output to the attenuation amplifiers 117 and 116, respectively. During operation, the attenuation amplifier 117 is driven by a control signal S8 (which is indicative of a sequence of control values and, thus, is also called a sequence S8 of control values) emitted from the element 115 ' of multiplication, and the attenuation amplifier 116 is actuated by a control signal S7 (which is indicative of a sequence of control values and, thus, It is also called sequence S7 of control values) emitted from the multiplication element 114 '. Elements 130, 131, 132, 134 and 135 of Fig. 3 are identical to (and function identically to) the elements with identical numbering in Fig. 1, and their description will not be repeated.
The procedure of Fig. 3 can be recognized as a side branch process. Following the signal path shown in Fig. 3, the N sub-signals generated in bank 302 for channel 102 without voice are each scaled by a member of a set of N gain values by the attenuation amplifier 117 and the N sub-signals generated in the Bank 303 for channel 103 without voice are scaled by a member of a set of N gain values by the attenuation amplifier 116. The derivation of these gain values will be described later. Next, the scaled sub-signals are recombined in a single audio signal. This can be done by a simple sum (by the sum circuit 313 for the channel 102 and by the sum circuit 314 for the channel 103). Alternatively, a synthesis filter bank that matches the analysis filter bank can be used. This procedure results in the signal R 'without modified voice (118) and the signal L' without modified voice (119).
Now, in relation to the description of the lateral branch path of the procedure of Fig. 3, each filter bank output is made available to a corresponding bank of N estimators (304, 305 and 306) of power. The resulting power spectra for channels 101 and 102 serve as inputs to an optimization circuit 307 which has as its output a C6 vector of N-dimensional gains. The resulting power spectra for channels 101 and 103 serve as inputs to an optimization circuit 308 which has as its output a C5 vector of N-dimensional gains. The optimization employs both an intelligibility prediction circuit (309 and 310) and a sound intensity calculation circuit (311 and 312) to find the gain vector that maximizes the sound intensity of each voiceless channel while maintaining a predetermined level of predicted intelligibility of the voice signal on channel 101. Suitable models for predicting intelligibility have been described with reference to Fig. 2. The sound intensity calculation circuits 311 and 312 can implement an adequate sound intensity prediction model according to the design options and commitments. Examples of suitable models are the American National Standard ANSI S3.4-2007 "Procedure for the Computation of Loudness of Steady Sounds" and the German standard DIN 45631 "Berechnung des Lautstarkepegels und der Lautheit aus dem Gerauschspektrum".
Depending on the computational resources available and the restrictions imposed, the shape and complexity of the optimization circuits (307, 308) can vary widely. According to one example, a restricted, multidimensional, iterative optimization of N free parameters is used. Each parameter represents the gain applied to one of the frequency bands of the voiceless channel. Standard techniques can be applied, such as following the steepest gradient in the N dimensional search space to find the maximum. In another embodiment, a computationally less demanding approach imposes a restriction on profit vs. profit functions. frequency so that they are members of a small set of possible profit vs. profit functions. frequency, such as a set of different spectral gradients or attenuating or limiting filters ("shelf"). With this additional restriction, the optimization problem can be reduced to a small number of one-dimensional optimizations. In yet another example, an exhaustive search is conducted on a very small set of possible gain functions. This last approach could be particularly desirable in real-time applications where a constant computational load and search speed is desired.
People with ordinary knowledge in the field will easily recognize the additional restrictions that may be imposed on the optimization and are also reflected by the respective embodiments of the dependent claims.
An example is to restrict the sound intensity of the channel without modified voice so that it is not greater than the sound intensity before the modification. Another example is to impose a limit on gain differences between adjacent frequency bands in order to limit the possibility of temporary overlap in the reconstruction filter bank (313, 314) or in order to reduce the possibility of ring modifications. objectionable Desirable restrictions depend on both the technical implementation of the filter bank and the commitment chosen between the improvement of intelligibility and the modification of the bell. For the sake of clarity in the illustration, these restrictions are omitted in Fig. 3.
The scaling of the N-dimensional attenuation gain control vector C6 in response to the attenuation gain control signal S2 according to the invention can be performed by multiplying (in element 115 ') each unprocessed gain control value of the vector C6 by a corresponding value between the scaled averaged difference values of the signal S2, to generate the N8-dimensional attenuation gain control vector S8. The scaling of the N-dimensional attenuation gain control vector C5 in response to the attenuation gain control signal S1 according to the invention can be performed by multiplying (in item 114 ') each unprocessed gain control value of the vector C5 by a corresponding value between the scaled averaged difference values of the signal S1, to generate the vector S7 of gain control of N-dimensional attenuation.
The system of Fig. 3 can be implemented in software by a processor (for example, processor 501 of Fig. 5) that has been programmed to implement the described operations of the system of Fig. 3. Alternatively, it can be implemented in hardware with connected circuit elements as shown in Fig. 3.
In variations in the example of Fig. 3, the scaling of the attenuation gain control vector C5 not processed in response to the attenuation gain control signal S1 according to the invention (to generate an attenuation gain control vector to drive amplifier 116) it can be performed in a non-linear manner. For example, said non-linear scaling can generate an attenuation gain control vector (which replaces vector S7) that does not cause an attenuation of amplifier 116 (i.e., the application of a unit gain by amplifier 116 and, of this way, a null attenuation of channel 103) when the current value of signal S1 is below a threshold, and causes the current values of the attenuation gain control vector (which replaces the vector S7) to be equal to the current values of the vector C5 (so that the signal S1 does not modify the current values of C5) when the current value of signal S1 exceeds the threshold. Alternatively, another linear or nonlinear scaling of vector C5 (in response to the attenuation gain control signal S1 of the invention) can be performed to generate an attenuation gain control vector to drive amplifier 116. For example, said scaling of vector C5 can generate an attenuation gain control vector (which replaces vector S7) that does not cause attenuation on the part of amplifier 116 (i.e., the application of a unit gain by amplifier 116 ) when the current value of signal S1 is below a threshold, and causes the current value of the attenuation gain control vector (which replaces vector S7) to be equal to the current value of vector C5 multiplied by the current value of signal S1 (or some other value determined from this product) when the current value of signal S1 exceeds the threshold.
Similarly, in variations of the example of Fig. 3, the scaling of the attenuation gain control vector C6 in response to the attenuation gain control signal S2 (to generate an attenuation gain control vector to operate amplifier 117) it can be performed in a non-linear manner. For example, said non-linear scaling can generate an attenuation gain control vector (which replaces vector S8) that does not cause attenuation on the part of amplifier 117 (ie, the application of a unit gain on the part of amplifier 117 and , in this way, a null attenuation of channel 102) when the current value of the signal S2 is below a threshold and causes the current values of the attenuation gain control vector (which replaces the vector S8) to be equal to the current values of the vector C6 (so that signal S2 does not change the current values of C6) when the current value of signal S2 exceeds the threshold. Alternatively, another linear or nonlinear scaling of vector C6 (in response to the attenuation gain control signal S2 of the invention) can be performed to generate an attenuation gain control vector to drive amplifier 117. For example, said scaling of vector C6 can generate an attenuation gain control vector (which replaces vector S8) that does not cause an attenuation on the part of amplifier 117 (i.e., the application of a unit gain by amplifier 117) when the current value of signal S2 is below a threshold, and causes the current value of the attenuation gain control vector (which replaces vector S8) to be equal to the current value of vector C6 multiplied by the current value of signal S2 (or some other value determined from this product) when the current value of signal S2 exceeds the threshold.
From this description, it will be evident to people with ordinary knowledge in the subject the way in which the system of Fig. 1, 1A, 2, 2A or 3 (and the variations in any of them) can be modified Filter a multichannel audio input signal that has a voice channel and any number of channels without voice. An attenuation amplifier (or a software equivalent thereof) is provided for each voiceless channel, and an attenuation gain control signal (for example, scaling an attenuation gain control signal) will be generated to drive each amplifier. of attenuation (or software equivalent to it).
As described, the system of Fig. 1, 1A, 2, 2A or 3 (and each of the many variations therein) are operable to carry out practical embodiments of the method of the invention to filter a signal. multichannel audio that has a voice channel and at least one voiceless channel to improve the intelligibility of the voice determined by the signal. In a first class of said embodiments, the procedure includes the steps of:
(a) determine at least one attenuation control value (for example, signal S1 or S2 of Fig. 1, 2 or 3, or signal V1, V2 or V3 of Fig. 1A or 2A) indicative of a measure of similarity between the voice-related content determined by the voice channel and the voice-related content determined by at least one voiceless channel of the audio signal; Y
(b) attenuate at least one voiceless channel of the audio signal in response to at least one attenuation control value (for example, in element 114 and amplifier 116, or in element 115 and amplifier 117, of Fig. 1, 1A, 2, 2A, or 3).
Typically, the attenuation step comprises scaling an unprocessed attenuation control signal (for example, the attenuation gain control signal C1 or C2 of Fig. 1 or 1A, or the C3 or C4 signal of Fig. 2 or 2A) for the voiceless channel in response to at least one attenuation control value. Preferably, the voiceless channel is attenuated to improve the intelligibility of the voice determined by the voice channel without undesirably attenuating the voice enhancing content determined by the voiceless channel. In some examples of embodiments in the first class, step (a) includes a step of generating an attenuation control signal (eg, signal S1 or S2 of Fig. 1, 2 or 3, or the signal V1, V2 or V3 of Fig. 1A or 2A) indicative of a sequence of attenuation control values, in which each of the attenuation control values is indicative of a measure of similarity between the voice related content determined by the voice channel and the voice related content determined by at least one voiceless channel of the audio signal at a different time (for example, at a different time interval) , and step (b) includes the steps of: scaling an attenuation gain control signal (for example, signal C1 or C2 of Fig. 1 or 1A, or signal C3 or C4) of Fig. 2 or 2A) in response to the attenuation control signal to generate a scaled gain control signal (for example, signal S3 or S4 of Fig. 1 or 1A, or signal S5 or S6 of Fig. 2 or 2A), and apply the gain gain control signal to attenuate the voiceless channel (for example, issue the gain gain control signal to the attenuation circuit 116 or 117, of Fig. 1, 1A, 2 or 2A, to control the attenuation of at least one voiceless channel through the attenuation circuit). For example, step (a) includes a step of comparing a first sequence of voice-related features (for example, the signal Q of Fig. one or 2) indicative of the voice related content determined by the voice channel with a second sequence of voice related features (for example, the signal P of Fig. one or 2) indicative of the voice related content determined by the voiceless channel to generate the attenuation control signal, and each of the attenuation control values indicated by the attenuation control signal is indicative of a similarity measure between the first sequence of voice-related features and the second sequence of voice-related features at a different time (for example, at a different time interval). In some examples of the embodiments, each attenuation control value is a gain control value.
In some embodiments in the first class, each attenuation control value is monotonically related to the probability that the voiceless channel is indicative of voice enhancing content that improves the intelligibility (or other perceived quality) of the voice content determined by the voice channel In some other examples related to embodiments in the first class, each attenuation control value is monotonically related to an expected voice enhancer value of the voiceless channel (for example, a measure of probability that the voiceless channel is indicative of content voice enhancer, multiplied by a measure of the perceived quality improvement that the voice enhancing content determined by the voiceless channel provides the voice content determined by the multichannel signal). For example, when step (a) includes a compare stage (for example, in element 134 or 135 of Fig. 1 or Fig. 2) a first sequence of voice related features indicative of the voice related content determined by the voice channel with a second sequence of voice related features indicative of the voice related content determined by the voiceless channel, the first voice related feature sequence can be a sequence of voice probability values, each of which indicates the probability at a different time (for example, at a different time interval) that the voice channel is indicative of voice (instead of audio content other than a voice), and the second sequence of voice-related features can also be a sequence of voice probability values, each of which indicates the probability at a different time (for example, at a different time interval) that the voiceless channel is indicative of voice.
As described, the system of Fig. 1, 1A, 2, 2A or 3 (and each of the many variations therein) is also operable to perform a second class of embodiments of the method of the invention to filter a multichannel audio signal that has a voice channel and at least one voiceless channel to improve the intelligibility of the voice determined by the signal. In the second class of embodiments, the procedure includes the steps of:
(a) compare a voice channel characteristic and a voiceless channel characteristic to generate at least one attenuation value (for example, values determined by signal C1 or C2 of Fig. 1, or by signal C3 or C4 of Fig. 2, or by the signal C5 or C6 of Fig. 3) to control the attenuation of the voiceless channel relative to the voice channel; Y
(b) adjust the at least one attenuation value in response to at least one voice improvement probability value (for example, signal S1 or S2 of Fig. 1, 2 or 3) to generate at least one value of adjusted attenuation (for example, determined values of signal S3 or S4 of Fig. 1, or by signal S5 or S6 of Fig. 2, or by signal S7 or S8 of Fig. 3) to control the attenuation of the voiceless channel in relation to the voice channel. Typically, the adjustment step is or includes scalar (for example, in element 114 or 115 of Fig. 1, 2 or 3) each of said attenuation values in response to one of said voice enhancement probability values to generate one of said adjusted attenuation values. Typically, each voice improvement probability value is indicative of (for example, is monotonically related to) the probability that the voiceless channel is indicative of voice enhancing content (content that improves intelligibility or other perceived quality of the content of voice determined by the voice channel). In related examples, the voice enhancement probability value is indicative of an expected voice enhancer value of the voiceless channel (for example, a measure of the probability that the voiceless channel is indicative of voice enhanced content multiplied by a measure of improvement of the perceived quality that the voice improvement content determined by the voiceless channel provides the voice content determined by the multichannel audio signal). In some embodiments in the second class, the voice enhancement probability value is a sequence of comparison values (eg, difference values) determined by a procedure that includes a step of comparing a first sequence of voice related features, indicative of the voice related content determined by the voice channel, with a second sequence of voice-related features, indicative of the voice-related content determined by the voiceless channel, and each of the comparison values is a measure of similarity between the first sequence of voice-related features and the second sequence of voice-related features at a time different (for example, in a different time interval). In typical embodiments in the second class, the method also includes the step of attenuating the voiceless channel (for example, in amplifier 116 or 117 of Fig. 1.2 or 3) in response to at least one attenuation value tight. Step (b) may comprise scaling the at least one attenuation value (for example, each attenuation value determined by the signal C1 or C2 of Fig. 1), or another attenuation value determined by an attenuation gain control signal or other unprocessed attenuation control signal) in response to at least one voice improvement probability value (for example, the corresponding value determined by signal S1 or S2 of Fig. 1).
During operation of the system of Fig. one To implement an embodiment in the second class, each attenuation value determined by signal C1 or C2 is a first factor indicative of an amount of attenuation of the channel without voice necessary to limit the ratio of signal power in the channel without Voice to signal strength in the voice channel so that it does not exceed a predetermined threshold, scaled by a second factor monotically related to the probability that the voice channel is indicative of voice. Typically, the adjustment step in these embodiments is (or includes) scaling each attenuation value C1 or C2 by a voice enhancement probability value (determined by signal S1 or S2) to generate an adjusted attenuation value (determined by signal S3 or S4), where the voice enhancement probability value is a factor monotically related to one of: a probability that the voiceless channel is indicative of voice-enhancing content (content that improves intelligibility or other perceived quality of the voice content determined by the multi-channel signal) and an expected voice-enhancing value of the voiceless channel (e.g., a measure of the probability that the voiceless channel is indicative of voice enhancement content multiplied by a measure of the perceived quality improvement that the voice enhancing content in the voiceless channel provides for the voice content determined by the multichannel signal) .
During operation of the system of Fig. 2 To carry out an embodiment in the second class, each attenuation value determined by signal C3 or C4 is a first factor indicative of an amount (for example, the minimum amount) of channel attenuation without voice sufficient to cause intelligibility of predicted voice determined by the voice channel in the presence of content determined by the voiceless channel exceeds a predetermined threshold value, scaled by a second factor monotically related to the probability that the voice channel is indicative of a voice. Preferably, the predicted speech intelligibility determined by the voice channel in the presence of the content determined by the voiceless channel is determined according to an intelligibility prediction model based on psycho-acoustics. Typically, the adjustment step in these embodiments is (or includes) scaling each of said attenuation values by one of said voice enhancement probability values (determined by signal S1 or S2) to generate an adjusted attenuation value ( determined by the signal S5 or S6), where the voice enhancement probability value is a factor monotonically related to one of: a probability that the voiceless channel is indicative of voice enhancement content and an expected voice enhancement value of the voiceless channel.
During operation of the system of Fig. 3 To implement an embodiment in the second class, each attenuation value determined by signal C1 or C2 is determined by steps that include determining (in element 301, 302 or 303) a power spectrum indicative of the power as a function of the frequency, of each of the voice channel 101 and the channels 102 and 103 without voice, and make a determination in the frequency domain of the attenuation value, determining attenuation in this way as a function of the frequency to be applied to the frequency components of the voiceless channel.
In one class of embodiments, the invention is a method and system for improving the voice determined by a multichannel audio input signal. In some examples thereof, the system of the invention includes an analysis module or subsystem (for example, elements 130-135, 104-109, 114 and 115 of Fig. 1, or elements 130-135, 201-204 , 114 and 115 of Fig. 2) configured to analyze the multichannel input signal to generate attenuation control values, and an attenuation subsystem (for example, amplifiers 116 and 117 of Fig. 1 or Fig. 2). The attenuation subsystem includes attenuation circuit (powered by at least some of the attenuation control values) coupled and configured to apply “ducking” to each voiceless channel of the input signal to generate an output signal from filtered audio The attenuation circuit is activated by control values in the sense that the attenuation applied to the voiceless channels is determined by the current values of the control values.
In some embodiments, a ratio of the power of the voice channel (eg, the center channel) to the power of the voiceless channel (eg, the side channel and / or the rear channel) is used to determine how much attenuation (" ducking ”) should be applied to each channel without voice. For example, in the example of Fig. 1, the profit applied by each of the attenuation amplifiers 116 and 117 are reduced in response to a decrease in a gain control value (emitted from element 114 or element 115) that is indicative of a lower power (within the Kmites) of voice channel 101 with relation to the power of a voiceless channel (left channel 102 or right channel 103) determined in the analysis module (i.e. an attenuation amplifier attenuates a voiceless channel in relation to the voice channel when the power of the voice channel decreases (within the limits) in relation to the power of the voiceless channel), assuming there are no changes in the probability ( as determined in the analysis module) that the voiceless channel includes voice enhancing content that improves the voice content determined by the voice channel.
In some alternative examples, a modified version of the analysis module of Fig. 1 or Fig. 2 individually processes each or more frequency subbands of each channel of the input signal. Specifically, the signal in each channel can be passed through a band pass filter bank, which provides three sets of n subbands: {L1, L2, ..., Ln}, {C1, C2, ... , Cn} and {R1, R2, ..., Rn}. The matching subbands are passed before the analysis module of Fig. 1 (or Fig. 2), and the filtered sub-signals (the outputs of the attenuation amplifiers for the voiceless channels, and the sub - voice channels not filtered) are combined by means of sum circuits to generate the filtered multichannel audio output signal. To carry out in each sub-band the operations performed by element 109 of Fig. 1, a separate threshold value On (corresponding to the threshold value O of element 109) can be selected for each sub-band. A good option is a set in which On is proportional to the average number of voice characteristics carried in the corresponding frequency region; that is, bands at the ends of the frequency spectrum are assigned lower thresholds than bands corresponding to the dominant voice frequencies. This implementation can offer a very good compromise between computational complexity and performance.
Fig. 4 is a block diagram of a system 420 (a configurable audio DSP) that has been configured to carry out an embodiment of the method of the invention. System 420 includes programmable circuitine 422 DSP (an active voice enhancement module of system 420) coupled to receive a multi-channel audio input signal. For example, the Lin and Rin channels without voice of the signal may correspond to channels 102 and 103 of the input signal described with reference to Figs. 1, 1A, 2, 2A and 3, the signal may also include additional voiceless channels (for example, the left rear and right rear channels), and the signal Cin channel of the signal may correspond to channel 101 of the signal of entry described with reference to Figs. 1, 1A, 2, 2A and 3. Circuit 422 is configured in response to control data from a control interface 421 to carry out an embodiment of the method of the invention, to generate a multi-channel audio output signal with improved voice in response to the input signal. audio To program the system 420, the appropriate software is activated from an external processor to control the control interface 421, and the interface 421 issues in response the appropriate control data to circuit 422 to configure circuit 422 to perform the procedure of the invention.
During operation, an audio DSP that has been configured to perform a voice enhancement according to the invention (for example, the system 420 of Fig. 4) is coupled to receive an N-channel audio input signal, and the DSP typically performs a variety of operations on the input audio (or a processed version thereof) in addition to (as well as) a voice enhancement. For example, the system 420 of Fig. 4 it can be implemented to perform other operations (on the output of circuit 422) in the processing subsystem 423.
According to yet another example, an audio DSP is operable to carry out an embodiment of the method of the invention after being configured (for example, programmed) to generate an output audio signal in response to an input audio signal. by performing the procedure on the input audio signal.
In some examples, the system of the invention is or includes a general purpose processor coupled to receive or generate input data indicative of a multichannel audio signal. The processor is programmed with software (or firmware) and / or if not configured (for example, in response to control data) to perform any of a variety of operations on the input data, including an embodiment of the procedure of the invention. The computer system of Fig. 5 is an example of such a system. The system of Fig. 5 includes a general purpose processor 501 that is programmed to perform any of a variety of operations on the input data, including an embodiment of the method of the invention.
The computer system of Fig. 5 also includes an input device 503 (for example, a mouse and / or a keyboard) coupled to the processor 501, a storage medium 504 coupled to the processor 501 and a display device 505 coupled to the processor 501. Processor 501 is programmed to implement the inventive procedure in response to the instructions and data entered by the user manipulation of the input device 503. The computer readable storage medium 504 (for example, an optical disk or other tangible object) has a computer code stored therein that is suitable for programming the processor 501 to perform an embodiment of the method of the invention. During operation, the processor 501 executes the computer code to process data indicative of a multichannel audio input signal according to the invention to generate output data indicative of a multichannel audio output signal.
The system of Fig. 1, 1A, 2, 2A or 3 described above could be implemented in the general purpose processor 501, in which the input signal channels 101, 102 and 103 are indicative data of the central audio (voice) input channels and left and right (without voice) (for example, of a surround sound signal), and in which the output signal channels 118 and 119 are output data indicative of left and right audio output channels emphasized with voice emphasized (for example, of a surround sound signal with enhanced voice). A conventional digital-to-analog converter (DAC) could operate on the output data to generate analog versions of the output audio channel signals for reproduction by physical speakers.
Aspects of the invention are a computer system programmed to perform any realization of the method of the invention, and a computer-readable medium that stores computer-readable code to implement any embodiment of the method of the invention.
Although specific embodiments of the present invention and applications of the invention have been described herein, it will be apparent to persons skilled in the art that many variations in the embodiments and applications described herein are possible without departing from the scope. of the described invention. and revindicated herein.
21 members in 9 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 311437P | United States of America | – | |
| 31143710 | United States of America | P | |
| 31143710 | United States of America | P | |
| 2011026505 | United States of America | W | |
| 2011026505 | United States of America | W | |
| 311437P | – | – | – |
| PCTUS2011026505 | – | – | – |
| US20100311437P | – | – | – |
| WO2011US26505 | – | – | – |
Members21
| Document | Office | Kind | |
|---|---|---|---|
| WO2011112382A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201215177A | Taiwan Province of China | A | |
| CN102792374A | China | A | |
| US2013006619A1 | United States of America | A1 | |
| EP2545552A1 | European Patent Office (EPO) | A1 | |
| JP2013521541A | Japan | A | |
| RU2012141463A | Russian Federation | A | |
| RU2520420C2 | Russian Federation | C2 | |
| TWI459828B | Taiwan Province of China | B | |
| JP5674827B2 | Japan | B2 | |
| CN102792374B | China | B | |
| CN104811891A | China | A | |
| US9219973B2 | United States of America | B2 | |
| US2016071527A1 | United States of America | A1 | |
| BR112012022571A2 | Brazil | A2 | |
| CN104811891B | China | B | |
| US9881635B2 | United States of America | B2 | |
| EP2545552B1 | European Patent Office (EPO) | B1 | |
| ES2709523T3This record | Spain | T3 | |
| BR122019024041B1 | Brazil | B1 | |
| BR112012022571B1 | Brazil | B1 |
Numbers
- Publication
- 2709523
- Publication, DOCDB
- 2709523
- Publication, EPODOC
- ES2709523T
- Application
- 11707537
- Application, DOCDB
- 11707537
- Application, EPODOC
- ES20110707537T
Titles2
- Spanish
- Procedimiento y sistema de escalado de atenuación de canales relevantes de voz en audio multicanal
- English
- Procedure and system for scaling attenuation of relevant voice channels in multichannel audio
Classification
- CPC, 8
- G10L21/0208
- G10L21/0364
- G10L21/0232
- H04S3/008
- H04S2400/09
- H04S2400/13
- H04S7/30
- G10L21/034
- IPC, 6
- H04S7 00
- G10L21 0208
- G10L21 0232
- G10L21 034
- G10L21 0364
- H04S3 00