Generating binaural audio in response to multi-channel audio using at least one feedback delay network.
Abstract
In some embodiments, the virtualization methods for generating a binaural signal in response to multichannel audio signal channels, which apply a binaural room response (BRIR) response to each channel that it includes when using at least one delayed-response network ( FDN) to apply a common late reverb to an audio mix of the channels. In some embodiments, the channels of the input signal are processed in a first processing path to apply to each channel a direct unreflective response portion of a single-channel BRIR for the channel, and the audio mix of the channels it is processed on a second processing path that includes at least one NDF that applies the common late reverb. Usually, The late common reverberation emulates macro collective attributes of late reverb portions of at least some of the single channel BRIRs. Other aspects are headset virtualizers configured to carry out any modality of the method.

Term
8.2 yearsleft in the term
Expires 18 December 2034.
- Priority
- Filed
- Granted
- Today
- Expires
44 claims: 23 independent, 21 dependent
- 1REIVINDICACIONES 1. Un método para generar una señal binaural en respuesta a un conjunto de canales de una señal de entrada de audio multicanal, gue incluye los pasos de:(a) aplicar una respuesta impulsiva binaural de sala, BRIR, a cada canal del conjunto, generando de esta manera señales filtradas, incluyendo mediante el uso de al menos una red de retardo realimentada (203, 204, 205, 220) para aplicar una reverberación tardía común a una mezcla de audio de los canales del conjunto;y (b) combinar las señales filtradas para generar la señal binaural, en donde en el paso (a) , la porción de reverberación tardía común emula macro atributos colectivos de porciones de reverberación tardía de BRIRs de un solo canal compartidas a través de al menos algunos canales del conjunto, el método también incluye un paso de afirmar valores de control a la red de retardo realimentada (203, 204, 205) para establecer al menos una ganancia de entrada, ganancias de tangue de reverberación, retrasos de tanque de reverberación, o parámetros de matriz de salida para dicha red de retardo realimentada (203, 204, 205), en donde los valores de control se afirman en tal manera que la porción de reverberación tardía común emula los macro atributos colectivos de las porciones de reverberación tardía de dichas BRIRs de un solo 106 IMPI instituto mexicano DE LA PROPIEDAD __... . 4 , , . , , , industrial _ canal compartidas a través de dichos al menos algunos cana del conjunto.
- 2El método de acuerdo con la reivindicación 1, caracterizado porque el paso (a) incluye un paso de aplicar a cada canal del conjunto una porción de respuesta directa y reflexión temprana de dicha BRIR de un solo canal para el canal.
- 3El método de acuerdo con la reivindicación 1 ó 2, caracterizado porque el paso (a) incluye un paso de utilizar un banco de redes de retardo realimentadas (203, 204, 205) para aplicar la reverberación tardía común a la mezcla de audio, con cada red de retardo realimentada (203, 204, 205) del banco aplicando reverberación tardía a una banda de frecuencia diferente de la mezcla de audio.
- 4El método de acuerdo con la reivindicación 3, caracterizado porque cada una de las redes de retardo realimentadas (203, 204, 205) se implementa en el dominio del filtro espejo en cuadratura complejo.
- 5El método de acuerdo con cualquiera delas reivindicaciones 1 a 4, caracterizado porque la mezclade audio de los canales del conjunto es una mezcla de audio monofónica de dichos canales del conjunto.
- 6El método de acuerdo con cualquiera delas reivindicaciones 1 a 5, caracterizado porque el paso(a) incluye un paso de generar la mezcla de audio en una manera 107 IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL que depende de una distancia de la 111Ίη hp los canales que se mezclan para generar dicha mezcla de audio, y del manejo de una porción de respuesta directa de la BRIR para cada uno de dichos canales los cuales se mezclan para generar dicha mezcla de audio, con el fin de mantener la relación de nivel y temporización apropiada entre la porción de respuesta directa de dicha BRIR y la reverberación tardía común.
- 7El método de acuerdo con cualquiera de las reivindicaciones 1 a 6, caracterizado porque el paso (a incluye un paso de utilizar una sola red de retardo realimentada (220) para aplicar la reverberación tardía común a la mezcla de audio de los canales del conjunto, en donde la red de retardo realimentada (220) se implementa en el dominio del tiempo.
- 8Un método para generar una señal binaural en respuesta a una señal de entrada de audio multicanal que tiene canales, al aplicar una respuesta impulsiva binaural de sala, BRIR, a cada canal de un conjunto de canales, que incluye al:(a) en una primera ruta de procesamiento, aplicar a cada canal del conjunto al menos una porción de respuesta directa de una respuesta impulsiva binaural de sala de un solo canal para el canal;y (b) en una segunda ruta de procesamiento en paralelo ιοθ IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL con la primera ruta de procesamiento, aplicar reverberación tardía común a una mezcla de audio de los canales del conjunto, donde la reverberación tardía común emula macro atributos colectivos de porciones de reverberación tardía de al menos algunas de las BRIRs de un solo canal compartidas a través de al menos algunos canales del conjunto, en donde la segunda ruta de procesamiento incluye al menos una red de retardo realimentada (203, 204, 205, 220), y el paso (b) incluye un paso de procesar la mezcla de audio en la red de retardo realimentada (203, 204, 205, 220), el método también incluye un paso de afirmar valores de control a la red de retardo realimentada (203, 204, 205, 220) para establecer al menos una ganancia de entrada, ganancias de tanque de reverberación, retrasos de tanque de reverberación, o parámetros de matriz de salida para dicha red de retardo realimentada (203, 204, 205, 220), en donde los valores de control se afirman en tal manera que la porción de reverberación tardía común emula los macro atributos colectivos de las porciones de reverberación tardía de dichas al menos algunas de las BRIRs de un solo canal compartidas a través de dichos al menos algunos canales del conjunto.
- 9El método de acuerdo con la reivindicación 8, caracterizado porque la segunda ruta de procesamiento incluye un banco de redes de retardo realimentadas (203, 204, 205) y 109 IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL el paso (b) incluye un paso de procesar la mezcla Ha Audio en el banco de redes de retardo realimentadas (203, 204, 205) de tal forma que cada red de retardo realimentada (203, 204, 205) del banco aplique reverberación tardía a una banda de frecuencia diferente de la mezcla de audio.
- 10El método de acuerdo con la reivindicación 9, caracterizado porque cada una de las redes de retardo realimentadas (203, 204, 205) se implementa en el dominio del filtro espejo en cuadratura complejo.
- 11El método de acuerdo con cualquiera de las reivindicaciones 8 a 10, caracterizado porque el paso (a) incluye un paso de aplicar la porción de respuesta directa y reflexión temprana de una BRIR de un solo canal diferente a cada canal diferente del conjunto.
- 12El método de acuerdo con cualquiera de las reivindicaciones 8 a 10, caracterizado porque la mezcla de audio de los canales del conjunto es una mezcla de audio monofónica de dichos canales del conjunto.
- 13El método de acuerdo con cualquiera de las reivindicaciones 8 a 10, caracterizado porque el paso (b) incluye un paso de generar la mezcla de audio en una manera que depende de una distancia de la fuente para cada uno de los canales que se mezclan para generar dicha mezcla de audio, y del manejo de una porción de respuesta directa de la BRIR para cada uno de dichos canales los cuales se mezclan 110 IMPI^^ INSTITUTO MEXICANO M LA MONEDAD para generar dicha mezcla de audio, con el fin cíe mantener la relación de nivel y temporización apropiada 1 U'flTFé 1Λ porción de respuesta directa de dicha BRIR y la reverberación tardía común.
- 14El método de acuerdo con cualquiera de las reivindicaciones 8 a 10, caracterizado porque la segunda ruta de procesamiento incluye una red de retardo realimentada (220), la red de retardo realimentada (220) se implementa en el dominio del tiempo, y el paso (b) incluye un paso de procesar la mezcla de audio en la red de retardo realimentada (220) .
- 15Un sistema configurado para generar una señal binaural en respuesta a una señal de entrada de audio multicanal que tiene canales, al aplicar una respuesta impulsiva binaural de sala a cada canal de un conjunto de los canales, dicho sistema incluye:una primera ruta de procesamiento acoplada y configurada para aplicar a cada canal del conjunto, al menos una porción de respuesta directa de una respuesta impulsiva binaural de sala, BRIR, de un solo canal, para el canal;y una segunda ruta de procesamiento, acoplada en paralelo con la primera ruta de procesamiento, y configurada para aplicar una reverberación tardía común a una mezcla de audio de los canales del conjunto, donde la reverberación tardía común emula macro atributos colectivos de porciones de 111 IMPI INSTITUTO MEXICANO DE LA PROPIEDAD industriai reverberación tardía de al menos algunas de las BRIRs de un solo canal compartidas a través de al menos algunos canales del conjunto, en donde la segunda ruta de procesamiento incluye al menos una red de retardo realimentada (203, 204, 205, 220), y la segunda ruta de procesamiento está configurada para procesar la mezcla de audio en dicha al menos una red de retardo realimentada (203, 204, 205, 220) para aplicar la reverberación tardía común a la mezcla de audio, el sistema también incluye: un subsistema de control (209) acoplado y configurado para afirmar valores de control a la red de retardo realimentada (203, 204, 205, 220) para establecer al menos una ganancia de entrada, ganancias de tanque de reverberación, retrasos de tanque de reverberación, o reverberación tardía común emula los macro atributos colectivos de las porciones de reverberación tardía de dichas al menos algunas de las BRIRs de un solo canal compartidas a través de dichos al menos algunos canales del conjunto.
- 16El sistema de acuerdo con la reivindicación 15, caracterizado porque la segunda ruta de procesamiento incluye un banco de redes de retardo realimentadas (203, 204, 205), y 112 IMPI INfmUTQMWtK^NÚ la segunda ruta de procesamiento está procesar la mezcla de audio en dicho banco de redAs cía . retardo realimentadas (203, 204, 205) de tal forma que cada red de retardo realimentada (203, 204, 205) del banco aplique reverberación tardía a una banda de frecuencia diferente de la mezcla de audio.
- 17El sistema de acuerdo con la reivindicación 16, caracterizado porque cada una de las redes de retardo realimentadas (203, 204, 205) se implementa en el dominio del filtro espejo en cuadratura complejo.
- 18El sistema de acuerdo con cualquiera de las reivindicaciones 15 a 17, caracterizado porque la primera ruta de procesamiento está configurada para generar señales filtradas en respuesta a cada uno de dichos canales del conjunto, la segunda ruta de procesamiento está configurada para generar señales filtradas adicionales en respuesta a la mezcla de audio, y en donde dicho sistema también incluye:un subsistema de combinación de señal (210), acoplado a la primera ruta de procesamiento y la segunda ruta de procesamiento, y configurado para generar la señal binaural al combinar las señales filtradas y las señales filtradas adicionales.
- 19El sistema de acuerdo con cualquiera de las reivindicaciones 15 a 18, caracterizado porque dicho sistema es un virtualizador de auriculares. 113 IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL
- 20El sistema de acuerdo con cualquiera de las reivindicaciones 15 a 18, caracterizado porque dicho sistema es un decodificador que incluye un subsistema virtualizador, y el subsistema virtualizador implementa la primera ruta de procesamiento y la segunda ruta de procesamiento.
- 21El sistema de acuerdo con cualquiera de las reivindicaciones 15 a 20, caracterizado porque la mezcla de audio de los canales del conjunto es una mezcla de audio monofónica de dichos canales del conjunto.
- 22El sistema de acuerdo con la reivindicación 15, caracterizado porque la segunda ruta de procesamiento incluye una red de retardo realimentada (220) , la red de retardo realimentada (220) se implementa en el dominio del tiempo, y la segunda ruta de procesamiento está configurada para procesar la mezcla de audio en el dominio del tiempo en dicha red de retardo realimentada (220) para aplicar la reverberación tardía común a dicha mezcla de audio.
- 23El sistema de acuerdo con la reivindicación 22, caracterizado porque la red de retardo realimentada (220) incluye:un filtro de entrada (400) que tiene una entrada acoplada para recibir la mezcla de audio, en donde el filtro de entrada (400) está configurado para generar una primera mezcla de audio filtrada en respuesta a la mezcla de audio;un filtro pasa todo (401), acoplado y configurado para IMPI» INSTITUTO MEXICANO DE LA PROPIEDAD generar una segunda mezcla de audio filtrada en ,N, ¥^^ues^S-? la primera mezcla de audio filtrada;। ..............— un subsistema de aplicación de reverberación que tiene una primera salida y una segunda salida, en donde el subsistema de aplicación de reverberación comprende un conjunto de tanques de reverberación, cada uno de los tanques de reverberación tiene un retraso diferente, y en donde el subsistema de aplicación de reverberación esta acoplado y configurado para generar un primer canal binaural no mezclado y un segundo canal binaural no mezclado en respuesta a la segunda mezcla de audio filtrada, para afirmar el primer canal binaural no mezclado en la primera salida, y para afirmar el segundo canal binaural no mezclado en la segunda salida;y una etapa de filtración de coeficiente de correlación cruzada interaural, IACC, y mezcla (424) acoplada al subsistema de aplicación de reverberación y configurada para generar un primer canal binaural mezclado y un segundo canal binaural mezclado en respuesta al primer canal binaural no mezclado y un segundo canal binaural no mezclado.
- 24El sistema de acuerdo con la reivindicación 23, caracterizado porque el filtro de entrada (400) se implementa como una cascada de dos filtros configurados para generar la primera mezcla de audio filtrada de tal manera que cada una de dichas BRIRs tenga una relación directa-a-tardía, DLR, que se adecúe, al menos sustancialmente, a una DLR objetivo.
- 25El sistema de acuerdo con la reivindicación 24/ caracterizado porque cada uno de los tanques de reverberación está configurado para generar una señal retrasada, e incluye un filtro de reverberación (406, 406A, 407, 407A, 408, 408A, 409, 409A) acoplado y configurado para aplicar una ganancia a una señal que se propaga en cada uno de dichos tanques de reverberación, para provocar que la señal retrasada tenga una ganancia gue se adecúe, al menos sustancialmente, a una ganancia decaída objetivo para dicha señal retrasada, con el fin de lograr un tiempo de decaimiento de reverberación objetivo característico de cada una de dichas BRIRs.
- 26El sistema de acuerdo con la reivindicación 25, caracterizado porque cada uno de dichos filtros de reverberación (406, 406A, 407, 407A, 408, 408A, 409, 409A) es un filtro limitador o una cascada de filtros limitadores.
- 27El sistema de acuerdo con cualquiera de las reivindicaciones 23 a 26, caracterizado porque el primer canal binaural no mezclado dirige el segundo canal binaural no mezclado, los tanques de reverberación incluyen un primer tanque de reverberación configurado para generar una primera señal retrasada que tenga un retraso más corto y un segundo tangue de reverberación configurado para generar una segunda señal retrasada que tenga un segundo retraso más corto, en donde el primer tanque de reverberación está configurado para 116 IMPI instituto mexicano DE LA PROPIEDAD INDUSTRIAL aplicar una primera ganancia a la primera señal retrasada, el segundo tanque de reverberación está configurado para aplicar una segunda ganancia a la segunda señal retrasada, la segunda ganancia es diferente que la primera ganancia, y la aplicación de la primera ganancia y la segunda ganancia resulta en atenuación del primer canal binaural no mezclado con relación al segundo canal binaural no mezclado.
- 28El sistema de acuerdo con cualquiera de las reivindicaciones 23 a 27, caracterizado porque el primer canal binaural mezclado y el segundo canal binaural mezclado son indicativos de una imagen estéreo re-centrada.
- 29El sistema de acuerdo con cualquiera de las reivindicaciones 23 a 28, caracterizado porque la etapa de filtración de IACC y mezcla (424) está configurada para generar el primer canal binaural mezclado y el segundo canal binaural mezclado de tal manera que dicho primer canal binaural mezclado y dicho segundo canal binaural mezclado tengan un IACC característico que se adecúe al menos sustancialmente a un IACC característico objetivo.
- 30Un sistema configurado para generar una señal binaural en respuesta a un conjunto de canales de una señal de entrada de audio multicanal, dicho sistema incluye:un subsistema de filtración acoplado y configurado para aplicar una respuesta impulsiva binaural de sala, BRIR, a cada canal del conjunto, generando de esta manera señales 17 IMPI INSTITUTO MEXICANO DE LA PROPIEDAD V INDUSTRIAL ** filtradas, incluyendo al generar una mezcla de audio de los canales del conjunto y procesar dicha mezcla de audio en al menos una red de retardo realimentada (203, 204, 205, 220) para aplicar una reverberación tardía común a dicha mezcla de audio;y un subsistema de combinación de señal (210), acoplado al subsistema de filtración, y configurado para generar la señal binaural al combinar las señales filtradas, en donde la reverberación tardía común emula macro atributos colectivos de porciones de reverberación tardía de BRIRs de un solo canal compartidas a través de al menos algunos canales del conjunto, el sistema también incluye un subsistema de control (209) acoplado al subsistema de filtración y configurado para afirmar valores de control a la red de retardo realimentada (203, 204, 205) para establecer al menos una ganancia de entrada, ganancias de tanque de reverberación, retrasos de tanque de reverberación, o parámetros de matriz de salida para dicha red de retardo realimentada (203, 204, 205), en donde los valores de control se afirman en tal manera que la porción de reverberación tardía común emula los macro atributos colectivos de las porciones de reverberación tardía de dichas al menos algunas de las BRIRs de un solo canal compartidas a través de dichos al menos algunos canales del conj unto. 118 IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL
- 31El sistema de acuerdo con la reivindicación 30, caracterizado porque el subsistema de filtración está configurado para aplicar a cada canal del conjunto una porción de respuesta directa y reflexión temprana de la BRIR de un solo canal para el canal.
- 32El sistema de acuerdo con la reivindicación 30 ó 31, caracterizado porque el subsistema de filtración incluye un banco de redes de retardo realimentadas (203, 204, 205) configuradas para aplicar la reverberación tardía común a la mezcla de audio, con cada red de retardo realimentada (203, 204, 205) del banco aplicando reverberación tardía a una banda de frecuencia diferente de la mezcla de audio.
- 33El sistema de acuerdo con la reivindicación 32, caracterizado porque cada una de las redes de retardo realimentadas (203, 204, 205) se implementa en el dominio del filtro espejo en cuadratura complejo.
- 34El sistema de acuerdo con cualquiera de las reivindicaciones 30 a 33, caracterizado porque dicho sistema es un virtualizador de auriculares.
- 35El sistema de acuerdo con cualquiera de las reivindicaciones 30 a 33, caracterizado porque dicho sistema es un decodificador que incluye un subsistema virtualizador, y el subsistema virtualizador implementa el subsistema de filtración y el subsistema de combinación de señal (210).
- 36El sistema de acuerdo con cualquiera de las 119 IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL reivindicaciones 30 a 35, caracterizado porque la mezcla de audio de los canales del conjunto es una mezcla de audio monofónica de dichos canales del conjunto.
- 37El sistema de acuerdo con la reivindicación 30 ó 31, caracterizado porque el subsistema de filtración incluye una red de retardo realimentada (220) implementada en el dominio del tiempo, y el subsistema de filtración está configurado para procesar la mezcla de audio en el dominio del tiempo en dicha red de retardo realimentada (220) para aplicar la reverberación tardía común a dicha mezcla de audio.
- 38El sistema de acuerdo con la reivindicación 37, caracterizado porque la red de retardo realimentada (220) incluye:un filtro de entrada (400) que tiene una entrada acoplada para recibir la mezcla de audio, en donde el filtro de entrada (400) está configurado para generar una primera mezcla de audio filtrada en respuesta a la mezcla de audio;un filtro pasa todo (401), acoplado y configurado para generar una segunda mezcla de audio filtrada en respuesta a la primera mezcla de audio filtrada;un subsistema de aplicación de reverberación que tiene una primera salida y una segunda salida, en donde el subsistema de aplicación de reverberación comprende un conjunto de tanques de reverberación, cada uno de los tanques «o IMPIOS instituto mexicano DE LA PROPIEDAD CWShJ INDUSTRIAL de reverberación tiene un retraso diferente, y en donde el subsistema de aplicación de reverberación esta acoplado y configurado para generar un primer canal binaural no mezclado y un segundo canal binaural no mezclado en respuesta a la segunda mezcla de audio filtrada, para afirmar el primer canal binaural no mezclado en la primera salida, y para afirmar el segundo canal binaural no mezclado en la segunda salida;y una etapa de filtración de coeficiente de correlación cruzada interaural, IACC, y mezcla (424) acoplada al subsistema de aplicación de reverberación y configurada para generar un primer canal binaural mezclado y un segundo canal binaural mezclado en respuesta al primer canal binaural no mezclado y un segundo canal binaural no mezclado.
- 39El sistema de acuerdo con la reivindicación 38, caracterizado porque el filtro de entrada (400) se implementa como una cascada de dos filtros configurados para generar la primera mezcla de audio filtrada de tal manera que cada una de dichas BRIRs tenga una relación directa-a-tardia, DLR, que se adecúe, al menos sustancialmente, a una DLR objetivo.
- 40El sistema de acuerdo con la reivindicación 38 ó 39, caracterizado porque cada uno de los tanques ' de reverberación está configurado para generar una señal retrasada, e incluye un filtro de reverberación (406, 406A, 407, 407A, 408, 408A, 409, 409A) acoplado y configurado para 121 IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL aplicar una ganancia a una señal que se propaga en cada uno de dichos tanques de reverberación, para provocar que la señal retrasada tenga una ganancia que se adecúe, al menos sustancialmente, a una ganancia decaída objetivo para dicha señal retrasada, con el fin de lograr un tiempo de decaimiento de reverberación objetivo característico de cada una de dichas BRIRs.
- 41El sistema de acuerdo con la reivindicación 40, caracterizado porque cada uno de dichos filtros de reverberación (406, 406A, 407, 407A, 408, 408A, 409, 409A) es un filtro limitador o una cascada de filtros limitadores.
- 42El sistema de acuerdo con cualquiera de las reivindicaciones 38 a 41, caracterizado porque el primer canal binaural no mezclado dirige el segundo canal binaural no mezclado, los tanques de reverberación incluyen un primer tanque de reverberación configurado para generar una primera señal retrasada que tenga un retraso más corto y un segundo tanque de reverberación configurado para generar una segunda señal retrasada que tenga un segundo retraso más corto, en donde el primer tanque de reverberación está configurado para aplicar una primera ganancia a la primera señal retrasada, el segundo tanque de reverberación está configurado para aplicar una segunda ganancia a la segunda señal retrasada, la segunda ganancia es diferente que la primera ganancia, la segunda ganancia es diferente que la primera ganancia, y la 122 IMPI INSTITUTO MEXICANO ^«*6 DE LA PROPIEDAD INDUSTRIAL aplicación de la primera ganancia y la segunda ganancia resulta en atenuación del primer canal binaural no mezclado con relación al segundo canal binaural no mezclado.
- 43El sistema de acuerdo con cualguiera de las reivindicaciones 38 a 42, caracterizado porque el primer canal binaural mezclado y el segundo canal binaural mezclado son indicativos de una imagen estéreo re-centrada.
- 44El sistema de acuerdo con cualguiera de las reivindicaciones 38 a 43, caracterizado porque la etapa de filtración de IACC y mezcla (424) está configurada para generar el primer canal binaural mezclado y el segundo canal binaural mezclado de tal manera que dicho primer canal binaural mezclado y dicho segundo canal binaural mezclado tengan un IACC característico que se adecúe al menos sustancialmente a un IACC característico objetivo. 123 IMPI 'NSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL
Independent claims44
463 paragraphs in 120 sections, as filed
(54) Title: BINAURAL AUDIO GENERATION IN RESPONSE TO MULTICHANNEL AUDIO USING AT LEAST ONE FEEDBACK DELAY NETWORK.
(54) Title: GENERATING BINAURAL AUDIO IN RESPONSE TO MULTI-CHANNEL AUDIO USING AT LEAST ONE FEEDBACK DELAY NETWORK.
(57) Summary
In some embodiments, virtualization methods for generating a binaural signal in response to channels of a multichannel audio signal, applying a Binaural Room Impulse Response (BRIR) to each channel including by using at least one feedback delay network ( FDN) to apply a common late reverb to an audio mix of channels. In some embodiments, the channels of the input signal are processed in a first processing path to apply to each channel an early thoughtless direct response portion of a single channel BRIR for the channel, and the audio mix of the channels. it is processed in a second processing path that includes at least one FDN that applies the common late reverb. Generally, common late reverb emulates collective macro attributes of late reverb portions of at least some of the single channel BRIRs. Other aspects are headphone virtualizers configured to carry out any modality of the method.
(57) Abstract ln some embodiments, virtualization methods for generating a binaural signal in response to channels of a multichannel audio signal, which apply a binaural room impulse response (BRIR) to each channel including by using at least one feedback delay network (FDN) to apply a common late reverberation to a downmix of the channels. In some embodiments, input signal channels are processed in a first Processing path to apply to each channel a direct response and early reflection portion of a single-channel BRIR for the channel, and the downmix of the channels is processed in a second Processing path including at least one FDN which applies the common late reverberation. Typically, the common late reverberation emulates collective macro attributes of late reverberation portions of at least some of the single-channel BRIRs. Other aspects are headphone virtualizers configured to perform any embodiment of the method.
<img file="MX352134B_D0001.tif" />
<img file="MX352134B_D0002.tif" />
PATENT TITLE No. 352134
Owner (s): DOLBY LABORATORIES LICENSING CORPORATION
D miciiio: 1275 Market Street, San Francisco, California, 94103, USA
D nomination: BINAURAL AUDIO GENERATION IN RESPONSE TO MULTICHANNEL AUDIO
USING AT LEAST ONE FEEDBACK DELAY NETWORK.
Classification!
CIP:
CPC:
H04S3 / 0O,
H04Sj / 60¿ '51 «5712:
Inventor (s)
KUAN-CMÉH VÉN; DIRK J<sub>s </sub>WlLSONy DAVID M. COOP ^ I, RT; 'JANG
DAVIDSON; RHONDA
Num,
MX / a / 2016/008696
Intel je 20 icional us
Number:
31/9^3,579
1410178258.0
Vieenciat ^ ugly '' l
Date of V ^ p <Aihiente | 18 d * cycigpOpening 2034
Exi date with
In accordance with the art from the date of presentation $ ¿Λ
Sejafle ^ le la? Ro | íef ^ lndustrial · íJZ _ ene »de ven * afifíjj ^ rorgables, counted to Sntener in force to the rights.
Who signs this title lo.A ”(Official Gazette of the Federation (trc 01/25/2006, 05/06/2009, 06/01/2010, ^ Regulation of the Mexican Institute of | fun <
the articles 1<sup>or</sup>, 3’, 4<sup>or</sup>, 5<sup>or</sup> fraction V subsection a), Wf 12/27/1999, amended 10/10/2002, 07/29/200 General Deputy Generals, Coordinator, Directors j / 1991 β / 06 ^ Jad la
Dr
Jo dieuestoec ermala on K 1. ^ 0872012 / articulate
I994, IpwW
1 (2012101115¾ reformed
Orgaatfñj lis lll 12/26/1
Departmental and other subordinates of the Mexican Institute of 08/04/2004 and 09/13/2007).
is ^ pcirW W ^ bstrial.
Λήβ'ΖΧβ the Industrial Property Law to <L ^ W1999, 01/26/2004, 06/16/2005, ώτί ^ Λώίο a), 4th and 12th sections I and III of & M8TO7 / 20O4, 07/28 / 2004 and 7/09/2007); Mexican Department of Industrial Property (DOF that delegates powers to the Directors n ^ ÍÍMs. Subdivisional Directors, Coordinators Ij. '12/5/1999. amended on 02/04/2000, 07/29/2004,
This document is signed with an advanced electronic signature (FIEL), based on articles 7 BIS 2 of the Industrial Property Law; 3 of its Regulations, and 1 section III, 2 section V, 26 BIS and 26 TER of the Agreement establishing the guidelines for the use of the Payment and Electronic Services Portal (PASE) of the Mexican Institute of Industrial Property, in the procedures indicated.
THE DIVISIONAL DIRECTOR OF PATENTS
NAHANNY CANAL REYES
<img file="MX352134B_D0003.tif" />
Original string.
NAHANNY MARISOL CANAL REYES | 00001000000403252793 | Administration Service
Tax | 1695 || MX / 2017/91961 | MX / a / 2016/008696 | PCT patent title | 1223 | GAGV | Page (s) | xZnS8IOBGpb6xAfaNYstGfl YXZc =
Digital stamp:
RV4SCIPM7ZOJDPjO + WQENMxKZy2nu9kKJBequx7ZxUoh95jnlls / 9MmrOBhy6 + Jw25y8h25ud + 8r3BxdkjSyAcCAaM + cfg8k6TGV1UdCAUpmQil8dDC5AO4D6YJd0XkNICq / kvnTe89aG2hrNavW0K4nFE3wu0OSk30Ch8fyiTVWreTbSGtl 8xDeUn / RgXzkjwmp1rKW L2VCulntfl44Ka9xcotcitAuqVZVXeNTKg7wkL5xwK49jYU05htq + + + N9Jf6kJoaoYF Afy
Q4JAOSWQp4z8C6YpFVUe227n / wZUWmBnnU2uqdDzG / HdUOEYV1uyEAslX98Pi312 / cwNZOXw ==
Arenal No, 550 Floor 1, Pueblo Santa María Tepepan. Xochimilco, 16020, Mexico City.
(55) 53340700 www.gob.mx/inipi
<img file="MX352134B_D0004.tif" />
55^37
IMPI
MEXICAN INSTITUTE
BINAURAL AUDIO GENERATION IN RESPONSE TO USING AT LEAST ONE DELAY NETWORK ΒΕΑ ^ ΤΜΕΝΤΆΠΑ .........
FIELD OF THE INVENTION
The invention relates to methods (sometimes referred to as headphone virtualization methods) and systems for generating a binaural signal in response to a multichannel audio input signal, by applying a Binaural Room Impulse Response (BRIR). ) to each channel of a channel set (e.g., all channels) of the input signal. In some embodiments, at least one Feedback Delay Network (FDN) applies a late reverb portion of an audio downmix BRIR to an audio mix of channels.
BACKGROUND OF THE INVENTION
Headphone virtualization (or binaural rendering or rendering) is a technology that aims to deliver an surround sound or invasive sound field experience using standard stereo headphones.
Early headphone virtualizers applied a Head-Related Transfer Function (HRTF) to transmit spatial information in binaural representation. An HRTF is a set<sup>2</sup> IMPI »
MUICANO INSTITUTE
OF THE PROPERTY
INDUSTRIAL of direction and distance dependent filter pairs that characterize how sound is transmitted from a specific point in space (sound source location) to both ears of a listener in an anechoic environment. Essential spatial signals such as Interaural Time Difference (ITD), Interaural Level Difference (ILD), head shadowing, spectral spikes and suppressions due to shoulder and pinna reflections, can be perceived in the binaural content rendering filtered by HRTF. Due to the restriction of the size of the human head, HRTFs do not provide sufficient or robust signals with respect to the distance from the source beyond approximately one meter. As a result, virtualizers based solely on HRTF generally do not achieve good outsourcing or perceived distance.
Most acoustic events in our daily lives occur in reverberant environments where, in addition to the direct path (from source to ear) modeled by HRTF, audio signals also reach the listener's ears through different reflection paths. Reflections introduce profound impact on auditory perception, such as distance, room size, and / or other attributes of the space. To transmit this information in binaural representation, a virtualizer needs to apply reverb
<img file="MX352134B_D0005.tif" />
IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL of the room in addition to the signs on the direct route HRTF.
A binaural room impulsive response (Bx ± k) characterizes the transformation of audio signals from a specific point in space to the listener's ears in a specific acoustic environment. In theory, BRIRs include all acoustic signals with respect to spatial perception.
Figure 1 is a block diagram of a type of conventional headphone virtualizer that is configured to apply a binaural room impulsive response (BRIR) to each full-range frequency channel (Xi, X<sub>N</sub>) of a multichannel audio input signal. Each of the X channels<sub>x</sub>, ..., X<sub>N</sub>, is a loudspeaker channel that corresponds to a different source direction relative to an assumed listener (that is, the direction of a direct path from an assumed position of a speaker corresponding to the assumed position of the listener), and each channel such it is convolved by the BRIR for the corresponding source address. The acoustic path from each channel needs to be simulated for each ear. Therefore, in the remainder of this document, the term BRIR refers to either an impulsive response, or a pair of impulsive responses associated with the left and right ears. Therefore, subsystem 2 is configured to convolve channel Xi with BRIRi (the BRIR for the corresponding source address), subsystem 4 is configured to
<img file="MX352134B_D0006.tif" />
convolve channel X<sub>N</sub> with BRIR<sub>N</sub> (the BRIR for the corresponding source address), and so on. The output of each BRIR subsystem (each of subsystems 2, ..., 4) is a time domain signal that includes a left channel and a right channel. The left channel outputs of the BRIR subsystems are mixed in edit item 6, and the right channel outputs of the BRIR subsystems are mixed in the add item 8. The output of item 6 is the left channel, L, of the binaural audio signal output from the virtualizer, and the output of item 8 is the right channel, R, of the binaural audio signal output from the virtualizer.
The multichannel audio input signal may also include a Low Frequency Effects (LFE) or subwoofer channel, identified in Figure 1 as the LFE channel. In a conventional way, the LFE channel does not convolve with a BRIR, but rather attenuates at gain stage 5 of Figure 1 (e.g., by -3dB or more) and the output of the Gain 5 is mixed equally (by items 6 and 8) on each of the channels of the virtualizer binaural output signal. An additional delay stage may be required in the LFE path in order to time align the output of stage 5 with the outputs of the BRIR subsystems (2, ..., 4). Alternatively, the LFE channel can simply be ignored s
MEXICAN USTtTUTO
OF THE PROPERTY
INDUSTRIAL (that is, not asserted or processed by the virtualizer). .
For example, the Figure 2 embodiment of the invention (to be described later) simply ignores any LFE channel of the multichannel audio input signal processed in this way. Many consumer headphones are not capable of accurately reproducing an LFE channel.
In some conventional virtualizers, the input signal undergoes transformation from time domain to frequency domain to Quadrature Mirror Filter (QMF) domain, to generate channels of QMF domain frequency components. These frequency components undergo leakage (e.g., in QMF domain implementations of subsystems 2, 4 of Figure 1) into the QMF domain and the resulting frequency components are generally then transformed back to the time domain ( e.g. in a final stage of each of the subsystems 2, ..., 4 of Figure 1) so that the audio output of the virtualizer is a time-domain signal (e.g., signal binaural time domain).
In general, each full-range frequency channel from a multichannel audio signal input to a headphone virtualizer is assumed to be indicative of the audio content output from a sound source at a known location relative to the listener's ears. The
ΙΜΡΙ «£ Mexican Institute
K THE PROPERTY Chcia
INDUSTRIAL ° headset virtualizer is configured to · erpilcal · · 'urn room binaural impulsive response (BRIR) to each such channel of the input signal. Each BRIR can be broken down into two parts: direct response and reflections. The direct response is the HRTF that corresponds to the Direction Of Arrival (DOA) of the sound source, adjusted with gain and appropriate due to the distance (between the sound source and the listener), and optionally increased with parallax effects for small distances.
The remaining portion of the BRIR models the reflections. Early reflections are generally primary or secondary reflections and have relatively little temporal distribution. The microstructure (eg, ITD and ILD) of each primary or secondary reflection is important. For late reflections (sound reflected from more than two surfaces before incident on the listener), the echo density increases with increasing number of reflections, and the micro attributes of individual reflections become difficult to observe. For later and later reflections, the macrostructure (eg, the decay rate of the reverb, interaural coherence, and spectral distribution of the total reverb) becomes more important. Because of this, reflections can be further segmented into two parts: early reflections and late reverbs.
<sup>7</sup> IMPIOUS"
MEXICAN INSTITUTE
M THE C— ™ PROPERTY
INDUSTRIAL ~ *
The direct response delay is the distance of the source from the listener divided by the speed of sound, and its level is (in the absence of walls or large surfaces near the location of the source) inversely proportional to the distance from the source . On the other hand, the delay and level of late reverbs are generally insensitive to the location of the source. Due to practical considerations, virtualizers may choose to time align direct responses from sources with different distances, and / or compress their dynamic range. However, the temporal and level relationship between direct response, early reflections, and late reverb must be maintained within a BRIR.
The effective length of a typical BRIR extends to hundreds of milliseconds or more in most acoustic environments. The direct application of BRIRs requires convolution with a filter of hundreds of shots, which is computationally expensive. Additionally, without parameterization, it would require a large memory space to store BRIRs from different positions of the source in order to achieve sufficient spatial resolution. Last but not least, sound source locations may change over time, and / or the position and orientation of the listener may vary over time. Precise simulation of such motion requires responses of
MEXICAN INSTITUTE
OF THE PROMITY
INDUSTRIAL impulse of BRIR variables in time. Proper interpolation and application of such time-varying filters can be challenging if the impulse responses of these filters have many taps.
A filter having the well known filter structure as a feedback delay network (FDN) can be used to implement a spatial reverb that is configured to apply simulated reverb to one or more channels of a multichannel audio input signal. The structure of a simple FDN. It comprises several reverb tanks (e.g. the reverb tank comprises the gain element gi and the delay line z '<sup>nl</sup>, in the FDN of Figure 4), each reverb tank has a delay and gain. In a typical FDN implementation, the outputs of all reverb tanks are mixed by means of a unity feedback matrix and the outputs of the matrix are fed back and summed with the inputs of the reverb tanks. Gain adjustments can be made to the reverb tank outputs, and the reverb tank outputs (or gain-adjusted versions of them) can be remixed appropriately for multi-channel or binaural playback. Natural resonant reverb can be generated and applied by an FDN with compact computational and memory sizes. Therefore FDNs have been used in
ΙΜΡΙ @>
Mexican iNjrmrro
OF PROPERTY C * .— CSSlf
INDUSTRIAL virtualizers to complement the direct response produced by the HRTF.
For example, the commercially available Dolby Mobile headphone virtualizer includes a reverb that has an FDN-based structure that is operable to apply reverb to each channel of a five-channel audio signal (which has front left, front right, center, surround channels left, and Right Surround) and to filter each reverb channel using a different filter pair from a set of five Head Relative Transfer Function (HRTF) filter pairs. The Dolby Mobile Headphone Virtualizer is also operable in response to a two-channel audio input signal, to generate a two-channel reverberated binaural audio output (a two-channel virtual surround sound output to which reverb has been applied). ). When the reverberated binaural output is rendered and played through a pair of headphones, it is perceived at the listener's eardrums as sound by HRTF, reverberated from five speakers in the front left, front right, center, rear left (surround), and rear positions. right (surround). The virtualizer separates (upmix) a downmixed two channel audio input (without using any spatial signal parameters received with the audio input) to generate five separate audio channels, apply <sup>10</sup> IMPI ^
MEXICAN INSTITUTE
OF ΙΑ PROPERTY Qua ™
INDUSTRIAL reverb to separate channels, and mixes the five reverb channel signals to generate the virtualizer's two-channel reverb output. The reverb for each separate channel is filtered on a different pair of HRTF filters.
In a virtualizer, an FDN can be configured to achieve a certain reverb decay time and echo density. However, FDN lacks the flexibility to simulate the microstructure of early reflections. Furthermore, in conventional virtualizers the adjustment and configuration of FDN has been mainly heuristic.
Headphone virtualizers that do not simulate all reflection paths (early and late) cannot achieve effective outsourcing. The inventors have recognized that virtualizers employing FDNs that attempt to simulate all reflection paths (early and late) generally have no more than limited success in simulating both early reflections and late reverb and applying both to an audio signal. The inventors have also recognized that virtualizers that employ FDNs but do not have the ability to properly control spatial acoustic attributes such as reverberation decay time, interaural coherence, and direct-to-late relationship, can achieve a degree of
<img file="MX352134B_D0007.tif" />
IMPI
INSTITUTE MEXICANO Dt LA PROPERTY INDUSTRIAL outsourcing but at the price of introducing distortion of timbre and reverberation.
BRIEF DESCRIPTION OF THE INVENTION
In a first class of embodiments, the invention is a method for generating a binaural signal in response to a set of channels (eg, each of the channels, or each of the full-range frequency channels) of a multichannel audio input signal, including the steps of: (a) applying a binaural room impulsive response (BRIR) to each channel of the ensemble (e.g., by convolving each channel of the ensemble with a BRIR corresponding to that channel), thereby generating filtered signals, including by means of the use of at least one feedback delay network (FDN) to apply a common late reverb to an audio mix (downmix ') (eg, a monophonic audio mix) of the channels of the ensemble; and (b) combining the filtered signals to generate the binaural signal. Generally, a bank of FDNs is used to apply the common late reverb to the audio mix (eg, with each FDN applying common late reverb to a different frequency band). Generally, step (a) includes a step of applying to each channel of the ensemble a direct response and early reflection portion of a single channel BRIR for the channel, and the common late reverb has<sup>12</sup> IMPI ^
MEXICAN INSTITUTE
OF THE PROPERTY been generated to emulate collective macro attributes<sup>-</sup>of late reverb portions of at least Sigilfiya (eg, all) single channel BRIRs.
A method of generating a binaural signal in response to a multichannel audio input signal (or in response to a set of channels of such a signal) is sometimes referred to herein as a headphone virtualization method, and a system configured to Carrying out such a method is sometimes referred to in this document as a headset virtualizer (or headset virtualization system or binaural virtualizer).
In typical embodiments in the first class, each of the FDNs is implemented in a filter bank domain (e.g., the Hybrid Complex Quadrature Mirror Filter (HCQMF) domain or the domain of the quadrature mirror filter (QMF), or other transform domain or sub-band that may include decimation), and in some embodiments such, the frequency-dependent spatial acoustic attributes of the binaural signal are controlled by controlling the setting of each FDN used to apply late reverb. Generally, a monophonic audio downmix of the channels is used as the input to the FDNs for efficient binaural representation of the audio content of the multichannel signal. Typical modalities in the first class<sup>13</sup> IMPI®
MEXICAN INSTITUTE
OF LJhINDUSTRIAL PROPERTY ”* include a step of adjusting FDN coefficients corresponding to frequency-dependent attributes (eg, reverberation decay time, interaural coherence, modal density, and direct-to-day relationship), for example, by assert control values to the feedback delay network to set at least one input gain, reverb tank gains, reverb tank delays, or output array parameters for each FDN. This enables better matching of acoustic environments and more natural sound outputs.
In a second class of embodiments, the invention is a method of generating a binaural signal in response to a multi-channel audio input signal having channels, by applying a room binaural impulsive response (BRIR) to each channel of a set of channels. of the input signal (e.g., each of the input signal channels or each full-range frequency channel of the input signal), including: processing each channel of the array into a first processing path configured to model, and applying to each of said channels, a direct response and early reflection portion of a single channel BRIR for the channel; and process an audio mix (downmix) (e.g., a mono channel mix (mono)) of the channels in the ensemble into a configured second render path (in parallel with the first render path)
<img file="MX352134B_D0008.tif" />
IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL to model, and apply a common late reverb to the audio mix. Generally, the common late reverb has been generated to emulate collective macro attributes of late reverb portions of at least some (eg, all) of the single channel BRIRs. Generally, the second processing path y includes at least one FDN (eg, one FDN for each of the multiple frequency bands). Generally, a mono audio mix is used as the input to all reverb tanks of each FDN implemented by the second processing path. Mechanisms for systematic macro-attribute control of each FDN are generally provided to better simulate acoustic environments and produce more natural-sounding binaural virtualization. Since most such macro attributes are frequency dependent, each FDN is generally implemented in the hybrid complex quadrature mirror filter (HCQMF) domain, the frequency domain, domain, or other domain of the filter bank, if used. a different or independent FDN for each frequency band. A primary benefit of implementing FDNs in a filterbank domain is to allow the application of reverb with frequency-dependent reverb properties. In different embodiments, FDNs are implemented in any of a wide variety of filterbank domains, using
<img file="MX352134B_D0009.tif" />
IMPI
MEXICAN INSTITUTE
M tA INDUSTRIAL PROPERTY any of a variety of filter banks, including, but not limited to, real or complex value quadrature mirror (QMF) filters, Finite-Impulse Response (FIR) filters, response filters Infinite-Impulse Response (IIR), Discrete Fourier Transforms (DFTs), cosine or sine (modified) transforms, wavelet transforms, or crossover filters. In a preferred implementation, the filter bank employed or transform includes decimation (eg, a decrease in the sampling rate of the frequency domain signal representation) to reduce the computational complexity of the FDN process.
Some modalities in the first class (and the second class) implement one or more of the following characteristics:
one. An implementation of FDN in the domain of the filter bank (e.g., hybrid complex quadrature mirror filter domain), or implementation of FDN in the domain of the hybrid filter bank and late reverb filter implementation in the domain time, which generally allows independent parameter setting and / or FDN adjustments for each frequency band (enabling simple and flexible control of frequency-dependent acoustic attributes), for example, by providing
<img file="MX352134B_D0010.tif" />
IMPI WfTTUTO MEXICAN
OF THE PROPERTY
INDUSTRIAL the ability to vary reverb tank delays in different bands to change modal density as a function of frequency;
2. The specific audio downmixing process, used to generate (from the multichannel input audio signal) the mixed audio signal (e.g. monophonic audio mixing) processed in the second processing path, it depends on the distance from the source of each channel and the handling of the direct response in order to maintain the appropriate level and timing relationship between the direct and late responses;
3. An All-Pass Filter (APF) is applied on the second processing path (e.g., at the input or output of the bank of FDNs) to introduce phase diversity and increased echo density without changing the spectrum and / or timbre of the resulting reverb;
Four. Fractional delays are implemented in the feedback path of each FDN in a complex, multi-rate valued structure to overcome problems related to delays quantized to the resolution reduction factor grid;
5. On the FDNs, the outputs of the reverb tangue are linearly mixed directly into the binaural channels, using output mix coefficients that are set based on interaural coherence.
I Μί Ρ I
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL ^> ^^ ”1 *!
desired in each frequency band. Optionally, the reverb tank mapping to the binaural output channels is alternating across frequency bands to achieve balanced delay between the binaural channels. Optionally as well, normalization factors are applied to the reverb tank outputs to equalize their levels while preserving fractional delay and overall power;
6. The frequency and / or modal density dependent reverb decay time is controlled by setting appropriate combinations of reverb tank delays and gains in each frequency band to simulate real rooms;
7. A scale factor per frequency band is applied (e.g., at any of the input or output of the relevant processing path), to:
control a frequency-dependent direct-to-late relationship (DLR, Direct-ToLate) that matches that of a real room (a simple model can be used to calculate the required scale factor based on the target DLR and the reverb decay time, e.g. T<sub>6</sub>q);
provide low-frequency attenuation to mitigate excess combing artifacts and / or low-frequency rumble; and / or applying diffuse field spectral shaping to the FDN responses;
<img file="MX352134B_D0011.tif" />
8. Simple parametric models are implemented to control essential frequency-dependent attributes of late reverb, such as reverb decay time, interaural coherence, and / or direct-evening relationship.
Aspects of the invention include methods and systems that carry out (or are configured to carry out, or support the performance of) binaural virtualization of audio signals (eg, audio signals whose audio content consists of channels speaker, and / or object-based audio signals).
In another class of embodiments, the invention is a method and system for generating a binaural signal in response to a set of channels of a multichannel audio input signal, including applying a room binaural impulsive response (BRIR) to each channel. of the set, thus generating filtered signals, including by using at least one feedback delay network (FDN) to apply a common late reverb to an audio mix (downmix) of the ensemble's channels; and combine the filtered signals to generate the binaural signal. The FDN is implemented in the time domain In some modalities, the time domain FDN includes:
an input filter that has an input coupled to
<img file="MX352134B_D0012.tif" />
IMPI
INSTITUTE
OF INDUSTRIAL AVERAGE receive the audio mix, where the fiTllU 'lie cnÍrada. is configured to generate a first filtered audio mix in response to the audio mix;
an all pass filter, coupled and configured for a second filtered audio mix in response to the first filtered audio mix;
a reverb application subsystem, having a first output and a second output, wherein the reverb application subsystem comprises a set of reverb tanks, each of the reverb tanks has a different delay, and wherein the subsystem reverb application is coupled and configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered audio mix, to assert the first unmixed binaural channel on the first output, and assert the second unmixed binaural channel on the second output; and an Interaural Cross-Correlation Coefficient (IACC) and mixing stage of filtering coupled to the reverb application subsystem and configured to generate a first mixed binaural channel and a second binaural mixed channel in response to the first binaural channel unmixed and a second unmixed binaural channel.
The input filter can be implemented to generate
IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL (preferably as a cascade of two filters configured to generate) the first audio mix filtered in such a way that each BRIR has a direct-to-late relationship (DLR) that conforms, at least substantially, to a target DLR.
Each reverb tank can be configured to generate a delayed signal, and can include a reverb filter (e.g., implemented as a shelf filter or a cascade of limiter filters) coupled and configured to apply a gain to a signal that propagates in each of said reverb tanks, to cause the delayed signal to have a gain that is at least substantially adequate at a target decay gain for that delayed signal, in an effort to achieve a target reverb decay time characteristic (eg, a characteristic Tso) of each BRIR.
In some embodiments, the first unmixed binaural channel drives the second unmixed binaural channel, the reverb tanks include a first reverb tank configured to generate a delayed first signal having a shorter delay and a second reverb tank configured to generate a delayed second signal having a shorter second delay, wherein the first reverb tank is configured to apply
2i IMPI 07
MEXICAN INSTITUTE
OF THE PROPERTY <'»» «™ (· βΙ<sub>β</sub>
INDUSTRIAL - a first gain to the first delayed signal, the second reverb tank is configured to apply a second gain to the second delayed signal, the second gain is different than the first gain, the second gain is different than the first gain, and applying the first gain and the second gain results in attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel. Generally, the first mixed binaural channel and the second mixed binaural channel are indicative of a re-centered stereo image. In some embodiments, the IACC filtering and mixing stage is configured to generate the first mixed binaural channel and the second mixed binaural channel in such a way that said first mixed binaural channel and said second mixed binaural channel have a characteristic IACC that matches the less substantially to a target characteristic IACC.
Typical embodiments of the invention provide a simple and unified frame of reference to support both input audio consisting of speaker channels, and object-based input audio. In modes where BRIRs are applied to the input signal channels that are object channels, the direct response and early reflection processing carried out on each object channel assumes a source address indicated by metadata.
<img file="MX352134B_D0013.tif" />
IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL provided with the audio content of the object channel. In modes in which BRIRs are applied to input signal channels that are horn channels, direct response and reflection processing early carried out on each horn channel assumes a source address that corresponds to the horn channel (i.e. , the direction of a direct path from an assumed position of a horn corresponding to the assumed listener position). Regardless of whether the input channels are speaker or object channels, late reverb processing is carried out on an audio mix (e.g., monophonic audio mix) of the input channels and does not assume any direction specific source for the audio content of the audio mix.
Other aspects of the invention are a headphone virtualizer configured (eg, programmed) to carry out any embodiment of the inventive method, a system (eg, a stereo, multichannel, or other decoder) that includes such a virtualizer. , and a computer-readable medium (eg, a disk) that stores code to implement any embodiment of the inventive method.
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1 is a block diagram of a conventional headset virtualization system.
<img file="MX352134B_D0014.tif" />
Figure 2 is a block diagram of a system that includes one embodiment of the inventive headset virtualization system.
Figure 3 is a block diagram of another embodiment of the inventive headset virtualization system.
Figure 4 is a block diagram of an FDN of a type included in a typical implementation of the system of Figure 3.
Figure 5 is a graph of the reverberation decay time (T<sub>60</sub>) in milliseconds as a function of frequency in Hz, which can be achieved by a mode of the inventive virtualizer for which the value T<sub>6</sub>or at each of two specific frequencies (f<sub>TO</sub> and f<sub>B</sub>) is stated as follows: T<sub>60</sub>, a = 320 ms in f<sub>TO</sub> = 10 Hz, and T<sub>60</sub>, b = 150 ms in f<sub>B</sub> = 2.4 kHz.
Figure 6 is a graph of interaural coherence (Coh) as a function of frequency in Hz, which can be achieved by one embodiment of the inventive virtualizer for which the control parameters Coh<sub>max</sub>, Coh<sub>min</sub>, and f<sub>c</sub> are set to have the following values: Coh<sub>max</sub> = 0.95, Cohmin = 0.05, and f<sub>c</sub> = 7 00 Hz.
Figure 7 is a direct-to-late relationship (DLR) plot with distance from the source of one meter, in dB, as a function of frequency in Hz, which can be achieved by an inventive virtualizer mode for which the
<img file="MX352134B_D0015.tif" />
IMPI
INSTITUTE MEXICANO DE LA PROHEBaD INDUSTRIAL DLR control parameters<sub>iK</sub>, DLR<sub>s</sub>i<sub>ope</sub>, DLR<sub>m</sub>i<sub>n</sub>, HPF<sub>s</sub>i<sub>ope</sub>, and f<sub>T</sub> are set to have the following values: DLR<sub>iK</sub> = 18 dB,
DLR<sub>s</sub>i<sub>op</sub>e = 6 dB / ΙΟχ frequency, DLR<sub>m</sub>i<sub>n</sub> = 18 dB, HPF<sub>slope</sub> = 6 dB / ΙΟχ frequency, and f<sub>T</sub> = 200 Hz.
Figure 8 is a block diagram of another embodiment of a late reverb processing subsystem of the inventive headset virtualization system.
Figure 9 is a block diagram of a time domain implementation of an FDN, of a type included in some embodiments of the inventive system.
Figure 9A is a block diagram of an example of an implementation of the filter 400 of Figure 9.
Figure 9B is a block diagram of an example of an implementation of the filter 406 of Figure 9.
FIG. 10 is a block diagram of one embodiment of the inventive headset virtualization system, in which the late reverb processing subsystem 221 is implemented in the time domain.
Figure 11 is a block diagram of one embodiment of elements 422, 423, and 424 of the FDN of Figure 9.
Figure 11A is a graph of the frequency response (Rl) of a typical implementation of the filter 500 of Figure 11, the frequency response (R2) of a typical implementation of the filter 501 of Figure 11, and the response of the filters 500 and 501 connected in parallel.
<sup>25</sup> ΙΜΡΙ »5
MEXICAN INSTITUTE
OF THE PROPERTY
Figure 12 is a graph of an example<sup>DUS</sup><3^<sup>TO THE</sup> a characteristic ta ^ C (curve I) that can be luyi'di <sup>1</sup> pbr an implementation of the FDN of Figure 9, and a characteristic IACC objective (curve I<sub>T</sub>)·
Figure 13 is a graph of a T<sub>60</sub> characteristic that can be achieved by an implementation of the FDN of Figure 9, by appropriately implementing each of the filters 406, 407, 408, and 409 implemented as a limiting filter.
Figure 14 is a graph of a T<sub>60</sub> characteristic that can be achieved by an implementation of the FDN of Figure 9, by appropriately implementing each of the filters 406, 407, 408, and 409 implemented as a cascade of two IIR limiting filters.
Notation and Nomenclature
Throughout this disclosure, including in the claims, the term "performing an operation on a signal or data (eg, filtering, scaling, transforming, or applying gain to, the signal or data") is used in a broad sense to denote carrying out the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performing the operation on it).
Throughout this disclosure, including in the claims, the term "system" is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a virtualizer can be referred to as a virtualizer system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) can also be referred to as a system virtualizer (or virtualizer).
Throughout this disclosure, including in the claims, the term processor is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to carry out operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chipset), a digital signal processor programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, a general-purpose processor or programmable computer, and a programmable microprocessor chip or chipset.
Throughout this disclosure, including in the claims, the term analysis filter bank
<img file="MX352134B_D0016.tif" />
IMPI INSTITUTE MEXICANO is used in a broad sense to denS ^ a ^ sT ^ (eg, a subsystem) configured for aoli transform (eg, a time domain to frequency domain transform) into a signal in the time domain to generate values (eg, frequency components) indicative of the content of the signal in the time domain, in each of a set of frequency bands. Throughout this disclosure, including in the claims, the term "filter bank domain" is used in a broad sense to denote the domain of frequency components generated by a transform or an analysis filter bank (eg. , the domain in which such frequency components are processed). Examples of filter bank domains include (but are not limited to) the frequency domain, the quadrature mirror filter domain (QMF), and the hybrid complex quadrature mirror filter domain (HCQMF). Examples of the transform that can be applied by an analysis filter bank include (but are not limited to) a Discrete Cosine Transform (DCT), Modified Discrete Cosine Transform (MDCT), Discrete Fourier Transform (DFT), and a wavelet transform. Examples of analysis filter banks include (but are not limited to) quadrature mirror (QMF) filters, finite impulse response filters
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL (FIR filters), infinite impulse response filters (IIR filters), crossover filters, and filters that have other suitable multi-rate structures.
Throughout this disclosure, including in the claims, the term metadata refers to data separate and different from the corresponding audio data (audio content of a bit stream that also includes metadata). The metadata is associated with audio data, and indicates at least one characteristic of the audio data (e.g., what type (s) of processing (s) has already taken place, or should ( n) carry out, on the audio data, or the path of an object indicated by the audio data). The association of metadata with audio data is synchronous in time. Therefore, the metadata present (most recently received or updated) may indicate that the corresponding audio data contemporaneously has an indicated characteristic and / or comprises the results of an indicated type of audio data processing.
Throughout this disclosure, including in the claims, the term coupled or coupled is used in the sense of a direct or indirect connection. Therefore, if a first device is coupled to a second device, that connection can be through a direct connection, or through an indirect connection through
<img file="MX352134B_D0017.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX352134B_D0018.tif" />
other devices and connections. ____—
Throughout this disclosure, including in the claims, the following expressions have the following definitions:
horn and loudspeaker are used synonymously to denote any sound-emitting transducer. This definition includes loudspeakers implemented as multiple drivers (eg, woofer and tweeter, woofer and tweeter respectively);
speaker feed: an audio signal to be applied directly to a loudspeaker, or an audio signal to be applied to an amplifier and loudspeaker in series;
channel (or audio channel): a monophonic audio signal. Such a signal can generally be represented in such a way as to be equivalent to applying the signal directly to a loudspeaker at a desired or nominal position. The desired position can be static, as is generally the case with physical speakers, or dynamic;
audio program: a set of one or more audio channels (at least one speaker channel and / or at least one object channel) and optionally also associated metadata (e.g. metadata describing a desired spatial audio presentation in a);
horn channel (or horn feed channel):
<img file="MX352134B_D0019.tif" />
IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL an audio channel that is associated with a designated loudspeaker (in a desired or nominal position), or with a designated speaker zone with a defined speaker configuration.
A horn channel is represented in such a way as to be equivalent to applying the audio signal directly to the designated loudspeaker (at the desired or nominal position) or to a horn in the designated horn area;
object channel: an audio channel indicative of sound emitted by an audio source (sometimes referred to as an audio object). Generally, an object channel determines a parametric audio source description (eg, metadata indicative of the parametric audio source description is included or provided with the object channel). The source description can determine the sound emitted by the source (as a function of time), the apparent position (e.g. 3D spatial coordinates) of the source as a function of time, and optionally at least one additional parameter (p. g., size or width of the apparent font) that characterizes the font;
object-based audio program: an audio program that comprises a set of one or more object channels (and optionally also comprises at least one speaker channel) and optionally also associated metadata (e.g. metadata indicative of a trajectory of an audio object that outputs sound indicated by an object channel, or metadata of <sup>31</sup> IMPI ^
MEXICAN INSTITUTE
OF THE PROñEDAt
INDUSTRY!
another form indicative of a desired spatial audio presentation of sound indicated by an object channel, or metadata indicative of an identification of at least one audio object that is a sound source indicated by an object channel); and render or render: the process of converting an audio program to one or more speaker feeds, or the process of converting an audio program to one or more speaker feeds and converting the speaker feed (s) to sound using one or more loudspeakers (in the latter case, the rendering is sometimes referred to in this document as rendering by the loudspeaker (s)). An audio channel can be represented trivially (in a desired position) by applying the signal directly to a physical speaker in the desired position, or one or more audio channels can be represented using one of a variety of virtualization techniques designed to be substantially equivalent (to the listener) to such trivial representation. In the latter case, each audio channel can be converted into one or more speaker feeds to be applied to the loudspeaker (s) at known locations, which are generally different from the desired position, such that the sound emitted by the speaker (s) in response to the power (s) will be perceived as emitting from the desired position.
IMPI
MEXICAN INSTITUTE
Examples of such virtualization techniques<sup>Dum</sup>they include binaural representation via headphones - (eg, using Dolby Headphone processing that simulates up to 7.1 channels of surround sound for the headphone user) and wavefield synthesis.
The notation that a multichannel audio signal is an xy or xyz channel signal denotes in this document that the signal has x full frequency horn channels (corresponding to speakers nominally positioned in the horizontal plane of the putative listener's ears) , and LFE (or subwoofer) channels, and optionally also z full-frequency top speaker channels (corresponding to speakers positioned above the head of the intended listener, e.g. on or near the ceiling of a room).
The expression IACC in this document denotes the interaural cross-correlation coefficient in its usual sense, which is a measure of the difference between the arrival times of the audio signal at a listener's ears, usually indicated by a number in a range from a first value that indicates that the arrival signals are equal in magnitude and exactly out of phase, to an intermediate value that indicates that the arrival signals have no similarity, to a maximum value indicating that identical arrival signals have the same amplitude and phase.
IMPI
MEXICAN INSTITUTE
OF THE PROPERTY Λ «~ 3ή zz INDUSTRIAL
DETAILED DESCRIPTION OF THE INVENTION
Many embodiments of the present invention are technologically possible. It will be apparent to those skilled in the art from the present disclosure how to implement them. The embodiments of the inventive system and method will be described with reference to Figures 2-14.
Figure 2 is a block diagram of a system (20) that includes one embodiment of the inventive headset virtualization system. The headset virtualization system (sometimes referred to as a virtualizer) is configured to apply a room binaural impulsive response (BRIR) to N full-range frequency channels (Xi, X<sub>N</sub>) of a multichannel audio input signal. Each of the channels Xi, X<sub>N</sub>, (which can be horn channels or object channels) corresponds to a specific source direction and distance relative to a supposed listener, and the system of Figure 2 is configured to convolve each such channel by a BRIR for the address and corresponding distance from the source.
System 20 may be a decoder that is coupled to receive a scrambled audio program, and includes a subsystem (not shown in Figure 2) coupled and configured to decode the program including retrieving the N full-range channels of <sup>34</sup> IMPI ^
INSTITUTE MEXICANO Ι ^ - ^ 'τηΤΖ
OF THE PROPERTY
INDUSTRIAL frequencies (Xi, ..., X<sub>N</sub>) thereof and provide them to elements 12, ..., 14, and 15 of the virtualization system (comprising elements, 12, ..., 14, 15, 16, and 18, coupled as shown). The decoder may include additional subsystems, some of which carry out functions unrelated to the virtualization function carried out by means of the virtualization system, and some of which may carry out functions related to the virtualization function. For example, the latter functions may include extraction of metadata from the encoded program, and provision of the metadata to a virtualization control subsystem that uses the metadata to control elements of the virtualizer system.
Subsystem 12 (with subsystem 15) is configured to convolve channel Xi with BRIRi (the BRIR for the corresponding direction and distance from the source), subsystem 14 (with subsystem 15) is configured to convolve channel X<sub>N</sub> with BRIR<sub>N</sub> (the BRIR for the corresponding source address), and so on for each of the other N-2 subsystems of BRIR. The output of each of subsystems 12, 14, and 15 is a time domain signal that includes a left channel and a right channel. Additional elements 16 and 18 are coupled to the outputs of elements 12, ..., 14, and 15. Additional element 16 is configured to combine (mix) the
<img file="MX352134B_D0020.tif" />
IMPI
INSTITUTE MEXICANO oe the INDUSTRIAL property left channel outputs of the BRIR subsystems, and additional element 18 is configured 'to combine (mix) the right channel outputs of the BRIR subsystems.
BRIR. The output of item 16 is the left channel, L, of the binaural audio signal output from the virtualizer of Figure 2, and the output of item 18 is the right channel, R, of the binaural audio signal output from the virtualizer in Figure 2.
Important features of typical embodiments of the invention are apparent from the comparison of the Figure 2 embodiment of the inventive headset virtualizer with the conventional headset virtualizer of Figure 1. For comparison purposes, we assume that the systems in Figure 1 and Figure 2 are configured such that, when the same multichannel audio input signal is asserted to each of them, the systems apply a BRIRj that has the same direct response and early reflection portion (that is, the relevant EBRIRi in Figure 2) to each full-range frequency channel, X<sub>if</sub> of the input signal (although not necessarily with the same degree of success). Each BRIRj applied by the system of Figure 1 or Figure 2 can be decomposed into two portions: a portion of direct response and early reflection (e.g. one of the portions EBIRi, ..., EBRIR<sub>N</sub> applied by subsystems 12-14 of the <sup>36</sup> IMPI ^
MEXICAN INSTITUTE
OF THE PROPERTY
Figure 2), and a late reverb portion. The embodiment of Figure 2 (and other typical embodiments of the invention assume that the late reverb portions of the single channel BRIRs, BRIRi, can be shared across source addresses and thus all channels, and therefore apply the same late reverb (that is, a common late reverb) to an audio mix (doivnmíx) of all full-range frequency channels of the input signal. This audio mix can be a monophonic (mono) audio mix of all input channels, but it can alternatively be a stereo or multichannel audio mix obtained from the input channels (e.g., a subset of input channels).
More specifically, subsystem 12 of Figure 2 is configured to convolve the input signal channel Xi with EBRIRi (the forward response and early reflection BRIR portion for the corresponding source direction), subsystem 14 is configured to convolve channel X<sub>N</sub> with EBRIR<sub>N</sub> (the direct response and BRIR portion of early reflection for the corresponding direction of the source), and so on. The late reverb subsystem 15 in Figure 2 is configured to generate a mono audio mix of all full-range frequency channels of the input signal, and to convolve
<img file="MX352134B_D0021.tif" />
IMPI
MEXICAN INSTITUTE
FROM 1.TO INDUSTRIAL PROPERTY mixing audio with LBRIR (a late reverb common to all channels undergoing audio mixing). The output of each BRIR subsystem of the virtualizer of Figure 2 (each of subsystems 12, ..., 14, and 15) includes a left channel and a right channel (of a binaural signal generated from the speaker channel corresponding or downmix). The left channel outputs of the BRIR subsystems are combined (mixed) in add-on 16, and the right channel outputs of BRIR subsystems are combined (mixed) in add-on 18.
The addition element 16 can be implemented to simply sum the corresponding Left binaural channel samples (the Left channel outputs of subsystems 12, ..., 14, and 15) to generate the Left channel of the binaural output signal, assuming the appropriate level settings and time alignments are implemented in subsystems 12, ..., 14 and 15. Similarly, the addition element 18 can also be implemented to simply sum the corresponding Right binaural channel samples (e.g., the Right channel outputs of subsystems 12, ..., 14, and 15) to generate the Right channel of the binaural output signal, again assuming that the appropriate level settings and time alignments are implemented in subsystems 12, ..., 14, and 15.
IMPI
INSTITUTE MEXICANO Of LA PROPIEDAD industrial
Subsystem 15 of Figure 2 can be implemented in any of a variety of ways, but generally includes at least one feedback delay network configured to apply the late reverb common to a monophonic audio mix of the asserted input signal channels. the same. Generally, when each of the subsystems 12, ..., 14 applies a direct response and early reflection portion (EBRIRi) of a single channel BRIR to the channel (Xi) it processes, the common late reverb has been generated. to emulate collective macro attributes of late reverb portions of at least some (e.g. all) of the single channel BRIRs (whose direct response and early reflection portions are applied by subsystems 12, ..., 14 ). For example, an implementation of subsystem 15 has the same structure as subsystem 200 of Figure 3, including a bank of feedback delay networks (203, 204, ..., 205) configured to apply a common late reverb to a monophonic audio mix of the input signal channels asserted to them.
The subsystems 12, ..., 14 of Figure 2 can be implemented in any of a variety of ways (in either a time domain or a filter bank domain), with the preferred implementation for any specific application depending on different considerations, such as (eg INDUSTRIAL computing, and memory. In an exemplary implementation, each of the subsystems 12, ..., 14 is configured to convolve the channel asserted thereto with an FIR filter corresponding to the direct and early responses associated with the channel, with gain and delay set appropriately. in such a way that the outputs of subsystems 12, ..., 14 can be combined simply and efficiently with those of subsystem 15.
Figure 3 is a block diagram of another embodiment of the inventive headset virtualization system. The modality of Figure 3 is similar to that of Figure 2, with two (left and right channel) time-domain signals being output from the direct response and early reflection processing subsystem 100, and two (left channel and right) signals in the time domain that are output from the late reverb processing subsystem 200. Adding element 210 is coupled to the outputs of subsystems 100 and 200. Element 210 is configured to combine (mix) the left channel outputs of subsystems 100 and 200 to generate the left channel, L, of the output of binaural audio signal from the virtualizer of Figure 3, and to combine (mix) the right channel outputs of subsystems 100 and 200 to generate the right channel, R, of the audio signal output
Binaural IMPI of the virtualizer in Figure 3.
can implement to simply sum the corresponding left channel samples that are output from subsystems 100 and 200 to generate the left channel of the binaural output signal, and to simply sum the corresponding right channel samples that are output from subsystems 100 and 200 to generate the right channel of the binaural output signal, assuming the appropriate level settings and time alignments are implemented in subsystems 100 and 200.
In the system of Figure 3, channels Xi of the multichannel audio input signal are directed to, and undergo processing on, two parallel processing paths: one through the early reflection and direct response processing subsystem 100; the other via late reverb processing subsystem 200. The system of Figure 3 is configured to apply a BRIRi to each channel, X¿. Each BRIRi can be broken down into two portions: a forward response and early reflection portion (applied by subsystem 100), and a late reverb portion (applied by subsystem 200). In operation, the direct response and early reflection processing subsystem 100 therefore generates the direct response and early reflections portions of the binaural audio signal that is output from the virtualizer, and the
<img file="MX352134B_D0022.tif" />
dog
IMPI
MEXICAN INSTITUTE
INDUSTRIAL PROPERTY reverb processing subsystem (late reverb generator) 200 therefore generates the late reverb portion of the binaural audio signal that is output from the virtualizer. The outputs of subsystems 100 and 200 are mixed (by adding subsystem 210) to generate the binaural audio signal, which is generally asserted from subsystem 210 to a rendering system (not shown) in which it experiences binaural rendering for playback. through headphones.
Generally, when rendered and played through a pair of headphones, a typical binaural audio signal that is output from Element 210 is perceived at the listener's eardrums as sound from N speakers (where N 2 and N is generally equal to 2.5 or 7) in any of a wide variety of positions, including positions in front of, behind, and above the listener. Reproduction of the output signals generated in the operation of the system of Figure 3 can give the listener the experience of sound coming from more than two (eg, five or seven) surround sources. At least some of these sources are virtual.
The early reflection and direct response processing subsystem 100 can be implemented in any of a variety of ways (in either the time domain or a filterbank domain), with the implementation
IMPI MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX352134B_D0023.tif" />
preferred for any specific application depending on different considerations, such as (for example) performance, computing, and memory. In an exemplary implementation, subsystem 100 is configured to convolve each channel asserted to it with an FIR filter that corresponds to the direct and early responses associated with the channel, with gain and delay set appropriately such that the outputs of subsystem 100 are can be combined simply and efficiently (in item 210) with those of subsystem 200.
As shown in Figure 3, the late reverb generator 200 includes the downmixing subsystem 201, the analysis filter bank 202, a bank of FDNs (FDNs 203, 204, ..., and 205 ), and synthesis filter bank 207, coupled as shown. Subsystem 201 is configured to mix the audio from the channels of the multichannel input signal into a mono audio mix, and the analysis filter bank 202 is configured to apply a transform to the mono audio mix to divide the mix of mono audio in K frequency bands, where K is an integer. The values in the domain of the filter bank (which are output from the filter bank 202) in each different frequency band are asserted to a different one of the FDNs 203, 204, ..., 205 (there are K of these FDNs, each one attached and configured to apply a
IMPI
Mexican INSTITUTE of LA PROPERTY portion of late reverberation from a BRIR to the values in the domain of the filter bank affirmed to them). The values in the domain of the filter bank are preferably designated in time to reduce the computational complexity of the FDNs.
In principle, each input channel (to subsystem 100 and subsystem 201 of Figure 3) can be processed into its own FDN (or bank of FDNs) to simulate the late reverb portion of its BRIR. Despite the fact that the late reverb portion of the BRIRs associated with different sound source locations are generally different in terms of mean square differences in impulse responses, their statistical attributes such as their average power spectrum, their energy decay structure, modal density, peak density among others are often very similar. Therefore, the late reverb portion of a set of BRIRs is generally perceptually quite similar across channels and consequently it is possible to use a common FDN or bank of FDNs (e.g. FDNs 203, 204, .. ., 205) to simulate the late reverb portion of two or more BRIRs. In typical embodiments, a common FDN (or bank of FDNs) is employed, and the input to it is comprised of one or more audio mixes constructed from the input channels. In the exemplary implementation of the<sup>44</sup> IMPW
INSTITUTE MEXICANO 1 ^ ««
FROM PROPERTY LV .: INDUSTRIAL Via »» Figure 2, the audio mix is a monophonic audio mix (asserted at the output of subsystem 201) of all input channels.
With reference to the modality of Figure 2, each one of the FDNs 203, 204, ..., 205, is implemented in the domain of the filter bank, and is coupled and configured to the process of a different frequency band of the values that are output from the analysis filter bank 202, to generate left and right reverberated signals for each band. For each band, the left reverb signal is a sequence of values in the filter bank domain, and the right reverb signal is another sequence of values in the filter bank domain. Synthesis filter bank 207 is coupled and configured to apply a frequency domain to time domain transform for the 2K sequences of values in the domain of the filter bank (e.g., frequency components in the domain of QMF) that are output from the FDNs, and to assemble the transformed values into a left channel time domain signal (indicative of the audio content of the mono audio mix to which late reverb has been applied) and a right channel time domain signal (Also indicative of the audio content of the mono audio mix to which the late reverb has been applied.) These left channel and right channel signals are output
<img file="MX352134B_D0024.tif" />
IMPI
MEXICAN INSTITUTE
FROM INDUSTRIAL PROPERTY to item 210.
In a typical implementation of each of the FDNs
203, 204, ..., and 205, is implemented in the QMF domain, and filter bank 202 transforms the mono audio mix from subsystem 201 to the QMF domain (e.g., mirror filter domain in Hybrid Complex Quadrature (HCQMF)), such that the asserted signal from filter bank 202 to one input of each of the FDNs 203, 204, ..., and 205 is a sequence of frequency components in the domain of QMF. In such an implementation, the asserted signal from filter bank 202 to FDN 203 is a sequence of frequency components in the QMF domain in a first frequency band, the asserted signal from filter bank 202 to FDN 204 is a sequence of frequency components in the QMF domain in a second frequency band, and the asserted signal from filter bank 202 to FDN 205 is a sequence of frequency components in the QMF domain in a K<sup>to</sup> frequency band. When the analysis filter bank 202 is implemented in this manner, the synthesis filter bank 207 is configured to apply a QMF domain time domain transform to the 2K frequency component sequences in the output QMF domain. of the FDNs, to generate the left channel and right channel late reverb time domain signals that are output to item 210.
IMPI
MEXICAN INSTITUTE
OF THE PROPERTY Q »ctAíÍ
INDUSTRIAL
For example, if K = 3 in the system of Figure 3, then there are six inputs to synthesis filter bank 207 (left and right channels, comprising samples in the frequency domain or QMF domain, which are output from each of FDNs 203, 204, and 205) and two 207 outputs (left and right channels, each consisting of time domain samples). In this example, the filter bank 207 would be generally implemented as two synthesis filter banks: one (to which the three left channels of the FDNs 203, 204, and 205 would be asserted) configured to generate the left channel signal in the time domain that is output from filter bank 207; and a second (to which the three right channels of the FDNs 203, 204, and 205 would be asserted) configured to generate the right channel signal in the time domain that is output from filter bank 207.
Optionally, control subsystem 209 is coupled to each of the FDNs 203, 204, ..., 205, and configured to assert control parameters to each of the FDNs to determine the late reverberation portion (LBRIR) to be applied by subsystem 200. Examples of such control parameters are described below. It is contemplated that in some implementations the control subsystem 209 is operable in real time (e.g., in response to user commands asserted thereto by means of a
IMPI
INSTITUTE MEXICANO M LA PROPIEDAD INDUSTRIAL input device) to implement real-time variation of the portion of reverb, late (LBRIR) applied by subsystem 200 to the monophonic audio mix of input channels.
For example, if the input signal to the system in Figure 2 is a 5.1 channel signal (whose full frequency range channels are in the following channel order: L, R, C, Ls, Rs), all channels full-range frequencies have the same distance from the source, and the 201 audio mixing subsystem can be implemented as the following audio mixing matrix, which simply adds the full-range frequency channels to form a mono audio mix :
£) = [! 1 1 1 1]
After the filtration passes everything (in element 301 in each of the FDNs 203, 204, ..., and 205), the channels are separated (upmix) of the mono audio mix to the four reverb tanks in one energy conservation way:
MEXICAN INSTITUTE
OF INDUSTRIAL PROPERTY
Alternatively (as an example), we can choose to move the left side channels to the first two reverb tanks, the right side channels to the last two reverb tanks, and the center channel to all reverb tanks. In this case, the audio mixing subsystem 201 would be implemented to form two channel mix signals:
o 1 / V2 1 o 0 1 1/72 o 1
In this example, the channel spacing (upmix) to the reverb tanks (in each of the FDNs 203, 204, ..., and 205) is:
Because there are two audio mix signals, the filtering pass all (at item 301 in each of the FDNs 203, 204, ..., and 205) needs to be applied twice. Diversity would be introduced for the late responses of (L, Ls), (R, Rs) and C even though they all have the same macro attributes. When the input signal channels have different distances from the source, they are still<sup>49</sup> IMPI
MEXICAN INSTITUTE »
OF INDUSTRIAL PROPERTY,,, would need to apply appropriate delay and gain audio mixing process. ————
Here we describe considerations for the specific implementations of the audio mixing subsystem 201, and the subsystems 100 and 200 of the virtualizer of Figure 3.
The audio mixing process implemented by means of subsystem 201 depends on the source distance (between the sound source and the assumed position of the listener) for each channel that will have audio mixing, and the handling of the direct response. The delay of the direct response is:
<img file="MX352134B_D0025.tif" />
t<sub>d</sub> = d / v<sub>s</sub> where d is the distance between the sound source and the listener and v<sub>s</sub> is the speed of sound. Also, the gain of the direct response is proportional to 1 / d. If these rules are preserved in handling the direct responses of channels with different distances from the source, subsystem 201 can implement direct audio mixing of all channels because the delay and level of late reverb is generally insensitive. to the location of the source.
Due to practical considerations, virtualizers (e.g., virtualizer subsystem 100 in Figure 3) <sup>50</sup> IMPI ^
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL - can be implemented to time align direct responses for input channels that have different distances from the source. In order to preserve the relative delay between the direct response and the late reverb for each channel, a channel with distance from the source d should be delayed by (dmax - d) / v<sub>s</sub> before mixing audio with other channels. Here dmax denotes the maximum possible distance from the source.
Virtualizers (eg, virtualizer subsystem 100 in Figure 3) can also be implemented to compress the dynamic range of direct responses. For example, the direct response for a channel with distance from the source d can be scaled by a factor of cT<sup>to</sup>, where 0 <a <1, instead of cT<sup>1</sup>. In order to preserve the level difference between direct response and late reverb, the audio mixing subsystem 201 may need to be implemented to scale a channel with distance from the source d by a factor of d<sup>1-a</sup> before mixing your audio with other scaled channels.
The feedback delay network of Figure 4 is an exemplary implementation of the FDN 203 (or 204 or 205) of Figure 3. Although the system of Figure 4 has four reverb tanks (each including a gain stage, gi, and a delay line, z ~<sup>neither</sup>, coupled to the output of the gain stage), variations in them the <sup>51</sup> IMPI ^
MEXICAN INSTITUTE
THE INDUSTRIAL PROPERTY system (and other FDNs employed in inventive virtualizer modes) employs more or less than four reverb tanks.
The FDN of Figure 4 includes input gain element 300, all pass filter (APF) 301 coupled to the output of element 300, addition elements 302, 303, 304, and 305 coupled to the output of APF 301, and four reverb tanks (each comprising a gain element, g ^ (one of the 306 elements), a delay line, z (one of the 307 elements) coupled thereto, and a gain element, 1 / g * (one of elements 309) attached thereto, where 0 k - 1 3) each coupled to the output of a different one of elements 302, 303, 304, and
305. The unit matrix 308 is coupled to the outputs of the delay lines 307, and is configured to assert a feedback output to a second input of each of the elements 302, 303, 304, and 305. The outputs of two of the Gain elements 309 (from the first and second reverb tanks) assert to the inputs of the addition element 310, and the output of the element 310 asserts to an input of the output mix matrix 312. The outputs of the other two of the gain elements 309 (from the third and fourth reverb tanks) are asserted to inputs of addition element 311, and the output of element 311 is asserted to the other input of the array
IMPI mix output 312.
Element 302 is configured to add the output of matrix 308 that corresponds to the delay line z '<sup>nl </sup>(that is, to apply feedback from the output of the delay line z '<sup>nl</sup> via matrix 308) at the inlet of the first reverb tank. Element 303 is configured to add the output of matrix 308 that corresponds to the delay line z '<sup>n2</sup> (that is, to apply feedback from the output of the delay line z '<sup>n2</sup> via matrix 308) at the inlet of the second reverb tank. Element 304 is set to add the output of matrix 308 that corresponds to the lag line z<sup>n3</sup> (that is, to apply feedback from the output of the delay line z ~<sup>n3</sup> via matrix 308) at the inlet of the third reverb tank. Element 305 is configured to add the output of matrix 308 that corresponds to the delay line z '<sup>n4</sup> (that is, to apply feedback from the output of the delay line z ~ <sup>n4</sup> via matrix 308) at the inlet of the fourth reverb tank.
The input gain element 300 of the FDN of Figure 4 is coupled to receive a frequency band of the transformed monophonic audio mix signal (a signal in the filter bank domain) that is output from the filter bank of analysis 202 of Figure 3. The element
<img file="MX352134B_D0026.tif" />
IMPI
MEXICAN INSTITUTE
FROM INDUSTRIAL PROPERTY input gain 300 apply a gain (scaling) factor, G<sub>in</sub>, to the signal in the domain of the filter bank asserted to it. Collectively, the Gi scale factors<sub>n </sub>(implemented by all FDNs 203, 204, ..., 205 of Figure 3) for all frequency bands they control the spectral shaping and level of the late reverberation. When setting the input gains, Gi<sub>n</sub>In all the FDNs in Figure 3, the virtualizer often takes the following goals into account:
a direct-to-late relationship (DLR), of the BRIR applied to each channel, which is adapted to real rooms;
low-frequency attenuation necessary to mitigate excess interlace artifacts and / or low-frequency rumble; and suitability of the diffuse field spectral envelope.
If we assume that the direct response (applied by subsystem 100 of Figure 3) provides unity gain in all frequency bands, a specific DLR (power ratio) can be achieved by setting G<sub>in</sub> so that 3Θ3 *
G<sub>in</sub> = sqrt (ln (10<sup>6</sup>) / (T<sub>60</sub> * DLR)), where T<sub>6</sub>o is the reverb decay time defined as the time it takes for the reverb to decay
<img file="MX352134B_D0027.tif" />
IMPI
MEX1CAN INSTITUTE OF INDUSTRIAL PROPERTY at 60 dB (determined by reverb delays and reverb gains discussed below), and
In denotes the natural logarithm function.
The input gain factor, G<sub>in</sub>, it may be dependent on the content that is being processed. One application of such content dependence is to ensure that the energy of the audio mix in each time / frequency segment is equal to the sum of the energies of the individual channel signals being mixed, regardless of any correlation that may occur. exist between the input channel signals. In that case, the input gain factor can be (or can be multiplied by) a term similar to or equal to:
ΣΣΧω and 'Σ / οί where i is an index through all the audio mix samples of a time / frequency or subband mosaic, and (i) are the audio mix samples for the mosaic, and Xí ( j) is the input signal (for channel XJ asserted to the input of the audio mixing subsystem 201.
In a typical QMF domain implementation of the FDN of Figure 4, the asserted signal from the output of the all-pass filter (APF) 301 to the inputs of the reverb tanks is a sequence of frequency components
<img file="MX352134B_D0028.tif" />
IMPI in the QMF domain.
INSTITUTE MEXICAN M LA PROmrMO ____
To generate FDN output<sup>IND</sup>For more natural sounding, the APF 301 is applied to the output of the gain element 300 to introduce phase diversity and increased echo density. Alternatively, or additionally, one or more pass-all-delay filters can be applied to: the individual inputs to audio mixing subsystem 201 (of Figure 3) before they are mixed in subsystem 201 and processed by the FDN; or in the reverb tank feed-back or feed-back paths shown in Figure 4 (e.g., in addition or replacement of delay lines z ~<sup>Mk</sup> in each reverb tank; or the outputs of the FDN (that is, the outputs of the output matrix 312).
In the implementation of the reverb tank delays, z ~<sup>neither</sup>, the reverb delays ni should be mutually prime numbers to prevent the reverb modes from aligning on the same frequency. The sum of the delays must be large enough to provide enough modal density to avoid artificial sound output. But the shorter delays should be short enough to avoid the excess time gap between the late reverb and the other components of the BRIR.
Generally, the reverb tank outputs are initially shifted to either the left or right binaural channel. Typically, sets of outputs
<img file="MX352134B_D0029.tif" />
IMPI of the displaced reverb tank
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL the two binaurals are equal in number and mutually exclusive.
You also want to balance the timing of the two binaural channels. So if the output of the reverb tank with the shortest delay goes to a binaural channel, the one with the second shortest delay would go to the other channel.
Reverb tank delays can be different across frequency bands to change modal density as a function of frequency. Generally, lower frequency bands require higher modal density, consequently longer reverb tank delays.
The amplitudes of the reverb tank gains, gi, and the reverb tank delays together determine the reverb decay time of the FDN of Figure 4:
Tgo - ~ 3rii / logio (lgríl) / Ffrm where F<sub>frm</sub> is the frame rate of filter bank 202 (from Figure 3). The reverb tank gains phases introduce fractional delays to overcome problems related to reverb tank delays that are quantized in the
IMPI
INSTITUTE MEXiCAN # DE LA PROPERTY INDUSTRIAL
<img file="MX352134B_D0030.tif" />
filter bank resolution reduction factor grid.
The unity feedback matrix 308 provides uniform mixing between the reverb tanks in the feedback path.
To equalize the levels of the reverb tank outputs, gain elements 309 apply a normalizing gain, l / lgil to the output of each reverb tank, to remove the level impact of the reverb tank gains while they preserve the fractional delays introduced by their phases.
The 312 output mix matrix (also identified as the M<sub>out</sub>) is a 2x2 matrix configured to mix the unmixed binaural channels (the outputs of elements 310 and 311, respectively) from the initial offset to achieve left and right binaural output channels (the L and R signals asserted at the output of the matrix 312) that have the desired interaural coherence. The unmixed binaural channels are close to being uncorrelated after the initial offset because they do not consist of any common reverb tank outputs. If the desired interaural coherence is Coh, where | Coh | ú 1, the output mix matrix 312 can be defined as:
<img file="MX352134B_D0031.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
M<sub>out</sub> eos β sin β sin β eos β where β = arcsin (Co ^) 77
Because the reverb tank delays are different, one of the unmixed binaural channels would drive the other constantly. If the combination of reverb tank delays and offset pattern is identical across the frequency bands, sound image drift would result. This drift can be mitigated if the offset pattern alternates through the frequency bands in such a way that the mixed binaural channels lead and follow each other in alternating frequency bands. This can be achieved by implementing the output mix matrix 312 to be shaped as stated in the previous paragraph in odd-numbered frequency bands (that is, in the first frequency band (processed by the FDN 203 of Figure 3 ), the third frequency band, and so on), and to have the following shape in even-numbered frequency bands (that is, in the second frequency band (processed by the FDN 204 of Figure 3), the fourth frequency band, and so on):
M<sub>out</sub>,<sub>ah</sub> = sin β eos / 7 eos β sin β
IMPI ^ faith
Mexican INSTITUTE
M THE PROPERTY VK.
INDUSTRIAL where the definition of β remains the same. It should be noted that matrix 312 can be implemented to be identical in the FDNs for all frequency bands, but the channel order of its inputs can be changed for alternating frequency bands (e.g., the output of the element 310 can be asserted to the first input of matrix 312 and the output of element 311 can be asserted to the second input of matrix 312 in odd frequency bands, and the output of element 311 can be asserted to the first input of matrix 312 and the output of element 310 can be asserted to the second input of matrix 312 in even frequency bands).
In the event that the frequency bands overlap (partially), the width of the frequency range through which the shape of the matrix 312 alternates can be increased (e.g., it could alternate once every two or three executive bands), or the value of β in the above expressions (for the 312 matrix shape) can be adjusted to ensure that the average coherence is equal to the desired value to compensate for the spectral overlap of consecutive frequency bands. '
If the previously defined target acoustic attributes Tgo, Coh, and DLR are known to the FDN for each specific frequency band in the inventive virtualizer, each of the FDNs (each of which may have the
<img file="MX352134B_D0032.tif" />
IMPI
INSTITUTE MEXICANO DE EA PROPERTY INDUSTRY !.
structure shown in Figure 4) can be configured to achieve the target attributes. Specifically, in some modes, the input gain (Gi<sub>n</sub>) and the reverb tank gains and delays (g¿ and n¿) and the parameters of the output matrix M<sub>ou</sub>t for each FDN can be set (eg, by means of control values asserted thereto by control subsystem 209 of Figure 3) to achieve target attributes in accordance with the relationships described in this document. In practice, setting the frequency-dependent attributes using models with simple control parameters is often sufficient to generate natural resonant late reverb that is suited to specific acoustic environments.
Below we describe an example of how a target reverb decay time (Tgo) can be determined for the FDN for each specific frequency band of an inventive virtualizer mode, by determining the target reverb decay time (Tgo) for each one of a small number of frequency bands. The FDN response level decays exponentially over time. T<sub>60</sub> is inversely proportional to the decay factor, df (defined as dB decay over a unit of time):
T<sub>6</sub>o = 60 / df.
IMPI
INSTITUTE MEXICANO DE IA INDUSTRIAL PROPERTY
<img file="MX352134B_D0033.tif" />
The decay factor, df, is frequency dependent and generally increases linearly against the logarithmic frequency scale, such that the reverberation decay time is also a function of frequency that generally decreases as the frequency increases. Therefore, if one determines (e.g., sets) the values of T<sub>60</sub> for two frequency points, the curve of T is determined<sub>6</sub>or for all frequencies. For example, if the reverb decay times for the frequency points f<sub>TO</sub> and f<sub>B</sub> are T<sub>6</sub>o, ay T<sub>6</sub>q, b, respectively, the curve of T<sub>60</sub> is defined as:
Figure 5 shows an example of a T curve<sub>6</sub>q that can be achieved by a mode of the inventive virtualizer for which the value of T<sub>6</sub>or at each of two specific frequencies (f<sub>TO</sub> and f<sub>B</sub>) is set: Tgo, A = 320 ms in f<sub>TO</sub> = 10 Hz, and T<sub>6</sub>or, b = 150 ms in f<sub>B</sub> = 2.4 kHz.
In the following we describe an example of how a target interaural coherence (Coh) for the FDN can be achieved for each specific frequency band of an inventive virtualizer mode by setting a small number of control parameters. The interaural coherence (Coh) of the late reverb follows the pattern of a diffuse field of
IMPIAS,
MEXICAN INSTITUTE ___. r, _ _. ___1 _ 1 _ __, ·, I »THE PROPERTY sound. It can be modeled by means of uneDusTfiuncSaSClSBe synchronization up to a frequency rmrp f<sub>c</sub>. and a constant above the crossover frequency. A simple model for the Coh curve is:
Coh {f) =, / </<sub>c</sub> . <sup>C</sup>°<sup>h</sup><sub>me</sub>n> f ^ fc where the parameters Coh<sub>min</sub> and Coh<sub>max</sub> satisfy -1 <Coh<sub>m</sub>in <Coh<sub>m</sub>ax 1, and control the range of Coh. The crossover frequency f<sub>c</sub> Optimal depends on the listener's head size. An F<sub>c</sub> Too high leads to an internalized sound source image, while too small a value leads to a scattered or divided sound source image. Figure 6 is an example of a Coh curve that can be achieved by means of an inventive virtualizer mode for which the Coh parameters<sub>max</sub>, Coh<sub>m</sub>i<sub>n</sub>, and f<sub>c</sub> are set to have the following values: Coh<sub>max</sub> - 0.95, Coh<sub>min</sub> = 0.05, and f<sub>c</sub> = 700 Hz.
Below we describe an example of how a target direct-to-late (DLR) ratio for the FDN can be achieved for each specific frequency band of an inventive virtualizer mode by setting a small number of control parameters. The direct-to-late ratio (DLR), in dB) generally increases linearly against the logarithmic frequency scale. Can be controlled by setting DLRi<sub>K </sub>(DLR in dB @ 1 kHz) and DLR<sub>s</sub>i<sub>OR</sub>p<sub>and</sub> (in dB per 10 frequency). Without
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX352134B_D0034.tif" />
However, low DLR in the lower frsnnRnce range often results in excess interlacing artifacts. In order to mitigate artifacts, two modifying mechanisms are added to DLR control:
a minimum floor of DLR, DLRmin (in dB); and a high-pass filter defined by a transition frequency, f<sub>T</sub>, and the slope of the attenuation curve below it, HPF<sub>s</sub>i<sub>ope</sub> (in dB per 10 frequency).
The resulting DLR curve in dB is defined as:
DLRIJ) = <sub>m</sub>ax (Di ^<sub>K</sub>+ Di ^ log,<sub>0</sub>(// 1000), DLR ,, .. ,,} + min (ffi> F ^ log ,, (///,), o)
It should be noted that the DLR changes with distance from the source even in the same acoustic environment. Therefore, both DLRi<sub>K</sub> and DLR<sub>min</sub> here are the values for a nominal source distance, such as 1 meter. Figure 7 is an example of a DLR curve for a source distance of 1 meter achieved by an inventive virtualizer mode with DLRi control parameters.<sub>K</sub>, DLR<sub>s</sub>i<sub>op</sub>ez DLR<sub>m</sub>i<sub>n</sub>, HPF<sub>s</sub>iope, and £ τ set to have the following values: DLRik = 18 dB, DLR<sub>s</sub>i<sub>ope</sub> = 6 dB / ΙΟχ frequency, DLR<sub>min</sub> = 18 dB, HPF<sub>s</sub>i<sub>OR</sub>p<sub>and</sub> = 6 dB / ΙΟχ frequency, and f<sub>T</sub> = 200 Hz.
Variations in the modalities disclosed in this document have one or more of the following characteristics:
INSTITUTE MEXICANO DE LA mOHEDAC 'η i Ί τ · i ·. INDUSTRIAL inventive virtualizer FDNs are either time domain implemented, or have non-compliant implementation with FDN based impulse response capture and FIR based signal filtering;
the inventive virtualizer is implemented to allow the application of power compensation as a function of frequency during performance of the audio mix step that generates the audio mix input signal (downmix) for the late reverb processing subsystem; and the inventive virtualizer is implemented to allow manual or automatic control of the applied late reverb attributes in response to external factors (that is, in response to the setting of control parameters).
For applications where the system latency is critical and the delay caused by the analysis and synthesis filter banks is prohibitive, the FDN structure in the domain of the filter bank of typical modalities of the inventive virtualizer can be translated into the domain of the inventive virtualizer. time, and each FDN structure can be implemented in the time domain in a class of virtualizer modes. In time-domain implementations, the subsystems that apply the input gain factor (Gin), reverb tank gains (gU, and normalization gains (1/1 |) are replaced by filters with <sup>65</sup> IMPIOS Mexican institute
OF THE PROPERTY
INDUSTRIAL similar amplitude responses in order to allow frequency-dependent controls. The matrix of 'mix output (M<sub>out</sub>) is also replaced by an array of filters. Unlike the other filters, the phase response of this filter array is critical as energy conservation and interaural coherence could be affected by the phase response. The reverb tank delays in a time domain implementation may need to be varied slightly (from their values in a filter bank domain implementation) to avoid sharing filter bank pass as a common factor. Due to different constraints, the performance of the inventive virtualizer FDNs time domain implementations might not exactly match that of the filterbank domain implementations thereof.
With reference to Figure 8, we now describe a hybrid implementation (filter bank domain and time domain) of the inventive late reverb processing subsystem of the inventive virtualizer. This hybrid implementation of the inventive late reverb processing subsystem is a variation of the late reverb processing subsystem 200 of Figure 4, implementing FDN-based impulse response capture and FIR-based signal filtering.
<sup>66</sup> IMPI
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL
The embodiment of Figure 8 includes the. elements 201, 202, 203, 204, 205, and 207 which are identical to the identically numbered elements of subsystem 200 of Figure 3. The previous description of these elements will not be repeated with reference to Figure 8. In the mode of Figure 8, pulse generator 211 is coupled to assert a pulse signal (a pulse) to analysis filter bank 202. A LBRIR filter 208 (mono input, stereo output) implemented as a FIR filter applies the appropriate late reverb portion of the BRIR (the LBRIR) to the mono audio mix output of subsystem 201. Therefore, elements 211 , 202, 203, 204, 205, and 207 are a chain from the process side to the LBRIR filter 208.
Whenever the setting of the late reverb portion LBRIR is to be changed, the pulse generator 211 is operated to assert a unit pulse to element 202, and the resulting output from filter bank 207 is captured and asserted to filter 208 ( to set filter 208 to apply the new LBRIR determined by the output of filter bank 207). To speed up the time span from the LBRIR setting change to the time the new LBRIR takes effect, samples of the new LBRIR can begin to replace the old LBRIR as they become available. To shorten the inherent latency of the FDNs, leading zeros can be discarded from the
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
LBRIR. These options provide f 1 ex-ibí UHa.H and allow hybrid implementation to provide potential performance improvement (relative to that provided by an implementation in the filter bank domain), at an aggregate computational cost of FIR filtering .
For applications where system latency is critical, but computing power is of less concern, the chain-side filterbank domain late reverb processor (e.g., implemented by elements 211, 202, 203, 204, ..., 205, and 207 of Figure 8) can be used to capture the effective FIR impulse response to be applied by filter 208. The FIR filter 208 can implement this captured FIR response and apply it directly to the mono audio mix of the input channels (during virtualization of the input channels).
The different FDN parameters and thus the resulting late reverb attributes can be manually tuned and subsequently wired to a mode of the inventive late reverb processing subsystem, for example by means of one or more adjustable presets (e.g. eg, by means of the operation control subsystem 209 of FIG. 3) by the user of the system. However, given the description of high level of late reverb, its relation to the FDN parameters,
IMPI ^
MEXICAN INSTITUTE
OF PROPERTY AND THE ABILITY TO MODIFY ITS BEHAVIOR, S ^ u ^^ eve'-HEa<sup>1 </sup>wide variety of methods for cont ^ aA ^ c .. various FDN-based late reverb processor modes, including (but not limited to) the following:
one. The end user can manually control the FDN parameters, for example through a user interface on a screen (e.g., implemented by one mode of the control subsystem 209 of Figure 3) or by changing presets using physical controls ( eg, implemented by one embodiment of the control subsystem 209 of Figure 3). In this way, the end user can tailor the room simulation according to taste, environment, or content;
2. The author of the audio content to be virtualized can provide desired settings or parameters that are conveyed with the content itself, for example by the metadata provided with the input audio signal. Such metadata can be analyzed and employed (eg, by one embodiment of the control subsystem 209 of FIG. 3) to control the relevant FDN parameters. The meta data can therefore be indicative of properties such as reverberation time, reverb level, direct-to-reverb ratio, among others, and these properties can be variable in time, indicated by metadata
IMPI variables in time;
3. A playback device can be aware of its location or environment, by means of one or more sensors. For example, a mobile device can use GSM networks, Global Positioning System (GPS), known Wi-Fi access points, or any other location service to determine where the device is. Subsequently, data indicative of location and / or environment (eg, by one embodiment, the control subsystem 209 of Figure 3) may be employed to control the relevant FDN parameters. Therefore, the FDN parameters can be modified in response to the location of the device, eg, to simulate the physical environment;
Four. Regarding the location of the playback device, a cloud service or social media can be used to derive the most common settings that consumers are using in a certain environment. Additionally, users can upload their current settings to a cloud service or social media, in association with the (known) location to make them available to other users, or to themselves;
5. A playback device may contain other sensors such as a camera, light sensor, microphone, gyroscope accelerometer, to determine the activity of the <sup>70</sup> IMPIAS
MEXICAN INSTITUTE
K THE INDUSTRIAL PROPERTY user and the environment in which the user is, to optimize the FDN parameters for that activity and / or particular environment;
6. The FDN parameters can be controlled by the audio content. Audio classification algorithms, or manually annotated content can indicate whether the audio segments comprise speech, music, sound effects, silence, and the like. The FDN parameters can be adjusted according to such labels. For example, the direct-to-reverb ratio can be lowered for dialogue to improve dialogue intelligibility. Additionally, video analytics can be used to determine the location of a current video segment, and the FDN parameters can be adjusted accordingly to closely simulate the environment depicted in the video; I
7. A solid state playback system can use different FDN settings like a mobile device,
eg, settings may be device dependent. A solid state system present in a living room can simulate a typical (quite reverberant) living room scenario with distant sources, while a mobile device can render content closer to the listener.
Some inventive virtualizer implementations include FDNs (e.g., an implementation of the FDN of the
<img file="MX352134B_D0035.tif" />
IMPI
MEXICAN INSTITUTE
OF INDUSTRIAL PROPERTY
Figure 4) that are configured to apply fractional delay as well as integer sample delay. For example, in such an implementation a fractional delay element is connected in each reverb tank in series with a delay line that applies integer delay equal to an integer number of sample periods (e.g., each fractional delay element is positioned after or otherwise in series with one of the lag lines). The fractional delay can be approximated by a phase change (complex unit multiplication) in each frequency band that corresponds to a fraction of the sample period: f = τ / Τ, where f is the delay fraction, τ is the delay desired for the band, and T is the sample period for the band. It is well known how to apply fractional delay! in the context of applying reverb in the QMF domain.
In a first class of embodiments, the invention is a headphone virtualization method for generating a binaural signal in response to a set of channels (e.g., each of the channels, or each of the full-range channels of frequencies) of a multichannel audio input signal, which includes the steps of: (a) apply a binaural room impulsive response (BRIR) to each channel of the ensemble (e.g., by convolving each channel of the ensemble with a BRIR corresponding to that channel, in the subsystems
<img file="MX352134B_D0036.tif" />
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL or in subsystems 12,.
IMPI
100 and 200 from Figure 3 to Figure 2), generating from eat'd<sup>1 1</sup> Inner filtered signals (e.g., the outputs of subsystems 100 and 200 of Figure 3, or the outputs of subsystems 12, ..., 14, and 15 of Figure 2), including by using al minus a feedback delay network (e.g., FDNs 203, 204, ..., 205 in Figure 3) to apply a common late reverb to an audio mix (downmix) (e.g., a mix of monophonic audio) of the ensemble channels; and (b) combining the filtered signals (eg, in subsystem 210 of Figure 3, or the subsystem comprising elements 16 and 18 of Figure 2) to generate the binaural signal. Generally, a bank of FDNs is used to apply the common late reverb to the audio mix (eg, with each FDN applying common late reverb to a different frequency band). Generally, step (a) includes a step of applying to each channel in the array a direct response and early reflection portion of a single channel BRIR for the channel (e.g., in subsystem 100 of Figure 3 or subsystems 12, ..., 14 of Figure 2), and the common late reverb has been generated to emulate collective macro attributes of late reverb portions of at least some (e.g. all) single BRIRs. channel.
In typical embodiments in the first class, each of the FDNs is implemented in the mirror filter domain in <sup>73</sup> IMPI
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL Hybrid Complex Quadrature (HCQMF) or Quadrature Mirror Domain Filter (QMF), and in some of such modalities, the frequency-dependent spatial acoustic attributes of the binaural signal are controlled (e.g., using the subsystem 209 of Figure 3) by controlling the setting of each FDN used to apply the late reverb. Generally, a monophonic audio mix of the channels (e.g., the audio mix generated by subsystem 201 of Figure 3) is used as the input to the FDNs for efficient binaural rendering of the audio content of the multichannel signal. . Generally, the audio mixing process is controlled based on a source distance for each channel (this is the distance between a presumed source of the channel's audio content and a presumed user position) and depends on the handling of responses. that correspond to the distances from the source in order to preserve the temporal and level structure of each BRIR (that is, each BRIR determined by the early thoughtless direct response portions of a single channel BRIR for a channel, along with the late reverb common for an audio mix that includes the channel). Although the channels to be mixed can be time aligned and scaled in different ways during audio mixing, the appropriate level and temporal relationship between the response portions must be maintained.<sup>74</sup> IMPI ^
MEXICAN INSTITUTE
FROM THE INDUSTRIAL PROPERTY „Direct, early reflection, and late reverb common to BRIR for each channel. In modes that use a single bank of FDN to generate the late reverb portion common to all channels that are mixed down (to generate an audio mix), appropriate gain and delay (to each channel that is mixed) need to be applied during the mixdown. generation of the audio mix.
Typical modalities in this class include a step of adjusting (e.g., using control subsystem 209 of Figure 3) the FDN coefficients that correspond to the frequency-dependent attributes (e.g., decay time of reverberation, interaural coherence, modal density, and direct-to-late relationship). This enables better matching of acoustic environments and more natural sound outputs.
In a second class of embodiments, the invention is a method of generating a binaural signal in response to a multichannel audio input signal, by applying binaural room impulsive response (BRIR) to each channel (e.g., by convolving each channel with a corresponding BRIR) of a set of input signal channels (e.g., each of the input signal channels or each full-range channel of the input signal), including: process each channel of the set in a first processing path (e.g., implemented by subsystem 100 of
<img file="MX352134B_D0037.tif" />
IMPI
INSTITUTE MEXICANO DE LA moriEBAD INDUSTRIAL Figure 3 or the subsystems 12, of Figure 2) which is configured to model, and apply to each of these channels, a portion of direct response early thoughtlessness (eg, the EBRIR applied by subsystem 12, 14, or 15 of Figure 2) of a single channel BRIR for the channel; and process an audio mix (e.g., a monophonic audio mix) of the channels in the ensemble in a second processing path (e.g., implemented by subsystem 200 of Figure 3 or subsystem 15 of Figure 2), in parallel with the first processing path. The second processing path is configured to model, and apply to the audio mix, a common late reverb (eg, the LBRIR applied by subsystem 15 of Figure 2). Generally, common late reverb emulates collective macro attributes of late reverb portions of at least some (eg all) of the single channel BRIRs. Generally, the second processing path includes at least one FDN (eg, one FDN for each of the multiple frequency bands). Generally, a mono audio mix is used as the input for all reverb tanks in each FDN implemented by the second processing path. Mechanisms (e.g., control subsystem 209 of Figure 3) are generally provided for systematic control of the macro attributes of each FDN in order to better simulate acoustic environments and
<img file="MX352134B_D0038.tif" />
IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL produce more natural-sounding binaural virtualization. Since most such macro attributes are frequency dependent, each FDN is generally implemented in the hybrid complex quadrature mirror filter (HCQMF) domain, the frequency domain, domain, or other domain of the filter bank, if a different FDN is used for each frequency band. A primary benefit of implementing FDNs in a filter bank domain is to allow the application of reverb with frequency-dependent reverb properties. In different embodiments, FDNs are implemented in any of a wide variety of filter bank domains, using any of a variety of filter banks, including, but not limited to, quadrature mirror filters (QMF), impulse response filters. finite (FIR filters), infinite impulse response filters (IIR filters), or crossover filters.
Some modalities in the first class (and the second class) implement one or more of the following characteristics:
one. An implementation of FDN in the domain of the filter bank (e.g., hybrid complex quadrature mirror filter domain) (e.g., the FDN implementation of Figure 4), or implementation of FDN in the domain of the Hybrid filter bank and reverb filter implementation <sub>77</sub>
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL late in the time domain (e.g., the structure described with reference to Figure 8), which generally allows independent adjustment of parameters and / or adjustments of the FDN for each frequency band (what which enables simple and flexible control of frequency-dependent acoustic attributes), for example, by providing the ability to vary reverb tank delays in different bands to change modal density as a function of frequency;
2. The audio mixing process (specific downmixingj, used to generate (from the multi-channel input audio signal) the mixed audio signal (e.g. monophonic audio mixing) processed in the second processing path, depends the distance from the source of each channel and the handling of the direct response in order to maintain the appropriate level and timing relationship between the direct and late responses;
3. An all-pass filter (APF 301 from Figure 4) is applied in the second processing path (e.g., at the input or output of the bank of FDNs) to introduce phase diversity and increased echo density without changing the spectrum. and / or timbre of the resulting reverb;
Four. Fractional delays are implemented in the feedback path of each FDN in a complex, multi-rate valued structure to overcome problems
INSTITUTE MEXICANO related to delays quantified to the resolution reduction factor; _
5. In FDNs, the reverb tank outputs are linearly mixed directly into the binaural channels (e.g., by matrix 312 in Figure 4), using output mixing coefficients that are set based on the desired interaural coherence. in each frequency band. Optionally, the reverb tank mapping to the binaural output channels is alternating across frequency bands to achieve balanced delay between the binaural channels. Optionally as well, normalization factors are applied to the reverb tank outputs to equalize their levels while preserving fractional delay and overall power;
6. Frequency-dependent reverb decay time is controlled (e.g., using control subsystem 209 in Figure 3) by setting appropriate combinations of reverb tank delays and gains in each frequency band to simulate real rooms ;
7. A scale factor (eg, by items 306 and 309 of Figure 4) is applied per frequency band (eg, at any of the input or output of the relevant processing path), to:
control a direct-to-late relationship (DLR) ”IMPI»
MEXICAN INSTITUTE
OF THE PROPERTY VAindustrial sgAarv frequency dependent to suit a real room (a simple model can be used to calculate the required scaling factor based on the target DLR and the reverberation decay time, e.g. ., T<sub>60</sub>) ;
provide low-frequency attenuation to mitigate excess interlace artifacts; and / or applying diffuse field spectral shaping to the FDN responses;
8. Simple parametric models are implemented (e.g., by control subsystem 209 of Figure 3) to control essential frequency-dependent attributes of late reverb, such as reverb decay time, interaural coherence, and / or ratio. direct-evening.
In some modalities (e.g., for applications in which system latency is critical and the delay caused by the analysis and synthesis filter banks is prohibitive), the FDN structures in the domain of the modal filter bank Typical of the inventive system (e.g., the FDN of Figure 4 in each frequency band) are replaced by FDN structures implemented in the time domain (e.g., the FDN 220 of Figure 10, which can be implemented as shown in Figure 9). In modalities in the time domain of the inventive system, the subsystems of the modalities in the domain of the
<img file="MX352134B_D0039.tif" />
IMPI
INSTITUTE MEXICANO filters that apply a gain factor of “enfessa ^ a to the gains of the reverb tank (gi), and pt d? normalization (1/1 g<sub>¿</sub> |) are replaced by time domain filters (and / or gain elements) in order to allow frequency dependent controls. The output mix matrix of a typical filterbank domain implementation (e.g., the output mix matrix 312 in Figure 4) is replaced (in typical time domain modes) by a set time-domain output filters (eg, elements 500503 of the implementation of Figure 11 of element 424 of Figure 9). Unlike the other typical time domain modality filters, the phase response of this set of output filters is generally critical (since energy conservation and interaural coherence could be affected by the phase response) . In some modes in the time domain, the reverb tank delays are varied (e.g., slightly varied) from their values in a domain implementation of the corresponding filter bank (e.g., to avoid sharing filter bank pass as a common factor).
Figure 10 is a block diagram of an embodiment of the inventive headset virtualization system similar to that of Figure 3, except that elements 202<sup>81</sup> IMPI »
MEXICAN INSTITUTE
OF THE PROPERTY
207 of the system of Figure 3 are replaced e ^ Fi '^ si ^ Topic of Figure 10 by a single FDN 220 which' SU iiiipleiiiLiiLa oh the time domain (e.g., the FDN 220 of Figure 10 can be implement as the FDN of Figure 9). In Figure 10, two signals in the time domain (left and right channel) are output from the direct response and early reflection processing subsystem 100, and two signals in the time domain (left and right channel) are output from the subsystem late reverb processing 221. Adder 210 is coupled to the outputs of subsystems 100 and 221. Element 210 is configured to combine (mix) the left channel outputs of subsystems 100 and 221 to generate the left channel, L, of the binaural audio signal output of the virtualizer of Figure 10, and to combine (mix) the right channel outputs of subsystems 100 and 221 to generate the right channel, R, of the binaural audio signal output of the virtualizer of Figure 10. Element 210 may be implemented to simply sum the corresponding left channel sample output from subsystems 100 and 221 to generate the left channel of the binaural output signal, and to simply sum the corresponding right channel sample output from the subsystems 100 and 221 to generate the right channel of the binaural output signal, assuming the level settings and
<img file="MX352134B_D0040.tif" />
IMPI
INSTITUTE MEXICANO ____ · 4- „J 4- · ___ J _ η, DE ί * PROPERTY appropriate time alignments in sub si © treman
221.
In the system of Figure 10, the multichannel audio input signal (which has channels, XJ are directed to, and undergoing processing in, two parallel processing paths: one through the direct response and early reflection processing subsystem 100 the other through the late reverb processing subsystem 221. The system of Figure 10 is configured to apply a BRIRi to each channel, Xi. Each BRIRi can be decomposed into two portions: a forward response and late reflection portion (applied by subsystem 100), and a late reverb portion (applied by subsystem 221). In operation, the direct response and early reflection processing subsystem 100 therefore generates the direct response and early reflections portions of the binaural audio signal that is output from the virtualizer, and the late reverb processing subsystem (reverb generator late) 221 thus generates the late reverb portion of the binaural audio signal that is output from the virtualizer. The outputs of subsystems 100 and 221 are mixed (via subsystem 210) to generate the binaural audio signal, which is generally asserted from subsystem 210 to a rendering system (not shown) in which you experiment
<img file="MX352134B_D0041.tif" />
binaural representation for headphones.
IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL reproduction medium of
The audio mixing subsystem 201 (of the late reverb processing subsystem 221) is configured to mix the audio of the channels of the multichannel input signal into a mono audio mix (which is a time-domain signal), and the FDN 220 is configured to apply the late reverb portion to the mono audio mix.
With reference to Figure 9, we now describe an example of a time domain FDN that can be used as the FDN 220 of the virtualizer of Figure 10. The FDN of Figure 9 includes the input filter 400, which it is coupled to receive a mono audio mix (eg, generated by subsystem 201 of the system of Figure 10) of all channels of a multi-channel audio input signal. The FDN of Figure 9 also includes the all pass filter (APF) 401 (corresponding to the APF 301 of Figure 4) coupled to the output of the filter 400, the input gain element 401A coupled to the output of the filter 401, adding elements 402, 403, 404, and 405 (corresponding to adding elements 302, 303, 304, and 305 of Figure 4) coupled to the outlet of element 401A, and four reverb tanks. Each reverb tank is coupled to the output of a different one of items 402, 403, 404, and <sup>84</sup> IMPI ^^
MEXICAN INSTITUTE
OF THE AVERAGE
405, and includes one of the reverb filters 40 ^ and ~ ^ 406A, 407 and 407A, 408 and 408A, and 409 and 409A, Uña de Ίar ^ ϊñ3as<sup>,</sup>”<sup>,</sup>'<sup>,, </sup>delay 410, 411, 412, and 413 (corresponding to delay lines 307 of Figure 4) coupled thereto, and one of the gain elements 417, 418, 419, and 420 coupled to the output of a of the delay lines.
The unit matrix 415 (corresponding to the unit matrix 308 of Figure 4, and generally implemented to be identical to the matrix 308) is coupled to the outputs of the delay lines 410, 411, 412, and 413. The matrix 415 is configured to assert a feedback output to a second input of each of elements 402, 403, 404, and 405.
When the delay (ni) applied by line 410 is shorter than that (n2) applied by line 411, the delay applied by line 411 is shorter than that (n3) applied by line 412, and the applied delay on line 412 is shorter than that (n4) applied by line 413, the outputs of gain elements 417 and 419 (of the first and third reverb tanks) are asserted to the inputs of the addition element 422, and the outputs of the gain elements 418 and 420 (from the second and fourth reverb tanks) are asserted to the inputs of the addition element 423. The output of element 422 is asserted to an input of IACC and the mix filter 424, and the other exit
<img file="MX352134B_D0042.tif" />
Filtering stage of element 423 asserts itself to the other IACC and mixes 424.
Examples of implementations of gain elements 417-420 and elements 422, 423, and 424 of Figure 9 will be described with reference to a typical implementation of elements 310 and 311 and the output mix matrix 312 of Figure 4. The output mix matrix 312 of Figure 4 (also identified as a Mout matrix) is a 2 x 2 matrix configured to mix the unmixed binaural channels (the outputs of elements 310 and
311, respectively) from initial offset to generate left and right binaural output channels (the left ear, L, and right ear, R signals asserted at the output of matrix 212) that have desired interaural coherence. This initial offset is implemented by elements 310 and 311, each of which combines two reverb tank outputs to generate one of the unmixed binaural channels, with the reverb tank output having the shortest delay asserting to a element 310 input and reverb tank output having the second shortest delay asserting to an element 311 input. Elements 422 and 423 of the embodiment of Figure 9 perform the same type of initial shift (on time domain signals asserted to their inputs) as elements 310 and
311 (in each frequency band) of the mode ofTTa Figure 4 are carried out in the component currents of the filter bank (in the relevant frequency band) affirmed at their inputs.
The unmixed binaural channels (output from elements 310 and 311 of Figure 4, or from elements 422 and 423 of Figure 9), which are close to being uncorrelated because they do not consist of any common reverb, can be mixed (via matrix 312 of Figure 4 or step 424 of Figure 9) to implement an offset pattern that achieves a desired interaural coherence for the left and right binaural output channels. However, because the reverb tank delays are different in each FDN (that is, the FDN in Figure 9, or the FDN implemented for each different frequency band in Figure 4), an unmixed binaural channel ( the output of one of the elements 310 and 311, or 422 and 423) constantly directs the other unmixed binaural channel (the output of the other of the elements 310 and 311, or 422 and 423).
Therefore, in the embodiment of Figure 4, if the combination of the reverb tank delays and the offset pattern is identical across all frequency bands, sound image drift would result. This deviation can be mitigated if the pattern of
IMPI offset alternates through the frequency bands in such a way that the mixed binaural output channels lead and follow each other on alternating frequency bands. For example, if the desired interaural coherence is Coh, where | Coh | 1, the output mix matrix 312 in odd-numbered frequency bands can be implemented to multiply the two inputs asserted to them by a matrix that has the following form:
eos β sin β, where β = arcsin (CoA) / 2 sin β eos β where the 312 output mix matrix in the odd-numbered frequency bands can be implemented to multiply the two inputs asserted to them by a matrix that has the following form:
M<sub>ouhah</sub> = sin β eos β eos β sin β where β = aresin (Coh) / 2.
Alternatively, the aforementioned sound image drift on the binaural output channels can be mitigated by implementing matrix 312 to be identical in the FDNs for all frequency bands, if the channel order of their inputs is changed for the alternating frequency bands (e.g., the output of element 310
<img file="MX352134B_D0043.tif" />
can be asserted to the first input of matrix 312 and the output of element 311 can be asserted to the second input of matrix 312 in odd frequency bands, and the output of element 311 can be asserted to the first input of the matrix 312 and the output of element 310 can be asserted to the second input of matrix 312 in even frequency bands).
In the embodiment of Figure 9 (and other modalities in the time domain of an FDN for the inventive system), it is not trivial to toggle the offset based on frequency to deal with the sound image drift that would otherwise result when the unmixed binaural channel output of item 422 constantly directs (or follows) the unmixed binaural channel output of item 423. This sound image deviation is dealt with in a typical time domain mode of an FDN of the inventive system in a different way than it is generally dealt with in a filter bank domain mode of an FDN of the inventive system. Specifically, in the embodiment of Figure 9 (and some other time-domain embodiments of an FDN of the inventive system), the relative gains of the unmixed binaural channels (e.g., those outputs of elements 422 and 423 of Figure 9) are determined by the gain elements (e.g., elements 417, 418, 419, and 420 of Figure 9) to compensate
IMPI
INSTITUTE MtXICAN · M LA INDUSTRIAL PROPERTY the deviation of sound image that would otherwise result due to the mentioned unbalanced timing. By implementing a gain element (e.g. item 417) to attenuate the signal that came up earlier (which has been shifted to one side, e.g. by item 422) and implementing a gain element (e.g. eg, item 418) to stimulate the next earliest signal (which has been shifted to the other side, eg by item 423), the stereo image is re-centered. Therefore, the reverb tank that includes the gain element 417 applies a first gain to the output of the element 417, and the reverb tank that includes the element 418 applies a second gain (different from the first gain) to the output. of element 418, such that the first gain and the second gain attenuate the first unmixed binaural channel (output of element 422) relative to the second unmixed binaural channel (output of element 423).
More specifically, in a typical implementation of the FDN of Figure 9, the four delay lines 410, 411, 412, and 413 have increasing length, with increasing delay values ni, n2, n3, and n4, respectively. In this implementation, filter 417 applies gain of gi. Therefore, the output of filter 417 is a delayed version of the input to delay line 410 to which it has been applied
9θ IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL ^ * - 55— a gain of gi. Similarly, filter 418 applies a gain of g<sub>2</sub>, filter 419 applies a gain of g<sub>3</sub>, and filter 420 applies a gain of g<sub>4</sub>. Therefore, the output of filter 418 is a delayed version of the input to delay line 411 to which a gain of g has been applied.<sub>2</sub>, and the output of filter 419 is a delayed version of the input to delay line 412 to which a gain of g has been applied<sub>3</sub>, and the output of filter 420 is a delayed version of the input to delay line 413 to which a gain of g has been applied<sub>4</sub>.
In this implementation, the choice of the following gain values may result in an undesirable drift of the output sound image (indicated by the output of the binaural channels of item 424) to the side (that is, to the left or right channel ): gi = 0.5, g<sub>2</sub> = 0.5, g<sub>3</sub> = 0.5, and g<sub>4</sub> = 0.5. According to one embodiment of the invention, the gain values gi, g<sub>2</sub>, g<sub>3</sub> , and g<sub>4</sub> (applied by elements 417, 418, 419, and 420, respectively) are chosen as follows to center the sound image: gi = 0.38, g<sub>2</sub> = 0.6, g3 = 0.5, and g<sub>4</sub> = 0.5. Therefore, the output stereo image is re-centered in accordance with one embodiment of the invention by attenuating the earliest arrival signal (which has been shifted to one side, by element 422 in this example) relative to the second earliest arrival signal (that is, choosing gi <g<sub>3</sub>), and stimulate the second <sup>91</sup> IMPI »Mexican institute
OF PROPERTY V ™
INDUSTRIAL earliest signal (that has been moved to the other side, by element 423 in the example), relative to the last arrival signal (that is, when choosing g<sub>4</sub> <g<sub>2</sub>).
Typical implementations of the FDN in the time domain of Figure 9 have the following differences and similarities to the FDN in the filterbank domain (CQMF domain) of Figure 4:
the same unitary feedback matrix, A (matrix 308 of Figure 4 and matrix 415 of Figure 9);
reverberation tank delays rii (that is, the delays in the CQMF implementation of Figure 4 can be η 64 = 17 * 64T<sub>S</sub> = 1088 * T<sub>s</sub>, n<sub>2</sub> = 21 * 64T<sub>S</sub> = 1344 * T<sub>S</sub>, n<sub>3</sub> = 26 * 64T<sub>S</sub> = 1664 * T<sub>S</sub>, and n<sub>4</sub> = 29 * 64T<sub>S</sub> = 1856 * T<sub>S</sub>, where 1 / T<sub>S</sub> is the sampling rate (1 / T<sub>S</sub> is generally equal to 48K Hz), while the implementation delays in the time domain can be: ηχ = 1089 * T<sub>s</sub>, n<sub>2</sub> = 1345 * T<sub>S</sub>, n<sub>3</sub> = 1663 * T<sub>S</sub>, and n<sub>4</sub> = 185 * T<sub>S</sub>. Note that in typical CQMF implementations there is a practical restriction that each delay is some integer multiple of the duration of a 64-sample block (the sample rate is generally 48K Hz), but in the time domain there is more flexibility in terms of to the choice of each delay and therefore more flexibility in terms of the choice of the delay of each reverb tank);
similar to filter implementations it passes everything (this
IMPI
MEXICAN INSTITUTE OF PROPERTY industrial is, similar implementations of filter 301 in Figure 4 and filter 401 in Figure 9). For example, the all pass filter can be implemented by cascading multiple all pass filters (eg three). For example, each filter passes all in cascade can be of the form
Λ - —— f where g <sup>=</sup> 0.6. ΕΣ 301 all-pass filter from Figure 4 can be implemented by three cascaded all-pass filters with adequate delays from what is shown (e.g., ni = 64 * T<sub>S</sub>, n<sub>2</sub>= 128 * T<sub>S</sub>, and n<sub>3</sub>= 196 * T<sub>S</sub>), while the 401 all-pass filter in Figure 9 (the time-domain all-pass filter) can be implemented by three cascaded all-pass filters with similar delays (e.g., ni = 61 * T<sub>S</sub>, n<sub>2</sub>= 127 * T<sub>S</sub>, and n<sub>3</sub>= 191 * T<sub>S</sub>) .
In some implementations of the FDN in the time domain of Figure 9, the input filter 400 is implemented in such a way that it causes the direct-to-late relationship (DLR) of the BRIR to be applied by means of the system of the Figure 9 to accommodate (at least substantially) a target DLR, and in such a way that the BRIR DLR is applied by means of a virtualizer that includes the system of Figure 9 (e.g., the virtualizer in Figure 10) can be changed by replacing filter 400 (or controlling a setting of filter 400). For example, in some embodiments, filter 400 is implemented as a cascade of filters (e.g., a first filter 400A and a second filter
<img file="MX352134B_D0044.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
400B, coupled as shown in Figure 9A) to implement the target DLR and optionally also to implement the desired DLR control. For example, the filters in the cascade are IIR filters (e.g., filter 400A is a first-order Butterworth high-pass filter (an IIR filter) configured to accommodate target low-frequency characteristics, and filter 400B is a second order low limiting IIR filter configured to suit target high frequency characteristics). As another example, the filters in the cascade are IIR and FIR filters (e.g., the 400A filter is a second-order Butterworth high-pass filter (an IIR filter) configured to accommodate the target low-frequency characteristics, and the filter 400B is a 14-order FIR filter configured to suit target high frequency characteristics). Generally, the forward signal is fixed, and filter 400 modifies the late signal to achieve the target DLR. The all pass filter (APF) 401 is preferably implemented to perform the same function as the APF 301 of Figure 4. That is, to introduce phase diversity and increased echo density to generate more natural sounding FDN output. The APF 401 generally controls the phase response while the input filter 400 controls the amplitude response.
In Figure 9, the filter 406 and the gain element 406A together implement a reverb filter, the filter <sup>94</sup> IMPI ^
MEXICAN INSTITUTE
OF THE PROPERTY O * —Ξ31. _ _<sub>Ίπ</sub> , , <sub>Λ</sub> ,. INDUSTRIAL - J
407 and the gain element 407A together implement another reverb filter, the filter 408 and the gain element 408A together implement another reverb filter, and the filter 409 and the gain element 409A together implement another reverb filter. Each of the filters 406, 407, 408, and 409 of Figure 9 is preferably implemented as a filter with a maximum gain value close to one (unity gain), and each of the gain elements 406A, 407A, 408A , and 409A is configured to apply a decay gain to the output of a corresponding one of filters 406, 407, 408, and 409 which accommodates the desired decay (after the relevant reverb tank delay, n<sub>Y</sub>) .
Specifically, the gain element 406A is configured to apply a decay gain (decaygaini) to the output of filter 406 to cause the output of element 406A to have a gain such that the output of delay line 410 (after delay of reverb tank, ni) has a first target decay gain, gain element 407A is configured to apply a decay gain<sub>2</sub>) to the output of filter 407 to cause the output of element 407A to have a gain such that the output of delay line 411 (after the reverb tank delay, n<sub>2</sub>) has a second target decayed gain, the element of
<img file="MX352134B_D0045.tif" />
IMPI
MEXICAN INSTITUTE
INDUSTRIAL PROPERTY gain 408A is configured to apply a decay gain (decaygain<sub>3</sub>) to the output of filter 408 to cause the output of element 408A to have such a gain that the output of delay line 412 (after the 5 reverb tank delay, n<sub>3</sub>) has a third target decay gain, and the gain element 409A is configured to apply a decay gain<sub>4</sub>) to the output of filter 409 to cause the output of element 409A to have such a gain that output 10 of delay line 413 (after the reverb tank delay, n<sub>4</sub>) has a fourth target decayed gain.
Each of the filters 406, 407, 408, and 409, and each of the elements 406A, 407A, 408A, and 409A of the system of Figure 9 is preferably implemented (with each of the filters 406, 407, 408 , and 409 preferably implemented as an IIR filter, e.g., a limiting filter or a cascade of limiting filters) to achieve a BRIR characteristic target T60 to be applied by a virtualizer including the system of Figure 9 (eg, the virtualizer of Figure 10), where T60 denotes the reverberation decay time (T<sub>6</sub>o). For example, in some embodiments each of the filters 406, 407, 408, and 409 is implemented as a limiting filter (eg, a limiting filter that has Q = 0.3 and a limiting frequency
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL CURRENCY (shelf frequency) of 500 Hz, para. at the characteristic Tñ, 0 „shown in Figure 13, in which T60 has units of seconds) or a cascade of two IIR limiting filters (e.g., having limiting frequencies of 100 Hz and 1000 Hz to achieve the characteristic T60 shown in Figure 14, in which T60 has units of seconds). The shape of each limiter filter is determined to suit the desired changing curve from low frequency to high frequency. When filter 406 is implemented as a limiting filter (or cascade of limiting filters), the reverb filter comprising filter 406 and gain element 406A is also a limiting filter (or cascade of limiting filters). In the same way, when each of the filters 407, 408, and 409 is implemented as a limiting filter (or cascade of limiting filters), each reverb filter comprising the filter 407 (or 408 or 409) and the element of Corresponding gain (407A, 408A, or 409A) is also a limiter filter (or cascade of limit filters).
Figure 9B is an example of filter 406 implemented as a cascade of a first limiter filter 406B and a second limiter filter 406C, coupled as shown in Figure 9B. Each of filters 407, 408, and 409 can be implemented as in the Figure 9B implementation of filter 406.
<img file="MX352134B_D0046.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
In some modalities, the decay gains (decaygaini) applied by items 406A, 407A, 408A, and
409A are determined as follows:
decaygaini = 10<sup>((</sup>-<sup>60</sup>*<<sup>ni / Fs) / T) / 20</sup>>, where i is the reverb tank index (that is, item 406A applies decaygaini, item 407 applies decaygain<sub>2</sub>, and so on), nor is the delay of the 1st reverb tank (e.g., nor is the delay applied by delay line 410), Fs is the sample rate, T is the reverb decay time desired (T<sub>6</sub>o) θη a predetermined low frequency.
Figure 11 is a block diagram of one embodiment of the following elements of Figure 9: elements 422 and 423, and the IACC (interaural cross-correlation coefficient) filtering and mixing stage 424. Element 422 is coupled and configured to sum the outputs of filters 417 and 419 (from Figure 9) and to assert the summed signal to the input of the low limiter filter 500, and element 422 is coupled and configured to sum the outputs of filters 418 and 420 (from Figure 9) and to assert the summed signal at the input of high pass filter 501. The outputs of filters 500 and 501 are summed (mixed) at element 502 to generate the signal from ear output
<img file="MX352134B_D0047.tif" />
IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL binaural left, and the outputs of the fg. ςηη and rm mix at element 502 (the output of filter 500 is subtracted from the output of filter 501) at element 502 to generate the binaural right ear output signal. Elements 502 and 503 mix (add and subtract) the filtered outputs of filters 500 and 501 to generate binaural output signals that achieve (within acceptable precision) the target characteristic IACC. In the embodiment of Figure 11, each of the low limiter filter 500 and high pass filter 501 is generally implemented as a first order IIR filter. In an example in which filters 500 and 501 have such an implementation, the Figure 11 modality of achieving the exemplary characteristic IACC plotted as curve I in Figure 12, which is a good match to the target characteristic IACC plotted as I<sub>T</sub> in Figure 12.
Figure 11A is a graph of the frequency response (Rl) of a typical implementation of the filter 500 of Figure 11, the frequency response (R2) of a typical implementation of the filter 501 of Figure 11, and the response of the filters 500 and 501 connected in parallel. It is apparent from Figure 11A, that the combined response is desirably flat throughout the 100 Hz 10,000 Hz range.
Therefore, in a class of modalities, the invention is a system (eg, that of Figure 10) and method for ”IMPIgfe ·
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL generating a binaural signal (e.g., the output of item 210 in Figure 10) in response to a set of channels of a multichannel audio input signal, including applying a room binaural impulsive response (BRIR) to each channel of the ensemble, thereby generating filtered signals, including by using at least one feedback delay network (FDN) to apply a common late reverb to an audio mix of the channels of the ensemble; and combining the filtered signals to generate the binaural signal. The FDN is implemented in the time domain. In some such embodiments, the time domain FDN (eg, FDN 220 of Figure 10, configured as in Figure 9) includes:
an input filter (e.g., filter 400 in Figure 9) having an input coupled to receive the audio mix, wherein the input filter is configured to generate a first mix of filtered audio in response to the audio mixing;
an all-pass filter (eg, all-pass filter 401 of Figure 9), coupled and configured for a second filtered audio mix in response to the first filtered audio mix;
a reverb application subsystem (e.g., all elements of Figure 9 other than elements 400, 401, and 424), having a first output (e.g., the
<img file="MX352134B_D0048.tif" />
element output 422)
100 and a second
IMPI INSTITUTE MEXICANO Dt LA PROPERTY INDUSTRIAL output (e.g., the output of item 423), where the reverb application subsystem comprises a set of reverb tanks, each of the reverb tanks has a different delay, and wherein the reverb application subsystem is coupled and configured to generate a first unmixed binaural channel and a second unmixed binaural channel in response to the second filtered audio mix, to assert the first unmixed binaural channel on the first output, and assert the second unmixed binaural channel on the second output; and an interaural cross-correlation coefficient (IACC) filtering and mixing step (eg, step 424 of Figure 9, which can be implemented as items 500, 501, 502, and 503 of Figure 11) coupled to the reverb application subsystem and configured to generate a first mixed binaural channel and a second mixed binaural channel in response to the first unmixed binaural channel and a second unmixed binaural channel.
The input filter can be implemented to generate (preferably as a cascade of two filters configured to generate) the first audio mix filtered such that each BRIR has a direct-to-late ratio (DLR) that matches, at least substantially to a target DLR.
<img file="MX352134B_D0049.tif" />
101
IMPI
INSTITUTE MEXICANO DE LA PROPERTY INDUSTRIA!
Each reverb tank can be configured to generate a delayed signal, and can include a reverb filter (e.g., implemented as a limiting filter or a cascade of limiting filters) coupled and configured to apply a gain to a signal that is propagates in each of said reverb tanks, to cause the delayed signal to have a gain that conforms, at least substantially, to a target decayed gain for said delayed signal, in an effort to achieve a characteristic target reverb decay time (e.g., a T<sub>60</sub> characteristic) of each BRIR.
In some embodiments, the first unmixed binaural channel drives the second unmixed binaural channel, the reverb tanks include a first reverb tank (e.g., the reverb tank in Figure 9 which includes delay line 410) configured to generate a delayed first signal having a shorter delay and a second reverb tank (e.g., the reverb tank of Figure 9 including delay line 411) configured to generate a second delayed signal having a second shorter delay, wherein the first reverb tank is configured to apply a first gain to the first delayed signal , the second reverb tank is configured to apply a second gain to the second delayed signal, the second
<img file="MX352134B_D0050.tif" />
102
<img file="MX352134B_D0051.tif" />
gain is different than gain is different than first <sup>FROM</sup> industrial gain, the first gain applying the first gain and the second gain results in attenuation of the first unmixed binaural channel relative to the second unmixed binaural channel. Generally, the first mixed binaural channel and the second mixed binaural channel are indicative of a re-centered stereo image. In some embodiments, the IACC filtering and mixing stage is configured to generate the first mixed binaural channel and the second mixed binaural channel in such a way that said first mixed binaural channel and said second mixed binaural channel have a characteristic IACC character that is conforms at least substantially to a target characteristic IACC.
Aspects of the invention include methods and systems (eg, the system 20 of Figure 2, or the system of Figure 3, or Figure 10) that carry out (or are configured to carry out, or support the performance of) binaural virtualization of audio signals (eg, audio signals whose audio content consists of speaker channels, and / or object-based audio signals).
In some embodiments, the inventive virtualizer is or includes a general purpose processor coupled to receive or to generate input data indicative of a multichannel audio input signal, and programmed with
103
<img file="MX352134B_D0052.tif" />
IMPI
MEXICAN INSTITUTE OF PROPERTY INDUSTRIAL software (or firmware and / or otherwise configured (e.g., in response to control data) to perform any of a variety of operations on the input data, including a mode of the Inventive method Such a general purpose processor would generally be coupled to an input device (eg, a mouse and / or keyboard), a memory, and a display device. For example, the system of Figure 3 (or the system 20 of Figure 2, or the virtualizer system comprising elements 12, ..., 14, 15, 16, and 18 of the system 20) could be implemented in a general purpose processor, with the inputs being audio data indicative of N channels of the audio input signal, and the outputs being audio data indicative of two channels of a binaural audio signal. A digital-to-analog converter (DAC) could operate on the output data to generate analog versions of the binaural signal channels for playback through speakers (e.g., a pair of headphones ).
While specific embodiments of the present invention and applications of the invention have been described herein, it will be apparent to those of skill in the art that many variations of the embodiments and applications described herein are possible without departing from the scope of the invention described and claimed in this document. It should be understood that while certain
<img file="MX352134B_D0053.tif" />
INSTITUTE MEXICANO DE LA PROPERTY INDUSTRIA!
<img file="MX352134B_D0054.tif" />
Forms of the invention have been shown and described, the invention should not be limited to the specific embodiments described and shown or the specific methods described.
<img file="MX352134B_D0055.tif" />
105
IMPI ixsTmrro maid DELAMortrtMP INOUSTRlAt
Contents120
75 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75
159 members in 14 offices
Members159
| Document | Office | Kind | |
|---|---|---|---|
| US2009200581A1 | United States of America | A1 | |
| TW200935601A | Taiwan Province of China | A | |
| US2010271133A1 | United States of America | A1 | |
| WO2010123712A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201044558A | Taiwan Province of China | A | |
| US7863645B2 | United States of America | B2 | |
| US2011063025A1 | United States of America | A1 | |
| US2011068376A1 | United States of America | A1 | |
| US7969243B2 | United States of America | B2 | |
| US2011215871A1 | United States of America | A1 | |
| TWI351100B | Taiwan Province of China | B | |
| EP2422447A1 | European Patent Office (EPO) | A1 | |
| KR20120030379A | Republic of Korea | A | |
| CN102414984A | China | A | |
| US8179197B2 | United States of America | B2 | |
| US8188540B2 | United States of America | B2 | |
| US2012205724A1 | United States of America | A1 | |
| JP2012525003A | Japan | A | |
| US8334178B2 | United States of America | B2 | |
| EP2422447A4 | European Patent Office (EPO) | A4 | |
| US8400222B2 | United States of America | B2 | |
| TWI405333B | Taiwan Province of China | B | |
| US2013248945A1 | United States of America | A1 | |
| KR101335202B1 | Republic of Korea | B1 | |
| JP5386034B2 | Japan | B2 | |
| EP2422447B1 | European Patent Office (EPO) | B1 | |
| US8928410B2 | United States of America | B2 | |
| US2015054038A1 | United States of America | A1 | |
| CN104766887A | China | A | |
| CN104768121A | China | A | |
| EP2892079A2 | European Patent Office (EPO) | A2 | |
| CA2935339A1 | Canada | A1 | |
| CA3043057A1 | Canada | A1 | |
| CA3148563A1 | Canada | A1 | |
| CA3170723A1 | Canada | A1 | |
| CA3226617A1 | Canada | A1 | |
| CA3242311A1 | Canada | A1 | |
| WO2015102920A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2892079A3 | European Patent Office (EPO) | A3 | |
| CN102414984B | China | B | |
| US9240402B2 | United States of America | B2 | |
| US2016111417A1 | United States of America | A1 | |
| AU2014374182A1 | Australia | A1 | |
| KR20160095042A | Republic of Korea | A | |
| CN105874820A | China | A | |
| CN105874820A8 | China | A8 | |
| EP3090573A1 | European Patent Office (EPO) | A1 | |
| US2016345116A1 | United States of America | A1 | |
| MX2016008696A | Mexico | A | |
| JP2017507525A | Japan | A | |
| US9621110B1 | United States of America | B1 | |
| US9627374B2 | United States of America | B2 | |
| EP3163747A1 | European Patent Office (EPO) | A1 | |
| US2017126179A1 | United States of America | A1 | |
| CN106656058A | China | A | |
| US2017187340A1 | United States of America | A1 | |
| BR112016014949A2 | Brazil | A2 | |
| JP6215478B2 | Japan | B2 | |
| MX352134BThis record | Mexico | B | |
| RU2637990C1 | Russian Federation | C1 | |
| CN105874820B | China | B | |
| JP2018014749A | Japan | A | |
| CN107750042A | China | A | |
| CN107770717A | China | A | |
| CN107770718A | China | A | |
| AU2014374182B2 | Australia | B2 | |
| CN107835483A | China | A | |
| AU2018203746A1 | Australia | A1 | |
| KR101870058B1 | Republic of Korea | B1 | |
| KR20180071395A | Republic of Korea | A | |
| EP3402222A1 | European Patent Office (EPO) | A1 | |
| EP3090573B1 | European Patent Office (EPO) | B1 | |
| HK1251757A | Hong Kong, China | A | |
| HK1251757A1 | Hong Kong, China | A1 | |
| RU2017138558A | Russian Federation | A | |
| ES2709248T3 | Spain | T3 | |
| CN104766887B | China | B | |
| MX365162B | Mexico | B | |
| HK1252865A | Hong Kong, China | A | |
| HK1252865A1 | Hong Kong, China | A1 | |
| CA2935339C | Canada | C | |
| US10425763B2 | United States of America | B2 | |
| JP6607895B2 | Japan | B2 | |
| US2019373397A1 | United States of America | A1 | |
| CN107750042B | China | B | |
| CN107770717B | China | B | |
| CN107770718B | China | B | |
| US10555109B2 | United States of America | B2 | |
| JP2020025309A | Japan | A | |
| AU2018203746B2 | Australia | B2 | |
| CN111065041A | China | A | |
| CN111065041A | China | A | |
| AU2020203222A1 | Australia | A1 | |
| KR102124939B1 | Republic of Korea | B1 | |
| KR20200075888A | Republic of Korea | A | |
| CN107835483B | China | B | |
| US2020245094A1 | United States of America | A1 | |
| US10771914B2 | United States of America | B2 | |
| EP3402222B1 | European Patent Office (EPO) | B1 | |
| JP6818841B2 | Japan | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 352134
- Application
- 8696
Titles2
- Spanish
- GENERACION DE AUDIO BINAURAL EN RESPUESTA A AUDIO MULTICANAL UTILIZANDO AL MENOS UNA RED DE RETARDO REALIMENTADA.
- English
- GENERATION OF BINAURAL AUDIO IN RESPONSE TO MULTI-CHANNEL AUDIO USING AT LEAST ONE FEEDBACK DELAY NETWORK.
Classification
- CPC, 10
- G10L19/008
- H04S3/004
- H04S7/306
- H04S7/307
- H04S2400/03
- H04S2400/13
- H04S2420/01
- G10K15/12
- H04S7/30
- H04S2400/01
- IPC, 2
- H04S3 00
- H04S5 00