Scalable compressed audio bit stream and codec using a hierarchical filterbank and multichannel joint coding
Abstract
A method of reconstructing a time domain output audio signal from an encoded bit stream, comprising: receive a scaled bit stream (599) having a predetermined data rate within a given interval as a sequence of frames, each frame containing at least one of the following and at least some of the frames containing all of the following (a) a plurality of quantized tonal components (2407) representing frequency domain contents at different frequency resolutions of an input signal, b) quantized residual time sample components (2403) representing a time domain residual formed from the difference between reconstructed tonal components and an input signal and c) scale factor grids (2404) representing signal energies of a residual signal formed from the difference between reconstructed tonal components and an input signal, wherein the scale factor grids (2404) extend at least partially over a frequency range of the input signal; receiving information (599) for each frame about the position of the quantized components and/or grids within the frequency interval; parsing the frames of the scaled bitstream into the components and grids (600); decoding any tonal component to form transform coefficients (2408); decode any time sample component and any grid (2401-2405); multiplying the time sample components by grid elements to form time domain samples (2406); and applying an inverse hierarchical filter bank (2400) to the transform coefficients (2407) and time domain samples (4002) to reconstruct a time domain output audio signal (614).

Term
Term ended
Projected expiry passed 16 June 2026, 0.3 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
35 claims: 4 independent, 31 dependent
- 1REIVINDICACIONES 1. Un procedimiento de reconstrucción de una señal de audio de salida de dominio de tiempo a partir de un flujo de bits codificado, que comprende:recibir un flujo (599) de bits escalado que tiene una tasa de datos predeterminada dentro de un intervalo dado como una secuencia de tramas, conteniendo cada trama al menos uno de los siguientes y conteniendo al menos algunas de las tramas todos los siguientes (a) una pluralidad de componentes (2407) tonales cuantificados que representan contenidos de dominio de frecuencia a diferentes resoluciones de frecuencia de una señal de entrada, b) componentes (2403) de muestra de tiempo residuales cuantificados que representan un residual de dominio de tiempo formado a partir de la diferencia entre componentes tonales reconstruidos y una señal de entrada y c) cuadrículas de factor de escala (2404) que representan energías de señal de una señal residual formada a partir de la diferencia entre componentes tonales reconstruidos y una señal de entrada, en el que las cuadrículas (2404) de factor de escala se extienden al menos parcialmente en un intervalo de frecuencia de la señal de entrada;recibir información (599) para cada trama acerca de la posición de los componentes cuantificados y/o cuadrículas dentro del intervalo de frecuencia;analizar las tramas del flujo de bits escalado en los componentes y cuadrículas (600);decodificar cualquier componente tonal para formar coeficientes (2408) de transformada;decodificar cualquier componente de muestra de tiempo y cualquier cuadrícula (2401-2405);multiplicar los componentes de muestra de tiempo por elementos de cuadrícula para formar muestras (2406) de dominio tiempo;y aplicar un banco (2400) de filtros jerárquico inverso a los coeficientes (2407) de transformada y muestras (4002) de dominio tiempo para reconstruir una señal (614) de audio de salida de dominio de tiempo.
- 2El procedimiento de la reivindicación 1, en el que las muestras de dominio tiempo se forman mediante, el análisis del flujo de bits en una cuadrícula (2404) de factor de escala G1 y los componentes (2403) de muestra de tiempo;la decodificación y cuantificación inversa de la cuadrícula de factor de escala de cuadrícula G1 para producir una cuadrícula (2405) de factor de escala G0;y la decodificación y cuantificación inversa de los componentes de muestra de tiempo, multiplicando esos valores de muestra de tiempo por valores (2406) de cuadrícula de factor de escala G0 para producir muestra (4002) de tiempo reconstruida.
- 3El procedimiento de la reivindicación 2, en el que la señal es una señal multicanal en la que los canales residuales se han agrupado y codificado, conteniendo también cada una de dichas tramas d) cuadrículas parciales que representan las relaciones de energía de señal de los canales de señal residuales dentro de grupos de canales comprendiendo además:analizar el flujo de bits en las cuadrículas (508) parciales;decodificar y cuantificar (2401) inversamente las cuadrículas parciales;y multiplicar las muestras de tiempo reconstruidas por la cuadrícula (508) parcial aplicada a cada canal secundario en un grupo de canales para producir las muestras de dominio tiempo reconstruidas.
- 4El procedimiento de la reivindicación 1, en el que la señal de entrada es multicanal en la que grupos de componentes tonales que contienen un canal primario y uno o más canales secundarios, conteniendo también cada una de dichas tramas e) una máscara de bits asociada con el canal primario en cada grupo en el que cada bit identifica la presencia de un canal secundario que se ha codificado conjuntamente con el canal primario, analizar el flujo de bits en las máscaras (3602) de bits;decodificar los componentes tonales para el canal primario en cada grupo (601);decodificar los componentes tonales conjuntamente codificados en cada grupo (601);para cada grupo, usar la máscara de bits para reconstruir los componentes tonales para cada uno de dichos canales secundarios a partir de los componente tonales de canal primario y los componentes (601) tonales conjuntamente codificados.
- 5El procedimiento de la reivindicación 4, en el que los componentes tonales de canal secundario se decodifican decodificando la información de diferencia entre las frecuencias primarias y secundarias, decodificándose por entropía y almacenándose amplitudes y fases para cada canal secundario en el que el componente tonal está presente.
- 6El procedimiento de la reivindicación 1, en el que el banco (2400) de filtros jerárquico inverso reconstruye la señal (614) de audio de salida transformando las muestras (4002) de dominio tiempo en coeficientes (2411) de transformada residuales, combinando (2412) los mismos con los coeficientes (2409) de transformada para un conjunto de componentes (2407) tonales a una resolución de frecuencia baja y transformado inversamente (2413) los coeficientes de transformada combinados para formar una señal (2415) de audio de salida parcialmente reconstruida, y repitiendo las etapas en esta señal de audio de salida parcialmente reconstruida con los coeficientes de transformada para otro conjunto de componentes tonales a la siguiente resolución de frecuencia más alta hasta que se reconstruya la señal (614) de audio de salida.
- 7El procedimiento de la reivindicación 6, en el que las muestras de dominio tiempo se representan como subbandas, reconstruyendo dicho banco de filtros jerárquico inverso la señal de audio de salida de dominio de tiempo mediante:a) la formación en ventanas de la señal o señales en cada una de las subbandas de dominio de tiempo de la trama de entrada para formar subbandas (2410) de tiempo de dominio formadas en ventana;b) la aplicación de una transformada de dominio de tiempo a frecuencia a cada una de las subbandas de tiempo de dominio formadas en ventana para formar coeficientes (2411) de transformada;c) la concatenación de los coeficientes de transformada resultantes para formar conjunto o conjuntos más grandes de los coeficientes (2411) de transformada residuales;d) sintetizar los coeficientes de transformada del conjunto de componentes (2409) tonales;e) la combinación de los coeficientes de transformada reconstruidos a partir de los componentes tonales y de dominio de tiempo en un único conjunto de coeficientes (2412) de transformada combinados;f) la aplicación de una transformada inversa a los coeficientes (2413) de transformada combinados, formando en ventanas y añadiendo solapamiento (2414) con la trama anterior para reconstruir una señal de dominio de tiempo parcialmente reconstruida (2415);y g) la aplicación de iteraciones sucesivas de las etapas (a) a (f) en la señal o señales de dominio de tiempo parcialmente reconstruida(s) usando el siguiente conjunto de componentes (2407) tonales hasta que se reconstruya la señal (614) de audio de salida de dominio de tiempo.
- 8El procedimiento de la reivindicación 6, en el que cada trama de entrada contiene Mi muestras de tiempo en cada una de las P subbandas, realizando dicho banco de filtros jerárquico inverso las siguientes etapas:a) en cada subbanda i, almacenar en memoria intermedia y concatenar las Mi muestras anteriores con las Mi muestras actuales para producir 2* Mi nuevas muestras (4004);b) en cada subbanda i, multiplicar las 2* Mi muestras de subbanda por una función de ventana de 2* Mi puntos (4006);c) aplicar una transformada de (2* Mi) puntos a las muestras de subbanda para producir Mi coeficientes de transformada para cada subbanda i (4008);d) concatenar los Mi coeficientes de transformada para cada subbanda i para formar un único conjunto de N/2 coeficientes (4010);e) sintetizar los coeficientes de transformada tonales a partir del conjunto decodificado e inversamente cuantificado de componentes tonales y combinar los mismos con los coeficientes concatenados de la etapa anterior para formar un único conjunto de coeficientes (2407, 2408, 2409, 2412) concatenados combinados;f) aplicar una transformada inversa de N puntos a los coeficientes concatenados combinados para producir N muestras (4012);g) multiplicar cada trama de N muestras por una función de ventana de N muestras para producir N muestras (4014) formadas en ventanas;h) añadir por superposición las muestras (4014) formadas en ventanas resultantes para producir N/2 nuevas muestras de salida en el nivel de subbanda dado como la señal (4016) de audio de salida parcialmente reconstruida;y i) repetir las etapas (a)-(h) en las N/2 nuevas muestras de salida usando el siguiente conjunto de componentes (2407) tonales hasta que se hayan procesado todas las subbandas y las N muestras de tiempo originales sean reconstruidas como la señal (614) de audio de salida.
- 9Un decodificador para reconstrucción de una señal de audio de salida de dominio de tiempo a partir de un flujo de bits codificado, que comprende:un analizador (600) de flujo de bits para analizar cada trama de un flujo de bits escalado en sus componentes de audio, conteniendo cada trama al menos uno de los siguientes y conteniendo al menos algunas de las tramas todos los siguientes (a) una pluralidad de componentes tonales cuantificados que representan contenidos de dominio de frecuencia a diferentes resoluciones de frecuencia de una señal de entrada, b) componentes de muestra de tiempo residuales cuantificados que representan un residual de dominio de tiempo formado a partir de la diferencia entre componentes tonales reconstruidos y una señal de entrada y c) cuadrículas de factor de escala que representan energías de señal de una señal residual formada a partir de la diferencia entre los componentes tonales reconstruidos y una señal de entrada;un decodificador (602) residual para codificar cualquier componente de muestra de tiempo y cualquier cuadrícula para reconstruir muestras de tiempo;un decodificador (601) tonal para decodificar cualquier componente tonal para formar coeficientes de transformada;y un banco (2400) de filtros jerárquico inverso que reconstruye la señal de salida transformando las muestras de tiempo en coeficientes de transformada residuales, combinando los mismos con los coeficientes de transformada para un conjunto de los componentes tonales a una resolución de frecuencia baja y transformado inversamente los coeficientes de transformada combinados para formar una señal de salida parcialmente reconstruida, y repitiendo las etapas en esta señal de salida parcialmente reconstruida con los coeficientes de transformada para otro conjunto de componentes tonales a la siguiente resolución de frecuencia más alta hasta que la señal de audio de salida se reconstruya.
- 10El decodificador de la reivindicación 9, en el que cada trama de entrada contiene Mi muestras de tiempo en cada una de las P subbandas, realizando dicho banco de filtros jerárquico inverso las siguientes etapas:a) en cada subbanda i, almacenar en memoria intermedia y concatenar las Mi muestras anteriores con las Mi muestras actuales para producir 2* Mi nuevas muestras (4004);b) en cada subbanda i, multiplicar las 2* Mi muestras de subbanda por una función (4006) de ventana de 2* Mi puntos;c) aplicar una transformada de (2* Mi) puntos a las muestras de subbanda para producir Mi coeficientes de transformada residuales para cada subbanda i (4008);d) concatenar los Mi coeficientes de transformada residuales para cada subbanda i para formar un único conjunto de N/2 coeficientes (4010);e) sintetizar los coeficientes de transformada tonales a partir del conjunto decodificado e inversamente cuantificado de componentes tonales y combinar los mismos con los coeficientes de transformada residuales concatenados para formar un único conjunto de coeficientes (2407, 2408, 2409, 2412) concatenados combinados;f) aplicar una transformada inversa de N puntos a los coeficientes concatenados combinados para producir N muestras (4012);g) multiplicar cada trama de N muestras por una función de ventana de N muestras para producir N muestras (4014) formadas en ventanas;h) añadir por superposición las muestras (4014) formadas en ventanas resultantes para producir N/2 nuevas muestras de salida en el nivel de subbanda dado como la señal (4016) de salida parcialmente reconstruida;y i) repetir las etapas (a)-(h) en las N/2 nuevas muestras de salida usando el siguiente conjunto de componentes (2407) tonales hasta que se hayan procesado todas las subbandas y las N muestras de tiempo originales sean reconstruidas como la señal (614) de salida.
- 11Un procedimiento de codificación de una señal de audio de entrada para formar un flujo (116) de bits escalable, que comprende:usar un banco (2101a,... 2101e) de filtros jerárquico (HFB) para descomponer una señal (100) de audio de entrada en una representación de tiempo/frecuencia de múltiples resoluciones;extraer componentes tonales en cada iteración del HFB en múltiples resoluciones de frecuencia de la representación (2109) de tiempo/frecuencia;extraer componentes residuales (2117, 2118, 2119) de la representación de tiempo/frecuencia eliminando componentes tonales de la señal de entrada para pasar una señal residual a la siguiente iteración del HFB y extrayendo los componentes residuales (2117, 2118, 2119) de la señal residual final;clasificando los componentes en base a su contribución relativa a la calidad (103, 107, 109) de señal decodificada;cuantificar y codificar los componentes (102, 107, 108);formar un flujo (126) de bits maestro que incluye los componentes (109) cuantificados clasificados, y escalar el flujo (126) de bits maestro eliminando un número suficiente de componentes (115) codificados de menor clasificación para formar el flujo (116) de bits escalado que tiene una tasa de datos menor que o aproximadamente igual a una tasa de datos deseada, en el que los componentes se clasifican y cuantifican con referencia a la misma función de enmascaramiento o diferentes criterios psicoacústicos, y en el que el flujo de bits escalado incluye información que indica la posición de los componentes en el espectro de frecuencia.
- 12El procedimiento de la reivindicación 11, en el que los componentes se clasifican agrupando primero los componentes tonales en al menos un subdominio (903, 904, 905, 906, 907) de frecuencia a diferentes resoluciones de frecuencia y agrupando los componentes residuales en al menos un subdominio (908, 909, 910) residual en diferentes escalas de tiempo y/o resoluciones de frecuencia, clasificando los subdominios en base a su contribución relativa a la calidad de señal decodificada y clasificando los componentes dentro de cada subdominio en base a su contribución relativa a la calidad de señal decodificada.
- 13El procedimiento de la reivindicación 12, que comprende adicionalmente:formar el flujo (126) de bits maestro en el que los subdominios y componentes dentro de cada subdominio se ordenan en base a su clasificación (109), eliminándose dichos componentes de baja clasificación comenzando con el componente de clasificación más baja en el subdominio de clasificación más baja y eliminando componentes en orden hasta que se consiga (115) la tasa de datos deseada .
- 14El procedimiento de la reivindicación 11, en el que el flujo (116) de bits escalado se graba en o se transmite a través de un canal que tiene la tasa de datos deseada como una restricción.
- 15El procedimiento de la reivindicación 14, en el que el flujo (116) de bits escalado es uno de los múltiples flujos de bits escalados y la tasa de datos de cada flujo de bits individual se controla independientemente, con la restricción de que la suma de las tasas de datos individuales no debe exceder una tasa de datos total máxima, controlándose dinámicamente cada una de dichas tasas de datos en tiempo de acuerdo con la calidad de señal decodificada a través de todos los flujos de bits.
- 16El procedimiento de la reivindicación 11, por el que componentes tonales que se eliminan para formar el flujo de bits escalado también se eliminan (2112) de la señal (2114) residual.
- 17El procedimiento de la reivindicación 11, en el que los componentes residuales incluyen componentes (2117) de muestra de tiempo y componentes (2118, 2119) de factor de escala que modifican los componentes de muestra de tiempo en diferentes escalas de tiempo y/o resoluciones de frecuencia.
- 18El procedimiento de la reivindicación 17, en el que los componentes de muestra de tiempo se representan mediante una cuadrícula G (2117) y los componentes de factor de escala comprenden una serie de una o más cuadrículas G0, G1 (2118, 2119) en múltiples escalas de tiempo y resoluciones de frecuencia que se aplican a los componentes de muestra de tiempo dividiendo la cuadrícula G por elementos de cuadrícula de G0, G1 en el plano de tiempo/frecuencia, teniendo cada cuadrícula G0, G1 un número diferente de factores de escala en tiempo y/o frecuencia.
- 19El procedimiento de la reivindicación 17, en el que los factores de escala se codifican (107) aplicando una transformación bidimensional a los componentes de factor de escala y cuantificando los coeficientes de transformada.
- 20El procedimiento de la reivindicación 19, en el que la transformada es una Transformada de Coseno Discreta bidimensional.
- 21Un procedimiento de la reivindicación 11, en el que el HFB descompone la señal de audio de entrada en coeficientes de transformada en niveles de resolución de frecuencia sucesivamente más bajos en iteraciones sucesivas, en el que dichos componentes tonales y residuales se extraen mediante:la extracción de componentes (2109) tonales de los coeficientes de transformada en cada iteración, cuantificando (2110) y almacenando los componentes tonales extraídos en una lista (2106) de tonos;la eliminación de los componentes (2111, 2112) tonales de la señal de audio de entrada para pasar una señal (2114) residual a la siguiente iteración del HFB;y la aplicación de una transformada (2115) inversa final con resolución de frecuencia relativamente menor que la iteración final del HFB a la señal (113) residual para extraer los componentes (2117) residuales.
- 22El procedimiento de la reivindicación 21, que comprende adicionalmente:eliminar algunos de los componentes (114) tonales de la lista de tonos después de la iteración final;y decodificar y cuantificar (104) inversamente localmente los componentes (114) tonales cuantificados eliminados, y combinar (105) los mismos con la señal (111) residual en la iteración final.
- 23El procedimiento de la reivindicación 22, en el que al menos algunos de los componentes tonales relativamente fuertes eliminados de la lista no se decodifican localmente ni se recombinan.
- 24El procedimiento de la reivindicación 21, en el que los componentes tonales en cada resolución de frecuencia se extraen (2109) mediante:la identificación de los componentes tonales deseados a través de aplicación de un modelo perceptual;la selección del más perceptualmente significativo de los coeficientes de transformada;el almacenamiento de parámetros de cada coeficiente de transformada seleccionado como el componente tonal, incluyendo dichos parámetros la amplitud, frecuencia, fase y posición en la trama del correspondiente coeficiente de transformada;y la cuantificación y codificación (2110) de los parámetros para cada componente tonal en la lista de tonos para inserción en el flujo de bits.
- 25El procedimiento de la reivindicación 21, en el que los componentes residuales incluyen componentes de muestra de tiempo representados como una cuadrícula G (2117), la extracción de los componentes residuales comprende además:construir una o más cuadrículas (2118, 2119) de factor de escala de diferentes resoluciones de tiempo/frecuencia, cuyos elementos representan valores de señal máximos o energías de señales en una región de tiempo/frecuencia;dividir los elementos de cuadrícula G de muestra de tiempo mediante correspondientes elementos de las cuadrículas de factor de escala para producir una cuadrícula G (2120) de muestra de tiempo escalada;y cuantificar y codificar la cuadrícula G (2122) de muestra de tiempo escalada y cuadrículas (2121) de factor de escala para inserción en el flujo de bits codificado.
- 26El procedimiento de la reivindicación 11, en el que la señal de audio de entrada se descompone y los componentes tonales y residuales se extraen mediante, (a) el almacenamiento en memoria intermedia de muestras de la señal de audio de entrada en tramas de N muestras (2900);(b) la multiplicación de las N muestras en cada trama por una función de ventana de N muestras (2900);(c) la aplicación de una transformada de N puntos para producir N/2 coeficientes de transformada originales (2902);(d) la extracción de componentes tonales de los N/2 coeficientes (2109) de transformada originales, cuantificando (2110) y almacenando los componentes tonales extraídos en una lista (2106) de tonos;(e) la resta de los componentes tonales cuantificando inversamente (2111) y restando los coeficientes de transformada tonales resultantes de los coeficientes de transformada originales (2112) para proporcionar N/2 coeficientes de transformada residuales;(f) la división de los N/2 coeficientes de transformada residuales en P grupos de M¡ coeficientes (2906), de tal p forma que la suma de los M¡ coeficientes es N/2 ( = N /2 ;) i=1 (g) para cada uno de los P grupos, la aplicación de una transformada inversa de (2* Mi) puntos a los coeficientes de transformada residuales para producir (2* Mi) muestras de subbanda desde cada grupo (2906);(h) en cada subbanda, la multiplicación de las 2* Mi muestras de subbanda por una función de ventana de 2* Mi puntos (2908);(i) en cada subbanda, la superposición con Mi muestras anteriores y la adición de los valores correspondientes para producir Mi nuevas muestras para cada subbanda (2910);(j) la repetición de las etapas (a)-(i) en una o más de las subbandas de Mi nuevas muestras usando tamaños N de transformada sucesivamente más pequeños (2912) hasta que la resolución de tiempo/transformada deseada se logre (29014);y (k) la aplicación de una transformada (2115) inversa final con resolución N de frecuencia relativamente menor a las Mi nuevas muestras para cada salida de subbanda en la iteración final para producir subbandas de muestras de tiempo en una cuadrícula G de subbandas y múltiples muestras de tiempo en cada subbanda.
- 27El procedimiento de la reivindicación 11, en el que la señal de audio de entrada es una señal de audio de entrada multicanal, codificándose juntos cada uno de dichos componentes tonales formando grupos de dichos canales y para cada uno de dichos grupos, seleccionar un canal primario y al menos un canal secundario, que se identifican a través de una máscara (3602) de bits, identificando cada bit la presencia de un canal secundario, cuantificar y codificar el canal primario (102, 108);y cuantificar y codificar la diferencia entre el canal primario y cada canal secundario (102, 108).
- 28El procedimiento de la reivindicación 27, en el que se selecciona un modo de canal conjunto de codificación de cada grupo de canales en base a una métrica que indica qué modo proporciona la menor distorsión percibida para la tasa de datos deseada en la señal de salida decodificada.
- 29El procedimiento de la reivindicación 11, en el que la señal de audio de entrada es una señal multicanal, comprendiendo adicionalmente:restar los componentes tonales extraídos de la señal de audio de entrada para cada canal para formar señales (2109a,...2109e) residuales;formar los canales de la señal residual en grupos determinados por criterios perceptuales y eficiencia de codificación (3702);determinar canales primarios y secundarios para cada grupo de señal residual (3704);calcular una cuadrícula (508) parcial para codificar información espacial relativa entre cada canal primario/secundario que se emparejan en cada grupo (502) de señales residuales;cuantificar y codificar componentes residuales para el canal primario en cada grupo como respectivas cuadrículas G (2110a);cuantificar y codificar la cuadrícula parcial para reducir la tasa (2110a) de datos requerida;e insertar la cuadrícula parcial codificada y la cuadrícula G para cada grupo en el flujo de bits escalado (3706).
- 30El procedimiento de la reivindicación 29, en el que los canales secundarios se construyen a partir de combinaciones lineales de uno o más canales (3704).
- 31Un codificador de flujo de bits escalable para codificar una señal de audio de entrada y formar un flujo de bits escalable, que comprende:un banco (2100) de filtros jerárquicos (HFB) que descompone la señal de audio de entrada en coeficientes (2108) de transformada en niveles de resolución de frecuencia sucesivamente más bajos y de vuelta en muestras (2114) de subbanda de dominio de tiempo en escalas de tiempo sucesivamente más finas en iteraciones sucesivas;un codificador (102) de tono que (a) extrae componentes (2109) tonales de los coeficientes de transformada en cada iteración, cuantifica (2110) y almacena los mismos en una lista (2106) de tonos, (b) elimina los componentes (2111, 2112) tonales de la señal de audio de entrada para pasar una señal (2114b) residual a la siguiente iteración del HFB y (c) clasifica todos los componentes tonales extraídos en base a su contribución relativa a calidad de señal decodificada;un codificador (107) residual que aplica una transformada (2115) inversa final con resolución de frecuencia relativamente menor que la iteración final del HFB (2101e) a la señal (113) residual final para extraer los componentes residuales (2117, 2118, 2119) y clasifica los componentes residuales en base a su contribución relativa a calidad de señal decodificada;en el que los componentes tonales extraídos y componentes residuales se clasifican y cuantifican con referencia a la misma función de enmascaramiento o diferentes criterios psicoacústicos;un formateador (109) de flujo de bits que ensambla los componentes tonales y residuales en una base de trama por trama para formar un flujo (126) de bits maestro;y un escalador (115) que elimina un número suficiente de los componentes codificados de menor clasificación de cada trama del flujo de bits maestro para formar un flujo (116) de bits escalado que tiene una tasa de datos menor que o aproximadamente igual a una tasa de datos deseada, en el que el flujo de bits escalado incluye información que indica la posición de los componentes en el espectro de frecuencia.
- 32El codificador de la reivindicación 31, en el que el codificador de tono agrupa los componentes tonales en subdominios de frecuencia a diferentes resoluciones (903, 904, 905, 906, 907) de frecuencia y clasifica los componentes con cada subdominio, el codificador residual agrupa los componentes residuales en subdominios residuales en diferentes escalas de tiempo y/o resoluciones (908, 909, 910) de frecuencia y clasifica los componentes con cada subdominio, y dicho formateador de flujo de bits clasifica los subdominios en base a su contribución relativa a calidad de señal decodificada.
- 33El codificador de la reivindicación 32, en el que el formateador de flujo de bits ordena los subdominios y los componentes dentro de cada subdominio en base a su clasificación, eliminando dicho escalador (115) dichos componentes de baja clasificación comenzando con el componente de clasificación más baja en el subdominio de clasificación más baja y eliminando componentes en orden hasta que se consiga la tasa de datos deseada.
- 34El codificador de la reivindicación 31, en el que la señal de audio de entrada es una señal de audio de entrada multicanal, codificando conjuntamente dicho codificador de tono cada uno de dichos componentes tonales formando grupos de dichos canales y para cada uno de dichos grupos, seleccionando un canal primario y al menos un canal secundario, que se identifican a través de una máscara (3602) de bits, identificando cada bit la presencia de un canal secundario;cuantificando y codificando el canal primario (102, 108);y cuantificando y codificando la diferencia entre el canal primario y cada canal secundario (102, 108).
- 35El codificador de la reivindicación 31, en el que la señal de entrada es una señal de audio multicanal, dicho codificador residual, formando los canales de la señal residual en grupos determinados por criterios perceptuales y eficiencia de codificación (3702);determinando canales primarios y secundarios para cada uno de dichos grupos de señales residuales (3704);calculando una cuadrícula (508) parcial para codificar información espacial relativa entre cada canal primario/secundario que se emparejan en cada grupo (502) de señales residuales;cuantificando y codificando componentes residuales para el canal primario en cada grupo como respectivas cuadrículas G (2110a);cuantificando y codificando la cuadrícula parcial para reducir la tasa (2110a) de datos requerida;e insertando la cuadrícula parcial codificada y la cuadrícula G para cada grupo en el flujo de bits escalado (3706). 36. El codificador de la reivindicación 31, en el que el codificador residual extrae componentes de muestra de tiempo representados por una cuadrícula G (2117) y una serie de una o más cuadrículas G0, G1 (2118, 2119) de factor de escala en múltiples resoluciones de tiempo y frecuencia que se aplican a los componentes de muestra de tiempo dividiendo la cuadrícula G por elementos de cuadrícula de G0, G1 en el plano (2120) de tiempo/frecuencia, teniendo cada cuadrícula G0, G1 un número diferente de factores de escala en tiempo y/o frecuencia.
Independent claims35
246 paragraphs in 1 section, as filed
DESCRIPTION
Encoding and decoding of scalable audio using a hierarchical filter bank
Background of the invention
Field of the Invention
The present invention relates to the scalable coding of an audio signal and more specifically to methods of performing this data rate scaling in an efficient manner for multichannel audio signals that include hierarchical filtering, joint coding of tonal components and joint coding. channel time domain components in the residual signal.
Description of the related technique
The main objective of an audio compression algorithm is to create a sonically acceptable representation of an input audio signal using as few digital bits as possible. This allows a low data rate version of the input audio signal to be distributed through limited bandwidth transmission channels, such as the Internet, and reduces the amount of storage required to store the input audio signal. for future reproduction. For those applications in which the data capacity of the transmission channel is fixed, and does not vary over time, or the amount, in terms of minutes, of audio that needs to be stored is known in advance and does not increase, procedures for Traditional audio compression set the data rate and therefore the level of audio quality at the time of compression coding. No further reduction in data rate can be made without either recording the original signal at a lower data rate or decompressing the compressed audio signal and then compressing this decompressed signal again at a lower data rate. These procedures are not "scalable" to address problems of variable channel capacity, storage of additional content in a fixed memory or extraction of data streams at variable data rates for different applications.
A technique used to create a bit stream with scalable features, and avoid the limitations described above, encodes the input audio signal as a high data rate data stream composed of subsets of low data rate data stream. These coded low data rate data streams can be extracted from the coded signal and combined to provide an output data stream whose data rate is adjustable over a wide range of data rates. One approach to implementing this concept is to first decode data at the lowest supported data rate, then encode an error between the original signal and a decoded version of this data flow of the lowest data rate. This encoded error is stored and also combined with the data stream of the lowest data rate supported to create a data stream of the second lowest data rate. The error between the original signal and a decoded version of this signal of the second lowest data rate is encoded, stored and added to the data stream of the second lowest data rate to form a data stream of the third lowest data rate. and so on. This procedure is repeated until the sum of the data rates associated with data flows of each of the error signals thus obtained and the data rate of the data stream of the lowest data rate supported is equal to the data stream of the highest data rate to support. The final scalable high data rate data flow consists of the data stream of the lowest data rate and each of the encoded error data streams.
A second technique, normally used to support a small number of different data rates between lower and higher widely spaced data rates, employs the use of more than one compression algorithm to create a "layered" scalable bit stream. The apparatus that performs the scaling operation in a bit stream encoded in this way chooses, depending on the output data rate requirements, which of the multiple data streams transported in the layered data stream to use as the output. of encoded audio. To improve coding efficiency and provide a wider range of scaled data rates, data transported in the lower rate data streams can be used by higher rate data streams to form higher quality and higher rate data streams.
The document DAUDET L ET AL: "Hybrid representations for audiophonic signal encoding", SIGNAL PROCESSING, ELSEVIER SCIENCE PUBLISHERS, vol. 82, No. 11, November 1, 2002, pages 1595-1617, XP004381254, ISSN: 0165-1684, DOI: 10.1016 / S0165-1684 (02) 00304-3, discloses a procedure in which transient, tonal and Stochastics of an audio signal are estimated and encoded using a strategy that includes transformation coding associated with excessive representation spaces. Between representation spaces, a union of local cosine and dyadic waves is used showing good separation properties for structural characteristics of audio signals, their tonal part and transient part, respectively. The separation of these two layers is improved by the use of structured representations. The approach described is not based on any previous segmentation of the audio signal.
US 2002/0176353 A1 describes a method and system for encoding and decoding an input signal in relation to the significantly more relevant aspects of the input signal. More particularly, a two-dimensional transformation is applied to the input signal to produce a matrix of magnitude and a phase matrix that can be inversely quantified by a decoder. A first column of coefficients of the magnitude matrix represents a function of average spectral density of the input signal. Relevant aspects of the mean spectral density function are coded at the beginning of a data packet for later use by a decoder to recreate the input signal, based on an encoding of the magnitude and phase matrices with the rest of the packet of data.
Summary of the invention
The invention provides a method of reconstructing a time domain output audio signal from a data stream encoded with the features of claim 1, a decoder for reconstructing a time domain output audio signal to from a data stream encoded with the features of claim 9, a method of encoding an input audio signal with the features of claim 11 and a scalable bitstream encoder encoding an input audio signal and forming a scalable bitstream with the features of claim 31 . In this context, it is noted that the invention is set forth in the aforementioned independent claims, and that all subsequent occurrences of the word "embodiment or embodiments", if they refer to combinations of characteristics other than those defined by the independent claims, are refer to examples that were originally presented but that do not represent embodiments of the presently claimed invention, in which these examples are still shown for illustration purposes only.
Accordingly, the present invention provides a method of encoding audio input signals to form a master bit stream that can be scaled to form a scaled bit stream having an arbitrarily prescribed data rate and decoding the scaled bit stream. to reconstruct the audio signals.
This is generally achieved by compressing the audio input signals and arranging them to form a master bit stream. The master bit stream includes quantified components that are classified based on their relative contribution to decoded signal quality. The input signal is suitably compressed by separating it into a plurality of tonal and residual components, and then classifying and then quantifying the components. The separation is done properly using a hierarchical filter bank. The components are properly classified and quantified with reference to the same masking function or different psychoacoustic criteria. The components can then be sorted based on their classification to facilitate efficient scaling. The master bit stream is scaled by eliminating a sufficient number of the low-ranking components to form the scaled bit stream that has a scaled data rate less than or approximately equal to a desired data rate. The scaled bitstream includes information that indicates the position of the components in the frequency spectrum. A scaled bit stream is properly decoded using a reverse hierarchical filter bank by arranging the quantized components based on the position formation, ignoring the missing components and decoding the components arranged to produce an output data stream.
In one embodiment, the encoder uses a hierarchical filter bank to decompose the input signal into a time / frequency representation of multiple resolutions. The encoder extracts tonal components at each iteration of the HFB at different frequency resolutions, removes those tonal components from the input signal to pass a residual signal to the next iteration of the HFB and then extracts residual components from the final residual signal. The tonal components are grouped into at least one frequency subdomain by frequency resolution and classified according to their psychoacoustic importance to the quality of the encoded signal. Residual components include time sample components (for example a G grid) and scale factor components (for example G0, G1 grids) that modify the time sample components. The time sample components are grouped into at least one time sample subdomain and classified according to their contribution to the quality of the decoded signal.
In the decoder, the reverse hierarchical filter bank can be used to extract both the tonal components and the residual components within an efficient filter bank structure. All components are quantified inversely and the residual signal is reconstructed by applying the scale factors to the time samples. The frequency samples are reconstructed and added to the reconstructed time samples to produce the output audio signal. Note that the reverse hierarchical filter bank can be used in the decoder regardless of whether the hierarchical filter bank was used during the coding procedure.
In an illustrative embodiment, the tonal components selected in a multichannel audio signal are encoded using differential coding. For each tonal component, a channel is selected as the primary channel. The channel number of the primary channel and its amplitude and phase are stored in the bit stream. A bit mask is stored indicating which of the other channels includes the indicated tonal component and, therefore, should be encoded as secondary channels. The difference between the primary and secondary amplitudes and phases is encoded by entropy and stored for each secondary channel in which the tonal component is present.
In an illustrative embodiment, the time sample and scale factor components that make up the residual signal are encoded using joint channel coding (JCC) extended to multichannel audio. A channel grouping procedure first determines which of the multiple channels can be coded together and all Channels are formed in groups with the last group being possibly incomplete.
Additional objects, features and advantages of the present invention are included in the following description of illustrative embodiments, the description of which should be read with the accompanying drawings. Although these illustrative embodiments pertain to audio data, it will be understood that video, multimedia and other types of data can also be processed in similar ways.
Brief description of the drawings
Figure 1 is a block diagram illustration of a scalable bitstream encoder using a residual coding topology according to the present invention;
Figures 2a and 2b are frequency and time domain representations of a Shmunk window for use with the hierarchical filter bank;
Figure 3 is an illustration of a hierarchical filter bank for the provision of a time / frequency representation of multiple resolutions of an input signal from which both tonal and residual components can be extracted with the present invention;
Figure 4 is a flow chart of the stages associated with the hierarchical filter bank;
Figures 5a to 5c illustrate an advantage of 'overlap-add' window formation;
Figure 6 is a graph of the hierarchical filter bank frequency response;
Figure 7 is a block diagram of an illustrative implementation of a bank of hierarchical analysis filters for use in the encoder;
Figures 8a and 8b are a simplified block diagram of a 3-stage hierarchical filter bank and a more detailed block diagram of a single stage;
Figure 9 is a bit mask for the extension of differential coding of tonal components to multichannel audio;
Figure 10 represents the detailed embodiment of the residual encoder used in an embodiment of the encoder of the present invention;
Figure 11 is a block diagram for joint channel coding for multichannel audio;
Figure 12 schematically represents a scalable frame of data produced by the scalable bitstream encoder of the present invention;
Figure 13 shows the detailed block diagram of an implementation of the decoder used in the present invention;
Figure 14 is an illustration of a reverse hierarchical filter bank for the reconstruction of time series data from both time and frequency sample components in accordance with the present invention;
Figure 15 is a block diagram of an illustrative implementation of a reverse hierarchical filter bank;
Figure 16 is a block diagram of the combination of tonal and residual components using a reverse hierarchical filter bank in the decoder;
Figures 17a and 17b are a simplified block diagram of a 3-stage inverse hierarchical filter bank and a more detailed single-stage block diagram;
Figure 18 is a detailed block diagram of the residual decoder;
Figure 19 is a correlation table of G1;
Figure 20 is a table of correction coefficients of base function synthesis; Y
Figures 21 and 22 are functional block diagrams of the encoder and decoder, respectively, illustrating an application of the multi-resolution time / frequency representation of the hierarchical filter bank in an audio encoder / decoder.
Description of illustrative embodiments
The present invention provides a method for compressing and encoding audio input signals to form a master bit stream that can be scaled to form a scaled bit stream having an arbitrarily prescribed data rate and decoding the scaled bit stream for Rebuild the audio signals. A hierarchical filter bank (HFB) provides a time / frequency representation of multiple resolutions of the input signal from which the encoder can efficiently extract both the tonal and residual components. For multichannel audio, joint coding of tonal components and joint channel coding of residual components in the residual signal is implemented. The components are classified on the basis of their relative contribution to decoded and quantified signal quality with reference to a masking function. The master bit stream is scaled by eliminating a sufficient number of the low-ranking components to form the scaled bit stream that has a scaled data rate less than or approximately equal to a desired data rate. The scaled bit stream is properly decoded using a reverse hierarchical filter bank by arranging the quantized components based on information position, ignoring the missing components and decoding the components arranged to produce an output data stream. In a possible application, the master bit stream is stored and then scaled down at a desired data rate for recording on other media or for transmission through a limited band channel. In another application, in which multiple scaled bit streams are stored in media, the data rate of each stream is independently and dynamically controlled to maximize the perceived quality while satisfying a aggregate data rate restriction on all bit streams.
As used herein the terms "Domain", "subdomain" and "component" describe the hierarchy of scalable elements in the bit stream. Examples will include:
<img file="ES2717606T3_D0001.tif" />
Scalable bitstream encoder with a residual encoding topology
As shown in Figure 1, in an illustrative embodiment a scalable bitstream encoder uses a residual encoding topology to scale the bitstream to an arbitrary data rate by selectively removing components with the lowest core classification (tonal components ) and / or residual components (time sample and scale factor). The encoder uses a hierarchical filter bank to efficiently decompose the input signal into a time / frequency representation of multiple resolutions from which the encoder can efficiently extract the tonal and residual components. The hierarchical filter bank (HFB) described herein for the provision of time / frequency representation of multiple resolutions can be used in many other applications where such a representation of an input signal is desired. A general description of the hierarchical filter bank and its configuration for use in the audio encoder are described below as well as the modified HFB used by the particular audio encoder.
The input signal 100 is applied to both the masking calculator 101 and the multi-order tone extractor 102. The masking calculator 101 analyzes the input signal 100 and identifies a masking level as a function of frequency below which frequencies present in the input signal 101 are not audible to the human ear. The multi-tone tone extractor 102 identifies frequencies present in the input signal 101 using, for example, multiple overlapping FFTs or as an MDCT based hierarchical filter bank is shown, which meet the psychoacoustic criteria that have been defined for tones , select tones according to this criterion, quantify the amplitude, frequency, phase and position components of these selected tones, and place these tones in a list of tones. At each iteration or level, the selected tones are removed from the input signal to pass a residual signal forward. Once completed, all other frequencies that do not meet the tone criteria are extracted from the input signal and emit from multi-tone tone extractor 102, specifically the last stage of the MDCT hierarchical filter bank (256), in the time domain on line 111 as the final residual signal.
Multi-tone tone extractor 102 uses, for example, five orders of overlapping transforms, starting from the largest and working to the smallest, to detect tones through the use of a base function. Size transforms are used: 8192, 4096, 2048, 1024 and 512 respectively, for an audio signal whose sampling rate is 44100 Hz. Other transform sizes could be chosen. Figure 7 graphically shows how the transforms overlap each other. The base function is defined by the equations:
<img file="ES2717606T3_D0002.tif" />
in which:
<img file="ES2717606T3_D0003.tif" />
Tones detected in each transform size are decoded locally using the same decoding procedure as used by the decoder of the present invention, which will be described later. These locally decoded tones are reversed phase and combined with the original input signal through time domain summing to form the residual signal that is passed to the next iteration or level of the HFB.
The masking level of the masking calculator 101 and the tone list of the multi-order tone extractor 102 are inputs to the tone selector 103. The tone selector 103 first classifies the list of tones provided thereto from the multi-order tone extractor 102 by means of relative power over the masking level provided by the masking calculator 101. He then uses an iterative procedure to determine which tonal components will fit in a frame of data encoded in the master bit stream. The amount of space available in a frame for tonal components depends on the predetermined data rate, before scaling, of the encoded master data stream. If the entire frame is assigned for tonal components then no residual coding is performed. In general, some portion of the available data rate is allocated for tonal components with the rest (less overhead) reserved for residual components.
Channel groups are properly selected for multichannel signals and primary / secondary channels identified within each channel group according to a metric such as contribution to perceptual quality. The selected tonal components are preferably stored using differential coding. For stereo audio, the two-bit field indicates the primary and secondary channels. The amplitude / phase and amplitude / differential phase are stored for the primary and secondary channels, respectively. For multichannel audio, the primary channel is stored with its amplitude and phase and a bit mask is stored (See Figure 9) for all secondary channels with amplitude / differential phase for the included secondary channels. The bit mask indicates which other channels are coded together with the primary channel and stored in the bit stream for each tonal component in the primary channel.
During this iterative procedure, some or all of the tonal components that are determined not to fit in a frame can be converted back into the time domain and combined with the residual signal 111. If, for example, the data rate is high enough, then usually all deselected tonal components are recombined. If, however, the data rate is lower, the relatively strong 'deselected' tonal components are adequately excluded from the residual. It has been found that this improves perceptual quality at lower data rates. The deselected tonal components represented by the signal 110 are decoded locally through the local decoder 104 to convert them back into the time domain on line 114 and combined with the residual signal 111 from the multi-order tone extractor 102 in the combiner 105 to form a combined residual signal 113. Note that the signals appearing in 114 and 111 are both time domain signals so that this combination procedure can be easily affected. The combined residual signal 113 is further processed by the residual encoder 107.
The first action performed by the residual encoder 107 is to process the combined residual signal 113 through a bank of filters that subdivides the signal into sub-bands of time domain frequency critically sampled. In a preferred embodiment, when the hierarchical filter bank is used to extract the tonal components, these time sample components can be read directly from the hierarchical filter bank thereby eliminating the need for a second filter bank dedicated to signal processing. residual. In this case, as shown in Figure 21, the combiner 104 operates at the output of the last stage of the hierarchical filter bank (MDCT (256)) to combine the 'deselected' and decoded tonal components 114 with the residual signal 111 before calculating the IMDCT 2106, which produces the subband time samples (see also Figure 7 steps 3906, 3908 and 3910). Additionally, the decomposition, quantification and arrangement of these subbands are carried out in a psychoacoustically relevant order. Residual components (time samples and scale factors) are adequately encoded using joint channel coding in which time samples are represented by a G grid and the scale factors by G0 and G1 grids (See Figure 11) . Joint coding of the residual signal uses partial grids, applied to groups of channels, representing the relationship of signal energies between primary channel and secondary channel groups. Groups are selected (dynamically or statically) through cross correlations or other metrics. More than one channel can be combined and used as a primary channel (for example, primary L + R, secondary C). The use of partial scale factor grids, G0, G1 in time / frequency dimensions is novel as it applies to these multichannel groups and more than one secondary channel can be associated with a given primary channel. Individual grid elements and time samples are sorted by frequency by ranking the lowest frequencies higher. The grids are classified according to the bit rate. The secondary channel information is classified with lower priority than the primary channel information. The code chain generator 108 takes an input of the tone selector 103, on line 120, and residual encoder 107 on line 122, and encodes values from these two inputs using entropy coding well known in the art in flow 124 of data. The bit stream formatter 109 ensures that psychoacoustic elements appear from the tone selector 103 and residual encoder 107, after being encoded through the code string generator 108, in the correct position in the master bit stream 126. The 'classifications' are implicitly included in the master bit stream by reordering the different components.
A scaler 115 removes a sufficient number of the coded components of lower classification from each frame of the master bit stream 126 produced by the encoder to form a scaled bit stream 116 having a data rate less than or approximately equal to a rate of desired data.
Hierarchical Filter Bank
Multi-tone tone extractor 102 preferably uses a 'modified' hierarchical filter bank to provide a time / frequency resolution of multiple resolutions from which both tonal components and residual components can be efficiently extracted. The HFB decomposes the input signal into transform coefficients at successively lower frequency resolutions and back into time domain subband samples at successively finer time scale resolution in each successive iteration. The tonal components generated by the hierarchical filter bank are exactly the same as those generated by multiple superimposed FFTs, however the calculation load is much smaller. The hierarchical filter bank addresses the problem of uneven time / frequency resolution modeling of the human auditory system by simultaneously analyzing the input signal in different time / frequency resolutions in parallel to achieve an almost arbitrary time / frequency decomposition. The hierarchical filter bank makes use of a window formation and superposition-addition stage in the interior transform not found in known decompositions. This stage and the novel design of the window function allow this structure to be placed in an arbitrary tree to achieve the desired decomposition, and could be done in an adaptive way to the signal.
As shown in Figure 21, a single channel encoder 2100 extracts tonal components from the transform coefficients in each iteration 2101a, ... 2101e, quantifies and stores the extracted tonal components in a list 2106 of tones. The joint coding of the tones and residual signals for multichannel signals is analyzed below. In each iteration the time domain input signal (residual signal) is formed in window 2107 and an MDCT of N points is applied 2108 to produce transform coefficients. The tones are extracted 2109 from the transform coefficients, quantify 2110 and added to the list of tones. The selected tonal components are decoded locally 2111 and subtracted 2112 from the transform coefficients before performing the inverse 2113 transform to generate the time domain subband samples that form the residual 2114 signal for the next iteration of the HFB. A final inverse transform 2115 with a frequency resolution relatively lower than the final iteration of the HFB is performed on the combined residual 113 and forms in a window 2116 to extract the residual components 2117 from G. As described above, any 'unselected' tone is decoded 104 locally and combines 105 with the residual signal 111 before the calculation of the final inverse transform. Residual components include time sample components (Grid) and scale factor components (G0, G1) that are extracted from G18 in 2118 and 2119. Grid is recalculated 2120 and G and G1 are recalculated. quantify 2121, 2122. The calculation of the G, G1 and G0 grids is described below. The quantized tones in the tone list, G grid and G1 scale factor grid are all coded and placed in the master bit stream. The elimination of the selected tones of the input signal in each iteration and the calculation of the final inverse transform are the modifications imposed on the HFB by the audio encoder.
A fundamental challenge in audio coding is the modeling of the time / frequency resolution of human perception. Transient signals, such as applause, require high resolution in the time domain, while harmonic signals, such as a horn, require high resolution in the frequency domain to be accurately represented by an encoded data stream. But it is a well known principle that time and frequency resolution are inverse with each other and no single transform can simultaneously convert high precision in both domains. The design of an effective audio codec requires balancing this compensation between time and frequency resolution.
Known solutions to this problem use window switching, adapting the transform size to the transient nature of the input signal (see K. Brandenburg et al., "The ISO-MPEG-Audio Codec: A Generic Standard for Coding of High Quality Digital Audio ", Journal of Audio Engineering Society, Vol. 42, No. 10, October, 1994). This adaptation of the analysis window size introduces additional complexity and requires a detection of transient events in the input signal. To manage algorithmic complexity, prior art window switching procedures usually limit the number of different window sizes to two. The hierarchical filter bank analyzed herein avoids this coarse adjustment to the signal / auditory characteristics by representing / processing the input signal by means of a filter bank that provides multiple time / frequency resolutions in parallel.
There are many filter banks, known as hybrid filter banks, that break down the input signal into a given time / frequency representation. For example, the MPEG Layer 3 algorithm described in ISO / IEC 11172-3 uses a bank of pseudo-quadrature mirror filters followed by an MDCT transform in each subband to provide the desired frequency resolution. In our hierarchical filter bank we use a transform, such as an MDCT, followed by the inverse transform (for example IMDCT) into groups of spectral lines to perform a flexible time / frequency transformation of the input signal.
Unlike the hybrid filter banks, the hierarchical filter bank uses results from two overlapping and consecutive exterior transforms to calculate 'superimposed' internal transforms. With the hierarchical filter bank it is possible to add more than one transform above the first transform. This is also possible with prior art filter banks (for example tree type filter banks), but it is not practical due to the rapid degradation of frequency domain separation with increasing number of levels. The hierarchical filter bank avoids this frequency domain degradation at the cost of some time domain degradation. This degradation of time domain can be controlled, however, through the correct selection of window forms or shapes. With the selection of the correct analysis window, the coefficients of the inner transform can also be made invariable at time shifts equal to the size of the inner transform (not the size of the outermost transform as in conventional approaches).
A suitable window W (x) referred to herein as the "Shmunk Window", for use with the hierarchical filter bank is defined by:
<img file="ES2717606T3_D0004.tif" />
Where x is the sample index of time domain (0 <x <= L), and L is the length of the window in samples.
The frequency response 2603 of the Shmunk window compared to the commonly used Kaiser-Bessel window 2602 is shown in Figure 2a. It can be seen that the two windows are similar in shape but the attenuation of lateral lobes is greater with the proposed window. The time domain response 2604 of the Shmunk window is shown in Figure 2b.
A hierarchical filter bank of general applicability for the provision of a time / frequency decomposition is illustrated in Figures 3 and 4. The HFB would have to be modified as described above for use in the audio codec. In Figure 3, the number on each dashed line represents the number of frequency containers equally spaced on each level (although not all of these containers are calculated). The down arrows represent an MDCT transform of N points resulting in N / 2 subbands. The up arrows represent an IMDCT that takes N / 8 subbands and transforms them to N / 4 time samples within a subband. Each square represents a subband. Each rectangle represents N / 2 subbands. The hierarchical filter bank performs the following stages:
(a) As shown in Figure 5a, samples 2702 of input signal are stored in buffer in frames of N samples 2704, and each frame is multiplied by a window function 2706 of N samples (Figure 5b) to produce N samples 2708 formed in windows (Figure 5c) (step 2900);
(b) As shown in Figure 3, a N-point transform (represented by arrow 2802 down in Figure 3) is applied to samples 2708 formed in windows to produce N / 2 transform coefficients 2804 (step 2902 );
(c) Optionally, ring reduction is applied to one or more of the transform coefficients 2804 by applying a linear combination of one or more adjacent transform coefficients (step 2904);
(d) The N / 2 transform coefficients 2804 are divided into P groups of M¡ coefficients, such that the
sum of the M<sup>i </sup>coefficients is<img file="ES2717606T3_D0005.tif" />
(e) For each of P groups, an inverse transform of (2 * M<sup>i</sup>) points (represented by date 2806 up in Figure 3) is applied to the transform coefficients to produce (2 * M<sup>i</sup>) subband samples from each group (step 2906);
(d) In each subband, the (2 * M<sup>i</sup>) Subband samples are multiplied by a 2706 window function of (2 * M<sup>i</sup>) points (step 2908);
(e) In each subband, the previous My samples overlap and add to corresponding current values to produce My new samples for each subband (step 2910);
(f) N is set equal to the previous Mi and selects new values for P and Mi, and
(g) The above steps are repeated (step 2912) in one or more of the new My subbands using successively smaller transform sizes for N until the desired time / transform resolution is achieved (step 2914). Note, the stages can be iterated in all subbands, only the smallest subbands or any desired combination thereof. If the stages are iterated in all subbands the HFB is uniform, otherwise it is not uniform.
The graph of the frequency response 3300 of an implementation of the filter bank of Figure 3 and described above is shown in Figure 6 in which N = 128, Mi = 16 and P = 4, and the stages are iterated in the Two minor subbands at each stage.
The potential applications for this hierarchical filter bank go beyond audio, for video processing and other types of signals (eg seismic, medical, other time series signals). Video coding and compression have similar requirements for time / frequency decomposition, and the arbitrary nature of the decomposition provided by the hierarchical filter bank can have significant advantages over current state of the art techniques based on Transform of Discrete cosine and particle breakdown. The filter bank can also be applied in the analysis and processing of seismic or mechanical measurements, processing of biomedical signals, analysis and processing of natural or physiological signals, voice and other time series signals. Frequency domain information can be extracted from the transform coefficients produced in each iteration at successively lower frequency resolutions. Likewise, time domain information can be extracted from the time domain subband samples produced in each iteration at successively finer time scales.
Hierarchical filter bank: uniformly separated subbands
Figure 7 shows a block diagram of an illustrative embodiment of the hierarchical filter bank 3900, which implements a uniformly separated subband filter bank. For a uniform filter bank Mi = M = N / (2 * P). The decomposition of the input signal into subband signals 3914 is described as follows:
one. Samples 3902 of input time are formed in a window at N points, 50% frames 3904 superimposed. 2. An MDCT of N points 3906 is made in each frame.
3. The resulting MDCT coefficients are grouped into P groups 3908 of M coefficients in each group.
Four. An IMDCT of (2 * M) dots 3910 is performed in each group to form (2 * M) 3911 subband time samples.
5. The resulting 3911 time samples are formed in a window at (2 * M) points, 50% overlapping frames and superimposed additions (OLA) 3912 to form M time samples in each subband 3914.
In an illustrative implementation, N = 256, P = 32 and M = 4. Note that different transform sizes and subband groupings represented by different choices for N, P and M can also be achieved to achieve a desired time / frequency decomposition.
Hierarchical filter bank: subbands not uniformly separated
Another embodiment of a bank 3000 of hierarchical filters is shown in Figures 8a and 8b. In this embodiment, some of the filter bank stages are incomplete to produce a transform with three different frequency intervals with the transform coefficients representing a different frequency resolution in each interval. The time domain signal is broken down into these transform coefficients using a series of single element filter banks in cascade. The detailed filter bank element can be iterated a number of times to produce a desired time / frequency decomposition. Note that the numbers for buffer sizes, transform sizes and window sizes and the use of the MDCT / IMDCT for the transform are for illustrative purposes only and do not limit the scope of the present invention. Other sizes of buffer, window and transform and other types of transform can also be used. In general, the Mi differ from each other but satisfy the restriction that the sum of the Mi is equal to N / 2.
As shown in Figure 8b, a single filter bank element stores 3022 input samples 3020 in buffer to form buffers of 256 samples 3024, which are formed in window 3026 by multiplying the samples by an advantage function of 256 samples . Samples formed in windows 3028 are transformed through a 25630 MDCT 3030 to form 128 transform coefficients 3032. Of these 128 coefficients, 3034 are selected the 96 highest frequency coefficients for output 3037 and are not further processed. The 32 lowest frequency coefficients are inversely transformed 3042 to produce 64 time domain samples, which are then formed in window 3044 in samples 3046 and add overlays 3048 with the previous output frame to produce 32 output samples 3050.
In the example shown in Figure 8a, the filter bank is composed of a filter bank element 3004 iterated once with an input buffer size of 256 samples followed by a filter bank element 3010 also iterated with an input buffer size of 256 samples. The last stage 3016 represents a single abbreviated filter bank element and is comprised of the buffer stages 3022, window formation 3026 and MDCT 3030 only to emit 128 frequency domain coefficients representing the lowest frequency range 0-1378 Hz.
Therefore, assuming an input 3002 with a sample rate of 44100 Hz, the filter bank shown produces 96 coefficients representing the frequency range 5513 to 22050 Hz at "Output1" 3008, 96 coefficients representing the frequency range 1379 at 5512 Hz at "Output2" 3014, and 128 coefficients representing the frequency range 0 to 1378 Hz at "Output3" 3018,
It should be noted that the use of MDCT / IMDCT for frequency inverse transform / transform is illustrative and other time / frequency transformations may be applied as part of the present invention. Other values for transform sizes are possible and other decompositions are possible with this approach, selectively expanding any branch in the hierarchy described above.
Multichannel joint coding of tonal and residual components
The tone selector 103 in Figure 1 takes as input data from the mask calculator 101 and the tone list from the multi-order tone extractor 102. The tone selector 103 first classifies the list of tones by relative power over the level of masking from the mask calculator 101, forming a sort by psychoacoustic importance. The formula used is provided by:
<img file="ES2717606T3_D0006.tif" />
in which:
Ak = spectral line amplitude
Mi, k = masking level for spectral line of k in the mask subframe of i
l = base function length in terms of mask subframes
The sum is made on the subframes in which the spectral component has a nonzero value.
The tone selector 103 then uses an iterative procedure to determine which tone components of the tone list classified for the same frame will fit in the bit stream. In stereo or multichannel audio signals, in which the amplitude of a tone is approximately the same in more than one channel, only the entire amplitude and phase are stored in the primary channel; the primary channel being the channel with the greatest amplitude for the tonal component. Other channels that have similar tonal characteristics store the difference of the primary channel.
The data for each transform size includes a number of subframes, covering the smallest transform size 2 subframes; the second 4 subframes; the third 8 subframes; the fourth 16 subframes; and the fifth 32 subplots. There are 16 subframes to 1 frame. Tone data is grouped by size of the transform in which the tone information was found. For each transform size, the following tonal component data is quantified, encoded by entropy and placed in the bit stream: subframe position encoded by entropy, spectral position encoded by entropy, quantified amplitude encoded by entropy and quantified phase.
In the case of multichannel audio, for each tonal component, a channel is selected as the primary channel. The determination of which channel should be the primary channel can be set or can be made based on the signal characteristics or perceptual criteria. The channel number of the primary channel and its amplitude and phase are stored in the bit stream. As shown in Figure 9, a 3602 bit mask is stored indicating which of the other channels include the indicated tonal component and, therefore, should be encoded as secondary channels. The difference between the amplitudes and primary and secondary phases is encoded by entropy and stored for each secondary channel in which the tonal component is present. This particular example assumes that there are 7 channels, and the main channel is channel 3. The 3602 bit mask indicates the presence of the tonal component in the secondary channels 1, 4 and 5. No bit is used for the primary channel.
The output 4211 of the multi-tone tone extractor 102 is composed of frames of MDCT coefficients in one or more resolutions. The tone selector 103 determines which tonal components can be retained for insertion into the bitstream output frame by the code chain generator 108, based on their relevance to the decoded signal quality. Those determined tonal components that do not fit the frame are the output 110 to the local decoder 104. The local decoder 104 takes the output 110 of the tone selector 103 and synthesizes all the tonal components by adding each scaled tonal component with 2000 synthesis coefficients of a query table (Figure 20) to produce frames of MDCT coefficients (See Figure 16). These coefficients are added to the output 111 of the multi-order tone extractor 102 in the combiner 105 to produce a residual signal 113 in the MDCT resolution of the last iteration of the hierarchical filter bank.
As shown in Figure 10, the residual signal 113 for each channel is passed to the residual encoder 107 as the MDCT coefficients 3908 of the hierarchical filter bank 3900, before the window formation and overlay-addition stages 3904 and IMDCT 3910 shown in Figure 7. The subsequent stages of IMDCT 3910, window formation and overlay-addition 3912 are performed to produce 32 critically sampled frequency subbands equally spaced 3914 in the time domain for each channel. The 32 subbands, which make up the time sample components, are called the G grid. Note that other embodiments of the hierarchical filter bank could be used in an encoder to implement different time / frequency decompositions than those described above and other transformations could be used to extract tonal components. If a hierarchical filter bank is not used to extract tonal components, another form of filter bank can be used to extract the subbands but at a higher calculation load.
For stereo or multichannel audio, several calculations are made in the channel selection block 501 to determine the primary and secondary channel for coding tonal components, as well as the procedure for coding tonal components (e.g., Left-Right or Middle -Side). As shown in Figure 11, a channel grouping procedure 3702 first determines which of the multiple channels can be coded together and all channels are formed in groups with the last group being possibly incomplete. Clusters are determined by perceptual criteria of a listener and coding efficiency, and groups of channels of combinations of more than two channels can be constructed (for example, a 5-channel signal composed of L, R, Ls, Rs and C channels can be grouped as {L, R}, {Ls, Rs}, {L + R, C}. Channel groups are then sorted as primary and secondary channels. In an illustrative multichannel embodiment, the selection of the primary channel is based on the relative power of the channels on the frame. The following equations define the relative powers:
<img file="ES2717606T3_D0007.tif" />
The grouping mode is also determined as shown in step 3704 of Figure 11. The tonal components can be encoded as a Left-Right or Medium-Lateral representation, or the output of this stage can result in a single primary channel only as It is shown by dashed lines. Representing Left-Right, the channel with the highest power for the subband is considered the primary and a single bit is set in bit stream 3706 for the subband if the right channel is the highest power channel. Medium-Lateral coding is used for a subband if the following condition for the subband is met:
<img file="ES2717606T3_D0008.tif" />
For multichannel signals, the above is done for each group of channels.
For a stereo signal, the grid calculation 502 provides a stereo pan grid in which the stereo pan can be reconstructed hard and applied to the residual signal. The stereo grid is 4 subbands for 4 time intervals, each subband in the stereo grid covers 4 subbands and 32 samples of the output of the 500 bank of filters, starting with frequency bands above 3 kHz. Other grid sizes, frequency covered subbands and time divisions could be chosen. Values in the cells of the stereo grid are the ratio of the power of the given channel to that of the primary channel, for the range of values covered by the cell. The relationship is then quantified to the same table as that used to encode tonal components. For multichannel signals, the previous stereo grid is calculated for each group of channels.
For multichannel signals, the grid calculation 502 provides multiple scale factor grids, one for each group of channels, which are inserted into the bitstream in order of their psychoacoustic importance in the spatial domain. It calculates the ratio of the power of the given channel to the primary channel for each group of 4 subbands per 32 samples. This ratio is quantified below and this quantified value plus logarithm sign of the power ratio is inserted into the bit stream.
The 503 calculation of the scale factor grid calculates the G1 grid, which is placed in the bit stream. The procedure for calculating G1 is now described. G0 is first obtained from G. G0 contains all 32 subbands but only half the time resolution of G. The contents of the cells in G0 are quantified values of the maximum of two neighboring values of a given subband of G. Quantification (referred to in the following equations as Quantify) is performed using the same modified logarithmic quantification table as used to encode the tonal components in the multi-order tone extractor 102. Each cell in G0 is therefore determined by:
<img file="ES2717606T3_D0009.tif" />
in which:
m is the number of subbands
n is the number of columns of G0
G1 is obtained from G0. G1 has 11 superimposed subbands and 1/8 of the time resolution of G0, forming a grid of 11 x 8 dimension. Each cell in G1 is quantified using the same table as the \ used for tonal components and found using the following formula :<img file="ES2717606T3_D0010.tif" />
in which: Wi is a weighting value obtained from Table 1 in Figure 19.
G0 is recalculated from G1 in decoder 506 of local grid. In the time sample quantification block 507, time output samples ("time sample components") are extracted from the hierarchical filter bank (Grid), which pass through the quantification level selection block 504 , scaled by dividing the time sample components by the respective values in the recalculated G0 of the local grid decoder 506 and quantified to the number of quantification levels, as a subband function, determined by block 504 of quantization level selection. These quantified time samples are then placed in the coded data stream together with the quantized G1 grid. In all cases, a model that reflects the psychoacoustic importance of these components is used to determine the priority for the bitstream storage operation.
In a further improvement step to improve coding gain for some signals, grids including G, G1 and partial grids can be further processed by applying a two-dimensional Discrete Cosine Transform (DCT) before quantification and coding. The corresponding inverse DCT is applied in the decoder after inverse quantization to reconstruct the original grids.
Scalable bitstream and scaling mechanism
Typically, each frame of the master bit stream will include (a) a plurality of quantized tonal components representing frequency domain contents at different frequency resolutions of the input signal, b) quantified residual time sample components representing the residual time domain formed from the difference between the reconstructed tonal components and the input signal, and c) scale factor grids representing the signal energies of the residual signal, which extend a frequency range of the input signal. For a multichannel signal each frame can also contain d) partial grids representing the signal energy ratios of the residual signal channels within groups of channels and e) a bit mask for each primary specifying the joint coding of secondary channels for components tonal Normally a portion of the data rate available in each frame is assigned from the tonal components (a) and a portion is assigned for the residual components (b, c). However, in some cases the entire available rate can be assigned to code the tonal components. Alternatively, the entire available rate can be assigned to code the residual components. In extreme cases, only scale factor grids can be encoded, in which case the decoder uses a noise signal to reconstruct an output signal. In most real applications, the scaled bitstream will include at least some frames that contain tonal components and some frames that include scale factor grids.
The structure and order of components located in the master bit stream, as defined by the present invention, provides scalability of wide bit range and fine granularity data flow. It is this structure and order that allows the bit stream to be smoothly scaled by external mechanisms. Figure 12 represents the structure and order of components based on the audio decompression codec of Figure 1 that decomposes the original data stream into a particular set of psychoacoustically relevant components. The scalable bit stream used in this example is composed of a number of data structures from the Resource Exchange File Format, or RIFF, called "fragments," although other data structures can be used. This file format, which is well known to those skilled in the art, allows the identification of the type of data transported by a fragment as well as the amount of data transported by a fragment. Note that any data flow format that carries information regarding the amount and type of data transported in its data flow data structures can be used to practice the present invention.
Figure 12 shows the design of a scalable data rate 900 fragment 900, together with sub-fragments 902, 903, 904, 905, 906, 906, 907, 908, 909, 910 and 912, comprising the psychoacoustic data that is they transport within the 900 fragment of plot. Although Figure 12 only represents fragment ID and fragment length for the frame item, sub- fragment ID and sub- fragment length data are included within each sub- fragment . Figure 12 shows the order of subfragments in a scalable bitstream frame. These subfragments contain the psychoacoustic components produced by the scalable bitstream encoder, with a unique subfragment used for each subdomain of the encoded data stream. In addition to the subfragments that are arranged in psychoacoustic importance, whether by a priori decision or calculation, the components within the subfragments are also arranged in psychoacoustic importance. The 911 fragment null, which is the last fragment in the plot, is used to package fragments in the case where the plot is required to have a constant or specific size. Therefore fragment 911 has no psychoacoustic relevance and is the least important psychoacoustic fragment. Fragment 910 of time samples 2 appears on the right side of the figure and the most important psychoacoustic fragment, fragment 902 of grid 1 appears on the left side of the figure. Operating to first remove data from the less relevant psychoacoustically fragment at the end of the bit stream, fragment 910, and working towards the elimination of increasingly relevant psychoacoustically components toward the beginning of the bit stream, fragment 902, the largest possible quality for each successive reduction in data rate. It should be noted that the highest data rate, together with the highest audio quality, capable of being supported by the bit stream, is defined in coding time. However, the lowest data rate after scaling is defined by the level of audio quality that is acceptable for use through an application or through the rate restriction located on the channel or media.
Each removed psychoacoustic component does not use the same number of bits. The scaling resolution for the current implementation of the present invention ranges from 1 bit for components of the least psychoacoustic importance to 32 bits for those components of the greatest psychoacoustic importance. The bitstream scaling mechanism does not need to delete entire fragments every time. As mentioned above, components within each fragment are arranged so that the most psychoacoustically important data is placed at the beginning of the fragment. For this reason, components at the end of the fragment, one component at a time, can be removed by a scaling mechanism while maintaining the best possible audio quality with each component removed. In one embodiment of the present invention, entire components are removed by the scaling mechanism, while in other embodiments, some or all of the components can be removed. The scaling mechanism removes components within a fragment as required, by updating the fragment length field of the particular fragment from which the components were removed, the length 915 of frame fragment and the sum 901 of frame control. As will be seen from the detailed description of the illustrative embodiments of the present invention, with updated fragment length for each scaled fragment, as well as updated frame fragment length and raster control sum information available to the decoder, the decoder can properly process the scaled bit stream, and automatically produce a fixed sample rate audio output signal for distribution to the DAC, even though there are fragments within the bit stream that are missing components, as well as fragments that are completely missing from the bit stream.
Scalable bitstream decoder for a residual coding topology
Figure 13 shows the block diagram for the decoder. The bit stream analyzer 600 reads initial secondary information consisting of: the hertz sample rate of the encoded signal before encoding, the number of audio channels, the original data rate of the stream and the encoded data rate. This initial secondary information allows it to reconstruct the full data rate of the original signal. Additionally, bit stream analyzer 600 analyzes components in data stream 599 and passes to the appropriate decoding element: tone decoder 601 or residual decoder 602. Components decoded through the tone decoder 601 are processed through the reverse frequency transform 604 that converts the signal back to the time domain. The overlay-add block 608 adds the values of the last half of the previously decoded frame to the values of the first half of the newly decoded frame that is the output of the reverse frequency transform 604. Components that the bit stream analyzer 600 determines to be part of the residual decoding process are processed through the residual decoder 602. The output of the residual decoder 602, which contains 32 frequency subbands represented in the time domain, is processed through the reverse filter bank 605. The reverse filter bank 605 recombines the 32 subbands in a signal to be combined with the add-over output 608 in the combiner 607. The output of the combiner 607 is the decoded output signal 614.
To reduce the calculation load, the inverse frequency transform 604 and the inverse filter bank 605 that convert the signals back to the time domain can be implemented with a reverse hierarchical filter bank, which integrates these operations with the combiner 607 to form the 614 decoded time domain output signal. The use of the hierarchical filter bank in the decoder is novel in the way in which the tonal components are combined with the residual in the hierarchical filter bank in the decoder. The residual signals are transformed forward using MDCT in each subband, and then the tonal components are reconstructed and combined before the last stage of IMDCT. The multi-resolution approach could be generalized for other applications (for example multiple levels, different decompositions would still be covered by this aspect of the invention).
Inverse hierarchical filter bank
To reduce the complexity of the decoder, the hierarchical filter bank can be used to combine the steps of the reverse frequency transform 604, reverse filter bank 605, add-overlay 608 and combiner 607. As shown in Figure 15, the output of the residual decoder 602 is passed to the first stage of the inverse hierarchy filter bank 4000 while the output of the tone decoder 601 is added to the residual samples in the higher frequency resolution stage before of the final reverse 4010 transform. The samples The resulting inverse transforms are added superimposed below to produce the linear output samples 4016.
Figure 22 shows the general operation of the decoder for a single channel using the HFB 2400. Additional steps for multichannel decoding of the tones and residual signals are shown in Figures 10, 11 and 18. The quantized grids G1 and G 'are read from the 599 bit stream by the data stream analyzer 600. The residual decoder 602 inversely quantifies (Q-1) 2401, 2402 the G'2403 and G1 2404 grids and reconstructs the G0 2405 grid from the G1 grid. The G0 grid is applied to the G 'grid by multiplying 2406 corresponding elements in each grid to form the scaled G grid, which consists of samples 4002 of time subband that are introduced in the next stage in the hierarchical filter bank 2401. For a multichannel signal, partial grid 508 will be used to decode the secondary channels.
The tonal components (T5) 2407 at the lowest frequency resolution (P = 16, M = 256) are read from the bit stream by the data stream analyzer 600. The tone decoder 601 inversely quantifies 2408 and synthesizes 2409 the tonal component to produce P groups of M frequency domain coefficients.
Grid time samples 4002 are formed in a window and add overlays 2410 as shown in Figure 15, transformed forward by P MDCT of (2 * M) points 2411 to form P groups of M frequency domain coefficients which are then combined 2412 with the P groups of M frequency domain coefficients synthesized from the tonal components as shown in Figure 16. The combined frequency domain coefficients are then concatenated and inversely transformed by an IMDCT of length N 2413, formed in windows and added overlapping 2414 to produce N output samples 2415 that are introduced in the next stage of the hierarchical filter bank.
The following tonal components (T4) of lower frequency resolution are read from the bit stream, and combined with the output of the previous stage of the hierarchical filter bank as described above, and then this continuous iteration P = 8, 4, 2, 1 and M = 512, 1024, 2048 and 4096 until all frequency components have been read from the bit stream, combined and rebuilt.
In the final stage of the decoder, the reverse transform produces N samples of full time bandwidth that are output as the decoded output 614. The above values of P, M and N are for an illustrative embodiment only and do not limit the scope of the present invention. Other sizes of buffer, window and transform and other types of transform can also be used.
As described, the decoder anticipates the reception of a frame that includes tonal components, time sample components and scale factor grids. However, if one or more of these of the scaled bit stream is missing, the decoder rebuilds the decoded output without problems. For example, if the frame includes only tonal components then the time samples in 4002 are zero and no residual 2403 is combined with the tonal components synthesized in the first stage of the reverse HFB. If one or more of the tonal components T5, ... T1 is missing, then 2403 combines a zero value in that iteration. If the frame includes only the scale factor grids, then the decoder replaces a noise signal for the G grid to decode the output signal. As a result, the decoder can seamlessly reconstruct the decoded output signal since the composition of each frame of the scaled bit stream can change due to the content of the signal, changing data rate restrictions, etc.
Figure 16 shows in more detail how the tonal components are combined within the reverse hierarchical filter bank of Figure 15. In this case, the residual subband signals 4004 are formed in a window and added overlapping 4006, transform 4008 forward and the resulting coefficients of all subbands are grouped to form the only plot of coefficients 4010. Each tonal coefficient is then combined with the residual coefficient plot by multiplying 4106 the envelope of amplitude tonal component 4102 by a group of synthesis coefficients 4104 (normally provided by a query table) and adding the results to the coefficients centered around the frequency 4106 of the given tonal component. The addition of these tonal synthesis coefficients is performed on the spectral lines of the same frequency region for the entire length of the tonal component. After all the tonal components are added in this way, the final IMDCT 4012 is performed and the results are formed in a window and added overlays 4014 with the previous frame to produce the output time samples 4016.
Figure 14 shows the general shape of the inverse hierarchical filter bank 2850 that is compatible with the hierarchical filter bank shown in Figure 3. Each input frame contains My time samples in each of P subbands, so that The sum of the My coefficients is N / 2:
p
'MS i = N / 2;
í = i
In Figure 14, the upward arrows represent an IMDCT transform of N points that takes N / 2 MDCT coefficients and transforms them to N time domain samples. Down arrows they represent an MDCT that takes N / 4 samples within a subband and transforms them to N / 8 MDCT coefficients. Each square represents a subband. Each rectangle represents N / 2 MDCT coefficients. The following stages are shown in Figure 14:
(a) In each subband, the previous My samples are stored in memory and concatenated with the current My samples to produce (2 * Mi) new samples for each subband 2828;
(b) In each subband, the (2 * Mi) subband samples are multiplied by a 2706 window function of (2 * Mi) dots (Figure 5a-5c);
(c) A transform of (2 * Mi) points (represented by down arrow 2826) is applied to produce My transform coefficients for each subband;
(d) The Mi transform coefficients for each subband are concatenated to form a single group 2824 of N / 2 coefficients;
(e) An inverse transform of N points (represented by date 2822 upwards) is applied to the concatenated coefficients to produce N samples;
(f) Each frame of N samples 2704 is multiplied by a window function of N samples 2706 to produce N samples 2708 formed in windows;
(g) Samples 2708 formed in resulting windows are added superimposed to produce N / 2 new output samples at the given subband level;
(h) The above stages are repeated at the current level and all subsequent levels until all subbands have been processed and the original 2840 time samples are reconstructed.
Inverse hierarchical filter bank: uniformly separated subbands
Figure 15 shows a block diagram of an illustrative embodiment of a reverse hierarchical filter bank 4000 compatible with the forward filter bank shown in Figure 7. The synthesis of the decoded output signal 4016 is described in more detail as Indicate below:
one. Each input frame 4002 contains M time samples in each of P subbands.
2. Store each subband 4004 in buffer memory, move new samples in M, apply dotted window (2 * M), 50% overlay-add (OLA) 4006 to produce M new subband samples.
3. An MDCT of (2 * M) points 4008 performed within each subband to form M MDCT coefficients in each of P subbands.
Four. The resulting MDCT coefficients are grouped to form a single frame 4010 of (N / 2) MDCT coefficients.
5. An IMDCT 4012 of N points made in each frame
6. The IMDCT output is formed in a window at N points, 50% overlapping frames and overlapping additions 4014 to form N / 2 new output samples 4016.
In an illustrative implementation, N = 256, P = 32 and M = 4. Note that different transform sizes and subband groupings represented by different choices for N, P and M can also be achieved to achieve a desired time / frequency decomposition.
Inverse hierarchical filter bank: subbands spaced unevenly
Another embodiment of the reverse hierarchical filter bank is shown in Figure 17a-b, which is compatible with the filter bank shown in Figure 8a-b. In this embodiment, some of the detailed filter bank elements are incomplete to produce a transform with three different frequency intervals with the transform coefficients representing a different frequency resolution in each interval. The reconstruction of the time domain signal from these transform coefficients is described as follows:
In this case, the first synthesis element 3110 omits the buffer stages 3122, window formation 3124 and the MDCT 3126 of the detailed element shown in Figure 17b. Instead, output 3102 forms a single set of coefficients that are inversely transformed 3130 to produce 256 time samples, which are formed in window 3132 and add overlays 3134 with the previous frame to produce output 3136 of 128 new time samples For this stage.
The output of the first element 3110 and 96 coefficients 3106 are introduced to the second element 3112 and combined as shown in Figure 17b to produce 128 time samples to introduce the third element 3114 of the filter bank. The second element 3112 and third element 3114 in Figure 17a implement the entire detailed element of Figure 17b, cascaded to produce 128 new time samples emitted from the filter bank 3116. Note that the buffer memory and transform sizes are provided as examples only and other sizes may be used. In particular, note that buffer storage 3122 at the input to the item detailed may change to accommodate different input sizes depending on where it is used in the hierarchy of the general filter bank.
Additional details regarding the decoder blocks will now be described.
600 data flow analyzer
The bit stream analyzer 600 reads IFF fragment information from the bit stream and passes elements of that information into the appropriate decoder, tone decoder 601 or residual decoder 602. It is possible that the bit stream may have escalated before reaching the decoder. Depending on the scaling procedure used, psychoacoustic data elements at the end of a fragment may not be valid due to missing bits. The tone decoder 601 and residual decoder 602 appropriately ignore data that is found to be invalid at the end of a fragment. An alternative to the tone decoder 601 and residual decoder 602 that ignores complete psychoacoustic data elements, when the element's bits are missing, is that these decoders recover as much of the element as possible by reading in the bits that come out and fill in the remaining missing bits with zeros. , random patterns or patterns based on previous psychoacoustic data elements. Although it is more intensive in calculation, the use of data based on previous psychoacoustic data elements is preferred because the resulting decoded audio may coincide more closely with the original audio signal.
601 tone decoder
Tone information found by the bit stream analyzer 600 is processed through the tone decoder 601. The resynthesis of tonal components is performed using the hierarchical filter bank as described above. Alternatively, a Fast Inverse Fourier Transform whose size is the same size as the smallest transform size that was used to extract the tonal components in the encoder can be used.
The following stages are performed for tonal decoding:
a) Initialize the domain of the subframe frequency with zero values.
b) Resynthesize the required portion of tonal components from the smallest transform size in the subframe frequency domain.
c) Resynthesize and add in the required positions, tonal components of the other four transform sizes in the same subframe. Resisthesis of these other four transform sizes can occur in any order.
The tone decoder 601 decodes the following values for each transform size grouping: quantized amplitude, quantized phase, spectral distance from the previous tonal component for grouping and the position of the component within the entire frame. For multichannel signals, the secondary information is stored as differences from the primary channel values and needs to be restored to absolute values by adding the values obtained from the bit stream to the value obtained for the primary channel. For multichannel signals, 'presence' per channel of the tonal component is also provided by the 3602 bit mask that is decoded from the bit stream. Additional processing is done on secondary channels regardless of the primary channel. If the tone decoder 601 is not able to fully acquire the elements necessary to reconstruct a tone from the fragment, that tonal element is discarded. The quantized amplitude is quantified using the inverse of the table used to quantify the value in the encoder. The quantified phase is quantified using the inverse of the linear quantification used to quantify the phase in the encoder. The absolute frequency spectral position is determined by adding the difference value obtained from the bit stream to the previously decoded value. Define Amplitude to be the quantized amplitude, Phase to be the quantified phase, and Frec to be the absolute frequency position, the following pseudocode describes the resynthesis of tonal components of the smallest transform size:
<b>Re [Frec] Amplitude * sin (2 * Pi * Phase / «?);</b>
<b>ImfFrec] =. Amplitude * cos (2 * Pi * Phase / 8);</b>
<b>Re [Freq 1] = Amplitude * sin (2 * Pi * Phase / 8);</b>
<b>Im [Freq 1] 4- = Amplitude * cos (2 * Pi * Phase / S);</b>
Resynthesis of longer base functions extend over more subframes therefore the amplitude and phase values necessary to be updated according to the frequency and length of the base function. The following pseudo code describes how this is done:
xFrec = Frec »(Group - 1);
CurrentPhase = Phase 2 * (2 * xFrec I);
for (i = 0; i <length; i = i 1)
{
CurrentPhase = 2 * (2 * Freq 1) / length .;
CurrentAmplitude = Amplitude, * Envolveute [Group] [i];
Re [i] [xFrec] 4 = CurreníAmplitude * sin (2 * Pi * CurrentPhase / 8);
Im [i] [xFrec] = CurreníAmplitude * cos (2 * Pi * CurrentPhase / 8);
Re [i] [xFrec + I] + = CurrentAmplitude <sup>* </sup>sin (2 * Pi * CurrentPhase / 8);
Im [i} [xFrec + l] = CurrentAmplitude * eos (2 * Pi <sup>* </sup>CurrentPhase / 8);
1
in which:
Amplitude, Frec and Phase are the same as defined above.
Group is a number that represents the base transform size function, 1 for the smallest transform and 5 for the longest.
Length is the subframes for the Group and is provided by:
Length = 2 A (Group -1).
>> is the right operator of displacement.
CurrentAmplitude and CurrentFase are stored for the next subframe. Envelope [Group] [i] is the triangular shaped envelope of appropriate length (length) for each group, having zero value at each end and having a value of 1 in the middle.
Resynthesis of lower frequencies in the three largest transform sizes through the procedure described above, causes audible distortion in the audio output, therefore the following empirically based correction is applied to spectral lines less than 60 in groups 3, 4 and 5:
xFrec = Frec >> (Group -1);
in which:
CurrentPhase = Phase - 2 * (2 * xFreq 1);
f_dlt = Frec - (xFrec «(Group - 1));
for (i = 0; i <length; i = i 1)
{
CurrentPhase = 2 * (2 * Freq 1) / length; .
CurreníAmplitude = Amplitude * Envolveute [Giupo] [i];
Re_Amp = CurrentAmplitude * sin (2 * Pi * CurrentPhase / 8);
Im ^ Amp - CurrentAmplitude * cos (2 <sup>* </sup>Pi * CurrentPhase / 8);
aO = Re_Amp * CorrCf [f_dlt] [Q];
bO = Im_Amp * CorrCf (f_dft] [0];
al ~ Re_Amp 4 CorrCf [f_dlt] [l];
bl = Im_Amp * CorrCf [f_dlt] [I];
a2 = Re_Amp * ConCf [f_dIt] [2];
b2 = Im_Amp * CorrCf {f_dlt] [2];
a3 - Re_Amp 4 CorrCf [f_dlt] [3];
b3 = Tra_Amp * CorrCf [f_dlt] [3];
a4 = Re_Amp * CorrCf [f_dlt] [4];
b4 = Im_Amp 4 CqrrCf [f_dlt] [4];
<img file="ES2717606T3_D0011.tif" />
Amplitude, Freq, Phase, Envelope [Group] [i], Group, and
Length are all as defined above.
CorrCf is provided by Table 2 (Figure 20).
abs (val) is a function that returns the absolute value of val
Since the bitstream does not contain any information such as the number of coded tonal components, the decoder only reads tone data for each transform size until data for that size is depleted. Therefore, the tonal components removed from the bit stream by external means have no effect on the ability of the decoder to process data still contained in the bit stream. Deleting elements from the bitstream only degrades audio quality by the amount of the data component removed. Tonal fragments can also be removed, in which case the decoder does not perform any tonal component reconstruction work for that transform size.
604 reverse frequency transform
The inverse frequency transform 604 is the inverse of the transform used to create the frequency representation domain in the encoder. The current embodiment employs the reverse hierarchical filter bank described above. Alternatively, a Fast Inverse Fourier Transform that is the inverse of the smallest FFT used to extract tones by the superimposed FFT provided encoder was used at coding time.
602 residual decoder
A detailed block diagram of the residual decoder 602 is shown in Figure 18. The data flow analyzer 600 passes elements of G1 from the bit stream to the inline grid decoder 702 610. The grid decoder 702 decodes G1 to recreate G0 which is 32 frequency subbands for 64 time slots. The bitstream contains quantified G1 values and the distances between those values. G1 values of the bit stream are quantified using the same quantification table as used to quantify tonal component amplitudes. Linear interpolation between bitstream values leads to 8 final G1 amplitudes for each subband of G1. Subbands 0 and 1 of G1 are initialized to zero, zero values being replaced when subband information is found for these two subbands in the bit stream. These amplitudes are then weighted in the G0 grid recreated using the correlation weights 1900 obtained from Table 1 in Figure 19. A general formula G0 is provided by:
<img file="ES2717606T3_D0012.tif" />
in which:
m is the number of subbands
W is the entry in Table 1
n is the number of columns of G0
k extends through 11 subbands of G1
700 quantizer
Time samples found by the data flow analyzer 600 are quantified in the quantizer 700. The quantizer 700 quantifies time samples of the bit stream using the reverse encoder procedure. Zero subband time samples are quantified at 16 levels, subbands 1 and 2 at 8 levels, subbands 11 to 25 at three levels and subbands 26 to 31 at 2 levels. Any missing or invalid time samples are replaced with a pseudorandom sequence of values in the range of -1 to 1 that has a spectral energy distribution of white noise. This improves the quality of scaled bitstream audio since a sequence of values of this type has characteristics that resemble the original signal more closely than the replacement with zero values.
701 channel demultiplexer
Secondary channel information in the bit stream is stored as the difference of the primary channel for some subbands, depending on flags set in the bit stream. For these subbands, channel demultiplexer 701 restores values in the secondary channel from the values in the primary channel and difference values in the bit stream. If the secondary channel information is missing the bit stream, secondary channel information can hardly be recovered from the primary channel by duplicating the primary channel information on secondary channels and using the stereo grid, to be analyzed later.
706 Channel Reconstruction
706 stereo reconstruction is applied to secondary channels when no secondary channel information (time samples) is found in the bit stream. The stereo grid, rebuilt by the grid decoder 702, is applied to the secondary time samples, retrieved by doubling the primary channel time sample information, to maintain the original stereo power ratio between channels.
Multi-channel reconstruction
Multichannel reconstruction 706 applies to secondary channels when no secondary information (whether time samples or grids) for the secondary channels is present in the bit stream. The procedure is similar to stereo reconstruction 706, except that the partial grid reconstructed by the grid decoder 702 is applied to the time samples of the secondary channel within each channel group, retrieved by duplicating primary channel time sample information to maintain the appropriate power level in the secondary channel. The partial grid is applied individually to each secondary channel in the reconstructed channel group after scaling by another grid or scale factor grids that include the G0 grid in step 703 of scaling by multiplying time samples of the G grid by corresponding elements of the partial grid for each secondary channel. G0 grid, partial grids can be applied in any order in accordance with the present invention.
While several illustrative embodiments of the invention have been shown and described, numerous variations and alternative embodiments will occur to those skilled in the art. Such variations and alternative embodiments are contemplated and can be made without departing from the scope of the invention as defined in the appended claims.
12 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12
42 members in 15 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 691558P | United States of America | – | |
| 69155805 | United States of America | P | |
| 69155805 | United States of America | P | |
| 452001 | United States of America | – | |
| 45200106 | United States of America | A | |
| 45200106 | United States of America | A | |
| 2006003986 | International Bureau of the World Intellectual Property Organization (WIPO) | W | |
| 2006003986 | International Bureau of the World Intellectual Property Organization (WIPO) | W | |
| 452001 | – | – | – |
| 691558P | – | – | – |
| PCTIB2006003986 | – | – | – |
| US20050691558P | – | – | – |
| US20060452001 | – | – | – |
| WO2006IB03986 | – | – | – |
Members42
| Document | Office | Kind | |
|---|---|---|---|
| US2007063877A1 | United States of America | A1 | |
| AU2006332046A1 | Australia | A1 | |
| CA2608030A1 | Canada | A1 | |
| CA2853987A1 | Canada | A1 | |
| WO2007074401A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2007074401A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1891740A2 | European Patent Office (EPO) | A2 | |
| KR20080025377A | Republic of Korea | A | |
| CN101199121A | China | A | |
| TR200806843T1 | Türkiye | T1 | |
| TR200708666T1 | Türkiye | T1 | |
| TR200806842T1 | Türkiye | T1 | |
| JP2008547043A | Japan | A | |
| HK1117655A1 | Hong Kong, China | A1 | |
| US7548853B2 | United States of America | B2 | |
| RU2008101778A | Russian Federation | A | |
| RU2402160C2 | Russian Federation | C2 | |
| NZ563337A | New Zealand | A | |
| IL187402A | Israel | A | |
| AU2006332046B2 | Australia | B2 | |
| AU2011205144A1 | Australia | A1 | |
| NZ590418A | New Zealand | A | |
| EP1891740A4 | European Patent Office (EPO) | A4 | |
| AU2011221401A1 | Australia | A1 | |
| NZ593517A | New Zealand | A | |
| CN101199121B | China | B | |
| JP2012098759A | Japan | A | |
| EP2479750A1 | European Patent Office (EPO) | A1 | |
| JP5164834B2 | Japan | B2 | |
| HK1171859A1 | Hong Kong, China | A1 | |
| JP5291815B2 | Japan | B2 | |
| KR101325339B1 | Republic of Korea | B1 | |
| EP2479750B1 | European Patent Office (EPO) | B1 | |
| AU2011221401B2 | Australia | B2 | |
| AU2011205144B2 | Australia | B2 | |
| PL2479750T3 | Poland | T3 | |
| CA2608030C | Canada | C | |
| CA2853987C | Canada | C | |
| TR200806842B | Türkiye | B | |
| EP1891740B1 | European Patent Office (EPO) | B1 | |
| ES2717606T3This record | Spain | T3 | |
| PL1891740T3 | Poland | T3 |
Numbers
- Publication
- 2717606
- Publication, DOCDB
- 2717606
- Publication, EPODOC
- ES2717606T
- Application
- 6848793
- Application, DOCDB
- 06848793
- Application, EPODOC
- ES20060848793T
Titles2
- Spanish
- Codificación y decodificación de audio escalable usando un banco de filtros jerárquico
- English
- Encoding and decoding of scalable audio using a hierarchical filter bank
Classification
- CPC, 9
- G10L19/0204
- G10L19/02
- G10L19/0212
- G10L19/022
- G10L19/035
- G10L19/24
- G10L25/18
- H03M7/28
- H03M7/30
- IPC, 2
- G10L19 24
- G10L19 02