Apparatus and method for generating audio subband values and apparatus and method for generating time-domain audio samples
23 claims: 13 independent, 10 dependent
- 1REIVINDICAÇÕES 1. Equipamento para a geração de valores de sub-banda de áudio em canais de sub-banda de áudio, compreendendo:um janelador de análise (110) para o janelamento de um frame (120) de amostras de entrada de áudio no domínio do tempo, sendo em uma sequência temporal que se estende de uma amostra prévia para uma amostra posterior usando uma função janela de análise (190), compreendendo uma sequência de coeficientes de janela para a obtenção de amostras janeladas, a função janela de análise compreendendo um primeiro número de coeficientes de janela obtidos de uma maior função janela compreendendo uma sequência de um maior segundo número de coeficientes de janela, caracterizado pelo fato de que os coeficientes de janela da função janela são obtidos por uma interpolação de coeficientes de janela da maior função janela;e onde o segundo número é um número par;e um calculador (170) para o cálculo dos valores de sub-banda de áudio usando as amostras janeladas.
- 2Equipamento (100), de acordo com a reivindicação 1, caracterizado pelo fato de que o equipamento (100) está adaptado para interpolar os coeficientes de janela da maior função janela para obter os coeficientes de janela da função janela.
- 3Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o equipamento (100) ou o janelador de análise (110) está adaptado de maneira que os coeficientes de janela da função janela sejam interpolados linearmente.
- 4Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o equipamento (100) ou o janelador de análise (110) está adaptado de maneira que os coeficientes de janela da função janela de análise são interpolados com base em dois coeficientes consecutivos de janela da maior função janela de acordo com a sequência dos coeficientes de janela da maior função janela, para obter um coeficiente de janela da função janela.
- 5Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o equipamento (100) ou o janelador de análise (110) está adaptado de maneira a obter os coeficientes de janela c(n) da função janela de análise com base na equação c(n) = (c 2 (2n) + c 2 (2n + 1)) , onde n é um número inteiro que indica um índice dos coeficientes de janela c(n), e C2(n) é um coeficiente de janela da maior função janela.
- 6Equipamento (100), de acordo com a reivindicação 5, caracterizado pelo fato de que o equipamento (100) ou o janelador de análise (110) está adaptado de maneira que os coeficientes de janela C2(n) da maior função janela obedeçam às relações dadas na tabela do Anexo 4.
- 7Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o janelador de análise (110) está adaptado de maneira que o janelamento compreende a multiplicação das amostras de entrada de áudio no domínio do tempo x(n) do frame (120) para obter as amostras janeladas z(n) do frame janelado com base na equação z(n) = x(n) · c(n) onde n é um número inteiro que indica um índice da sequência de coeficientes de janela na faixa de 0 a T · N-l, onde c(n) é o coeficiente de janela da função janela de análise correspondente ao índice n, onde x(N · T-l) é a última amostra de entrada de áudio no domínio do tempo de um frame (120) de amostras de entrada de áudio no domínio do tempo, onde o janelador de análise (110) está adaptado de maneira que o frame (120) de amostras de entrada de áudio no domínio do tempo compreende uma sequência de T blocos (130) de amostras de entrada de áudio no domínio do tempo que se prolonga da mais anterior até a mais posterior das amostras de entrada de áudio no domínio do tempo do frame (120), cada bloco compreendendo N amostras de entrada de áudio no domínio do tempo, e onde T e N são inteiros positivos e T é maior que 4.
- 8Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o janelador de análise (110) está adaptado de maneira que a função janela de análise (190) compreende um primeiro grupo (200) de coeficientes de janela compreendendo uma primeira porção da sequência de coeficientes de janela e um segundo grupo (210) de coeficientes de janela compreendendo uma segunda porção da sequência de coeficientes de janela, onde a primeira porção compreende menos coeficientes de janela que a segunda porção, onde um valor de energia dos coeficientes de janela na primeira porção é maior que um valor de energia dos coeficientes de janela da sequnda porção, e onde o primeiro grupo de coeficientes de janela é usado para o janelamento posterior das amostras no domínio do tempo e o segundo grupo de coeficientes de janela é usado para o janelamento das amostras mais anteriores no domínio do tempo.
- 9Equipamento, de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o equipamento (100) está adaptado para usar uma função janela de análise (190) sendo uma versão de tempo reverso ou do índice reverso da função de janela de síntese (370) a ser usado para os valores de sub-banda de áudio.
- 10Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que um janelador de análise (110) está adaptado de maneira que a maior função janela é assimétrica com referência à sequência de coeficientes de janela.
- 11Equipamento (300) para a geração de amostras de áudio no domínio do tempo, compreendendo:um calculador (310) para o cálculo da sequência (330) de amostras intermediárias no domínio do tempo a partir de valores de sub-banda de áudio nos canais de sub-banda de áudio, a sequência compreendendo amostras mais anteriores intermediárias no domínio do tempo e amostras mais posteriores no domínio do tempo;um janelador de síntese (360) para o janelamento da seqíiência (330) de amostras intermediárias no domínio do tempo usando uma função de janela de síntese (370) compreendendo uma sequência de coeficientes de janela para obter amostras janeladas intermediárias no domínio do tempo, a função de janela de síntese compreendendo um primeiro número de coeficientes de janela obtidos de uma maior função janela compreendendo uma sequência de um maior segundo número de coeficientes de janela, caracterizado pelo fato de que os coeficientes de janela da função janela são obtidos por uma interpolação de coeficientes de janela da maior função janela;e onde o segundo número é par;e um estágio de saída de adição de sobrepasso (400) para o processamento das amostras janeladas intermediárias no domínio do tempo de maneira a obter amostras no domínio do tempo.
- 12Equipamento (300), de acordo com a reivindicação 11, caracterizado pelo fato de que o equipamento (300) é adaptado para interpolar os coeficientes de janela da maior função janela para obter os coeficientes de janela da função janela.
- 13Equipamento (300), de acordo com qualquer uma das reivindicações 11 ou 12, caracterizado pelo fato de que o equipamento (300) está adaptado de maneira que os coeficientes de janela da função de janela de síntese sejam interpolados linearmente.
- 14Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 13, caracterizado pelo fato de que o equipamento (300) está adaptado de maneira que os coeficientes de janela da função de janela de síntese são interpolados com base em dois coeficientes consecutivos de janela da maior função janela de acordo com a sequência de coeficientes de janela da maior função janela para obter um coeficiente de janela da função janela.
- 15Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 14, caracterizado pelo fato de que o equipamento (300) está adaptado para obter os coeficientes de janela c(n) da função de janela de síntese com base na equação c(n) = y (c 2 (2n) + c 2 (2n + 1)) , onde C2(n) são coeficientes de janela de uma maior função janela correspondente ao índice n.
- 16Equipamento (300), de acordo com a reivindicação 15, caracterizado pelo fato de que o equipamento (300) está adaptado de maneira que os coeficientes de janela C2(n) obedecem as relações dadas na tabela do Anexo 4.
- 17Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 16, caracterizado pelo fato de que a janela de síntese (360) está adaptada de maneira que o janelamento compreende a multiplicação das amostras intermediárias no domínio do tempo g(n) da sequência de amostras intermediárias no domínio do tempo para obter as amostras janeladas z(n) do frame janelado (380) com base na equação z(n) = g(n) · c(t · N - 1 - n) para n = 0,..., T · N - 1.
- 18Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 17, caracterizado pelo fato de que o janelador de síntese (360) está adaptado de maneira que a função de janela de síntese (370) compreende um primeiro grupo (420) de coeficientes de janela compreendendo uma primeira porção da sequência de coeficientes de janela e um segundo grupo (430) de coeficientes de janela compreendendo uma segunda porção da sequência de coeficientes de janela, a primeira porção compreendendo menos coeficientes de janela que a segunda porção, onde um valor de energia dos coeficientes de janela na primeira porção é maior que um valor da energia dos coeficientes de janela da segunda porção, e onde o primeiro grupo de coeficientes de janela é usado para o janelamento posterior das amostras intermediárias no domínio do tempo e o segundo grupo de coeficientes de janela é usado para o janelamento prévio das amostras intermediárias no domínio do tempo.
- 19Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 18, caracterizado pelo fato de que o equipamento (300) está adaptado de maneira para usar uma função de janela de síntese (370) sendo uma versão de tempo reverso ou de índice reverso de uma função janela de análise (190) usada para a geração de valores de sub-banda de áudio.
- 20Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 19, caracterizado pelo fato de que o janelador de síntese (360) está adaptado de maneira que a maior função janela é assimétrica com referência a uma sequência de coeficientes de janela.
- 21Método para a geração de valores de sub-banda de áudio em canais de sub-banda de áudio, compreendendo:janelamento de um frame de amostras de entrada de áudio no domínio do tempo estando em uma sequência temporal que se estende de uma amostra prévia para uma amostra posterior usando uma função janela de análise para obter amostras janeladas, a função janela de análise compreendendo um primeiro número de coeficientes de janela obtidos a partir de uma maior função janela compreendendo uma sequência de um maior segundo número de coeficientes de janela, caracterizado pelo fato de que os coeficientes de janela da função janela são obtidos por uma interpolação pelos coeficientes de janela da maior função janela;e onde o segundo número é um número par;e calcular os valores de sub-banda de áudio usando as amostras janeladas.
- 22Método para a geração de amostras de áudio no domínio do tempo, compreendendo:calcular uma sequência de amostras intermediárias no domínio do tempo a partir de valores de sub-banda de áudio em canais de sub-banda de áudio, a sequência compreendendo amostras prévias intermediárias no domínio do tempo e amostras posteriores intermediárias no domínio do tempo;janelamento da sequência de amostras intermediárias no domínio do tempo usando uma função de janela de síntese para obter amostras janeladas no domínio do tempo, a função de janela de síntese compreendendo um primeiro número de coeficientes de janela obtidas a partir de uma maior função janela compreendendo a sequência de um maior segundo número de coeficientes de janela, caracterizado pelo fato de que os coeficientes de janela da função janela são obtidos por uma interpolação de coeficientes de janela da maior função janela;e onde o segundo número é par;e adição de sobrepasso das amostras janeladas no domínio do tempo para obter as amostras no domínio do tempo.
- 23Programa com um código de programas para a execução, caracterizado pelo fato de que opera em um processador, de um 5 método de acordo com a reivindicação 21 ou de acordo com a reivindicação 22. RESUMO EQUIPAMENTO E MÉTODO PARA A GERAÇÃO DE VALORES DE SUB-BANDA DE ÁUDIO E EQUIPAMENTO E MÉTODO PARA A GERAÇÃO DE AMOSTRAS DE ÁUDIO NO DOMÍNIO DO TEMPO Uma configuração de um equipamento (100) para a geração de valores de sub-banda de áudio em canais de sub-banda de áudio compreende um janelador de análise (110) para o janelamento de um frame (120) de amostras de entrada de áudio no domínio do tempo, estando em uma sequência temporal que se estende de uma amostra prévia até uma amostra posterior usando uma função janela de análise (190), compreendendo uma sequência de coeficientes de janela para obter amostras janeladas. A função janela de análise compreende um primeiro número de coeficientes de janela obtidos a partir da maior função janela compreendendo uma sequência de um maior segundo número de coeficientes de janela, onde os coeficientes de janela da função janela são obtidos por uma interpolação de coeficientes de janela da maior função janela. O equipamento (100) ainda compreende um calculador (170) para o cálculo dos valores de sub-banda de áudio usando as amostras janeladas. REIVINDICAÇÕES 1. Equipamento para a geração de valores de sub-banda de áudio em canais de sub-banda de áudio, compreendendo:um janelador de análise (110) para o janelamento de um frame (120) de amostras de entrada de áudio no dominio do tempo, sendo em uma sequência temporal que se estende de uma amostra prévia para uma amostra posterior usando uma função janela de análise (190), compreendendo uma sequência de coeficientes de janela para a obtenção de amostras janeladas, a função janela de análise compreendendo um primeiro número de coeficientes de janela obtidos de uma maior função janela compreendendo uma sequência de um maior segundo número de coeficientes de janela, caracterizado pelo fato de que os coeficientes de janela da função janela são obtidos por uma interpolação de coeficientes de janela da maior função janela;e onde o segundo número é um número par;e um calculador (170) para o cálculo dos valores de sub-banda de áudio usando as amostras janeladas. 2. Equipamento (100), de acordo com a reivindicação 1, caracterizado pelo fato de que o equipamento (100) está adaptado para interpolar os coeficientes de janela da maior função janela para obter os coeficientes de janela da função janela. 3. Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o equipamento (100) ou o janelador de análise (110) está adaptado de maneira que os coeficientes de janela da função janela sejam interpolados linearmente. 4. Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o equipamento (100) ou o janelador de análise (110) está adaptado de maneira que os coeficientes de janela da função janela de análise são interpolados com base em dois coeficientes consecutivos de janela da maior função janela de acordo com a sequência dos coeficientes de janela da maior função janela, para obter um coeficiente de janela da função janela. 5. Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o equipamento (100) ou o janelador de análise (110) está adaptado de maneira a obter os coeficientes de janela c(n) da função janela de análise com base na equação c(n) = 3 (c 2 (2n) + c 2 (2n + 1)) , onde n é um número inteiro que indica um índice dos coeficientes de janela c(n), e c 2 (n) é um coeficiente de janela da maior função janela. 6. Equipamento (100), de acordo com a reivindicação 5, caracterizado pelo fato de que o equipamento (100) ou o janelador de análise (110) está adaptado de maneira que os coeficientes de janela c 2 (n) da maior função janela obedeçam às relações dadas na tabela do Anexo 4. 7. Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o janelador de análise (110) está adaptado de maneira que o janelamento compreende a multiplicação das amostras de entrada de áudio no domínio do tempo x(n) do frame (120) para obter as amostras janeladas z(n) do frame janelado com base na equação z(n) = x(n) · c(n) onde n é um número inteiro que indica um índice da sequência de coeficientes de janela na faixa de 0 a T · N-l, onde c(n) é o coeficiente de janela da função janela de análise correspondente ao índice n, onde x(N · T-l) é a última amostra de entrada de áudio no domínio do tempo de um frame (120) de amostras de entrada de áudio no domínio do tempo, onde o janelador de análise (110) está adaptado de maneira que o frame (120) de amostras de entrada de áudio no domínio do tempo compreende uma sequência de T blocos (130) de amostras de entrada de áudio no domínio do tempo que se prolonga da mais anterior até a mais posterior das amostras de entrada de áudio no domínio do tempo do frame (120), cada bloco compreendendo N amostras de entrada de áudio no domínio do tempo, e onde T e N são inteiros positivos e T é maior que 4. 8. Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que o janelador de análise (110) está adaptado de maneira que a função janela de análise (190) compreende um primeiro grupo (200) de coeficientes de janela compreendendo uma primeira porção da sequência de coeficientes de janela e um segundo grupo (210) de coeficientes de janela compreendendo uma segunda porção da sequência de coeficientes de janela, onde a primeira porção compreende menos coeficientes de janela que a segunda porção, onde um valor de energia dos coeficientes de janela na primeira porção é maior que um valor de energia dos coeficientes de janela da segunda porção, e onde o primeiro grupo de coeficientes de janela é usado para o janelamento posterior das amostras no domínio do tempo e o segundo grupo de coeficientes de janela é usado para o janelamento das amostras mais anteriores no domínio do tempo. 9. Eguipamento, de acordo com gualguer uma das reivindicações anteriores, caracterizado pelo fato de que o equipamento (100) está adaptado para usar uma função janela de análise (190) sendo uma versão de tempo reverso ou do indice reverso da função de janela de síntese (370) a ser usado para os valores de sub-banda de áudio. 10. Equipamento (100), de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que um janelador de análise (110) está adaptado de maneira que a maior função janela é assimétrica com referência à sequência de coeficientes de janela. 11. Equipamento (300) para a geração de amostras de áudio no domínio do tempo, compreendendo: um calculador (310) para o cálculo da sequência (330) de amostras intermediárias no domínio do tempo a partir de valores de sub-banda de áudio nos canais de sub-banda de áudio, a sequência compreendendo amostras mais anteriores intermediárias no domínio do tempo e amostras mais posteriores no domínio do tempo;um janelador de síntese (360) para o janelamento da sequência (330) de amostras intermediárias no domínio do tempo usando uma função de janela de síntese (370) compreendendo uma sequência de coeficientes de janela para obter amostras janeladas intermediárias no domínio do tempo, a função de janela de síntese compreendendo um primeiro número de coeficientes de janela obtidos de uma maior função janela compreendendo uma sequência de um maior segundo número de coeficientes de janela, caracterizado pelo fato de que os coeficientes de janela da função janela são obtidos por uma interpolação de coeficientes de janela da maior função janela;e onde o segundo número é par;e um estágio de saída de adição de sobrepasso (400) para o processamento das amostras janeladas intermediárias no domínio do tempo de maneira a obter amostras no domínio do tempo. 12. Equipamento (300), de acordo com a reivindicação 11, caracterizado pelo fato de que o equipamento (300) é adaptado para interpolar os coeficientes de janela da maior função janela para obter os coeficientes de janela da função janela. 13. Equipamento (300), de acordo com qualquer uma das reivindicações 11 ou 12, caracterizado pelo fato de que o equipamento (300) está adaptado de maneira que os coeficientes de janela da função de janela de síntese sejam interpolados linearmente. 14. Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 13, caracterizado pelo fato de que o equipamento (300) está adaptado de maneira que os coeficientes de janela da função de janela de síntese são interpolados com base em dois coeficientes consecutivos de janela da maior função janela de acordo com a sequência de coeficientes de janela da maior função janela para obter um coeficiente de janela da função janela. 15. Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 14, caracterizado pelo fato de que o equipamento (300) está adaptado para obter os coeficientes de janela c(n) da função de janela de síntese com base na equação c(n) = j (c 2 (2n) + c 2 (2n + 1)) onde c 2 (n) são coeficientes de janela de uma maior função janela correspondente ao índice n. 16. Equipamento (300), de acordo com a reivindicação 15, caracterizado pelo fato de que o equipamento (300) está adaptado de maneira que os coeficientes de janela c 2 (n) obedecem as relações dadas na tabela do Anexo 4. 17. Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 16, caracterizado pelo fato de que a janela de síntese (360) está adaptada de maneira que o janelamento compreende a multiplicação das amostras intermediárias no domínio do tempo g(n) da sequência de amostras intermediárias no domínio do tempo para obter as amostras janeladas z(n) do frame janelado (380) com base na equação z(n) = g(n) · c(t · N - 1 — n) para n = 0,..., T · N - 1. 18. Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 17, caracterizado pelo fato de que o janelador de síntese (360) está adaptado de maneira que a função de janela de síntese (370) compreende um primeiro grupo (420) de coeficientes de janela compreendendo uma primeira porção da sequência de coeficientes de janela e um segundo grupo (430) de coeficientes de janela compreendendo uma segunda porção da sequência de coeficientes de janela, a primeira porção compreendendo menos coeficientes de janela que a segunda porção, onde um valor de energia dos coeficientes de janela na primeira porção é maior que um valor da energia dos coeficientes de janela da segunda porção, e onde o primeiro grupo de coeficientes de janela é usado para o janelamento posterior das amostras intermediárias no domínio do tempo e o segundo grupo de coeficientes de janela é usado para o janelamento prévio das amostras intermediárias no domínio do tempo. 19. Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 18, caracterizado pelo fato de que o equipamento (300) está adaptado de maneira para usar uma função de janela de síntese (370) sendo uma versão de tempo reverso ou de índice reverso de uma função janela de análise (190) usada para a geração de valores de sub-banda de áudio. 20. Equipamento (300), de acordo com qualquer uma das reivindicações de 11 a 19, caracterizado pelo fato de que o janelador de síntese (360) está adaptado de maneira que a maior função janela é assimétrica com referência a uma sequência de coeficientes de janela. 21. Método para a geração de valores de sub-banda de áudio em canais de sub-banda de áudio, compreendendo: janelamento de um frame de amostras de entrada de áudio no domínio do tempo estando em uma sequência temporal que se estende de uma amostra prévia para uma amostra posterior usando uma função janela de análise para obter amostras janeladas, a função janela de análise compreendendo um primeiro número de coeficientes de janela obtidos a partir de uma maior função janela compreendendo uma sequência de um maior segundo número de coeficientes de janela, caracterizado pelo fato de que os coeficientes de janela da função janela são obtidos por uma interpolação pelos coeficientes de janela da maior função janela;e onde o segundo número é um número par;e calcular os valores de sub-banda de áudio usando as amostras janeladas. 22. Método para a geração de amostras de áudio no domínio do tempo, compreendendo: calcular uma sequência de amostras intermediárias no domínio do tempo a partir de valores de sub-banda de áudio em canais de sub-banda de áudio, a sequência compreendendo amostras prévias intermediárias no domínio do tempo e amostras posteriores intermediárias no domínio do tempo;janelamento da sequência de amostras intermediárias no domínio do tempo usando uma função de janela de síntese para obter amostras janeladas no domínio do tempo, a função de janela de síntese compreendendo um primeiro número de coeficientes de janela obtidas a partir de uma maior função janela compreendendo a sequência de um maior segundo número de coeficientes de janela, caracterizado pelo fato de que os coeficientes de janela da função janela são obtidos por uma interpolação de coeficientes de janela da 5 maior função janela;e onde o segundo número é par;e adição de sobrepasso das amostras janeladas no dominio do tempo para obter as amostras no domínio do tempo. 23. Programa com um código de programas para a execução, caracterizado pelo fato de que opera em um processador, de um 10 método de acordo com a reivindicação 21 ou de acordo com a reivindicação 22.
Independent claims23
453 paragraphs in 1 section, as filed
DEWCRIÇÃO
EQUIPMENT AND METHOD FOR THE GENERATION OF AUDIO SUB-BAND VALUES AND EQUIPMENT AND METHOD FOR THE GENERATION OF TIME SAMPLES OF AUDIO
Technical Field
The configurations of the present invention refer to an equipment and method for the generation of audio subband values, an equipment and a method for the generation of audio samples in the time domain and systems comprising any of the aforementioned equipment, which can , for example, be implemented in the field of modern audio coding, audio decoding or other applications related to audio transmission.
Modern digital audio processing is typically based on encoding schemes that allow a significant reduction in terms of bit rates, transmission bandwidths and storage space when compared to direct transmission or storage of the respective audio data. This is achieved by encoding the audio data on the sender side and decoding the encoded data on the container side before, for example, providing the decoded audio data to the listener or other signal processing.
These digital audio processing systems can be implemented with reference to a wide range of parameters, which typically influence the quality of transmitted or otherwise processed audio data, on the one hand, and computational efficiency, bandwidth and others performance-related parameters, on the other hand. Very generally, higher qualities require higher bit rates, greater computational complexity and a higher storage requirement for the corresponding encoded audio data. Thus, depending on the desired application, factors such as allowable bit rates, acceptable computational complexity and acceptable amounts of data must be balanced with a desirable and obtainable quality.
Another particularly important parameter in real-time applications such as bidirectional or monodirectional communication, the delay imposed by the different coding schemes can also play an important role. As a consequence, the delay imposed by audio encoding and decoding places another restriction in terms of the parameters mentioned above when balancing the needs and costs of different encoding schemes having a specific desired field of application. Thus, digital audio systems can be applied in many different fields of use, ranging from ultra-low quality transmission to high-end transmission, with different parameters and different restrictions being very generally imposed on the respective audio systems. In some applications, a lower delay may, for example, require a higher bit rate and thus a higher transmission bandwidth when compared to an audio system with greater delay, as of comparable level of quality.
However, in many cases, compromises must be made in terms of different parameters such as bit rate, computational complexity, memory requirements, quality and delay.
summary
According to a configuration of the present invention, an equipment for the generation of audio subband values in audio subband channels comprises an analysis window for windowing a frame of audio input samples in the field of audio. time being in a sequence of time that extends from a previous sample to a later sample using a window analysis function that comprises a sequence of window coefficients to obtain windowed samples, the analysis window function comprising a first number of window coefficients obtained from a larger window function comprising a sequence of a greater second number of window coefficients, in which the window coefficients of the window function are obtained by means of of an interpolation of window coefficients of the largest window function, in which the second number is an even number and a calculator for calculating the audio subband values using windowed samples. According to a configuration of the present invention, an equipment for the generation of audio samples in the time domain comprises a calculator for calculating a sequence of intermediate samples in the time domain of the subband audio values in subband channels audio, the sequence comprising pre-intermediate samples in the time domain and later samples in the time domain, a synthesis window for windowing the sequence of intermediate samples in the time domain using a window synthesis function comprising a sequence of window coefficients to obtain intermediate windowed samples in the time domain, the synthesis window function comprising a first number of window coefficients obtained from a larger window function, comprising a sequence of a greater second number of window coefficients, in which the window coefficients of the window function are obtained by an interpolation of window coefficients of the largest window function, and in which the second number is even and an overpass addition output stage for the processing of the intermediate window samples in the domain time to obtain samples in the time domain.
Brief Description of Drawings
The configurations of the present invention will now be described, with reference to the accompanying drawings.
Fig. 1 shows a block diagram of an equipment configuration for the generation of audio subband values;
Fig. 2a shows a block diagram of an equipment configuration for generating audio samples in the time domain;
Fig. 2b illustrates a functional principle according to a configuration of the present invention in the form of a device for the generation of samples in the time domain;
Fig. 3 illustrates the concept of interpolation window coefficients according to a configuration of the present invention;
Fig. 4 illustrates the concept of interpolation window coefficients in the case of the sine window function;
Fig. 5 shows a block diagram of a configuration of the present invention comprising an SBR decoder and an SBR encoder;
Fig. 6 illustrates the delay sources of an SBR system;
Fig. 7a shows a flowchart of a method configuration for generating audio subband values;
Fig. 7b illustrates a step in configuring the method shown in Fig. 7a;
Fig. 7c shows a flowchart of a method configuration for the generation of audio subband values;
Fig. 8a shows a flow chart of a comparative example of a method for generating samples in the time domain;
Fig. 8b shows a flow chart of a comparative example of a method for generating samples in the time domain;
Fig. 8c shows a flowchart of a method configuration for generating samples in the time domain;
Fig. 8d shows a flowchart of another configuration of a method for generating samples in the time domain;
Fig. 9a shows a possible implementation of a comparative example of a method for generating audio subband values;
Fig. 9b shows a possible implementation of a method configuration for the generation of audio subband values;
Fig. 10a shows a possible implementation of a comparative example of a method for generating samples in the time domain;
Fig. 10b shows another possible implementation of a method configuration for generating samples in the time domain;
Fig. 11 shows a comparison of a synthesis window function according to a configuration of the present invention and a sine window function;
Fig. 12 shows a comparison of a synthesis window function according to a configuration of the present invention and a prototype SBR QMF filter function;
Fig. 13 illustrates the different delays caused by the window function and the prototype filter function shown in Fig. 12;
Fig. 14a shows a table illustrating the different contributions to the delay of a conventional AAC-LD + SBR codec and an AAC-ELD codec comprising a configuration of the present invention;
Fig. 14b shows another table comprising details regarding the delay of different components of different codecs;
Fig. 15a shows a comparison of a frequency response of equipment based on a window function according to a configuration of the present invention and an equipment based on the sine window function;
Fig. 15b shows a close view of the frequency response shown in Fig. 15a;
Fig. 16a shows a comparison of the frequency response of 4 different window functions;
Fig. 16b shows a close view of the frequency response shown in Fig. 16a;
Fig. 17 shows a comparison of a frequency response of two different window functions, a window function according to the present invention and a window function being a symmetric window function;
Fig. 18 schematically shows the general temporal masking property of the human ear; and
Fig. 19 illustrates a comparison of an original audio timing signal, a timing signal generated based on a HEAAC codec and a timing signal based on a codec comprising a configuration of the present invention.
Detailed Description of Settings
Figs. 1 to 19 show block diagrams and other diagrams describing the functional properties and characteristics of the different equipment configurations and methods for generating audio subband values, equipment and methods for generating time domain samples and systems comprising at least one of the aforementioned equipment or methods. However, before describing a first configuration of the present invention in more detail, it should be noted that the configurations of the present invention can be implemented in hardware and software. Thus, the implementations described in terms of block diagrams of hardware implementations of the respective configurations can also be considered as flowcharts of an appropriate configuration of a corresponding method. Also, a flowchart describing a configuration of the present invention can be considered to be a block diagram of a corresponding hardware implementation.
Implementations of filter banks, which can be implemented as an analysis filter bank or a synthetic filter bank, will be described below. An analysis filter bank is an equipment for the generation of subband audio values in subband audio channels based on samples (input) of audio in the time domain being in a temporal sequence that extends to from a previous sample to a later sample. In other words, the term analysis filter bank can also be used as a synonym for a configuration of the present invention in the form of an equipment for generating audio subband values. Likewise, a synthesis filter bank is a filter bank for generating audio samples in the time domain of audio subband values in audio subband channels. In other words, the term synthesis filter bank can be used as a synonym for a configuration according to the present invention in the form of a device for generating audio samples in the time domain.
Both the analysis filter bank and the synthesis filter bank, which are also named to summarize how filter banks can, for example, be implemented as modulated filter banks. Modulated filter banks, examples and configurations of which will be highlighted in greater detail below, are based on oscillations with frequencies that are based on or that are obtained from central frequencies of corresponding sub-bands in the frequency domain. The term modulated refers in this context to the fact that the aforementioned oscillations are used in the context with a window function or a prototype filter function, depending on the concrete implementation of this modulated filter bank. Modulated filter banks can, in principle, be based on oscillations of real value as on a harmonic oscillation (sine oscillation or cosine oscillation) or on corresponding oscillations of complex values (complex exponential oscillations). Likewise, modulated filter banks are referred to as actual modulated filter banks or complex modulated filter banks, respectively.
In the following description, the configurations of the present invention will be described in greater detail in the form of complex modulated low-delay filter banks and modulated real low-delay filter banks and corresponding methods and software implementations. One of the main applications of these low-delay modulated filter banks is an integration in a low-delay spectral band (SBR) replication system, which is currently based on the use of a complex QMF filter bank with a symmetric prototype filter (QMF = Quadrature Mirror Filter).
As will become apparent in the structure of the present description, an implementation of low-delay filter banks according to the configurations of the present invention will give the advantage of a better decision between computational complexity, frequency response, spread of temporal noise and quality (reconstruction). In addition, a better decision is made between the delay and the quality of reconstruction based on an approach to use the so-called zero wait techniques in order to prolong the filter impulse response of the respective filter banks without introducing additional delays. . A shorter delay at a predefined quality level, a better quality at the predefined delay level or a simultaneous improvement of both the delay and the quality, can be achieved by employing a bank of analysis filters or a bank of synthesis filters according to a configuration of the present invention.
The configurations of the present invention are based on the finding that these improvements can be obtained by employing an interpolation scheme to obtain a window function having a first number of window coefficients based on a window function having a second larger number of window coefficients. Using an interpolation scheme, a better distribution of energy values of the window coefficients of the window functions can be obtained. This leads, in many cases, to a better level of aliasing and an improvement in audio quality. For example, when the largest window function comprises an even number of window coefficients, an interpolation scheme can be useful.
The computational complexity only increases a little with the use of an interpolation scheme. However, this small increase is not only ruled out by the improvement regarding quality, but also by the resulting savings related to the reduced use of memory when comparing the situation with two separate window functions being stored independently. While interpolation can be done in one or a few cycles of a processor's clock signal in an implementation, in many cases leading to negligible delay and increased computational complexity, the requirement for additional memory can be extremely important in many applications. For example, in the case of mobile applications, memory can be limited, especially when long window functions are used with a significant number of window coefficients.
Furthermore, the configurations according to the present invention can be used in the context of a new window function for any of the two filter banks described above, further improving the aforementioned decisions. The quality and / or delay can be further improved in the case of a bank of analysis filters, employing a window analysis function comprising a sequence of window coefficients, which comprises a first group having a first consecutive portion of the sequence of window coefficients. window and the second group of window coefficients comprising a second consecutive portion of the sequence of window coefficients. The first portion and the second portion comprise all window coefficients of the window function. Furthermore, the first portion comprises fewer window coefficients than the second portion, but an energy value of the window coefficients in the first portion is greater than an energy value of the window coefficients of the second portion. The first group of window coefficients is used for later window sampling in the time domain and the second group of window coefficients is used for early windowed time domain samples. This form of the window function gives the opportunity to process samples in the time domain with window coefficients having higher energy values previously. This is a result of the described distribution of window coefficients for the two portions and their applications in the sequence of audio samples in the time domain. As a consequence, the use of this window function can reduce the delay introduced by the filter bank by a constant quality level or allows a better quality level based on a constant delay level.
Likewise, in the case of a configuration of the present invention in the form of a device for generating audio samples in the time domain and corresponding method, a synthesizer can use the synthesis window function, which comprises a sequence of window coefficients ordered correspondingly in a first (consecutive) portion and (consecutive) second portion. Also in the case of a synthesis window function, the energy value or a total energy value of a window coefficient in the first portion is greater than the energy value of a total energy value of a second window coefficient. portion, wherein the first portion comprises fewer window coefficients than the second portion. Due to this distribution of the window coefficients between the two portions and the fact that the synthesizer pane uses the first portion of the pane coefficients to window later samples in the time domain and the second portion of window coefficients for window pane previews in the time domain, the effects and advantages described above also apply to a synthesis filter bank or to a corresponding method configuration.
The detailed descriptions of the synthesis window function and the analysis window functions employed in the structure of some configurations of the present invention will be described below in greater detail. In many configurations of the present invention, the sequence of window coefficients of the synthesis window function and / or of the analysis window function comprises exactly the first group and the second group of window coefficients. Furthermore, each of the window coefficients in the sequence of window coefficients belongs exactly to one of the window coefficients of the first and second groups.
Each of the two groups comprises exactly a portion of the sequence of window coefficients in a consecutive manner. In the present description, a portion comprises a consecutive set of window coefficients according to the sequence of window coefficients. In configurations according to the present invention, each of the two groups (first and second group) comprises exactly a portion of the sequence of the window coefficients in the aforementioned manner. The respective groups of window coefficients do not comprise any of the window coefficients, which do not belong to exactly a portion of the respective group. In other words, in many configurations of the present invention, each of the first and the second group of window coefficients comprises only the first portion and the second portion of window coefficients without understanding other window coefficients.
In the structure of the present description, a consecutive portion of the sequence of window coefficients should be seen as a linked set of window coefficients in the mathematical sense, where the set has no lack of window coefficients compared to the sequence of window coefficients, which would be located in a range (for example, index range) of the window coefficients of the respective portion. As a consequence, in many configurations of the present invention, the sequence of window coefficients is divided exactly into two linked portions of window coefficients, each forming the first or the second group of window coefficients. In such cases, each window coefficient comprised in the first group of window coefficients can be arranged before or after each of the window coefficients of the second group of window coefficients with reference to the total sequence of window coefficients.
In other words, in many configurations according to the present invention, the sequence of window coefficients is divided exactly into two groups or portions without leaving out any window coefficients. According to the sequence of the window coefficients, which also represents such an order, each of the two groups or portions comprises all the window coefficients up to (but excluding) or starting from (including) a border window coefficient. As an example, the first portion or the first group may comprise window coefficients with indexes from 0 to 95 and from 96 to 639 in the case of a window function comprising 640 window coefficients (with indexes from 0 to 639). Here, the boundary window coefficient would be the one corresponding to the index 96. Of course, other examples are also possible (for example, 0 to 543 and 544 to 639).
The detailed example implementation of an analysis filter bank described below provides a filter length covering 10 blocks of input samples causing a system delay of only 2 blocks, which is the corresponding delay introduced by an MDCT (discrete modified transform of cosine) or an MDST (modified discrete sine transform). One difference is due to the longer filter length covering 10 blocks of input samples compared to an implementation of an MDCT or MDST in which the overpass is increased from 1 block in the case of MDCT and MDST to an overpass and 9 blocks. However, other implementations can also be made covering a different number of input sample blocks, which are also called audio input samples. Furthermore, other decisions can also be considered and implemented.
Fig. 1 shows a block diagram of an analysis filter bank 100 as an equipment configuration for generating audio subband values in audio subband channels. The analysis filter bank 100 comprises an analysis window 110 for window frame 120 of the audio input samples in the time domain. Frame 120 comprises T blocks 130-1,. .., 130-T blocks of audio (input) samples in the time domain, where T is a positive integer and equal to 10 in the case of the configuration shown in Fig. 1. However, frame 120 can also comprise a different number of blocks 130.
Both frame 120 and each of blocks 130 comprise audio input samples in the time domain in a time sequence that extends from a previous sample to a later sample according to a timeline, as indicated by arrow 140 in the Fig. 1. In other words, in the illustration shown in Fig. 1, the more the audio sample in the time domain, which in this case is also a sample of audio input in the time domain, is to the right, the more later the corresponding audio sample in the time domain will be with reference to the sequence sample audio in the time domain.
Analysis window 110 generates, based on a sequence of audio samples in the time domain, windowed samples in the time domain, which are arranged in a frame 150 of windowed samples. According to the time domain audio input sample frame 120, also the windowed sample frame 150 comprises T windowed sample blocks 160-1. .., 160T. In the preferred configurations of the present invention, each of the windowed sample blocks 160 comprises the same number of windowed samples as the number of time domain audio input samples of each time domain audio input sample block 130. Thus, when each of the blocks 130 comprises N time domain input audio samples, frame 120 and frame 150 each comprise T · N samples. In this case, N is a positive integer, which can, for example, have the values of 32 or 64. For T = 10, frames 120, 150 each comprise 320 and 640, respectively, in the case above.
The analysis window 110 is coupled to a calculator 170 for the calculation of the audio subband values based on the windowed samples provided by the analysis window 110. The audio subband values are provided by the calculator 170 as a block 180 of audio subband values, where each of the audio subband values corresponds to an audio subband channel. In a preferred configuration, also the block 180 of audio subband values comprises N subband values.
Each of the audio subband channels corresponds to a characteristic central frequency. The central frequencies of the different audio subband channels can, for example, be equally distributed or evenly spaced with respect to the frequency bandwidth of the corresponding audio signal described by the audio input samples in the time domain provided to the bank analysis filters 100.
Analysis window 110 is adapted to window audio input samples in the time domain of frame 120 based on an window analysis function comprising a sequence of window coefficients having a first number of window coefficients to obtain the windowed samples of frame 150. The analysis window 110 is adapted to window the audio samples frame in the time domain 120, multiplying the values of the audio samples in the time domain by the window coefficients of the analysis window function. In other words, the window comprises and multiplies the audio samples in the time domain by the corresponding window coefficient by the element. Like both, the frame 120 of audio samples in the time domain and the window coefficients comprise a corresponding sequence, the multiplication in elements of the window coefficients and the audio samples in the time domain are made according to the respective sequences, for example, as indicated by a sample and window coefficient index.
In configurations of the present invention, the window functions used to window the audio input sample frame in the time domain are generated based on a larger window function comprising a greater second number of window coefficients employing an interpolation scheme such as, for example, highlighted in the context of Fig. 3 and 4. The largest window function typically comprises an even number of window coefficients and can, for example, be asymmetric with reference to the sequence of window coefficients. Symmetric window functions can also be used.
The window function 190 used for window frame 120 of time domain input samples is, for example, obtained by analysis window 110 or filter bank 100 by interpolating the window coefficients of the largest window function. In configurations according to the present invention, this is done, for example, by interpolating consecutive window coefficients of the largest window function. Here, a polynomial or spline linear interpolation scheme can be used.
When, for example, each window coefficient of the largest window function is used once to generate a window coefficient of the window function and the second number is even, the number of window coefficients of the window function 190 (first number) is half of the second number. This interpolation can be based on a linear interpolation, an example of which will be highlighted later in the context of equation (15). However, as noted, other interpolation schemes can also be used.
In configurations of the present invention in the form of an analysis filter bank 100 as shown in Fig. 1, the analysis window function, as well as the synthesis window function in the case of a synthesis filter bank can, for example , only understand windowed coefficients of real value. In other words, each of the window coefficients assigned to a window coefficient index is a real value.
The window coefficients together form the respective window function, an example of which is shown in Fig. 1 as an analysis window function 190. Next, window functions will be considered, which allow a reduction of the delay when used in the context of described filter banks. However, the configurations of the present invention are limited to these low-delay window functions.
The sequence of window coefficients forming the analysis window function 190 comprises a first group 200 and a second group 210 of window coefficients. The first group 200 comprises a first consecutive, linked portion of the window coefficients of the sequence of window coefficients, whereas the second group 210 comprises a second consecutive, linked portion of a window coefficient. Together with the first portion in the first group 200, they form the entire sequence of window coefficients of the analysis window function 190. Furthermore, each window coefficient of the sequence of window coefficients either belongs to the first portion or the second portion of coefficients window, so that the entire analysis window function 190 is composed of the window coefficients of the first portion and the second portion. The first portion of window coefficients is thus identical to the first group 200 of window coefficients and the second portion is identical to the second group 210 of window coefficients, as indicated by the corresponding arrows 200, 210 in Fig. 1.
The number of window coefficients in the first group 200 of the first portion of window coefficients is less than the number of window coefficients in the second group of the second portion of window coefficients. However, an energy value or a total energy value of the window coefficients in the first group 200 is greater than an energy value or a total energy value of the window coefficients in the second group 210. As will be highlighted later, an energy value for a set of window coefficients is based on the sum of the squares of the absolute values of the corresponding window coefficients.
In configurations according to the present invention, the analysis window function 190, as well as the corresponding synthesis window function, can therefore be asymmetrical with reference to the sequence of window coefficients or an index of a window coefficient. Based on the set of definition of the window coefficient indices in which the analysis window function 190 is defined, the analysis window function 190 is asymmetric, when for all real numbers there is another real number in, so that the value absolute value of the window coefficient corresponding to the window coefficient of the window coefficient index (n<sub>0</sub> - n) is not equal to the absolute value of the window coefficient corresponding to the window coefficient index (n<sub>0</sub> + n), when (n<sub>0</sub> - n) and (n<sub>0 </sub>+ n) belong to the definition set.
Furthermore, as also shown schematically in Fig. 1, the analysis window function 190 comprises changes in signals in which the product of two consecutive window coefficients is negative. More details and other characteristics of the possible window functions according to the configurations of the present invention will be discussed in more detail in the context of Figs. 11 to 19.
As previously indicated, the windowed sample frame 150 comprises a block structure similar to individual blocks 160-1. .., 160-T as frame 120 of the individual time domain input samples. As the analysis window 110 is adapted for the window of the audio input samples in the time domain by multiplying these values by the window coefficients of the analysis window function 190, the frame 150 of windowed samples is also in the time domain. The calculator 170 calculates the audio subband values, or to be more exact, the block 180 of audio subband values using the windowed sample frame 150 and performs a transfer from the time domain to the frequency domain. Therefore, the calculator 170 can be considered as a time / frequency converter, capable of providing the block 180 subband audio values as a spectral representation of the frame 150 of the windowed samples.
Each audio subband value of block 180 corresponds to a subband with a characteristic frequency. The number of audio subband values comprised in block 180 is also sometimes referred to as the band number.
In many configurations according to the present invention, the number of audio subband values in block 180 is identical to the number of audio input samples in the time domain of each of blocks 130 of frame 120. In the case where the windowed sample frame 150 comprises the same block structure as the frame 120, so that each of the windowed sample blocks 160 also comprises the same number of windowed samples as the block of these audio input samples in the time domain 130, block 180 of audio subband values also naturally comprises the same number as block 160.
Frame 120 can be optionally generated, based on a block of new audio input samples in time domain 220 by changing blocks 130-1,. .., 130- (Tl) by a block in the opposite direction to the arrow 140 indicating the direction in time. Therefore, a frame 120 of audio input samples in the time domain to be processed is generated by changing the last (Tl) blocks of a frame 120 directly preceding audio samples in the time domain by a block in the direction of the audio samples. in the time domain and by adding the new block 220 of new audio samples in the time domain as the new block 130-1 comprising the last audio samples in the time domain of the current frame 120. In Fig. 1 this is also indicated by a series of dashed arrows 230 indicating the change of blocks 130-1, ..., 130- (Tl) in the opposite direction to arrow 140.
Due to this change of blocks 130 in the opposite direction of time as indicated by arrow 140, the current frame 120 to be processed, comprises block 130- (Tl) of frame 120 directly preceding it as the new block 130-T. Likewise, blocks 130- (Tl), ..., 130-2 of the current frame 120 to be processed are the same as block 130- (T-2), ..., 130-1 of the directly preceding frame 120 Block 130-T of the directly preceding frame 120 is discarded.
As a consequence, each audio sample in the time domain of the new block 220 will be processed T times in the structure of T consecutive processing of T consecutive frames 120 of the audio input samples in the time domain. Thus, each audio input sample in the time domain of the new block 220 contributes not only to different T frames 120, but also to different T frames 150 of windowed samples and T blocks 180 of audio subband values. As previously indicated, in a preferred configuration according to the present invention, the number of T blocks in frame 120 is equal to 10, so that each time domain audio sample supplied to the analysis filter bank 100 contributes to 10 different blocks 180 of audio subband values.
In the beginning, before a single frame 120 is processed by the analysis filter bank 100, frame 120 can be initialized to a small absolute value (below a predetermined limit), for example, the value 0. As will be explained in greater detail details below, the format of the analysis window function 190 comprises a central point or center of mass, which typically corresponds to or is located between two window coefficient indices of the first group 200.
As a consequence, the number of new blocks 220 to be inserted in frame 120 is small, before frame 120 is filled to a point where the portions of frame 120 are occupied by non-volatile values (that is, non-zero values) ) that correspond to window coefficients having a significant contribution in terms of their energy values. Typically, the number of blocks to be inserted in frame 120 before a significant process can be started, is 2 to 4 blocks, depending on the format of the analysis window function 190. Thus, the analysis filter bank 100 is capable of provide blocks 180 faster than the corresponding filter bank is employing, for example, a symmetric window function. As the new blocks 220 are typically supplied to the analysis filter bank 100 as a set, each of the new blocks corresponds to a recording or sampling time, which is essentially given by the length of the block 220 (that is, the number of blocks). audio input samples in the time domain comprised in block 220) and the sampling rate or sampling frequency. Therefore, the analysis window function 190, as incorporated in a configuration of the present invention, leads to a reduced delay before the first and the following blocks 180 of audio subband values can be provided or sent by the filter bank 100 .
As another option, equipment 100 may be able to generate a signal or incorporate a piece of information regarding the analysis window function 190 used in the generation of frame 180 or regarding the synthesis window function to be used in the structure of a database. synthesis filters. Therefore, the analysis filter function 190 can, for example, be a reverse version in time or in the index of the synthesis window function to be used by the synthesis filter bank.
Fig. 2a shows a block diagram of an equipment configuration 300 for generating audio samples in the time domain based on the audio subband value block. As already explained, a configuration of the present invention in the form of an equipment 300 for generating audio samples in the time domain is generally also called a synthesis filter bank 300, since the equipment is capable of generating audio samples in the time domain, which can in principle be reproduced, based on audio subband values that comprise spectral information referring to an audio signal. Thus, the synthesis filter bank 300 is capable of synthesizing audio samples in the time domain based on audio subband values, which can, for example, be generated by a corresponding analysis filter bank 100.
Fig. 2a shows a block diagram of the synthesis filter bank 300 comprising a calculator 310 which is provided with a block 320 of audio subband values (in the frequency domain). The calculator 310 is capable of calculating a frame 330 comprising a sequence of intermediate samples in the time domain from the audio subband values of block 320. The frame 330 of intermediate samples in the time domain comprises in many configurations according to the present invention also a block structure similar, for example, to the frame 150 of windowed samples of the analysis filter bank 100 of Fig. 1. In such cases , frame 330 comprises blocks 340-1, ..., 340-T blocks of intermediate samples in the time domain.
The sequence of intermediate samples in the time domain of frame 330, as well as each block 340 of intermediate samples in the time domain comprise an order according to time, as indicated by arrow 350 in Fig. 2a. As a consequence, frame 330 comprises a previous intermediate sample in the time domain in block 340-T and a last intermediate sample in the time domain in block 340-1, which represent the first and last intermediate sample in the time domain of the frame 330, respectively. Also, each of the blocks 340 comprises a similar order. As a consequence, in configurations of a synthesis filter bank, the terms frame and sequence can generally be used interchangeably.
The calculator 310 is coupled to a synthesis window 360, for which the frame 330 of intermediate samples in the time domain is provided. The synthesis window is adapted for
4 window the sequence of intermediate samples in the time domain using a synthesis window function 370 shown schematically in Fig. 2a. As an output, the synthesis pane 360 provides a frame 380 of intermediate windowed samples in the time domain, which can also comprise a block structure of blocks 390-1, ..., 390-T.
Frames 330 and 380 can comprise T blocks 340 and 390, respectively, where T is a positive integer. In a preferred configuration according to the present invention, in the form of a synthesis filter bank 300, the number of T blocks is equal to 10. However, in different configurations, also different numbers of blocks can be comprised in one of the frames . To be more precise, in principle, the number of T blocks can be greater than or equal to 3, or greater than or equal to 4, depending on the circumstances of the implementation and the previously explained decisions of the configurations according to the present invention comprising the structure in block of frames for both the synthesis filter bank 100 and the synthesis filter bank 300.
The synthesis pane 360 is coupled to an overpass addition 400 output stage, to which frame 380 of intermediate windowed samples in the time domain is provided. The overlapping addition output stage 400 is capable of processing the intermediate windowed samples in the time domain to obtain a block 410 of samples in the time domain. Block 410 of the time domain (output) samples can then, for example, be supplied to other components for further processing, storage or transformation into audible audio signals.
The calculator 310 for calculating the sequence of samples in the time domain comprised in frame 330 is capable of transferring data from the frequency domain in the time domain. Therefore, the calculator 310 may comprise a time / frequency converter capable of generating a signal in the time domain of the spectral representation comprised in block 320 of audio subband values. As explained in the context of the calculator 170 of the analysis filter bank 100 shown in Fig. 1, each of the audio subband values of block 320 corresponds to an audio subband channel having a characteristic center frequency.
In contrast, the intermediate samples in the time domain comprised in frame 330 represent, in principle, information in the time domain. The synthesis winder 360 is capable and adapted for windowing the sequence of intermediate samples in the time domain comprised in frame 330 using the synthesis window function 370 as shown schematically in Fig. 2a.
As already highlighted in the context of Fig. 1, the synthesis window 360 also uses a synthesis window function 370, which is obtained by interpolating a larger window function comprising a second number of window coefficients. The second number is, therefore, greater than a first number of window coefficients of the synthesis window function 370 used for the windowing of the intermediate samples in the time domain of frame 330.
The synthesis window function 370 can, for example, be obtained by the synthesis window 360 or by the filter bank 300 (the equipment) that performs one of the interpolation schemes highlighted above. The window coefficients of the synthesis window function can, for example, be generated based on linear, polynomial or spline interpolation. Furthermore, in configurations according to the present invention, interpolation can be based on the use of consecutive window coefficients of the largest window function. When each window coefficient of the largest window function is used exactly once, the window function 370 comprising the (lowest) first number of window coefficients can, for example, comprise exactly half the number of window coefficients of the largest window function, when the second number is even. In other words, in this case, the second number can be double the first number. However, also other interpolation scenarios and schemes can be implemented in the structure of the configurations of the present invention.
In the following, the case of a so-called low-delay window function will be considered more closely. As previously indicated, the configurations according to the present invention are not limited to these window functions. Other window functions can also be used, such as symmetric window functions.
The synthesis window function 370 comprises a sequence of window coefficients, which also comprises a first group 420 and a second group 430 of window coefficients as explained above in the context of the window function 190 with a first group 200 and a second group 210 of window coefficients.
The first group 420 of window coefficients of the synthesis window function 370 comprises a first consecutive portion of the sequence of window coefficients. Similarly, the second group 430 of coefficients also comprises a second consecutive portion of the sequence of window coefficients, in which window the total value of the first portion comprises less coefficients the second portion and where an energy or energy value of the coefficients of the window in the first portion is greater than the corresponding energy value of the window coefficients of the second portion. Other characteristics and properties of the synthesis window function 370 may be similar to the corresponding characteristics and properties of the analysis window function 190 as shown schematically in Fig. 1. As a consequence, reference is made at present to the corresponding description in the structure of the window function of analysis 190 and the other description of the window functions with reference to Figs. 11 to 19, where the first group 200 corresponds to the first group 420 and the second group 210 corresponds to the second group 430.
For example, the portions comprised in the two groups 420, 430 of window coefficients, typically form each consecutive and connected set of window coefficients together comprising all the window coefficients of the window function sequence of the window function 370. In many configurations according to the present invention, the analysis window function 190 as shown in Fig. 1 and the synthesis window function 370 as shown in Fig. 2a are based on each other. For example, the analysis window function 190 may be a reverse-time or reverse index version of the synthesis window function 370. However, other relationships between the two window functions 190, 370 may also be possible. It may be advisable to employ a synthesis window function 370 in the structure of the synthesis window 360, which is related to the analysis window function 190, which was used in the course of the generation (optionally before other modifications) of block 320 of values of audio subband supplied to the synthesis filter bank 300.
As highlighted in the context of Fig. 1, the synthesis filter bank 300 in Fig. 2a can be optionally adapted so that the input block 320 can comprise other signals and new pieces of information relating to the window functions. As an example, block 320 can comprise information regarding the analysis window function 190 used to generate block 320 or regarding the synthesis window function 370 to be used by the synthesis window 360. Thus, the filter bank 300 can be adapted to isolate the respective information and provide it to the synthesis window 360.
The overlapping addition output stage 400 is capable of generating the block 410 of samples in the time domain by processing the intermediate windowed samples in the time domain comprised in frame 380. In different configurations according to the present invention, the output stage overpaste addition 4000 can comprise a memory for temporarily storing previously received frames 380 of intermediate windowed samples in the time domain. Depending on the implementation details, the overlapping addition output stage 400 may, for example, comprise T different storage positions comprised in memory for storing a total number of T frames 380 of intermediate windowed samples in the time domain. However, also a different number of storage positions can be comprised in the overpaste addition exit stage 400, as needed. Furthermore, in different configurations according to the present invention, the overpaste addition output stage 400 may be able to provide block 410 of samples in the time domain based solely on a frame 380 of intermediate samples in the time domain. The configurations of different synthesis filter banks 300 will be explained later in greater detail.
Fig. 2b illustrates a functional principle according to a configuration of the present invention in the form of a synthesis filter bank 300, in which the generation of the window function 370 by interpolation does not focus only on simplicity.
The block 320 of audio subband values is first transferred from the frequency domain to the time domain by the calculator 310, which is illustrated in Fig. 2b by an arrow 440. The resulting frame 320 of intermediate samples in the time domain comprising blocks 340-1, ..., 340-T of intermediate samples in the time domain is then windowed by the synthesis window 360 (not shown in Fig. 2b) by multiplying the sequence of intermediate samples in the time domain of frame 320 by the sequence of window coefficients of the synthesis window function 370 to obtain frame 380 of intermediate windowed samples in the time domain. Frame 380 again comprises blocks 390-1, ..., 390-T of intermediate windowed samples in the time domain, together forming frame 380 of intermediate windowed samples in the time domain.
In the configuration shown in Fig. 2b of a synthesis filter bank of the invention 300, the overpaste addition output stage 400 is then able to generate block 410 of time domain output samples by adding for each index value of the samples of audio in the time domain of block 410, the intermediate windowed samples in the time domain of a block 390 of different frames 380. As illustrated in Fig. 2b, the audio samples in the time domain of block 410 are obtained by adding for each audio sample index an intermediate sample in the windowed time domain of block 390-1 of frame 380, processed by the synthesis window 360 at the current time and as previously described, the corresponding intermediate sample in the time domain of the second block 3902 of a frame 380-1 processed immediately before frame 380 and stored in a storage position in the overflow addition output stage 400. As illustrated in Fig. 2b, other corresponding intermediate windowed samples in the time domain of other blocks 390 (for example, block 390-3 of frame 380-2, block 390-4 of frame 380-3, block 390-5 of frame 380-4) processed synthesis filter bank 300 before they can be used. Frames 380-2, 380-3, 380-4 and optionally other frames 380 have been processed by the synthesis filter bank 300 on previous occasions. Frame 380-2 was immediately processed before frame 380-1 and, likewise, frame 380-3 was generated immediately before frame 380-2 and so on.
The output stage of overpaste addition 400 as used in the configuration is capable of adding different blocks 390-1, ..., 390-T of intermediate windowed samples for each index of block 410 of (outgoing) time domain T in the time domain of different T frames 380, 380-1, ..., 380 (Tl). Thus, apart from the first T blocks processed, each of the samples (outgoing) time domain of block 410 is based on different T blocks 320 of audio subband values.
As in the case of the configuration of the present invention of an analysis filter bank 100 described in Fig. 1, due to the shape of the synthesis window function 370, the synthesis filter bank 300 offers the possibility of quickly providing block 410 of samples (outgoing) time domain. This is also a consequence of the shape of the window function 370. As the first group 420 of window coefficients corresponds to a higher energy value and comprises fewer window coefficients than the 430 sequence, the synthesis window 360 is capable of providing significant frames 380 of windowed samples when frame 330 of intermediate samples in the time domain it is filled in such a way that at least the window coefficients of the first group 420 contribute to frame 380. The window coefficients of the second group 430 demonstrate a lower contribution due to their lower energy values.
Therefore, when at the beginning, the synthesis filter bank 300 is initialized to 0, the provision of blocks 410 can, in principle, be initiated when only a few blocks 320 of audio subband values have been received by the filter bank. synthesis filter 300. Therefore, the synthesis filter bank 300 also allows a significant delay reduction compared to the synthesis filter bank having, for example, a symmetric synthesis window function.
As previously indicated, calculators 170 and 310 of the configurations shown in Figs. 1 and 2a can be implemented as real value calculators, generating or being able to process audio subband values of real value blocks 180 and 320, respectively. In such cases, the calculators can, for example, be implemented as real value calculators based on harmonic oscillating functions such as the sine function or the cosine function. However, complex value calculators such as calculators 170, 310 can also be implemented. In such cases, calculators can, for example, be implemented based on complex exponential functions or other complex value harmonic functions. The frequency of oscillations of real values or complex values usually depends on the index of the audio subband value, sometimes also called the band index or subband index of the specific subband. Furthermore, the frequency can be identical or depend on the central frequency of the corresponding subband. For example, the frequency of the oscillation can be multiplied by a constant factor, changed with reference to the central frequency of the corresponding subband or it can depend on a combination of both modifications.
A complex value calculator 170, 310 can be built or implemented based on real value calculators. For example, for a complex value calculator, an efficient implementation can, in principle, be used for both the sine and cosine modulated part of a filter bank representing the real and imaginary part of a complex value component. . This means that it is possible to implement both the cosine-modulated part and the sine-modulated part based on, for example, modified structures DCT-IV- and DST-IV. Furthermore, other implementations may employ the use of an FFT (FFT = Fast Fourier Transform) optionally being implemented together for both the real part and the part of the complex modulated calculators, using an FFT or using a separate FFT stage for each transformed.
Mathematical Description
The following sections describe an example of an analysis filter bank and synthesis filter bank configurations with multiple 8-block overrides for the part, which no longer causes delay, as explained above, and one block for the future, which causes the same delay as for an MDCT / MDST structure (MDCT = Discrete Modified Cosine Transform;
MDST = Modified Sine Discrete Transform). In other words, in the following example, parameter T is equal to 10.
First, a description of a modulated complex low-delay analysis filter bank will be given. As shown in Fig. 1, the analysis filter bank 100 comprises the steps of transforming an analysis window performed by the analysis window 110 and an analysis modulation performed by the calculator 170. The analysis window is based on the equation<sup>z</sup>i<sub>r</sub>n “w (102V - 1 - n) · x<sub>ifn</sub> for <n <10 · N <sub>r</sub> (1) where, z<sub>i <n</sub> is the windowed sample (of real value) corresponding to the block index ie the sample index n of frame 150 shown in Fig. 1. The x value<sub>i <n</sub> is the time input sample (of real value) corresponding to the same block index ie sample index η. The analysis window function 190 is represented in equation (1) by its window coefficients of real value w (n), where n is also the window coefficient index in the range indicated in equation (1). As previously explained, parameter N is the number of samples in a block 220, 130, 160, 180.
From the analysis window function w (10N-ln) it can be seen that the analysis window function represents an inverted version or a time-reversed version of the synthesis window function, which is actually represented by the window coefficient w (n).
The analysis modulation performed by the calculator 170 in the configuration shown in Fig. 1, is based on two equations <sup>2N</sup>~<sup>r</sup> (π (1Ϋ = 2 · y z. cos - (n + n<sub>n</sub> 1 k Hn = -8N V<sup>V</sup> \
2^-<sup>1</sup> Λ <sub>π</sub> Λ jai <sup>x</sup><sub>Im</sub> ag.ik = <sup>2</sup> · Σ <sup>sin</sup> - (n + nj k + n = -8N (2) (3) for the spectral coefficient index or k band index being an integer in the range of (4)
The XR values<sub>and the</sub>i, i, ke Xim<sub>ag</sub>, iA represent the real part and the imaginary part of the audio subband complex value value corresponding to the block index ie the spectral coefficient index k of block 180. The parameter n<sub>0</sub> represents an index option, which is equal to n<sub>0</sub> = —N / 2 + 0.5. (5)
The corresponding low-delayed modulated complex synthesis filter bank comprises the steps of transforming a synthesis modulation, a synthesis window, an overlapping addition being described.
Synthesis modulation is based on the equation
<img file="PT1994530E_D0001.tif" />
7V-1 k = Q <n <10 N <sup>π</sup> í cos - \ n +
<img file="PT1994530E_D0002.tif" />
Nl j Imag, i, kk = Q sin - (η + ηΛ k [N + -
2JJ (6) where x'i,<sub>n</sub> is an intermediate sample in the time domain of frame 330 corresponding to the sample index n and the block index i. Again, parameter N is an integer that indicates the length of the block 320, 340, 390, 410, which is also called the transform block length or, due to the block structure of frames 330, 380, as a deviation from the previous block. The other variables and parameters were also introduced above, such as the spectral coefficient index k and the deviation n<sub>0</sub>.
The synthesis window performed by the 360 synthesis window in the configuration shown in Fig. 2a is based on the equation <sup>zl</sup>i<sub>f</sub>n <sup>=</sup> ‘ <sup>X</sup>'i, n °<sup>r</sup> 0 <n <10 · N, (7) where z'i,<sub>n</sub> is the value of the intermediate sample in the windowed time domain corresponding to the sample index n and the block index i of frame 380.
The transformation pattern of the overpaste addition is based on the equation <sup>0U</sup>h, n “ι_γ ^<sub>η + Ν</sub>ΡΖ i- ^ n + lN + Z, <sup>Z</sup> i- ^ n + 4N + Z <sup>Z</sup> i-6, n + 6N ~ ^ ~<sup>Z</sup> i-7, n + 7N ~ ^ ~<sup>Z</sup> iS, n + SN ^ ~<sup>Z</sup> i-9, n + 9N (β) is 0 <n <N where outi,<sub>n</sub> represents the sample in the time domain (output) corresponding to the sample index n and the block index i. Equation (8), thus illustrates the overpaste addition operation as done in the overpaste addition exit stage 400 illustrated at the bottom of Fig. 2b.
However, the configurations according to the present invention are not limited to complex modulated low-delay filter banks, allowing audio signal processing with one of these filter banks. A real-time implementation of a low-delay filter bank can also be made for an extended, low-delay audio encoding. As a comparison, for example, equations (2) and (6) in terms of a cosine part reveals, the cosine contribution of the analysis modulation and the synthesis modulation shows a comparable structure when considering that of an MDCT . Although the design method in principle allows an extension of the MDCT in both directions regarding time, only an extension of E (= T-2) blocks to the past is applied here, where each of the T blocks comprises N samples. The frequency coefficient Χι, κ of the band k and block i within an analysis filter bank with N channels or N bands can be summarized by w<sub>The</sub>(n) · x (n) · cosi— (n + | - f) (k + |) \ N (9) for the spectral coefficient index k as defined by equation (4). Here again, n is a sample index and w<sub>The</sub> is an analysis window function.
To guarantee the totality, the mathematical description previously provided of the modulated complex low delay analysis filter bank can be given in the same summarized form of equation (9), exchanging the cosine function for the exponential function of complex value. To be more precise, with the definition and variables given above, equations (1), (2), (3) and (5) can be summarized and extended according to<sup>x</sup>'ik = -<sup>2</sup> Σ% k + U η = -Ε · Ν \ N where in contrast to equations (2) and (3), the extension of 8 blocks in the past has been replaced by the variable E (= 8).
The steps of the synthesis modulation and the synthesis window, as described for the complex case in equations (6) and (7), can be summarized in the case of a real value synthesis filter bank. Frame 380 of intermediate windowed samples in the time domain, which is also called a demodulated vector, is given by, 1 (71 j z'i, n = - Σ '<sup>x</sup>i, k <sup>WAISTBAND</sup> (- (Λ + | - f) (k + |) N <sub>k = 0</sub> k N) r \) where z'i<sub>zn</sub> is the intermediate sample in the windowed time domain corresponding to the band index ie the sample index η. The sample index n is again an integer in the range of <n <n (2 + e) = N · T (12) ew<sub>s</sub>(n) is the synthesis window, which is compatible with the analysis window w<sub>The</sub>(n) of equation (9).
The transformation step of overpaste addition is then given by 0 <sup>— 2</sup> ί + Ι, η-1-Ν (1 Q \ where χ'ί, η θ the reconstructed signal, or a time domain sample of block 410 as provided by the overpaste addition output stage 400 shown in Fig. 2a.
For the complex value synthesis filter bank 300, equations (6) and (7) can be summarized and generalized with reference to the extension of E (= 8) blocks for the path according to <sup>z</sup>'vn = - Σ <sup>w</sup>3<sup>n></sup>) <sup>Re x</sup>i, k <sup>ex</sup>P - j · (- (η + I - f) (k + |) <sub>(14)</sub> where j = VI is the imaginary unit. Equation (13) represents the generalization of equation (8), and is also valid for the case of complex value.
As the direct comparison of equation (14) with equation (7) shows, the window function w (n) of equation (7), that is, the same synthesis window function as w<sub>s</sub>(n) of equation (14). As previously mentioned, the similar comparison of equation (10) with the window analysis function coefficient w<sub>The</sub>(n) with equation (1) shows that the analysis window function is the time reverse version of the synthesis window function in the case of equation (1).
As in both, an analysis filter bank 100 is shown in Fig. 1 and a synthesis filter bank 300 as shown in Fig. 2a offers a significant improvement in terms of the decision between the delay, on the one hand, and the quality the audio process; on the other hand, filter banks 100, 300 are generally referred to as low-delay filter banks. Its complex value version is sometimes called a complex low-delay filter bank, abbreviated by CLDFB. In some circumstances, the term CLDFB is not only used for the complex value version, but also for the real value version of the filter bank.
As the previous discussion of the mathematical history showed, the structure used to implement the proposed low-delay filter banks uses a structure of the MDCT or IMDCT type (IMDCT = inverse MDCT), as shown by the MPEG-4 Standard, using an extended overpass. Additional overlapping regions can be attached en bloc on the left side as well as on the right side of the MDCT type core. Here, only the extension for the right side (for the synthesis filter bank) is used, which only works from past samples and therefore does not cause any further delay.
Inspection of equations (1), (2) and (14) showed that the processing is very similar to that of MDCT or IMDCT. With only minor modifications comprising a modified analysis window function and synthesis window function, respectively, the MDCT or IMDCT is extended to a modulated filter bank that can operate multiple overpasses and is very flexible with respect to its delay. As for example, equations (2) and (3) demonstrated that the complex version is, in principle, obtained by the simple addition of a modulated sine to the given cosine modulation.
Interpolation
As highlighted in the context of Figs. 1 and 2a, both the analysis window 110 and the synthesis window 360 or the respective filter banks 100, 300 are adapted to window the respective sample frames in the time domain by multiplying each of the respective audio samples in the domain time by the individual window coefficient. Each of the samples in the time domain is, in other words, multiplied by a window coefficient (individual), such as equations (1), (7), (9), (10), (11), and (14) demonstrated. As a consequence, the number of window coefficients for the respective window function is typically identical to the number of the respective audio samples in the time domain.
However, under certain implementation circumstances, it may be advisable to implement the window function having a higher second number of window coefficients compared to the actual window function with a lower first number of coefficients, which is actually used during the windowing of the respective frame or sequence of audio samples in the time domain. This may, for example, be advisable in the case when the memory requirements of a specific implementation can be more valuable than computational efficiency. Another scenario in which a sub-sampling of the window coefficients can be useful is the case of the so-called dual index approach, which is, for example, used in the structure of SBR systems (SBR = Spectral Band Replication). The concept of SBR will be explained in more detail in the context of Figs. 5 and 6.
In this case, the analysis window 110 or the synthesis window 360 can also be adapted in such a way that the respective window function used for windowing audio samples in the time domain provided for the respective window window 110, 360 is obtained by an interpolation of window coefficients of the largest window function having a higher second number of window coefficients.
The interpolation can, for example, be done by a linear polynomial or spline-based interpolation. For example, in the case of linear interpolation, but also in the case of a polynomial or spline-based interpolation, the respective window pane 100, 360 may then be able to interpolate the window coefficients of the window function used for the window based on two consecutive coefficients window of the largest window function, according to a sequence of the window coefficients of the largest window function to obtain a window coefficient of the window function.
Especially in the case of an even number of audio samples in the time domain and window coefficients, an implementation of an interpolation as previously described, results in a significant improvement in audio quality. For example, in the case of an even number N · T of audio samples in the time domain in one of frames 120, 330, not using an interpolation, for example, a linear interpolation, will result in serious aliasing effects during the rest of the processing of the respective audio samples in the time domain.
Fig. 3 illustrates an example of a linear interpolation based on a window function (an analysis window function or a synthesis window function) to be used in the context with frames comprising N · T / 2 audio samples in the time domain . Due to memory constraints or other implementation details, the window function's own window coefficients are not stored in memory, but in a larger window function comprising N · T window coefficients being stored during the availability of suitable memory or otherwise. form. Fig. 3 illustrates in the upper graph, the corresponding window coefficients c (n) as a function of the window coefficient indices n in the range between 0 and N · Tl.
Based on a linear interpolation of two consecutive window coefficients of the window function having the largest number of window coefficients as shown in the upper graph in Fig. 3, an interpolated window function is calculated based on the equation ci [n] = | (c [2n] + c [2n + 1]) for 0 <η <N · T / 2<sub>Ί</sub> .
The number of window coefficients ci (n) interpolated from the window function to be applied to the frame having N · T / 2 audio samples in the time domain comprises half the number of window coefficients.
To better illustrate the fact, in Fig. 3 window coefficients 450-0,. .., 450-7 are shown at the top of Fig. 3 corresponding to a window coefficient c (0), c (7). Based on these window coefficients and the other window coefficients of the window function, an application of equation (15) leads to the window coefficients ci (n) of the interpolated window function shown at the bottom of Fig. 3. For example, based on the window coefficients 450-2 and 450-3, the window coefficient 460-1 is generated based on equation (15), as illustrated by arrows 470 in Fig. 3. Likewise, the coefficient of window 460-2 of the interpolated window function is calculated based on the window coefficient 450-4, 450-5 of the window function shown at the top of Fig. 3. Fig. 3 shows the generation of other window coefficients ci ( n).
To illustrate the aliasing cancellation achieved by the interpolated subsampling of the window function, Fig. 4 illustrates the interpolation of the window coefficients in the case of the sine window function, which can, for example, be used in an MDCT. Simply put, the left half of the window function and the right half of the window function are drawn overlapping. Fig. 4 shows a simplified version of a sine window, comprising only 2 · 4 window coefficients or points for an MDCT having the length of 8 samples.
Fig. 4 shows four window coefficients 480-1, 480-2, 480-3 and 480-4 of the first half of the sine window and four window coefficients 490-1, 490-2, 490-3 and 390-4 of the second half of the sine window. The window coefficient 490-1, ..., 4904 corresponds to the window coefficient indices 5, ..., 8. The window coefficients 490-1,. .., 490-4 correspond to the second half of the length of the window function so that the given indices = '= 4 must be added to obtain the real indices.
To reduce or even achieve the cancellation of the aliasing effects described above, the window coefficient must comply with the condition w (n) · (Ν'-l - n) = w ^ N ^ n) · w (2N'-1 - n) (16) in the best possible way. The better the relationship (16) is obeyed, the better is the alias suppression or alias cancellation.
Assuming the situation in which a new window function having half the number of window coefficients is determined for the left half of the window function, the following problem arises. Due to the fact that the window function comprises an even number of window coefficients (subsampling of even numbers), without employing an interpolation scheme as highlighted in Fig. 3, the window coefficients 480-1 and 480-3 or 480-2 and 480-4 correspond to only one aliasing value of the original window function or the original filter.
This leads to an unbalanced proportion of the spectral energy and leads to an asymmetric redistribution of the central point (center of mass) of the corresponding window function. Based on the interpolation equation (15) for the window coefficient w (n) of Fig. 4, the interpolated values li and I2 obey the aliasing relationship (16) much better, and will thus lead to a better improvement regarding quality. of processed audio data.
However, the use of an even more elaborate interpolation scheme, for example, a spline-based or similar interpolation scheme, may even result in window coefficients, which even better obey the relationship (16). Linear interpolation is, in most cases, sufficient and allows for quick and efficient implementation.
The situation in the case of a typical SBR system that employs an SBR-QMF filter bank (QMF = Quadrature Mirror Filter), linear interpolation or other interpolation scheme does not need to be implemented since the SBR-QMF prototype filter comprises a odd number of prototype filter coefficients. This means that the SBR-QMF prototype filter comprises a maximum value with reference to which the subsampling can be implemented so that the symmetry of the SBR-QMF prototype filter remains intact.
In Figs. 5 and 6, a possible application of the configurations according to the present invention will be described in the form of an analysis filter bank and a synthesis filter bank. An important field of application is the SBR system or SBR tool (SBR = Spectral Band Replication). However, other configurations applications according to the present invention may come from other fields, in which there is a need for spectral modifications (for example, modifications or gain equalizations), such as spatial audio object coding, parametric stereo coding of low delay, low delay spatial / surround encoding, hiding frame loss, echo cancellation or other corresponding applications.
The basic idea behind the SBR is the observation that there is usually a strong correlation between the characteristics of a high frequency range of a signal, which will be called a high band signal, and the characteristics of the low band frequency range, still called low bandwidth or low bandwidth signals of the same signal. Therefore, a good approximation of the representation of the original high band input signal can be obtained by transposing the low band to the high band.
In addition to transposition, the reconstruction of the high band incorporates the conformation of the spectral envelope, which includes an adjustment of the qanhos. This process is typically controlled by a transmission of the high-band spectral envelope of the original input signal. Other guidance information can also be sent by the synthesis modules of the encoder control, such as reverse filtering, an addition of noise and sine to adapt to the audio material when only the transposition may not be sufficient. The corresponding parameters comprise the high noise band parameters for the addition of noise and the high tone band parameter for the addition of sine. This guidance information is usually referred to as SBR data.
The SBR process can be combined with any conventional waveform or codec through pre-processing on the encoder side and post-processing on the decoder side. The SBR encodes the high frequency portion of an audio signal at a very low cost, whereas the audio codec is used to encode the lower frequency portion of the signal.
On the encoder side, the original input signal is analyzed, the high band spectral envelope and its characteristics with respect to the low band are encoded and the resulting SBR data is multiplexed with a bit stream from the codec to the low band. On the decoder side, the SBR data is first demultiplexed. The decoding process is usually organized in stages. First, the core decoder generates the low band, and second, the SBR decoder operates a post processor using the decoded SBR data to guide the spectral band replication process. It is then obtained as a total bandwidth output signal.
In order to obtain the highest possible coding efficiency, and to keep computational complexity low, extended SBR codecs, such as so-called dual index systems, are generally implemented. dual index means that the limited bandwidth core codec is operating at half the external audio sampling index. In contrast, the SBR part is processed at the total sampling frequency.
Fig. 5 shows a schematic block diagram of an SBR 500 system. The SBR 500 system comprises, for example, an AAC-LD encoder (AAC-LD = Advanced Low Delay Audio Codec) 510 and an SBR 520 encoder, to which the audio data to be processed are supplied in parallel. The SBR 520 encoder comprises a bank of analysis filters 530, which is shown in Fig. 5 as a bank of analysis filters QMF. The analysis filter bank 530 is capable of providing subband audio values corresponding to the subband based on the audio signals supplied to the SBR 500 system. These subband audio values are then sent to an extraction module of SBR 540 parameters, which generates the SBR data as previously mentioned, for example, comprising the high band spectral envelope, the high band noise parameter and the high band tone parameter. This SBR data is then sent to the AAC-LD 510 encoder.
The AAC-LD 510 encoder is shown in Fig. 5 as a dual index encoder. In other words, the encoder 510 operates at half the sampling frequency compared to the sampling frequency of the audio data provided to the encoder 510. To facilitate this, the AAC-LD 510 encoder comprises a sub-sampling stage 550, which can optionally understand a low-pass filter to avoid distortions caused, for example, a violation of Theory
Nyquist-Shannon. The sub-sampled audio data as sent by the sub-sampling stage 550 is then supplied to an encoder 560 (analysis filter bank) in the form of an MDCT filter bank. The signals provided by the encoder 560 are then quantized and encoded in the quantization and encoding stage 570. Furthermore, the SBR data provided by the SBR 540 parameter extraction module is also encoded to obtain a bit stream, which will then be sent by the ACC-LD 510 encoder. The quantization and encoding stage 570 can, for example, quantize the data according to the hearing properties of the human ear.
The bit stream is then sent to an AAC-LD 580 Decoder, which is part of the decoder side to which the bit stream is transported. The AAC-LD decoder comprises a 590 decoding and de-quantizing stage, which extracts the SBR data from the bit stream and de-quantized or re-quoted audio data in the frequency domain representing the low band. The low band data is then sent to a synthesis filter bank 600 (reverse MDCT filter bank). The reverse MDCT stage (MDCT<sup>-1</sup>) 600 converts the signals provided to the MDCT stage inverse of the frequency domain in the time domain to provide the time signal. This time-domain signal is then supplied to the SBR 610 decoder, which comprises an analysis filter bank 620, which is shown in Fig. 5 as a QMF analysis filter bank.
The analysis filter bank 620 makes a spectral analysis of the time signal supplied to the analysis filter bank 620 which represents the low band. This data is then supplied to a 630 high frequency generator, which is also called an HF generator. Based on the SBR data provided by the AAC-LD 580 encoder and its 590 decoding and dequantization stage, the HF 630 generator generates the high band based on the low band signals provided by the 620 analysis filter bank. Both the low band and high band signals are then fed to a synthesis filter bank 640, which transfers the low band and high band signals from the frequency domain to the time domain to provide a form of audio output in the time domain for the SBR 500 system.
For totality, it should be noted that in many cases the SBR 500 system as shown in Fig. 5 is not implemented in this way. To be more exact, the AAC-LD 510 encoder and the SBR 520 encoder are usually implemented on the encoder side, which is usually implemented separately on the decoder side comprising the AAC-LD 580 decoder and the SBR 610 decoder. the system 500 shown in Fig. 5 essentially represents the connection of two systems, that is, of an encoder comprising the aforementioned encoders 510, 520 and a decoder comprising the aforementioned decoders 580, 610.
The configurations according to the present invention in the form of analysis filter banks 100 and synthesis filter banks 300 can, for example, be implemented in the system 500 shown in Fig. 5, as a replacement for the analysis filter bank 530, analysis filter bank 620 and synthesis filter bank 640. In other words, synthesis filter banks or SBR component analysis system 500 can, for example, be replaced by the corresponding configurations according to the present invention. Furthermore, the MDCT 560 and the MDCT inverse 600 can also be replaced by low-delay synthesis and analysis filter banks, respectively. In this case, if all the described substitutions have been implemented, the so-called low delayed extended AAC codec (codec will be completed.
encoder-decoder)
The extended low-delay AAC (AAC-ELD) aims to combine the low-delay characteristics of an AAC-LD (Advanced Audio Codec - Low-delay) with high HE coding efficiency -AAC (High Efficiency Advanced Audio Codec) using SBR with AAC-LD. The SBR 610 decoder acts in this scenario as a post-processor, which is provided after the core 580 decoder including a complete analysis filter bank and a 640 synthesis filter bank. Therefore, the components of the SBR 610 decoder add another decoding delay. , which is illustrated in Fig. 5, by shading the components 620, 630, 540.
In many implementations of SBR 500 systems, the lower frequency part or lower band ranges typically from 0 kHz to typically 5-15 kHz are encoded using a waveform encoder, called the core codec. The core codec can, for example, be one of the MPEG audio codec family. In addition, the reconstruction of the high frequency or high band part is achieved by a transition from the low band. The combination of SBR with a core encoder is, in many cases, implemented as a dual index system, where the underlying AAC encoder / decoder is operated at half the sampling rate of the SBR encoder / decoder.
Most of the control data is used for the representation of the spectral envelope, which has varying resolution of time and frequency in order to be able to control the SBR process as best as possible with the least possible bit rate overhead. The other control data mainly seek to control the tonal-to-noise index of the high band.
As shown in Fig. 5, the output of the underlying AAC decoder 580 is typically analyzed with a 32-channel QMF filter bank 620. Then, the HF 630 generator module recreates the high band by patching QMF sub-bands from from the existing low band to the high band. In addition, reverse filtering is done on a subband basis, based on the control data obtained from the bit stream (SBR data). The envelope adjuster modifies the spectral envelope of the regenerated high band and adds other components such as noise and sine waves are added according to the control data in the bit stream. Since all operations are done in the frequency domain (also known as QMF or subband domain), the final step of decoder 610 is a QMF 640 synthesis to retain the time domain signal. For example, in the case where the QMF analysis on the encoder side is done in a 32 QFM subband system for 1024 samples in the time domain, the high frequency reconstruction results in 64-QMF subband in which the synthesis is done producing 2048 samples in the time domain, in order to obtain an oversampling by a factor of 2.
In addition, the delay of the core 510 encoder is doubled by operating at half the original sampling rate in the dual mode index, which gives rise to other sources of delay in both, in the process of encoding and decoding an AACLD in combination with SBR. These sources of delay are then examined and their associated delays are minimized.
Fig. 6 shows a simplified block diagram of the system 500 shown in Fig. 5. Fig. 6 concentrates the delay sources in the encoder / decoder process using SBR and low delay filter banks for encoding. Comparing Fig. 6 with Fig. 5, the MDCT 560 and the inverse MDCT 600 have been replaced by the delay-optimized modules, the so-called low-delay MDCT 560 '(LD MDCT) and the low-delay inverse MDCT 600' (LD IMDCT). Furthermore, the HF 630 generator has also been replaced by the optimized delay module 630 '.
Apart from the low-delay MDCT 560 'and the low-delay inverse MDCT 600', a modified SBR framing and a modified HF generator 630 'are employed in the system shown in Fig. 6. To avoid framing delay other than an encoder / decoder core 560, 600 and the respective SBR modules, the SBR framing is adapted to fit the framing length of 480 or 512 samples of the AAC-LD. In addition, the variable time grid of the HF 630 generator, which implies 384 delay samples, is restricted with reference to the dissemination of SBR data in the adjacent AC-LD frames. Therefore, the only remaining sources of delay in the SBR module are the filter banks 530, 620 and 640.
According to the situation shown in Fig. 6, which represents a partial implementation of the AAC-ELD codec, some delay optimizations have already been implemented, including the use of a low delay filter bank in the AAC-LD core and the removal of a previously mentioned SBR overpass. For further delay improvements, the remaining modules should be investigated. Fig. 6 shows the delay sources in the encoder / decoder process using SBR and the low delay filter banks named LD-MDCT and LD-IMDCT in the present. Compared with Fig. 5, in Fig. 6 each box represents a delay source, where the delay optimizer modules are drawn in shading. So far, similar models have not been optimized for low delay.
Fig. 7a illustrates a flowchart comprising a pseudo C- or C + + code to illustrate a configuration according to the present invention in the form of an analysis filter bank or a corresponding method for generating audio subband values in subband audio channels. For even more precise, Fig. 7a represents a flow chart of a complex value analysis filter bank for 32 bands.
As pointed out earlier, the analysis filter bank is used to divide the signal in the time domain, for example, sent from the core encoder into N = 32 subband signals. The output of the filter bank and the subband samples or audio subband values have, in the case of a complex value analysis filter bank, complex values and therefore super-sampled by a factor of 2, compared to a bank of real value filters. The filtering involves and comprises the following steps, where a set x (n) comprises exactly 320 samples in the time domain. The higher the index of samples n in the set, the older the samples are.
After starting the method settings in step S100, first, the samples in the set x (n) are changed from 32 positions in step S110. The 32 oldest samples are discarded and 32 new samples are saved in positions 31 to 0 in step S120. As shown in Fig. 7a, the audio samples in the input time domain are stored in positions corresponding to a decreasing n index in the range of 31 to 0. This results in a time reversal of the samples stored in the corresponding frame or vector, so that the reversal of the window function index, to obtain an analysis window function based on the (equally long) synthesis window function, has already been done.
During step S130, window coefficients ci (j) are obtained by linear interpolation of the coefficients c (j) based on equation (15). The interpolation is based on a block size (block length or number of subband values) of N = 64 values and is based on a frame comprising T = 10 blocks. Thus, the index of the window coefficients of the interpolated window function are in the range between 0 and 319 according to equation (15). The window coefficients c (n) are given in the table in Annex 1 of the description. However, depending on the implementation details, to obtain the window coefficients based on the values given in the tables in Annexes 1 and 3, other changes in signs with reference to the window coefficients corresponding to indexes 128 to 255 and 384 to 511 (multiplication by factor (-1)) must be considered.
In such cases, the window coefficients w (n) or c (n) to be used can be obtained according to w (n) = w<sub>table</sub>(n) s (n) (16a) with the signal change function s (n) according to s (n) = + 1
V for 128 <n <255 and 384 <n <51 f or (16b) for n = 0 to 639, where w<sub>table</sub>(n) are the values given in the tables in the Annexes.
However, the window coefficients do not have the necessary implementation according to the table in Annex 1 to obtain, for example, the reduction already described for the delay. To obtain this delay reduction, maintaining the quality level of the processed audio data, or to obtain another decision, the window coefficients c (n) for the window coefficient index n in the range between 0 and 639, can obey a of the sets of relationships given in one of Annexes 2 to 4. Furthermore, it should be noted that other window coefficients c (n) can also be employed in configurations in accordance with the present invention. Of course, also other window functions comprising a different number of window coefficients that 320 or 640 can be implemented, although the tables in Annexes 1 to 4 only apply to window functions having 640 window coefficients.
A linear interpolation according to S130 leads to a significant improvement in quality and reduction or cancellation of aliasing effects in the case of the window function comprising an even number of window coefficients. It should also be noted that the complex unit is not j as in equations (1), (2) and (16), but indicated by i = V<sup>—</sup> 1 .
In step S140, the samples of the set x (n) then have their elements multiplied by the coefficients ci (n) of the interpolated window.
In step S150, the windowed samples are added according to the equation given in the flowchart in Fiq. 7a to create the set of 64 elements u (n). In step S160, 32 new subband samples or audio subband values W (k, l) are calculated according to the Mu matrix operation, where the elements of the M matrix are given by
M (k, n) = 2 · exp i · π · (k + 0.5) · (2 · n - 95t [0 <k <32 (17) where exp () indicates the complex exponential function and, as mentioned earlier, i is the imaginary unit. Before the end of the loop of a flowchart with step S170, each of the subband values W (k, l) (= W [k] [1]) can be sent, which corresponds to the sample of sub-band 1 in the sub-band with the index k. In other words, each loop in the flowchart shown in Fig. 7a produces 32 subband values of complex value, each representing the output of a subband filter bank.
Fiq. 7b illustrates the step S150 of the collapse of frame 150 of the audio samples windowed in the time domain comprising 10 blocks 160-1, ..., 160-10 of audio samples windowed in the time domain z (n) for the vector u (n) 5 times, with the sum of blocks of frame 150 each. The collapse or retraction is done based on the elements, so that the audio samples windowed in the time domain corresponding to the same sample index within each of the blocks 160-1, 160-3, 160-5, 160- 7 and 160-9 are added to obtain the corresponding value in the first blocks 650-1 of the vector u (n). Likewise, based on blocks 160-2, 160-4, 160-6, 160-8 and 160-10 the corresponding elements of the vector u (n) in block 160-2 are generated in step S150.
Another configuration according to the present invention in the form of an analysis filter bank can be implemented as a complex 64-band low-delay filter bank. The processing of this complex low-delay filter bank as an analysis filter bank is basically similar to the analysis filter bank described in the context of Fig. 7a. Due to the similarities and basically the same processing described in the context of Fig. 7a, the differences between the complex analysis filter bank described for 32 bands of Fig. 7a and the complex analysis filter bank for 64 sub-bands will be highlighted here.
In contrast to the 32 sub-bands comprising the analysis filter bank shown in Fig. 7a, the frame vector x (n) comprises, in the case of a 64-band analysis filter bank 640 elements having indexes of 0- 639. Thus, step S110 is modified so that the samples in the set x (n) are changed in 64 positions, where the oldest 64 samples are discarded. In step S120 instead of 32 new samples, 64 new samples are saved in positions 63 to 0. As shown in Fig. 7c, the audio samples in the input time domain are stored in positions corresponding to a decreasing index n in the range of 63 to 0. This results in a time reversal of the samples saved in the corresponding frame or vector, so that the reversal of the window function index to obtain an analysis window function based on the (equally long) synthesis window function has already been done.
Since the window c (n) used for windowing the elements of the vector of frame x (n), typically comprises 640 elements, the step S130 of linear interpolation of the window coefficients to obtain the interpolated windows ci (n) can be omitted .
Then, during step S140, the samples in the set x (n) are multiplied or windowed using the sequence of window coefficients c (n), which are based once again on the values in the table in Annex 1. In the case where the window coefficients c (n) are those of the synthesis window function, a window or multiplication of the set x (n) is made by the window c (n), according to the equation z (n) = x (n) · C (n) (18) for n = 0,. .., 639. Again, to obtain the low-delay properties of the window function, it is not necessary to implement the window function exactly according to the window coefficients based on the values given in the table in Annex 1. For many applications, an implementation in which the coefficients of the window obey each set of relations given in the tables in Annexes 2 to 4 will be sufficient to obtain an acceptable decision between quality and the significant reduction of the delay. However, depending on the implementation details, to obtain the window coefficients based on the values given in the tables in Annexes 1 and 3, other signal changes with reference to the window coefficients corresponding to indexes 128 to 255 and 384 to 511 (multiplication by factor (-1)) must be considered according to equations (16a) and (16b).
The step S150 of the flowchart shown in Fig. 7a is then replaced by the sum of the samples of the vector of frame z (n) according to the equation = Σ U + J · 128) <sub>(19) </sub>j = 0
To create the set of 128 u (n) elements.
The step S160 of Fig. 7a is then replaced by a step in which 64 new subband samples are calculated according to the matrix operation Mu, where the matrix elements of the matrix M are given by. . (i · π · (k + 0.5) · (2 · η - 19 lh 0 <k <64
M (k, n) = 2 · exp -----------------— -------------, í <128) 0 <n < 128 (20) where exp () indicates the complex exponential function and i is, as explained, the imaginary unit.
Fig. 7c illustrates a flow chart according to a configuration of the present invention in the form of a real value analysis filter bank for 32 subband channels. The configuration shown in Fig. 7c does not differ significantly from the configuration shown in Fig. 7a. The main difference between the two configurations is that step S160 of the calculation of the new 32 audio values of subband of complex value is replaced in the configuration shown in Fig. 7c by a step S162 in which the 32 real value subband audio samples are calculated according to the M matrix operation<sub>r</sub>u, where the elements of the matrix M<sub>r</sub> are given by, <sub>λ o</sub> (π · (k + 0.5) · (2 · n - 9 5) ú% <k <3 2
M (k, n) = 2 · cos -------------—------------ L <sup>r</sup> 64) 0 <n <64 (21)
As a consequence, each flowchart loop produces 32 subband samples of real value where W (k, l) corresponds to subband 1 audio samples of subband k.
The real value analysis filter bank can, for example, be used in the structure of a low power mode of an SBR system, as shown in Fig. 5. The low power mode of the SBR tool differs from the high power SBR tool quality, mainly with reference to the fact that banks of real value filters are used. This reduces computational complexity and computational work by a factor of 2, so that the number of operations per unit of time is essentially reduced by a factor of 2, since no imaginary part must be calculated.
The new filter banks proposed in accordance with the present invention are fully compatible with the low power mode of SBR systems. Therefore, with filter banks according to the present invention, SBR systems can still operate in both normal and high quality mode with complex filter banks and in low power mode with real value filter banks. The real-value filter bank can, for example, be obtained from the complex filter bank using only the real values (contributions modulated by cosine) and omitting the imaginary values (contributions modulated by sine).
Fig. 8a shows a flow chart according to a comparative example of the present invention in the form of a complex value synthesis filter bank for 64 subband channels. As mentioned earlier, the synthesis filtering of the subband signals processed by SBR is done using a synthesis filter bank with 64 sub-bands. The output of the filter bank is a sample block in the real-time time domain as highlighted in the context of Fig. 1. The process is illustrated by the flowchart in Fig. 8a, which also illustrates a comparative example in the form of a method for generating audio samples in the time domain.
The synthesis filtering comprises after the start (step S200), the following steps, where a set v comprises 1280 samples. In step S210, the samples in set v are changed in 128 positions, where the oldest 128 samples are discarded. In step S220, the 64 new complex sub-band audio values are multiplied by an N matrix, where the elements of matrix N (k, n) are given by
N ^ k, n) = - · expf 6 4 '· (k + 0.5) · (2 · n - 63) j í 0 <k <64
128 / [The <n <128 (22) where exp () indicates the complex exponential function and i is the imaginary unit. The actual part of the output from this operation is stored at position 0-127 of set v, as shown in Fig. 8a.
In step S230, the samples, which are now in the time domain, are extracted from the set v according to the equation given in Fig. 8a to create a set of 640 elements g (n). In step S240, the real value samples in the time domain of the set g are multiplied by the window coefficient c (n) to produce a set w, where the window coefficients are again the window coefficients based on the values given in the table of the Annex 1.
However, as noted earlier, the window coefficients need not be based exactly on the values given in the table in Annex 1. It is sufficient, in different comparative examples, if the window coefficients satisfy one of the sets of relationships given in the tables in Annexes 2 to 4, to obtain the desired low-delay property of the synthesis filter bank. Furthermore, as explained in the context of the analysis filter bank, other window coefficients can also be used in the structure of the synthesis filter bank. However, depending on the implementation details, in order to obtain the window coefficients based on the values stored in the tables in Annexes 1 and 3, other changes in signs with reference to the window coefficients corresponding to indexes 128 to 255 and 384 to 511 (multiplication by factor (-1)) must be considered.
In step S250, 64 new output samples are calculated by summing the samples from the set w (n) according to the last step and the formula given in the flowchart of Fig. 8a, before a loop of a flowchart ends in step S260. In the flowchart as shown in Fig. 8a, X [k] [1] (= X (k, l)) corresponds to the value of audio subband 1 in the subband having the index k. All new loops shown in Fig. 8a produce 64 real-time audio samples of real value as output.
The implementation as shown in Fig. 8a of a 64-band complex value analysis filter bank does not require an overlap / addition buffer comprising various storage positions as explained in the context of the configuration shown in Fig. 2b. Here, the overlap / addition buffer is hidden in the veg vectors, which is calculated based on the values stored in the vector v. The overlap / addition buffer is implemented in the structure of these vectors with these indexes being greater than 128, so that the values correspond to the values of the previous or past blocks.
Fig. 8b illustrates a flow chart of a real value synthesis filter bank for 64 real value audio subband channels. The real value synthesis filter bank according to Fig. 8b can also be implemented in the case of a low power SBR implementation as a corresponding SBR filter bank.
The flowchart of Fig. 8b differs from the flowchart of Fig. 8a, mainly with reference to step S222, which replaces S220 of Fig. 8a. In step S222, the 64 new real-value audio subband values are multiplied by an N matrix<sub>r</sub>, where the elements of the matrix N<sub>r</sub>(k, n) are given by
N<sub>r</sub>(k, n) = (π · (k + 0.5) · (2 · n - 63) 'cos ---------------------- k 128.
'0 <k <64 <n <128 (23) where the output of this operation is again stored in the positions
0-127 of set v.
In addition to these modifications, the flowchart as shown in Fig. 8b in the case of a real value synthesis filter bank for low power SBR mode, does not differ from the flowchart shown in Fig. 8a of the complex value synthesis filter bank for high quality SBR mode.
Fig. 8c illustrates a flow chart according to a configuration of the present invention in the form of a bank of sub-sampled complex value synthesis filters and the suitable method that can, for example, be implemented in a high quality SBR implementation. . To be more precise, the synthesis filter bank described in Fig. 8c refers to a complex value synthesis filter bank capable of processing complex value audio subband values for 32 subband channels.
The sub-sampled synthesis filtering of the sub-band signals from the SBR process is done using a 32-channel synthesis filter bank as illustrated in Fig. 8c. The output of the filter bank is a block of samples in the real-time time domain. The process is given in the flow chart of Fig. 8c. The synthesis filtering comprises after the start (step S300), the following steps, where the set v comprises 640 samples in the real time domain.
In step S310, the samples in set v are changed in 64 positions, where the oldest 64 samples are discarded. Then, in step S320, the new 32 complex subband samples or complex audio subband values are multiplied by an N matrix, the elements of which are given by, x 1 íi π · (k + 0.5) · (2 n - 31) j Í0 <N (k, n) = - · exp ---------------—---------- L
Λ 64 0 <
k <32 n <64 (24) where exp () indicates the complex exponential function and i is again from that operation is then saved in positions 0-63 of set v.
In step S330, samples are extracted from vector v according to the equation given in the flowchart of Fig. 8c to create a set of 320 g elements. In step S340, the window coefficients ci (n) of an interpolated window function are obtained by a linear interpolation of the coefficients c (n) according to equation (15), where the index n is again in the range between 0 and 319 ( N = 64, T = 10 for equation (15)). As illustrated earlier, the coefficients of the window c (n) function are based on the values given in the table in Appendix 1. Furthermore, to obtain the low-delay property as previously obtained, the window coefficients c (n) need not be exactly the data provided in the table in Annex 1. It is sufficient that the window coefficients c (n) obey at least one set of relationships given in Annexes 2 to 4. However, depending on the implementation details, to obtain the window coefficients based on the values given in the tables in Annexes 1 and 3, other changes of signs with reference to the window coefficients corresponding to the indexes 128 to 255 and 384 to 511 (multiplication by factor (-1)) must be considered according to equations (16a) and (16b). Furthermore, also different window functions comprising different window coefficients c (n) can of course be employed in configurations of the present invention.
In step S350, the samples of the set g are multiplied by the interpolated window coefficient ci (n) of the interpolated window function to obtain the sample windowed in the time domain w (n).
Then, in step S360, 32 new output samples are calculated by a sum of samples from the set w (n) according to the last step S360, before the final step S370 in the flowchart of Fig. 8c.
As previously indicated, in the flowchart of Fig. 8c, X ([k] [l]) (= x (k, l)) corresponds to an audio subband value 1 in the audio subband channel k. Furthermore, all new loops in a flowchart as shown in Fig. 8c produce 32 samples in the real time domain as an output.
Fig. 8d shows a flowchart of a configuration according to the present invention in the form of a sub-sampled real-value synthesis filter bank that can, for example, be used in the case of a low-level SBR filter bank. power. The configuration and flowchart shown in Fig. 8d differ from the flowchart shown in Fig. 8c of the sub-sampled complex value synthesis filter bank only with reference to step S320, which is replaced in the flowchart shown in Fig. 8d by step S322.
In step S322, the 32 new real-value audio subband values, or subband samples, are multiplied by the matrix N<sub>r</sub>, where the elements of the matrix N<sub>r</sub> are given by
N<sub>r</sub>(k, n) (25) where the output of this operation is stored at position 0 to 64 of set v.
Fig. 9a shows an implementation of a comparative example in the form of a method corresponding to a complex value analysis filter bank for 64 sub-bands. Fig. 9a shows an implementation as a MATLAB implementation, which outputs a vector y and a state vector. The function defined in this inscription shown in Fig. 9a is called LDFB80 for which a vector x comprising new audio samples and the state vector is provided as an output. The function name LDFB80 is an abbreviation for low delay filter bank for 8 blocks extending in the past and 0 blocks in the future.
In the MATLAB programming language, the percent sign (%) indicates observations, which are not made, but which only serve to comment and illustrate the source code. In the following description, different segments of the source code will be explained with reference to their functions.
In a S400 code sequence, the buffer that is represented by the state vector is updated so that the contents of the state vector having indexes 577 to 640 are replaced by the content of vector x comprising the new audio input samples in the time domain . In an S410 code sequence, the window coefficients of the analysis window function as stored in the LDFB80_win variable are transferred to the win_ana vector.
In step S420, which assumes that the last samples are aligned on the right side of the buffer, the actual window is made. In block S420, the content of the state vector is multiplied in elements (. *) By the elements of the win_ana vector comprising the analysis window function. The product of this multiplication is then stored in the x_win_orig vector.
In step S430, the content of the x_win_orig vector is reshaped to form a matrix with a dimension of 128 · 5 elements called x_stack. In step S440, the signal change of the x_stack stack is made with reference to the second and fourth columns of the x_stack matrix.
In step S450, the x_stack stack collapses or is retracted by adding the elements of x_stack with reference to the second index and simultaneously reversing the order of the elements and transposing the result before again storing the result in the various X-stacks.
In the S460 code segment, the transformation of the time domain into the frequency domain is done by computing a complex Fast Fourier transformation (FFT) of the content multiplied by elements of the x_stack stack multiplied by the complex exponential function for which the argument is provided (—I · π · n / 128), by the indexes and in the range 0 to -127 and the imaginary unit i.
In the code segment S470, a post-turn is performed by defining the variable m = (64 + l) / 2 and calculating the block comprising the audio subband values as a vector y according to the equation yt) = 2 · Temp (k) exp (- 2i π ((k - 1 + ^)). <sub>(26)</sub>
The k index covers the range of integers 1-64 in the implementation shown in Fig. 9a. The vector y is then produced as the vector or block comprising the audio subband values 180 of Fig. 1. The bar above the second factorization equation (26), as well as the coding segment conj () of the function S417 in Fig. 9a refer to the complex conjugate of the argument of the respective complex number.
In a final S480 segment code, the state vector is changed into elements. The state vector in its altered form can then be supplied to the LDFB80 function as input again in another function loop.
Fig. 9b shows a MATLAB implementation according to a configuration of the present invention in the form of a method corresponding to a complex value analysis filter bank for 32 sub-bands. Likewise, the defined function is called LDFB80_32, indicating that the implementation represents a low-delay filter bank for 32 sub-bands based on another overlap of 8 blocks in the past and 0 blocks in the future.
The implementation of Fig. 9b differs from the implementation shown in Fig. 9a, only with reference to a few code strings, as will be highlighted in the description below. The sequence of codes S400, S430, S460, S470 and S480 is replaced by a corresponding sequence of codes S400 ', S430', S460 ', S470' and S480 'taking into account mainly the fact that the number of sub-bands, or the number of subband value outputs by the LDFB80_32 function, is reduced by a factor of 2. Likewise, step S400 'refers to the state vector being updated with reference to the last 32 entries corresponding to indices 289 to 320 with the corresponding 32 audio input samples in the time domain of the new block 220 as shown in Fig. 1 .
However, the main difference between the implementations as shown in Figs. 9a and 9b appears in a code sequence S410 of Fig. 9a, which is replaced by a code sequence S412 in the implementation shown in Fig. 9b. The code sequence for S412 in Fig. 9b first comprises a copy of the 640 window coefficients comprising windows stored in the LDFB80_win vector in the local win_ana vector. Then, an interpolation occurs according to equation (15), in which two consecutive window coefficients represented by the vector elements of the win_ana vector are added and divided by 2 and then saved back in the win_ana vector.
The next code sequence S420 is identical to the code sequence S420 shown in Fig. 9a, which performs the actual multiplication in elements (. *) Of a window of the values, or elements, of the state vector by the elements of the win_ana vector comprising the coefficients interpolated windows of the interpolated window function. The product of this operation is stored in the x_win_orig vector. However, the difference between the S420 code sequence of Fig. 9b and the corresponding S420 code sequence of Fig. 9a, is that in the case of Fig. 9b, not 640, but only 320 multiplications are made in the window structure.
In the code sequence 3430 'in the replacement of the code sequence S430, the stack x_stack is prepared by reconformation of the vector X — win_orig. However, since the X_win_orig vector only comprises 320 elements, compared to the corresponding vector in Fig. 9a comprising 640 elements, the x_stack matrix is a matrix of only 64 · 5 elements.
The code sequence 3440 of the signal exchange and the code sequence S450 of the battery collapse are identical in both implementations according to Figs. 9a and 9b, not considering the reduced number of elements (320 compared to 640).
In a code sequence S460 ', the replacement in a code sequence S460 of a fast complex Fourier Transform (FFT) is odd of the window data, which is very similar to the transform of the code sequence 3460 of Fig. 9a. However, again, due to the reduced number of audio subband values output, the temp vector is provided with the result of a Fast Fourier Transform, the multiplication in elements by the elements of the x_stack stack and the complex exponential function of the argument (-i · π · n / 64), where the index n is in the range between 0 and 63.
Then, in the modified code sequence S470 ', the post-rotation is done by defining the variable m = (32 + 1) / 2 and generating the output vector y according to equation (26), where the index k only covers the range 1 to 32 and where the number 128 that appears in the arch of the complex exponential function is replaced by the number 6 4.
In the final sequence code S480 ', the state of the buffer is changed by 32 elements in the case of the implementation shown in Fiq. 9b, where in the corresponding code sequence S480, the buffer is changed by 64 elements.
Fig. 10a shows a MATLAB inscription illustrating an implementation according to a comparative example in the form of a method corresponding to a complex value synthesis filter bank for 64 sub-bands. The inscription shown in Fig. 10a defines the function ILDFB80 to which the vector x representing the block 320 of audio subband values of Fig. 2a and the state state vector are provided as input parameters. The name ILDFB80 indicates that the defined function is a low inverse delay filter bank corresponding to 8 blocks of audio data from the past and 0 blocks of the future. The function provides a vector y and a new or redefined state state vector as an output, where vector y corresponds to block 410 of audio samples in the time domain of Fig. 2a.
In a sequence of code S500, a pre-rotation is performed, in which a variable m = (64 + l) / 2 is defined, as well as a temp vector. The temp (n) elements of the temp vector are defined according to the equation temp (n) = j · x (n) · exp (2i · π (η - 1 + f) ·<sub>2?)</sub> where the bar above the vector element x (n) and the function conj () represents the complex conjugate, exp () represents the complex exponential function, i represents the imaginary unit and n is an index in the range 1 - to 64.
In a code sequence S510, the temp vector is spent in the matrix comprised in the first column of the elements of the temp vector and in the second column, the complex conjugate of the reverse temp vector with reference to the order of the elements defined by the vector index. Thus, in an S510 code sequence, an odd symmetry of the temp matrix is established based on the temp vector.
In an S520 code sequence, a unique Fast Fourier Transform (FFT) is made based on the temp matrix. In this code sequence, having the real part of the multiplication in elements of the result of the inverse Fourier Transform of the temp matrix by the exponential function having the argument of (i · π / 128) and sent to a vector y_knl, where the index n is in range 0 to 127.
In an S530 code sequence, an extension of the data and an alternate signal change are made. Therefore, the order of the elements of the vector y_knl is reversed and at the same time a change of sign is made. Then, the tmp matrix is defined, comprising the first, third and fifth columns of the vector y_knl, where the second and fourth columns comprise a vector y_knl of exchanged sign.
In a sequence of code S540, the window coefficients stored in the vector LDFB80_win are first copied to the vector win_ana. Then, the synthesis window coefficients are determined based on the analysis window coefficients and stored in the win_ana vector generating a reverse time version of the analysis window function according to win _ syn (n) = win _ ana ^ NT - n) (28) where N · T is the total number of window coefficients and n is the index of the window coefficients.
In an S550 code sequence, the synthesis window is applied to the vector tmp by a multiplication in elements of the vector by the synthesis window function. In a sequence of code S560, the buffer is updated by establishing the elements of the state vector at indexes 577 to 640 to 0 and adding the contents of the tmp vector with the state vector.
In an S570 code sequence, the output vector y that comprises the audio samples in the time domain is extracted from the state vector by extracting elements from the state vector, extracting the elements from the state vector with indexes 1 to 64.
In an S580 code sequence, the final code sequence of the function shown in Fig. 10a, the state state vector is changed by 64 elements, so that elements with indexes from 65 to 640 are copied from the first 576 elements of the state vector.
Fig. 10b shows a MATLAB inscription of an implementation according to a configuration of the present invention in the form of a complex value synthesis filter bank for 32 subband values. The name of the function defined by the inscription shown in Fig. 10b illustrates this as a defined function called ILDFB80_32 indicating that the defined function is a low delay inverse filter bank for 32 bands with 8 overlapping blocks from the past and 0 overpassing blocks from the future.
As discussed with reference to the comparison of the implementation shown in Figs. 9a and 9b, the implementation according to the inscription of Fig. 10b is also closely related to the implementation of the synthesis filter bank with 64 sub-bands according to Fig. 10a. As a consequence, the same vectors are provided for the function and are sent by the function, which, however, comprises only half the number of elements compared to the implementation of Fig. 10a. The implementation of a bank of synthesis filters from 32 bands to 32 bands differs from the version of 64 sub-bands illustrated in Fig. 10a, mainly with reference to two aspects. The code sequence S500, S510, S520, S53B, S560, S570 and S580 is replaced by the code sequence in which the number of elements to be considered and the number of parameters relating to the elements are divided by 2. Furthermore, the S540 code sequence for the generation of a synthesis function is replaced by an S542 code sequence, in which the synthesis function is generated as a linearly interpolated synthesis function according to equation ( 15).
In a code sequence S500 'that replaces a code sequence S500, the variable m is defined as equal to m = (32 + 1) / 2 and the vector temp is defined according to equation (27), where the index n only covers the range from 1 to 32 and where the factor of 1/128 is replaced by the factor 1/64 in the exponential function argument.
Likewise, in a 3510 'code sequence that replaces a 3510 code sequence, the index range only covers the indexes of the 32 elements comprising the temp vector. In other words, the index only covers the values from 1 to 32. Likewise, in a code sequence S520 'that replaces a code sequence 3520, the exponential function argument is replaced by (i · π · n / 64 ), where the index n is in the range 0 to 63. In the structure of a code sequence S530 ', the index range is also reduced by a factor of 2 when compared to a code sequence 3530.
The code sequence S542 that replaces a code sequence S540 of Fig. 10a also copies a window function like the one stored in the vector LDFB80_win to vector win_ana and generates a reverse version in time win_syn according to equation (28). However, the S542 code sequence of the implementation shown in Fig. 10b further comprises an interpolation step according to equation (15), in which for each element of the redefined vector win_syn comprising the window coefficients of the synthesis window function, a linear interpolation of two consecutive window coefficients of the window function of original synthesis.
The 3550 code sequence for applying the window to the tmp vector and replacing the tmp elements with its windowed version is identical in terms of code to a direct comparison of the respective code sequence in Figs. 10a and 10b. However, due to the smaller size of the tmp vector in the implementation of Fig. 10b, during an implementation, only half the number of multiplications is done.
Also in the structure of a code sequence 3560 ', 3570' and S580 'replacing the code sequence S560, S570 and S580, respectively, indexes 640 and 64 are replaced by 320 and 32, respectively. Therefore, these three final code sequences only differ from the code sequence of the implementation shown in Fig. 10a with reference to the size of the states vector tmp and y.
As the configurations have illustrated so far, the analysis window as well as the synthesis window are adapted to the window of the respective samples in the time domain included in the respective frames by multiplying these in the base of elements by the window coefficients of the window function.
Before describing the window function, which can be used, for example, as a synthesis window function and as an analysis window function in its reverse version over time, the advantages of the configurations according to the present invention will be emphasized. in greater detail, especially in view of an implementation in the structure of an SBR tool or system as shown in Figs. 5 and 6.
Among the advantages, configurations according to the present invention and systems comprising more than one configuration according to the present invention can offer a significant reduction of the delay according to other filter banks. However, this low-delay property will be seen in the context of Figs. 13 and 14 in greater detail. An important aspect of this context is to note that the length of the window function, in other words, the number of window coefficients to be applied to a frame or sample block in the time domain is independent of the delay.
The configurations according to the present invention offer another advantage of improving the quality of (reconstructed) audio data. The interpolation employed in configurations according to the present invention offers a significantly reduced aliasing compared to other reduction schemes referring to the number of window coefficients.
Furthermore, as will be highlighted in the context of Figs. 17 and 18 in more detail, in terms of psychoacoustics, the configurations according to the present invention generally use the temporal masking properties of the human ear better than many other filter banks. Furthermore, as will be better emphasized in the context of Figs. 15, 16 and 19, the configurations according to the present invention offer an excellent frequency response.
Also, in many filter banks according to a configuration of the present invention, perfect reconstruction is obtainable if the analysis filter bank and the synthesis filter bank are interconnected. In other words, the configurations according to the present invention not only offer an indistinctly audible output compared to the input of an interconnected set of an analysis filter bank and a synthesis filter bank, as (not considering quantization errors, computational rounding effects and other effects caused by the necessary discretization), an identical output compared to the input.
An integration into the SBR module of filter banks according to the present invention can easily be done. Although SBR modules typically operate in double index mode, low-delay filter banks of complex value according to the configurations of the present invention are able to provide a perfect reconstruction in single index mode, while SBR QMF filter banks originals are able to provide only an almost perfect reconstruction. In dual index mode, the 32-band version of the impulse response is obtained by linear interpolation also called sub-sampling of two adjacent taps or window coefficients of the 64-band impulse response or window function as explained in the context of Fig 3.
In the case of a complex value implementation of a filter bank, a significantly reduced analysis (or synthesis) delay can be obtained for the critically sampled filter banks, where the sampling or processing of the frequency corresponds to the limit frequency of according to the Nyquist-Shannon theory. In the case of a real value implementation of a filter bank, an efficient implementation can be obtained using optimized algorithms as, for example, illustrated in the context of the MATLAB implementation shown in Figs. 9 and 10. These implementations can, for example, be used in the low power mode of the SBR tool as described in the context of Figs. 5 and 6.
As highlighted in the context of Figs. 5 and 6, it is possible to obtain a greater reduction regarding the delay in the case of an SBR system using a low-delay filter bank of complex value according to a configuration of the present invention. As pointed out earlier, in the SBR 610 decoder as shown in Fig. 5, the QMF 620 analysis filter bank is replaced by a complex low delay filter bank (CLDFB) according to a configuration of the present invention. This substitution can be done by computer, maintaining the number of bands (64), the length of the impulse response (640) and using complex modulation. The delay obtained by this tool is thus minimized in order to obtain a total delay sufficiently low for bidirectional communication without sacrificing an obtainable level of quality.
Compared, for example, to a system comprising an MDCT and an MDST to form a complex value MDCT type system, a configuration according to the present invention provides a much better frequency response. Compared to the QMF filter bank, for example, used today in the MPEG-4 SBR, a system comprising one or more filter banks according to the configurations of the present invention provides significantly less delay.
Even compared to a low-delay QMF filter bank, the configurations according to the present invention offer the advantage of perfect reconstruction combined with a shorter delay. The advantages that arise from the property of perfect reconstruction in contrast to the almost perfect reconstruction of the QMF filter bank are the following. For almost perfect reconstruction, high cut-band attenuation is required to attenuate aliasing to a sufficiently low level. This restricts the possibility of obtaining a very low delay in the filter design. In contrast, the use of a configuration according to the present invention now has the possibility to independently design the filter, so that no high cut-band attenuation is required to attenuate aliasing at sufficiently low levels. The cut band attenuation should be low enough to allow a reduced enough aliasing for the desired signal processing application. Therefore, a better decision with less delay in the design of the filter can be achieved.
Fig. 11 shows a comparison of the window function 700 as it can, for example, be used in a configuration according to the present invention, together with the sine window function 710. The window function 700, which is also called the synthesis CMLDFB window ( CMLDFB = complex low modulated delay filter bank), comprising 640 window coefficients based on the values given in the table in Annex 1. With regard to the magnitude of the window functions, it should be noted that the following general amplification factors or damping factors are not considered for adjusting the amplitude of the windowed signal. The window functions can, for example, be normalized with reference to a value corresponding to the delay center, as highlighted in the context of Fig. 13, or with reference to a value η = N, n = N-loun = N + l, where N is the length of the block and n is the index of the window coefficients. In comparison, the sine window function 710 is only defined in 128 samples and, for example, it is used in the case of an MDCT or MDST module.
However, depending on the implementation details, to obtain the window coefficients based on the values given in the tables in Annexes 1 and 3, other changes in signs with reference to the window coefficients corresponding to indexes 128 to 255 and 384 to 511 (multiplication by factor (-1)) must be considered according to equations (16a) and (16b).
Before discussing the differences of the two window functions 700, 710, it should be noted that both window functions comprise only window coefficients of real value. Furthermore, in both cases, the absolute value of the window coefficient corresponding to an index n = 0 is less than 0.1. In the case of a CMLDFB 700 window, the respective value is even less than 0.02.
Considering the two window functions 700, 710 with reference to their definition sets, several significant differences are evident. Considering that the sine window function 710 is symmetric, the window function 700 shows an asymmetric behavior. To define the fact more clearly, the sine window function is symmetric when there is a real value value in with reference to all real n numbers, so that a window function 710 is defined for (n<sub>0</sub>+ n) and (n<sub>0</sub>-n), and the relationship | J<sup>n</sup>o - Ú = p (<sup>n</sup>o + d (29) is true in a desired margin (ε> 0; the absolute value of the difference in terms on both sides of equation (29) is less than or equal to ε), where w (n) represents the window coefficient corresponding to index n. In the case of the sine window, the respective index n<sub>0</sub> it is exactly half of the two largest window coefficients. In other words, for the sine window 710, the index is n<sub>0</sub> = 63.5. The sine window function is defined for the indices n = 0, ..., 127.
In contrast, the window function 700 is defined in the set of indices n = 0, ..., 639. The window function 700 is clearly asymmetric in the sense that for all numbers n<sub>0</sub> of real values at least one real number always exists, so that (n<sub>0</sub>+ n) and (n<sub>0</sub>-n) belong to the window function definition set, for which the inequality
M ^<sub>The</sub> - d * kto + d (30)
Have an (almost deliberately) margin of definition (ε> 0; the absolute value of the difference under both sides of equation (29) is greater than or equal to ε), where again w (n) is the window coefficient corresponding to index n.
Other differences between the two window functions, which relate to the block sizes of N = 64 samples, is that the maximum value of the window function 700 is greater than 1 and has its indexes in the range of
N <η <2N (31) for a synthesis window. In the case of the window function 700 shown in Fig. 11, the maximum value obtained is greater than 1.04, obtained in the sample index n = 77. In contrast, the maximum values of the sine window 710 are less than or equal to 1, obtained emn = 63en = 64.
However, the window function 700 also obtains a value of approximately 1 with sample rates around η = N. To be more precise, the absolute value or the value of the window coefficient w (Nl) corresponding to the index η = Nl is less than 1, considering that the absolute value or the value of the window coefficient w (N) corresponding to the index n = N is greater than 1. In some configurations according to the present invention, these two window coefficients obey the relationships
0.99 <w (Nl) <1.0, 0 1.0 <w (N) <1.01 (32) which is the result of optimizing the quality of the audio filter bank according to the configurations of the present invention. In many cases, it is desirable to have a window coefficient w (0) comprising the smallest possible absolute value. In this case, a determinant of the window coefficients | w (0) · w (2N - 1) - w (N - 1) · w (n) | ® 1 (33) must be as close as possible to 1 to obtain a audio quality, which is optimized with reference to possible parameters. The sign of the determinant given by equation (33) is, however, of free choice. As a consequence of the window coefficient w (0) being less than, or approximately 0, the product of w (Nl) · w (N) or its absolute values must be as close as possible to +/- 1. In this case, the window coefficient w (2N-l) can then be chosen almost freely. Equation (33) is the result of employing the zero-delay matrix technique described in New Framework for Modulated Perfect Reconstruction Filter Banks by GDT Schuller and MJT Smith, IEEE Transactions on Signal Processing, Vol. 44, No. 8, August 1996.
In addition, as will be highlighted in more detail in the context of Fig. 13, the window coefficients corresponding to the Nl and N indices are comprised in half of the modulation core and therefore correspond to the sample having a value of approximately 1.0 and which coincides with the delay of the filter bank defined by the prototype filter function or the window function.
The synthesis window function 700 as shown in Fig. 11, furthermore shows an oscillating behavior with increasing strictly monotonic window coefficients of the window coefficients of the sequence of window coefficients corresponding to the index (n = 0) used for windowing the window. last audio sample in the time domain up to the window coefficient comprising the highest absolute value of all the window coefficients of the synthesis window function 700. Naturally, in the case of the reverse time analysis window function, the oscillating behavior comprises a strictly monotonic reduction of the window coefficients of the window coefficient which comprises the highest absolute value of all window coefficients of a corresponding (reverse time) window analysis function for the window coefficients of a sequence of coefficients of window corresponding to an index (n = 639) used for winding the last audio sample in the time domain.
As a consequence of the oscillating behavior, the development of the synthesis window function 700 starts with a window coefficient corresponding to the index n = 0 with an absolute value less than 0.02 and an absolute value of the window coefficient corresponding to the index η = 1 less than 0.03, acquiring a value of about 1 in an index η = N, acquiring a maximum value of more than 1.04 in an index according to equation (31), acquiring another value of approximately 1 in an index n = 90 and 91, a first exchange of signals in index values of n = 162 and n = 163, acquiring a minimum value less than -0.1 or 0., 12755 in an index of approximately η = 3N and another signal exchange with index values of n = 284 and n = 285. However, the synthesis window function 700 can still comprise other signal exchanges in other index values n. When comparing the window coefficients with the values given in the tables in Annexes 1 and 3, the other changes in signs with reference to the window coefficients corresponding to indexes 128 to 255 and 384 to 511 (multiplication by factor (-1)) must be considered according to equations (16a) and (16b).
The oscillating behavior of the synthesis window function 700 is similar to that of a strongly damped oscillation, which is illustrated by the maximum value of about 1.04 and the minimum value of about -0.12. As a consequence, more than 50% of all window coefficients comprise absolute values being less than or equal to 0.1. As highlighted in the context of the configurations described in Figs. 1 and 2a, the development of the window function comprises a first group 420 (or 200) and the second group 430 (or 210), where the first group 420 comprises a first consecutive portion of window coefficients and the second group 430 comprises a second portion consecutive window coefficients. As already pointed out before, the window window coefficient sequence comprises only the first group 420 of window coefficients and the second group of window functions 430, where the first group 420 of window coefficients comprises exactly the first consecutive sequence of window coefficients. window, and where the second group 430 comprises exactly the second consecutive portion of window coefficients. Thus, the terms first group 420 and first portion of window coefficients, as well as the terms second group 430 and second portion of window coefficients can be used interchangeably.
More than 50% of all window coefficients with absolute values less than or equal to 0.1 are included in the second group or in the second portion 430 of window coefficients as a consequence of the strongly damped oscillatory behavior of the window 700 function. more than 50% of all window coefficients comprised in the second group or second portion 430 of window coefficients comprise absolute values less than or equal to 0.01.
The first portion 420 of window coefficients comprises less than one third of all window coefficients in the window coefficient sequence. Likewise, the second portion 430 of window coefficients comprises more than two thirds of window coefficients. In the case of a total number of T blocks to be processed in one of frames 120, 150, 330, 380 of more than four blocks, the first portion typically comprises 3/2 · N window coefficients, where N is the number of samples in the time domain of a block. Likewise, the second portion comprises the rest of the window coefficients or, to be more precise, (T-3/2) N window coefficients. In the case of T = 10 blocks per frame as shown in Fig. 11, the first portion comprises 3/2 · N window coefficients, whereas the second portion 210 comprises 8.5 · N window coefficients. In the case of a block size of N = 64 audio samples in the time domain per block, the first portion comprises 96 window coefficients, whereas the second portion comprises 544 window coefficients. The synthesis window function 700 as shown in Fig. 11 acquires the value of approximately 0.96 at the limit of the first portion and the second portion with an index of about n = 95 or 96.
Despite the number of window coefficients comprised in the first portion 420 and the second portion 430, the energy value or a total energy value of the corresponding window coefficients differ significantly from each other. The energy value is defined by £ = ΣΚΗ<sup>2</sup><sub>r (34)</sub> n
where w (n) is a window coefficient and the index n at which the sum in equation (34) is evaluated corresponds to the indexes of the respective portions 420, 430, the entire set of window coefficients or any set of window coefficients to which the respective energy values E correspond. Despite the significant difference in window coefficients, the energy value of the first portion 420 is equal to or greater than 2/3 of the total energy value of all window coefficients. Likewise, the energy value of the second portion 430 is less than or equal to 1/3 of the total energy value of all window coefficients.
To illustrate the fact, the energy value of the first portion 420 of the window coefficients of the window function 700 is approximately 55.85, while the energy value of the window coefficients of the second portion 430 is approximately 22.81. The total energy value of all window coefficients of the window function 700 is approximately 78.03, so the energy value of the first portion 420 is approximately 71.6% of the total energy value, while the energy value of the second portion 430 is approximately 28.4% of the total energy value of all window coefficients.
Naturally, equation (34) can be presented in a normalized version, dividing the energy value E by a normalization factor E<sub>The</sub>, which in principle can have any energy value. The normalization factor E<sub>The</sub> it can, for example, be the total energy value of all window coefficients in a sequence of window coefficients calculated according to equation (34).
Based on the absolute values of the window coefficients or on the basis of the energy values of the respective window coefficients, a central point or center of mass of a sequence of window coefficients can also be determined. The center of mass or center point of a sequence of window coefficients is a real number and is typically within the range of the indices of the first portion 420 of the window coefficients. In the case of the respective frames comprising more than four blocks of audio samples in the time domain (T> 4), the center of mass n<sub>here</sub> based on the absolute values of the window coefficients or the center of mass n<sub>ce</sub> based on the energy values of the window coefficients it is less than 3/2 · N. In other words, in the case of T = 10 blocks per frame, the center of mass is well within the region of the indices of the first portion 200.
The center of mass n<sub>here</sub> based on the absolute values of the window coefficients w (n) is defined according to
Yn · Mn) and the center of mass n<sub>ce</sub> in view of the energy values of the window coefficients w (n) is defined according to
Σ n · Md<sup>2 </sup>n<sub>ce</sub> = ------ Σ Md<sup>2</sup> n = 0 (36) where N and T are positive integers that indicate the number of audio samples in the time domain per block and the number of blocks per frame, respectively. Of course, the central points according to equations (35) and (36) can also be calculated with reference to a limited set of window coefficients, replacing the limits of the sums above in the same way.
For the window function 700 as shown in Fig. 1, the center of mass n<sub>here</sub> based on the absolute values of the window coefficients w (n) is equal to a value of n<sub>here</sub> ~ 87.75 and the center point or center of mass n<sub>ce</sub> with reference to the energy values of the window coefficients w (n) is n<sub>ce</sub> ~ 80.04. Since the first portion 200 of window coefficients of the window function 700 comprises 96 (= 3/2 · N; N = 64) window coefficients, both central points are well within the first portion 200 of the window coefficients, as highlighted previously.
The window coefficients w (n) of the window 700 function are based on the values given in the table in Appendix 1. However, to obtain, for example, the low delay property of the filter bank as pointed out before, it is not necessary to implement the function window as exactly as given by the window coefficients in the table in Annex 1. In many cases, it is more than enough for the window coefficients of the window function comprising 640 window coefficients to comply with any of the relationships or equations given in the tables in Annexes 2 to 4. The window coefficients or filter coefficients given in the table in Annex 1 represent preferred values, which can be adapted according to equations (16a) and (16b) in some implementations. However, as indicated, for example, by the other tables given in the other Annexes, the preferred values may vary from the second, third, fourth, fifth digit after the decimal point, so that the resulting filters or window functions still have advantages of configurations. according to the present invention. However, depending on the implementation details, to obtain the window coefficients based on the values given in the tables in Annexes 1 and 3, other changes in signs with reference to the window coefficients corresponding to indexes 128 to 255 and 384 to 511 (multiplication by factor (-1)) should be considered according to equations (16a) and (16b).
Of course, other window functions comprising a different number of window coefficients can also be defined and used in the structure of the configurations in accordance with the present invention. In this context, it should be noted that both the number of audio samples in the time domain per block and the number of blocks per frame, as well as the distribution of blocks with reference to past and future samples can vary over a wide range of parameters.
Fig. 12 shows a comparison of a low-delay modulated filter bank with complex window (CMLDFB window) 700 shown in Fig. 11 and the original prototype filter SBR QMF 720 as used, for example, in the SBR tool according to standards MPEG. As shown in Fig. 11, the CMLDFB 700 window is again a synthesis window according to a configuration of the present invention.
Although the window function 700 according to a configuration of the present invention is clearly asymmetric as defined in the context of equation (30), the original prototype filter SBR QMF 720 is symmetrical with reference to the indices n = 319 and 320, as the window function 700, as well as the SBR QMF 720 prototype filter are defined with reference to 640 indexes each. In other words, with reference to equation (29), the index value in which the index of the center of symmetry represents is given by no = 319.5 in the case of the prototype filter SBR QMF 720.
Furthermore, due to the symmetry of the prototype filter SBR QMF 720, also the central point n<sub>here</sub> en<sub>ce</sub> according to equations (35) and (36), respectively, they are identical to the center of symmetry n<sub>0</sub>. The energy value of the SBR QMF 720 prototype filter is 64.00, since the prototype filter is an orthogonal filter. In contrast, the symmetrical asymmetric window function 700 comprises an energy value of 78.0327 as noted above.
In the following sections of the description, SBR systems will be considered as highlighted in the context of Figs. 5 and 6, wherein the SBR 610 decoder comprises configurations according to the present invention in the form of an analysis filter bank such as the filter bank 620 and a configuration according to the present invention in the form of an analysis bank. synthesis filters for the synthesis filter bank 640. As will be highlighted in more detail, the total delay of a bank of analysis filters according to the present invention that employs a window function 700 as shown in Figs. 11 and 12 comprises a total delay of 127 samples, considering that the SBR tool based on the original SBR QMF prototype filter results in a total delay of 640 samples.
The replacement of the QMF filter banks in the SBR module, for example, in the SBR 610 decoder, with a low value complex filter bank (CLDFB) results in a delay reduction from 42 ms to 31.3 ms without introducing any degradation of audio quality or other computational complexity. With the new filter bank, both the standard SBR mode (high quality mode) and the low power mode that employs only real value filter banks are supported, as the description of the configurations according to the present invention showed. with reference to Figs. 7 to 10.
Especially in the field of telecommunications and two-way communication, a low delay is of great importance. Although the extended low-delay AAC is already capable of achieving a delay low enough for 42 ms communications applications, its algorithmic delay is even greater than that of the low-delay AEC corecodec, which is capable of achieving a delay as low as 20 ms and that of other telecommunications codecs. In the SBR 610 decoder, the QMF analysis and synthesis stages still cause a 12 ms reconstruction delay. An approach that promises to reduce this delay is to use a low-delay filter bank technique according to a configuration of the present invention and replace the current QMF filter banks with the respective low-delay version according to the configurations of the present invention. In other words, another delay reduction is achieved by simply replacing the regular filter banks used in the SBR 610 module with a complex, low-delay filter bank according to the configurations of the present invention.
For use in the SBR 610 module, the new filter banks according to the configurations of the present invention, which are also called CLDFBs, are designed to be as similar to the QMF filter banks originally used as possible. This includes, for example, the use of 64 sub-bands or bands, of an equal length of the impulse responses and a compatibility with the dual index modes used in the SBR systems.
Fig. 13 illustrates the comparison of the CLDFB 700 window format according to a configuration of the present invention and the original prototype filter SBR QMF 720. In addition, it illustrates the delay of the modulated filter banks, which can be determined by analyzing the delay of overpass introduced by the prototype filter or window function, in addition to the framing delay of the modulation core with a length of N samples in the case of a system based on DCT-IV. The situation shown in Fig. 13 refers again to the case of a synthesis filter bank. The window function 700 and the prototype filter function 720 also represent impulse responses from the prototype synthesis filters of the two filter banks involved.
With reference to the delay analysis of both the SBR QMF filter bank and the proposed CLDFB, according to a configuration of the present invention, in the analysis and synthesis only the overpass to the right and left sides of the modulation core, respectively , add delay.
For both filter banks, the modulation core is based on a DCT-IV that introduces a delay of 64 samples, which is marked in Fig. 13 as the delay 750. In the case of the prototype filter SBR QMF 720 due to symmetry, the delay of the modulating core 750 is arranged symmetrically with reference to the center of mass or the central point of the respective prototype filter function 720 as shown in Fig. 13. The reason for this behavior is that the buffer of the SBR QMF filter bank must be filled to a point where the 720 prototype filter function will be considered in the processing, having the most significant contribution in terms of the respective energy values of the prototype filter values. . Due to the shape of the prototype filter function 720, this requires the buffer to be filled to at least the central point or center of mass of the respective prototype filter function.
To better illustrate the fact, starting from all the initialized buffers of the corresponding SBR QMF filter bank, the buffer must be filled to the point that the data processing will result in the processing of significant data, which require that the respective window function or function of prototype filter has a significant contribution. In the case of the prototype filter function SBR QMF, the symmetrical shape of the prototype filter 720 produces a delay, which is in the order of the center of mass or the central point of the prototype filter function.
However, as the delay introduced by the modulation core of the N = 64 DCT-IV based system for samples is always present and the system also comprises the delay of a block, it can be observed that the SBR QMF synthesis prototype introduces a delay overpaste of 288 samples.
As previously indicated, in the case of the synthesis filter banks referred to in Fig. 13, this new overlap on the left side 760 causes the delay, while the overlap on the right side 770 refers to the past samples and, therefore, does not introduce any new delay in the case of a synthesis filter bank.
In contrast, starting with all the initialized buffers of the CLDFB according to a configuration of the present invention, the synthesis filter bank, as well as the analysis filter bank, are able to provide significant data earlier, compared to the database. SBR QMF filters due to the shape of the window function. In other words, due to the shape of the analysis or synthesis window function 700, samples processed by the window functions indicative of the significant contribution are possible. As a consequence, the CLDFB synthesis prototype or synthesis window function only introduces an overlap delay of 32 samples, taking into account the delay already introduced by the 750 modulation core. The first portion 420 or first group 420 of window coefficients of the window function 700 according to a configuration of the present invention comprises, in a preferred configuration according to the present invention, the 96 window coefficients corresponding to the delay caused by the overlap on the left side 760 in conjunction with the delay of the modulation core 750.
The same delay is introduced by the analysis filter bank or by the analysis prototype function. The reason is that the analysis filter bank is based on the time-reverse version of the synthesis window function or prototype function. Therefore, the overpass delay is introduced on the right side, comprising the same overpass size as that of the synthesis filter bank. Thus, in the case of an original QMF prototype filter bank, a delay of 288 samples is also introduced, whereas for an analysis filter bank according to a configuration of the present invention only 32 samples are introduced as a delay.
The table shown in Fig. 14a gives an overview of the delay with different stages of modification assuming a frame length of 480 samples and a sample rate of 48 kHz. In a standard configuration, comprising an AACLD codec together with a standard SBR tool, the MDCT and IMDCT filter banks in double index mode cause a 40 ms delay. Furthermore, the QMF tool itself causes a delay of 12 ms. Furthermore, due to an overlapping SBR, another delay of 8 ms is generated, so that the general delay of this codec is in the range of 60 ms.
In comparison, an AAC-ELD codec comprising low-delay versions of MDCT and IMDCT generates a 30 ms delay in the dual index approach. Compared to the original QMF filter bank of a SBR tool, the use of a low delay filter bank of complex value according to a configuration of the present invention will result in a delay of only 1 ms compared to the 12 ms of the original QMF tool. By avoiding the SBR overlap, the additional 8 ms overlap of a forward combination of an AAC-LD with the SBR tool can be completely avoided. Therefore, an extended low-delay AAC codec is capable of a total algorithmic delay of 31 ms, instead of 60 ms from the forward combination highlighted earlier. Therefore, it can be seen that the combination of the described delay reduction methods really results in the saving of the total delay of 29 ms.
The table in Fig. 14b gives another view of the general codec delay caused by the original and proposed filter bank versions in a system as shown in Figs. 5 and 6. The data and values presented in Fig. 14b are based on the sampling rate of 48 kHz and a core encoder frame size of 480 samples. Due to the dual index approach of an SBR system as shown and discussed in Figs. 5 and 6, the core encoder operates effectively at a sample rate of 24 kHz. As the framing delay of 64 samples of the modulation core is already introduced by the core encoder, it can be subtracted from the standalone delay values of the two filter banks as described in the context of Fig. 13.
The table in Fig. 14b highlights that it is possible to reduce the overall delay of the extended low-delay AAC codec comprising the low-delay versions of MDCT and IMDCT (LD MDCT and LD IMDCT). Although a total algorithmic delay of 42 ms is obtainable only by the use of low-delay versions of MDCT and IMDCT, as well as of the original QMF filter banks, using complex low-delay filter banks according to the configurations of the present invention instead of the conventional QMF filter banks, the total algorithmic delay can be significantly reduced to just 31.3 ms.
To assess the quality of the filter banks according to the configurations of the present invention and systems comprising one or more filter banks, hearing tests must be carried out, from which it can be concluded whether filter banks according to the configurations of the The present invention maintains the audio quality of AAC-ELD at the same level and introduces no degradation, either for the complex SBR mode or for the low power and real value SBR mode. Therefore, the delayed filter banks optimized according to the configurations of the present invention do not introduce any load on the audio quality, although they are able to reduce the delay by more than 10 ms. For transient items, it can even be observed that some small improvements, despite not having statistical significance, can be achieved. The aforementioned improvements were observed during hearing tests of castanets and vibraphones.
In order to better verify that the sub-sampling in the case of a 32-band filter bank according to a configuration of the present invention works equally well for the filter banks according to the present invention compared to the QMF filter banks, following evaluation. First, a sine-logarithmic scan was analyzed with a 32-band sub-sampled filter bank, where the upper 32 bands, initialized with zeros, were added. Then, the result was synthesized by a 64-band filter bank, again sub-sampled and compared to the original signal. Using a conventional SBR QMF prototype filter results in a signal to noise ratio (SNR) of 59.5 dB. A filter bank according to the present invention, however, achieves an SNR value of 78.5 dB, which illustrates that the filter banks according to the configurations of the present invention also operate in the subsampled version at least as well as the banks of filters. original QMF filters.
To show that this optimized, non-symmetric delay filter bank approach as used in configurations according to the present invention provides an additional value compared to a classic filter bank with a symmetrical prototype, asymmetric prototypes will be compared with symmetrical prototypes having the same delay to follow.
Fig. 15a shows a comparison of a frequency response in a far field illustration of a filter bank according to the present invention using a low delay window (graph 800) compared to the frequency response of a filter bank that employs a sine window with a length of 128 taps (graph 810). Fig. 15b shows an enlargement of the frequency response in the field next to the same filter banks that employ the same window functions as highlighted above.
A direct comparison of the two graphs 800, 810 shows that the frequency response of the filter bank employing a low-delay filter bank according to a configuration of the present invention is significantly better than the corresponding frequency response of a filter bank which employs a sine window of 128 taps having the same delay.
Also, Fig. 16a shows a comparison of different window functions with an overall delay of 127 samples. The filter bank (CLDFB) with 64 bands comprises a general delay of 127 samples including the framing delay and the overpass delay. A filter bank modulated with a symmetrical prototype and the same delay would, therefore, have a prototype with a length of 128, as already illustrated in the context of Figs. 15a and 15b. For these filter banks with 50% overlap, such as, for example, the MDCT, sine windows or Kaiser-Bessel-derived windows generally provide a good choice of prototypes. Thus, in Fig. 16a an overview of the frequency response of a filter bank that employs a low-delay window as a prototype according to a configuration of the present invention is compared with the frequency responses of the alternative symmetric prototypes with the same delay. Fig. 16a shows, in addition to the frequency response of the filter bank according to the present invention (graph 800) and the frequency response of a filter bank that employs a sine window (graph 810), as already shown in Figs. 15a and 15b, in addition to the two KBD windows based on the parameters oc = 4 (graph 820) and OC = 6 (graph 830). Both, Fig. 16a and the approach to Fig. 16a shown in Fig. 16b, clearly show that a much better frequency response can be obtained with a filter bank according to a configuration of the present invention having a non-symmetric window function or the prototype filter function with the same delay.
To illustrate this advantage more generally, in Fig. 17 two prototypes of filter banks are compared with delay values different from the filter bank previously described. Although the filter bank according to the present invention, which was considered in Figs. 15 and 16, have a general delay of 127 samples, which corresponds to an overlap of 8 blocks in the past and 0 blocks in the future (CLDFB 80), Fig. 17 shows a comparison of the frequency response of two different filter bank prototypes with the same 383 sample delay. To be more precise, Fig. 17 shows the frequency response of a non-symmetric prototype filter bank (graph 840) according to a configuration of the present invention, which is based on an overlap of 6 blocks of samples in the time domain in the past and 2 blocks of samples in the time domain in the future (CLDFB 62) . Furthermore, Fig. 17 it also shows the frequency response (graph 850) of a corresponding symmetric prototype filter function, also having a delay of 383 samples. It can be seen that the same delay value of a non-symmetric prototype or window function achieves a better frequency response than a filter bank with the symmetric window function or prototype filter. This shows the possibility of a better decision between delay and quality, as previously indicated.
Fig. 18 illustrates the temporal masking effect of the human ear. When a sound or tone appears at a time in the time indicated by a line 860 in Fig. 18, a masking effect appears regarding the frequency of the tone or the sound and neighboring frequencies approximately 20 ms before the start of the actual sound. This effect is called pre-masking, being an aspect of the psychoacoustic properties of the human ear.
In the situation illustrated in Fig. 18, the sound remains audible for approximately 200 ms until the time shown in line 870. During this time, the human ear mask is active, which is also called simultaneous masking. After the sound stops (illustrated by line 870), the frequency masking at the neighboring frequency of the tone decays slowly over a period of approximately 150 ms as illustrated in Fig. 18. This psychoacoustic effect is also called post-masking.
Fig. 19 illustrates a comparison of a pre-echo behavior of a conventional HE-AAC encoded signal and a HE-AAC encoded signal that are based on a filter bank that employs a low-delay filter bank (CMLDFB) according to with a configuration of the present invention. Fig. 19a illustrates the original time signal of the castanets, which was processed with a system comprising a HE-AAC codec (HE-AAC = advanced high efficiency audio codec). The HE99-based system output
Conventional AAC is illustrated in Fig. 19b. A direct comparison of the two signals, the original timing signal and the HE-AAC codec output signal shows that before the castanets sound in the area illustrated by arrow 880, the HE-AAC codec output signal comprises important effects pre-echo.
Fig. 19c illustrates an output signal from a system comprising a HE-AAC based on filter banks comprising CMLDFB windows according to a configuration of the present invention. The same original timing signal shown in Fig. 19a and processed using filter banks according to a configuration of the present invention shows a significantly reduced appearance of the pre-echo effects just before the start of a castanets signal as indicated by an arrow 890 in Fig. 19c. Due to the pre-masking effect described in the context of Fig. 18, the pre-echo effect indicated by arrow 890 in Fig. 19c will be much better masked than the pre-echo effects indicated by arrow 880 in the case of the conventional HEAAC codec. Therefore, the pre-echo behavior of the filter banks according to the present invention, which are also the result of the significantly reduced delay compared to conventional filter banks, makes the output much better adapted to the time masking properties and to the psychoacoustics of the human ear. As a consequence, as already indicated when describing the hearing tests, the use of filter banks according to a configuration of the present invention can even lead to an improvement in quality due to the reduced delay.
The configurations according to the present invention do not increase the computational complexity compared to conventional filter banks. Low-delay filter banks use the same filter length and modulation mode as,
100 for example, the QMF filter banks in the case of SBR systems so that the computational complexity does not increase. In terms of memory requirements due to the asymmetric nature of the prototype filters, the read-only memory requirement of the synthesis filter bank increases by approximately 320 words in the case of a filter bank based on N = 64 samples per block and T = 10 blocks per frame. Furthermore, in the case of an SBR system, the memory requirement increases by another 320 words if the analysis filter is stored separately.
However, as the current ROM requirements for an AACELD core is approximately 2.5 k words (kilo words) and for the SBR implementation another 2.5 k words, the ROM requirement increases only moderately by around 10%. As a possible decision between memory and complexity, if a small memory consumption is necessary, a linear interpolation can be used to generate the analysis filter from the synthesis filter as highlighted in the context of Fig. 3 and in equation (15) . This interpolation operation increases the number of instructions required by only approximately 3.6%. Therefore, when replacing conventional QMF filter banks in the SBR module structure with low delay filter banks according to the configurations of the present invention, the delay can be reduced in some configurations by more than 10 ms without any degradation of the quality of noticeable increase in complexity.
The configurations according to the present invention, therefore, relate to an analysis window or equipment or window method. Furthermore, an analysis or synthesis filter bank or method for the analysis or synthesis of a signal using the window is described. Naturally, the computer program
101 that implements one of the above methods is also revealed.
The implementation according to the configurations of the present invention can be made in hardware implementations, software implementations or a combination of both. The data, vectors and variables generated, received or stored to be processed can be stored in different types of memories such as random access memories, buffers, read memories, non-volatile memories (for example, EEPROMs, flash memories) or other memories like magnetic or optical memories. A storage position can, for example, be one or more memory units required to store or save the respective amounts of data, such as variables, parameters, vectors, matrices, window coefficients or other pieces of information and data.
Software implementations can be operated on different computers, computer-type systems, processors, ASICs (application-specific integrated circuits) or other integrated circuits (ICs).
Depending on certain requirements for implementing the method configurations of the invention, the configurations of the methods of the invention can be implemented in hardware, software or a combination of both. The implementation can be done using a digital storage medium, in particular a CD disc, a DVD or another disc having a stored electronic reading control signal, which cooperates with a programmable computer system, processor or integrated circuit, in order to a configuration of the method of the invention can be carried out. In general, a configuration of the present invention is, therefore, a computer program product with a program code stored in a machine-readable carrier, the
102 program code being operated to perform a configuration of the methods of the invention when the computer program product operates on a computer, processor or integrated circuit. In other words, the configurations of the methods of the invention are, therefore, a computer program having a program code to carry out at least one configuration of the methods of the invention when the computer program operates on a computer, processor or integrated circuit.
An equipment for generating audio sub-band values in audio sub-band channels according to the configurations of the present invention comprises an analysis window (110) for the window of a frame (120) of input samples of audio in the time domain being in a temporal sequence that extends from a previous sample to a later sample using an analysis window function (190) comprising a sequence of window coefficients to obtain the window function sampled window analysis (190) comprising a first group (200) of window coefficients comprising a first portion of a sequence of window coefficients and a second group (210) of window coefficients comprising a second portion of a sequence of window coefficients window, the first portion comprising fewer window coefficients than the second portion, where an energy value of the window coefficients in the first portion is greater than an energy value of the window coefficients of the second portion, where the first group of window coefficients is used for the subsequent window pane of samples in the time domain and the second a group of window coefficients is used for the early windowing of samples in the time domain, and a calculator (170) for the calculation of the audio subband values using the windowed samples.
103
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the analysis window function (190 ) is asymmetric with respect to the sequence of window coefficients.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted so that an energy value of the window coefficients of the first portion is equal to or greater than 2/3 of an energy value of all window coefficients of a sequence of window coefficients and an energy value of the window coefficients of the second portion of window coefficients is less than or equal to 1/3 of an energy value of all window coefficients of a sequence of window coefficients.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted so that the first portion of window coefficients comprises 1/3 or less than 1/3 of the total number of window coefficients in a sequence of window coefficients and the second portion comprises 2 / 3 or more than 2/3 of the total number of window coefficients in a sequence of window coefficients.
In an equipment for the generation of subband audio values in subband audio channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that a central point of the
104 window coefficients of the analysis window function (190) corresponds to a real value in a range of the index of the first portion of window coefficients.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the analysis window function (190) comprises a strictly monotonic reduction of the window coefficient comprising the highest absolute value of all the window coefficients of the analysis window function (190) of a window coefficient of a sequence of window coefficients used for the window of the last audio sample in the time domain.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the analysis window function (190 ) comprises oscillating behavior.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the window coefficient corresponding to a index n = (Tl) · N comprises an absolute value in the range of 0.9 to 1.1, where an index of a sequence of window coefficients is an integer in the range of 0 to N • Tl, where the window coefficient used for the window of the last audio input sample in the time domain of frame 120 is the window coefficient corresponding to the N • Tl index, where the analysis window (110) is adapted so that the frame (120) of audio input samples on the
105 time domain comprises a sequence of T blocks (130) of audio input samples in the time domain extending from the most recent to the last audio input samples in the time domain of the frame (120), each block comprising N audio input samples in the time domain, and where T and N are positive integers and T is greater than 4.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the window coefficient corresponding to the index of the window coefficients η = N · T 1 comprises an absolute value less than 0.02.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted so that the window comprises the multiplication of the audio input samples in the time domain x (n) of the frame (120) to obtain the windowed samples z (n) of the windowed frame based in the equation z (n) = x (n) · c (n) where n is an integer that indicates an index of a sequence of window coefficients in the range 0 to T · Nl, where c (n) is the coefficient window of the analysis window function corresponding to the index n, where x (N · Tl) is the last audio input sample in the time domain of a frame (120) of audio input samples in the time domain, where the analysis window (110) is adapted so that the frame (120) of audio input samples in the time domain comprises a sequence of T blocks (130) of samples of
106 audio input in the time domain extending from the most recent to the last audio input samples in the frame time domain (120), each block comprising N audio input samples in the time domain, and where T and N are positive integers and T is greater than 4.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the window coefficients c (n ) obey the relationships given in the table in Annex 4.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the equipment (100) is adapted to use an analysis window function (190) being a reverse-time or reverse-index version of the synthesis window function (370) to be used for the audio subband values.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted so that the first portion of the window function of The analysis comprises a window coefficient having an absolute maximum value greater than 1.
In an equipment for the generation of subband audio values in subband audio channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that all the window coefficients of an sequence of window coefficients are real value window coefficients.
107
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted so that the frame (120) of samples audio input in the time domain comprises a sequence of T blocks (130) of audio input samples in the time domain extending from the most recent to the last audio input samples in the time domain of the frame (120) , each block comprising N samples of audio input in the time domain, where T and N are positive integers and T is greater than 4.
In an equipment for the generation of subband audio values in subband audio channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the window comprises an elementary multiplication of the samples of audio input in the time domain of the frame (120) by the window coefficients of a sequence of window coefficients.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted so that each sample of audio input in the time domain is multiplied elementally by a window coefficient of the analysis window function according to the sequence of audio input samples in the time domain and the sequence of window coefficients.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted so that for each audio input sample in the time domain of the frame (120) of samples of
108 audio input in the time domain is generated exactly a windowed sample.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the window coefficient corresponding to a index of window coefficients η = (T-3) • N comprises a value less than -0.1, where the index of a sequence of window coefficients is an integer in the range 0 to N · T - 1, and where the window coefficient used for windowing the last sample of audio input in the time domain is the window coefficient corresponding to the N · T - 1 index.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the first portion of window coefficients comprises 3/2 · N window coefficients and the second portion of window coefficients comprises (T-3/2) · N window coefficients of a sequence of window coefficients.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the window coefficients c (n ) obey the relationships given in the table in Annex 3.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window
109 (110) is adapted so that the window coefficients c (n) obey the relationships given in the table in Annex 2.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the analysis window (110) is adapted in such a way that the window coefficients c (n ) comprise the values given in the table in Annex 1.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the equipment (100) is adapted so that the present frame (120) of audio input samples in the time domain to be processed is generated by changing (Tl) subsequent blocks of a directly preceding frame (120) of input samples audio in the time domain by a block in the direction of the previous audio input samples in the time domain and adding a block (220) of new audio samples in the time domain as the block comprising the last samples of time audio input in the time frame of the current frame (120).
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the equipment (100) is adapted in such a way that the present frame (120) of samples of audio input in the time domain x (n) to be processed is generated based on the change of the audio input samples in the time domain x<sub>prev</sub>(n) of the directly preceding frame (120) of audio input samples in the time domain based on the equation x (n - 32) = x<sub>prev</sub>(n)
110 for a time or sample index n = 32, ...<sub>r</sub> 319, and where the equipment (100) is still adapted to generate the audio input samples in the time domain x (n) of the current frame (120) of audio input samples in the time domain by including the next 32 input samples in the time domain according to the order of the audio input samples in the decreasing time time domain or in sample n indexes of the audio input samples in the time domain x (n) of the current frame (120) starting at time index or sample n = 31.
In an equipment for the generation of subband audio values in subband audio channels according to the configurations of the present invention, the calculator (170) comprises a time / frequency converter adapted to generate the subband values. audio band so that all subband values based on a windowed sample frame (150) represent a spectral representation of the windowed samples of the windowed sample frame (150).
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the time / frequency converter is adapted to generate complex values or real subband values of audio.
In an equipment for generating audio subband values in audio subband channels according to the configurations of the present invention, the calculator (170) is adapted to calculate an audio subband value for each sample of audio input in the time domain of a block (130) of audio input samples in the time domain, where the calculation of each audio subband value or each one
111 of the audio input samples in the time domain of a block (130) of audio input samples in the time domain is based on the windowed samples of the windowed frame (150).
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the calculator (170) is adapted to calculate the audio subband values based on multiplying the windowed samples (150) by the harmonic oscillation function of each subband value and adding the multiplied windowed samples, where the frequency of the harmonic oscillation function is based on a central frequency of a corresponding subband of the subband values.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the calculator (170) is adapted in such a way that the harmonically oscillating function is a complex exponential function , a sine function or a cosine function.
In an equipment for generating audio subband values in audio subband channels according to the configurations of the present invention, the calculator (170) is adapted to calculate the audio subband values w<sub>ki</sub> based on equation 4 u<sub>n</sub> = X z (n + j · 6 4)
J = 0 for n = 0, ..., 63 and <sup>63</sup> (τι = E<sup>u</sup>n <sup>2</sup> Osc - · t + 0.5) (2n - 95) n = 0 V4
112 for k = 0,. .., 31, where z (n) is a windowed sample corresponding to an index n, where k is a subband index, where 1 is an index of a block (180) of audio subband values and where F<sub>the C</sub>(x) is an oscillating function depending on the real value variable x.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the calculator (170) is adapted in such a way that the oscillating function f<sub>the C</sub>(X and
CD<sup>x</sup>) = e * pU · *)
OR
CscW = cos (x) or
OscU) = Sin (x), where i is an imaginary unit.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the equipment (100) is adapted to process a frame (120) of input samples from audio in the real time domain.
In an equipment for the generation of audio subband values in audio subband channels according to the configurations of the present invention, the equipment (100) is adapted to provide a signal indicative of the synthesis window function ( 370) to be used with the audio subband values or indicative of the analysis window function (190) used
113 to generate the audio subband values.
An equipment for the generation of audio samples in the time domain according to the configurations of the present invention comprises a calculator (310) for the calculation of a sequence (330) of intermediate samples in the time domain of the subband values of audio in subband audio channels, a sequence comprising pre-intermediate samples in the time domain and later samples in the time domain, the synthesis winder (360) for winding the sequence (330) of intermediate samples in the time domain using the synthesis window function (370) comprising the sequence of window coefficients to obtain intermediate windowed samples in the time domain, the synthesis window function (370) comprising a first group (420) of window coefficients comprising a first portion of a sequence of window coefficients and a second group (430) of window coefficients comprising a second portion of a sequence of window coefficients window coefficients, the first portion comprising fewer window coefficients than the second portion, where an energy value of the window coefficients in the first portion is greater than an energy value of the window coefficients of the second portion, where the first group of window coefficients is used for the subsequent winding of intermediate samples in the time domain and the second group of window coefficients is used for pre-winding intermediate samples in the time domain, and an overpaste addition output stage (400) for processing the intermediate windowed samples in the time domain in order to obtain the samples in the time domain.
In an equipment for the generation of audio samples in the time domain according to the configurations of this
114 invention, the synthesis window (360) is adapted so that an energy value of the window coefficients of the first portion of window coefficients is greater than or equal to 2/3 of an energy value of all function window coefficients synthesis window (370) and an energy value of the second portion of window coefficients is less than or equal to 1/3 of the energy value of all window coefficients of the synthesis function.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesizer window (360) is adapted so that the first portion of window coefficients comprises 1/3 or less than 1/3 of the total number of all window coefficients in a sequence of window coefficients and the second portion of window coefficients comprise 2/3 or more than 2/3 of the total number of window coefficients in a sequence of window coefficients.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the central point of the window coefficients of the synthesis window function (370 ) corresponds to a real value in a range of the index of the first portion of window coefficients.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the synthesis window function comprises a strictly monotonic increase of the window coefficient of a sequence of window coefficients used for the window of the last intermediate sample in the time domain for the
115 window coefficient comprising the highest absolute value of all the window coefficients of the synthesis window function.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the synthesis window function (370) comprises an oscillating behavior.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the window coefficient corresponding to an index n = N comprises an absolute value in the range between 0.9 and 1.1, where the index n of a sequence of window coefficients is an integer in the range 0 to T · N - 1, where the window coefficient used for the window of the last intermediate sample in the time domain is the window coefficient corresponding to the index n = 0, where T is an integer greater than 4 indicating the number of blocks comprised in the frame (330) of intermediate samples in the time domain, where the equipment (300) is adapted to generate a block (410) of audio samples in the time domain, the block (410) of audio samples in the time domain comprising N audio samples in the time domain, where N is a positive integer.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the window coefficient corresponding to the index n = 0 comprises a smaller absolute value or equal to 0.02.
In an equipment for the generation of audio samples in the time domain according to the configurations of this
116 invention, the synthesis window (360) is adapted so that the window coefficient corresponding to an index η = 3N is less than -0.1, where the equipment (300) is adapted to generate a block (410) of samples audio in the time domain, the block (410) of audio samples in the time domain comprising N audio samples in the time domain, where N is a positive integer.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the window comprises the multiplication of the intermediate samples in the time domain g (n ) of a sequence of intermediate samples in the time domain to obtain the windowed samples z (n) of the windowed frame (380) based on the equation z (n) = g (n) · c (t · N - 1 - n) for n = 0, ..., T · N - 1.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the window coefficient c (n) obeys the relations given in the table of the Annex 4.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the equipment (300) is adapted to use a synthesis window function (370) being a time-reverse or index-reverse version of an analysis window function (190) used to generate the audio subband values.
117
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the equipment (300) is adapted to generate a block (410) of audio samples in the time domain, the block (410 ) of audio samples in the time domain comprising N audio samples in the time domain, where N is a positive integer.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the equipment (300) is adapted to generate a block (410) of audio samples in the time domain, based on a block (320) of audio subband values comprising N audio subband values and where the calculator (310) is adapted to calculate a sequence (330) of intermediate audio samples in the time domain comprising T · N intermediate audio samples in the time domain, where T is a positive integer.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the synthesis window function is asymmetric with reference to a coefficient sequence of window.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted so that the first portion comprises a maximum value of all the window coefficients of the function synthesis window having an absolute value greater than 1.
In an equipment for the generation of audio samples in the
118 time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the first portion comprises 3/2-N window coefficients and the second portion of window coefficients comprises (T-3 / 2) -N window coefficients, where T is an index greater than or equal to 4 indicating a number of blocks 340 comprised in the frame (330) of intermediate samples in the time domain.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the window of the sequence of intermediate samples in the time domain comprises an elementary multiplication of intermediate samples in the time domain by a window coefficient.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesizer (360) is adapted so that each intermediate sample in the time domain is multiplied elementally by the window coefficient of the function synthesis window (370) according to the sequence of intermediate samples in the time domain and a sequence of window coefficients.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted so that the window coefficients of the synthesis window function (370) are values real.
In an equipment for the generation of audio samples in the time domain according to the configurations of this
119 invention, the synthesis window (360) is adapted in such a way that the window coefficient c (n) obeys the relationships given in the table in Annex 3.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the window coefficients c (n) obey the relationships given in the table of the Annex 2.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the synthesis window (360) is adapted in such a way that the window coefficients c (n) comprise the values given in the table of the Annex 1.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the calculator (310) is adapted to calculate the intermediate samples in the time domain of a sequence of intermediate samples in the time domain based on multiplying the audio subband values by the harmonic oscillation function and adding the multiplied audio subband values, where the frequency of the harmonic oscillation function is based on a central frequency of the corresponding subband.
In a device for the generation of audio samples in the time domain according to the configurations of the present invention, the calculator (310) is adapted in such a way that the harmonic oscillation function is a complex exponential function, a sine function or a function cosine.
In an equipment for the generation of audio samples in the
120 time domain according to the configurations of the present invention, the calculator (310) is adapted to calculate intermediate samples of real value in the time domain based on the audio subband values of complex values or real values.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the calculator (310) is adapted to calculate a sequence of intermediate real value samples in the time domain z (i, n) based on the equation
Ί N 1 / \ Λ
z. = - V Re X.. F - 0 + 7-7 -) - (/: + 7)<sup>1, n</sup> N "V <sup>V 2 2 2</sup>0J for an integer n in the range 0 to N · Tl, where Re (x) is the real part of the complex value number x, π = 3.14 ... is the circular number and f<sub>the C</sub>(x) is a function of harmonic oscillation, where f<sub>the C</sub>(x) = exp (í · x) when the audio subband values provided for the calculator are complex values, where I is the imaginary unit, and where
CscU) = cos (x), when the audio subband values provided for the calculator (310) are real values.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the calculator (310) comprises a time / frequency converter adapted to generate a sequence of samples
121 intermediate in the time domain, so that the audio subband values provided for the calculator (310) represent a spectral representation of a sequence of intermediate samples in the time domain.
In a device for generating audio samples in the time domain according to the configurations of the present invention, the time / frequency converter is adapted to generate a sequence of intermediate time / domain samples based on the audio subband values complex values or real values.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the calculator (310) is adapted to calculate a sequence of intermediate samples in the time domain g (n) from the values of audio subband X (k) based on equation v (n) = v<sub>prev</sub>{n - 2N) for an integer n in the range of 20N - 1 and 2N,
Nl f η v (n) = Σ <sup>Re</sup> '77 k = ok • exp i <sub>π</sub> ι --- (k + -0 · (2n - (n - 1)) for
2N the integer n in the range of 0 and 2N-1 eg (2N · j + k) = v (^ Nj + k) g (2N · j + N + k) = v (4Nj + 3N + k) for a number integer j in the range of 0 and 4 and for an integer k in the range of 0 and Nl, where N is an integer that indicates the number of subband audio values and the number of audio samples in the time domain , where v is a real value vector, where v<sub>prev</sub> is a real value vector v of generation
122 directly anterior of audio samples in the time domain, where i is the imaginary unit and π is the circular number.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the calculator (310) is adapted to calculate a sequence of intermediate samples in the time domain g (n) from the values of audio subband X (k) based on equation = <sup>v</sup><sub>FOR</sub>King<sup>n</sup> ~ 9N) for an integer n in the range of 20N - 1 and 2N,
N__1 Ί (TT = Σ 77 cos - (k + |) (2n - (n - 1)) k = o ó 2, for an integer n in the range of 0 and 2N-1 eg (2N · j + k) = v (4Nj + k) g (2N · j + N + k) = v (^ Nj + 3N + k) for an integer j in the range of 0 and 4 and for an integer k in the range of 0 and Nl , where N is an integer that indicates the number of subband audio values and the number of audio samples in the time domain, where v is a real value vector, where v<sub>prev</sub> is a vector of real value v of the directly previous generation of audio samples in the time domain and where π is the circular number.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the overpaste addition output stage (400) is adapted to process the intermediate windowed samples in the time domain in an overpaste manner. , based on the T
123 consecutively provided blocks (320) with audio subband values.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the overpaste addition output stage (400) is adapted to supply the samples in the time domain outi (n), where n is an integer that indicates a sample index based on the equation
Tl out ^ n) = k = 0 where Zi<sub>zn</sub> is an intermediate sample in the windowed time domain corresponding to a sample index in a frame or sequence index 1 in the range 0 to T - 1, where 1 = 0 corresponds to the last frame or sequence and values less than 1 of the frames or sequences previously generated.
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the overpaste addition output stage (400) is adapted to supply the samples in the time domain out (k) based in equation 9 out (k) = w (n · n + k), k = 0 where w is a vector comprising the intermediate windowed samples in the time domain and k is an integer that indicates an index in the range between 0 and (Nl ).
In an equipment for the generation of audio samples in the time domain according to the configurations of the present invention, the equipment (300) is adapted to receive a signal
124 indicative of the analysis window function (190) used to generate the audio subband values, or indicative of the synthesis window function (370) to be used for the generation of audio samples in the time domain.
According to the configurations of the present invention, an encoder (510) comprises equipment (560) for generating audio subband values in audio subband channels according to a configuration of the present invention.
According to the configurations of the present invention, an encoder (510) further comprises a quantizer and encoder (570) coupled to the equipment (560) for the generation of audio subband values and adapted to quantize and encode the sub values - audio band produced by the equipment (560) and producing the encoded quantized audio subband values.
According to the configurations of the present invention, a decoder (580) comprises equipment (600) for the generation of audio samples in the time domain according to a configuration of the present invention.
According to the configurations of the present invention, a decoder (580) further comprises a decoder and quantizer (590) adapted to receive encoded and quantized audio subband values, coupled to the equipment (600) for the generation of audio samples in the time domain and adapted to provide decoded and dequantized audio subband values as the audio subband values for the equipment (600).
According to the configurations of the present invention, an SBR encoder (520) comprises equipment (530) for the
125 generation of audio subband values in audio subband channels, based on a frame of audio input samples in the time domain provided for the SBR encoder (520) and an SBR parameter extraction module ( 540) coupled to the equipment (530) for the generation of audio subband values and adapted for the extraction and generation of SBR parameters based on the audio subband values.
According to the configurations of the present invention, a system (610) comprises equipment (620) for generating audio subband values from a frame of audio input samples in the time domain provided for the system (610); and an equipment (640) for the generation of audio samples in the time domain based on the audio subband values generated by the equipment (640) for the generation of audio subband values.
According to the configurations of the present invention, a system (610) is an SBR decoder.
According to the configurations of the present invention, the system further comprises an HF generator (630) interconnected between the equipment (620) for the generation of audio subband values and the equipment (640) for the generation of audio samples in the time domain and adapted to receive SBR data adapted to modify or add audio subband values based on the SBR data and the audio subband values of the equipment (620) for the generation of subband values of audio.
With reference to all equipment and methods according to the configurations of the present invention, depending on the implementation details, to obtain the window coefficients based on the values given in the tables in Annexes 1 and 3, can be
126 other changes of signs were implemented with reference to the window coefficients corresponding to indexes 128 to 255 and 384 to 511 (multiplication by factor (-1)). In other words, the window coefficients of the window function are based on the window coefficients given in the table in Annex 1. To obtain the window coefficients of the window function shown in the figures, the window coefficients in the table corresponding to the indexes 0 to 127, 256 to 383 and 512 to 639 must be multiplied by (+1) (that is, without changing the sign) and the window coefficients corresponding to indexes 128 to 255 and 384 to 511 must be multiplied by (-1) (that is, a sign change) to obtain the window coefficients of the shown window function. In the same way, the relationships given in the table in Annex 3 must be treated in the same way.
It should be noted that in the structure of the present application according to an equation based on an equation, it includes an introduction of additional delays, factors, other coefficients and the introduction of other simple functions. Then, simple constants, constant addendums, etc. can be abandoned. Furthermore, algebraic transformations, equivalence transformations and approximations (for example, a Taylor approximation) are also included, without altering the result of the equation in any way or in any significant way. In other words, both small changes as well as transformations that lead essentially in terms of an identical result are included in the event that an equation or expression is based on an equation or expression.
Although the foregoing has been particularly shown and described with reference to its particular configurations, it will be understood by those skilled in the art that various other changes in shape and details can be made without abandoning
127 its spirit and scope. It should be understood that several changes can be made in adapting to different configurations without abandoning the broader concept now revealed and contained by the following claims.
128
31 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31
138 members in 26 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 86295406 | United States of America | P | |
| 86295406 | United States of America | P | |
| 862954P | – | – | – |
| US20060862954P | – | – | – |
Members138
| Document | Office | Kind | |
|---|---|---|---|
| AU2007308415A1 | Australia | A1 | |
| AU2007308416A1 | Australia | A1 | |
| CA2645618A1 | Canada | A1 | |
| CA2667505A1 | Canada | A1 | |
| WO2008049589A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2008049590A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200836166A | Taiwan Province of China | A | |
| TW200837719A | Taiwan Province of China | A | |
| MX2008011898A | Mexico | A | |
| KR20080102222A | Republic of Korea | A | |
| EP1994530A1 | European Patent Office (EPO) | A1 | |
| AR063394A1 | Argentina | A1 | |
| AR063400A1 | Argentina | A1 | |
| HK1119824A | Hong Kong, China | A | |
| HK1119824A1 | Hong Kong, China | A1 | |
| CN101405791A | China | A | |
| MX2009004477A | Mexico | A | |
| KR20090058029A | Republic of Korea | A | |
| EP1994530B1 | European Patent Office (EPO) | B1 | |
| EP2076901A1 | European Patent Office (EPO) | A1 | |
| AT435480T | Austria | T | |
| ATE435480T1 | Austria | T1 | |
| NO20084012L | Norway | L | |
| NO20091951L | Norway | L | |
| NO20170452A1 | Norway | A1 | |
| DE602007001460D1 | Germany | D1 | |
| JP2009530675A | Japan | A | |
| DK1994530T3 | Denmark | T3 | |
| PT1994530EThis record | Portugal | E | |
| EP2109098A2 | European Patent Office (EPO) | A2 | |
| ES2328187T3 | Spain | T3 | |
| CN101606194A | China | A | |
| IL197976A0 | Israel | A0 | |
| IL197976D0 | Israel | D0 | |
| US2009319283A1 | United States of America | A1 | |
| ZA200810308B | South Africa | B | |
| PL1994530T3 | Poland | T3 | |
| US2010023322A1 | United States of America | A1 | |
| JP2010507820A | Japan | A | |
| RU2008137468A | Russian Federation | A | |
| ZA200902199B | South Africa | B | |
| KR100957711B1 | Republic of Korea | B1 | |
| AU2007308416B2 | Australia | B2 | |
| AU2007308415B2 | Australia | B2 | |
| RU2009119456A | Russian Federation | A | |
| MY142520A | Malaysia | A | |
| RU2411645C2 | Russian Federation | C2 | |
| RU2420815C2 | Russian Federation | C2 | |
| BRPI0709310A2 | Brazil | A2 | |
| KR101056253B1 | Republic of Korea | B1 | |
| IL193786A | Israel | A | |
| TWI355649B | Taiwan Province of China | B | |
| CN101405791B | China | B | |
| TWI357065B | Taiwan Province of China | B | |
| JP4936569B2 | Japan | B2 | |
| CN101606194B | China | B | |
| JP5083779B2 | Japan | B2 | |
| CA2645618C | Canada | C | |
| US8438015B2 | United States of America | B2 | |
| US8452605B2 | United States of America | B2 | |
| MY148715A | Malaysia | A | |
| US2013238343A1 | United States of America | A1 | |
| IL197976A | Israel | A | |
| US8775193B2 | United States of America | B2 | |
| CA2667505C | Canada | C | |
| EP2076901B1 | European Patent Office (EPO) | B1 | |
| BRPI0716315A2 | Brazil | A2 | |
| EP2109098A3 | European Patent Office (EPO) | A3 | |
| EP2076901B8 | European Patent Office (EPO) | B8 | |
| PT2076901T | Portugal | T | |
| ES2631906T3 | Spain | T3 | |
| PL2076901T3 | Poland | T3 | |
| NO341567B1 | Norway | B1 | |
| NO341610B1 | Norway | B1 | |
| EP3288027A1 | European Patent Office (EPO) | A1 | |
| NO342691B1 | Norway | B1 | |
| HK1251073A | Hong Kong, China | A | |
| HK1251073A1 | Hong Kong, China | A1 | |
| BRPI0709310B1 | Brazil | B1 | |
| EP2109098B1 | European Patent Office (EPO) | B1 | |
| PT2109098T | Portugal | T | |
| PL2109098T3 | Poland | T3 | |
| EP3288027B1 | European Patent Office (EPO) | B1 | |
| ES2834024T3 | Spain | T3 | |
| PT3288027T | Portugal | T | |
| EP3848928A1 | European Patent Office (EPO) | A1 | |
| PL3288027T3 | Poland | T3 | |
| ES2873254T3 | Spain | T3 | |
| EP3848928B1 | European Patent Office (EPO) | B1 | |
| FI3848928T3 | Finland | T3 | |
| PT3848928T | Portugal | T | |
| DK3848928T3 | Denmark | T3 | |
| EP4207189A1 | European Patent Office (EPO) | A1 | |
| PL3848928T3 | Poland | T3 | |
| ES2947516T3 | Spain | T3 | |
| EP4207189B1 | European Patent Office (EPO) | B1 | |
| EP4207189C0 | European Patent Office (EPO) | C0 | |
| EP4300824A2 | European Patent Office (EPO) | A2 | |
| EP4300825A2 | European Patent Office (EPO) | A2 | |
| EP4325723A2 | European Patent Office (EPO) | A2 |
Numbers
- Publication, DOCDB
- 1994530
- Publication, EPODOC
- PT1994530E
- Application
- 7819260
- Application, DOCDB
- 07819260
- Application, EPODOC
- PT20070819260T
Titles2
- English
- APPARATUS AND METHOD FOR GENERATING AUDIO SUBBAND VALUES AND APPARATUS AND METHOD FOR GENERATING TIME-DOMAIN AUDIO SAMPLES
- Portuguese
- EQUIPAMENTO E MÉTODO PARA A GERAÇÃO DE VALORES DE SUB-BANDA DE ÁUDIO E EQUIPAMENTO E MÉTODO PARA A GERAÇÃO DE AMOSTRAS DE ÁUDIO NO DOMÍNIO DO TEMPO
Classification
- CPC, 8
- G10L19/022
- G10L19/0204
- G10L19/02
- H03H17/0266
- G10L21/038
- G10L25/45
- G11B20/10
- H03M7/30
- IPC, 2
- G10L19 02
- H03H17 02
