Audio splicing concept.
30 claims: 7 independent, 23 dependent
- 1REIVINDICACIONES 1. Un flujo de datos de audio empalmadle (40), que comprende:una secuencia de paquetes de carga útil (16), cada uno de los paquetes de carga útil pertenecen a una respectiva de una secuencia de unidades de acceso (18) en las que el flujo de datos de audio empalmable está particionado, cada unidad de acceso está asociada con una respectiva de las tramas de audio (14) de una señal de audio (12) que está codificada en el flujo de datos de audio empalmable en unidades de las tramas de audio;y un paquete de unidad de truncamiento (42;58) insertado en el flujo de datos de audio empalmable y ajustable con el fin de indicar, para una unidad de acceso predeterminada, una porción de extremo (44;56) de una trama de audio con la que la unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión.
- 2El flujo de datos de audio empalmable de acuerdo con la reivindicación 1, en donde la porción de extremo es una porción de extremo delantero (56) y la unidad de acceso predeterminada tiene codificada en su interior la trama de audio asociada respectiva de una manera tal que la reconstrucción de los mismos en el lado de decodificación es independiente de la unidad de acceso inmediatamente anterior a la unidad de acceso predeterminada, lo que de ese modo permite una emisión inmediata.
- 3El flujo de datos de audio empalmable de acuerdo con la reivindicación 1, en el que el flujo de datos de audio empalmable además comprende:un paquete de unidad de truncamiento adicional (58) insertado en el flujo de datos de audio empalmable y ajustable con el fin de indicar una unidad de acceso predeterminada adicional, una porción de extremo (44;56) de una trama de audio adicional con el que la unidad de acceso predeterminada adicional está asociada, como para ser descartados durante la emisión.
- 4El flujo de datos de audio empalmable de acuerdo con la reivindicación 3, en el que la unidad de acceso predeterminada tiene codificada en su interior la trama de audio asociada respectiva de una manera tal que una reconstrucción de los mismos en el lado de decodificación es dependiente de una unidad de acceso inmediatamente anterior a la unidad de acceso predeterminada, y la unidad de acceso predeterminada adicional tiene codificada en su interior la trama de audio asociada respectiva de una manera tal que la reconstrucción de los mismos en el lado de decodificación es independiente de la unidad de acceso inmediatamente anterior a la unidad de acceso predeterminada adicional, lo que de ese modo permite una emisión inmediata.
- 5El flujo de datos de audio empalmable de acuerdo con la reivindicación'4, en donde una mayoría de las unidades de acceso tiene codificada en su interior la trama de audio asociada respectiva de una manera tal que la reconstrucción de los mismos en el lado de decodificación es dependiente de la unidad de acceso inmediatamente anterior respectiva.
- 6El flujo de datos de audio empalmable de acuerdo con la reivindicación 4 o 5, en donde el paquete de unidad de truncamiento (42) y el paquete de unidad de truncamiento adicional (58) comprenden un elemento de sintaxis de empalme de salida (50), respectivamente, que indica si el respectivo del paquete de unidad de truncamiento o el paquete de unidad de truncamiento adicional se refiere a una unidad de acceso de empalme de salida o no, en el que el elemento de sintaxis de empalme de salida (50) comprendido por el paquete de unidad de truncamiento indica que el paquete de unidad de truncamiento se refiere a una unidad de acceso de empalme de salida y el elemento de sintaxis compuesto por el paquete de unidad de truncamiento adicional indica que el paquete de unidad de truncamiento adicional no se refiere a una unidad de acceso de empalme de salida.
- 7El flujo de datos de audio empalmable de acuerdo con la reivindicación 4 o 5, en donde el paquete de unidad de truncamiento (42) y el paquete de unidad de truncamiento adicional (58) comprenden un elemento de sintaxis de empalme de salida, respectivamente, que indica si el respectivo del paquete de unidad de truncamiento o el paquete de unidad de truncamiento adicional se refiere a una unidad de acceso de empalme de salida o no, en el que el elemento de sintaxis (50) compuesto por el paquete de unidad de truncamiento indica que el paquete de unidad de truncamiento se refiere a una unidad de acceso de empalme de salida y el elemento de sintaxis de empalme de salida compuesto por el paquete de unidad de truncamiento adicional indica que el paquete de unidad de truncamiento adicional se refiere a una unidad de acceso de empalme de salida, también, en el que el paquete de unidad de truncamiento adicional además comprende un elemento de sintaxis de truncamiento de extremo delantero/trasero (54) y un elemento de longitud de truncamiento (48), en el que el elemento de sintaxis de truncamiento de extremo delantero/trasero es para indicar si la porción de extremo de la trama de audio adicional es una porción de extremo trasero (44) o una porción de extremo delantero (56) y el elemento de longitud de truncamiento es para indicar una longitud (Át) de la porción de extremo de la trama de audio adicional.
- 8Flujo de datos de audio empalmado, que comprende:una secuencia de paquetes de carga útil (16), cada uno de los paquetes de carga útil pertenecen a una respectiva de una secuencia de unidades de acceso (18) en la que el flujo de datos de audio empalmado está particionado, cada unidad de 100 acceso está asociada con una respectiva de las tramas de audio (14) ;un paquete de unidad de truncamiento (42;58;114) insertado en el flujo de datos de audio empalmado y que indica una porción de extremo (44;56) de una trama de audio con la que una unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión, en el que en una primera subsecuencia de paquetes de carga útil de la secuencia de paquetes de carga útil, cada paquete de carga útil pertenece a una unidad de acceso (AU#) de un primer flujo de datos de audio que tiene codificada en su interior una primera señal de audio en unidades de tramas de audio de la primera señal de audio, y las unidades de acceso del primer flujo de datos de audio, que incluyen la unidad de acceso predeterminada, y en una segunda subsecuencia de paquetes de carga útil de la secuencia de paquetes de carga útil, cada paquete de carga útil pertenece a unidades de acceso (AU'#) de un segundo flujo de datos de audio que tiene codificada en su interior una segunda señal de audio en unidades de tramas de audio del segundo flujo de datos de audio, en el que la primera y la segunda subsecuencia de paquetes de carga útil son inmediatamente consecutivas una con respecto a la otra y hacen tope entre si en la unidad de acceso predeterminada y la porción de extremo es una porción de 101 extremo trasero (44) en el caso de la primera subsecuencia que precede a la segunda subsecuencia y una porción de extremo delantero (56) en el caso de la segunda subsecuencia que precede a la primera subsecuencia;en donde la primera subsecuencia precede la segunda subsecuencia y el flujo de datos de audio empalmado además comprende un paquete de unidad de truncamiento adicional (58) insertado en el flujo de datos de audio empalmado y que indica una porción de extremo delantero (58) de una trama de audio adicional con la que una unidad de acceso predeterminada adicional está asociada, como para ser descartados durante la emisión, en el que en una tercera subsecuencia de paquetes de carga útil de la secuencia de paquetes de carga útil, cada paquete de carga útil pertenece a unidades de acceso (AU#) del primer flujo de datos de audio, después de las unidades de acceso del primer flujo de datos de audio a la que pertenecen los paquetes de carga útil de la primera subsecuencia, en el que las unidades de acceso del primer flujo de datos de audio incluyen la unidad de acceso predeterminada adicional.
- 9El flujo de datos de audio empalmado de acuerdo con la reivindicación 8, en el que una mayoría de las unidades de acceso del flujo de datos de audio empalmado que incluyen la unidad de acceso predeterminada tiene codificada en su interior la trama de audio asociada respectiva de una manera tal que una reconstrucción de los mismos en el lado de decodificación 102 es dependiente de una respectiva unidad de acceso inmediatamente anterior,
- 10El flujo de datos de audio empalmado de acuerdo con la reivindicación 9, en donde la unidad de acceso inmediatamente posterior a la unidad de acceso predeterminada y que forma un inicio de las unidades de acceso del segundo flujo de datos de audio tiene codificada en su interior la trama de audio asociada respectiva en una manera tal que la reconstrucción de la misma es independiente de la unidad de acceso predeterminada, lo que de ese modo permite una emisión inmediata, y la unidad de acceso predeterminada adicional tiene codificada en su interior la trama de audio adicional en una manera tal que la reconstrucción de la misma es independiente de la unidad de acceso inmediatamente anterior a la unidad de acceso predeterminada adicional, lo que de ese modo permite una emisión inmediata, respectivamente.
- 11El flujo de datos de audio empalmado de acuerdo con la reivindicación 8 o 10, en el que el flujo de datos de audio empalmado además comprende un paquete de unidad de truncamiento incluso adicional (114) insertado en el flujo de datos de audio empalmado y que indica una porción de extremo trasero (44) de una trama de audio incluso adicional con la que la unidad de acceso inmediatamente anterior a la unidad de acceso predeterminada adicional está asociada, como para ser descartados durante la emisión, en el que el flujo de datos de 103 audio empalmado comprende información de marca de tiempo (24) que indica para cada unidad de acceso del flujo de datos de audio empalmado una marca de tiempo correspondiente en la que la trama de audio con la que la unidad de acceso respectiva está asociada, se va a emitir, en el que una marca de tiempo de la unidad de acceso predeterminada adicional es igual a la marca de tiempo de la unidad de acceso inmediatamente anterior a la unidad de acceso predeterminada adicional más una longitud temporal de la trama de audio con la que la unidad de acceso inmediatamente anterior a la unidad de acceso predeterminada adicional está asociada, menos la suma de una longitud temporal de la porción de extremo delantero de la trama de audio adicional y la porción de extremo trasero de la trama de audio incluso adicional.
- 12El flujo de datos de audio empalmado de acuerdo con la reivindicación 10, en el que una marca de tiempo temporal de la unidad de acceso inmediatamente posterior a la unidad de acceso predeterminada es igual a la marca de tiempo de la unidad de acceso predeterminada más una longitud temporal de la trama de audio con la que la unidad de acceso predeterminada está asociada, menos una longitud temporal de la porción de extremo trasero de la trama de audio con la que la unidad de acceso predeterminada está asociada.
- 13Empalmador de flujo para el empalme de flujos de datos de audio, que comprende:104 una primera interfaz de entrada de audio (102) para la recepción de un primer flujo de datos de audio (40) que comprende una secuencia de paquetes de carga útil (16), cada uno de los cuales pertenece a una respectiva de una secuencia de unidades de acceso (18) en el que el primer flujo de datos de audio está particionado, cada unidad de acceso del primer flujo de datos de audio está asociada con una respectiva de las tramas de audio (14) de una primera señal de audio (12) que está codificada en el primer flujo de datos de audio en unidades de tramas de audio de la primera señal de audio;una segunda interfaz de entrada de audio (104) para la recepción de un segundo flujo de datos de audio (110) que comprende una secuencia de paquetes de carga útil, cada uno de los cuales pertenece a una respectiva de una secuencia de unidades de acceso en la que el segundo flujo de datos de audio está particionado, cada unidad de acceso del segundo flujo de datos de audio está asociada con una respectiva de las tramas de audio de una segunda señal de audio que está codificada en el segundo flujo de datos de audio en unidades de tramas de audio de la segunda señal de audio;un regulador del punto de empalme;y un multiplexor de empalme, en donde el primer flujo de datos de audio además comprende un paquete de unidad de truncamiento (42;58) insertado en el primer flujo de datos de audio y ajustable con 105 el fin de indicar una unidad de acceso predeterminada, una porción de extremo (44;56) de una trama de audio con la que la unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión, y el regulador del punto de empalme (106) está configurado para establecer el paquete de unidad de truncamiento (42;58) de una manera tal que el paquete de unidad de truncamiento indique una porción de extremo (44;56) de la trama de audio con la que la unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión, o el regulador del punto de empalme (106) está configurado para insertar un paquete de unidad de truncamiento (42;58) en el primer flujo de datos de audio y configura el mismo con el fin de indicar una unidad de acceso predeterminada, una porción de extremo (44;56) de una trama de audio con la que la unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión;y en donde el multiplexor de empalme (108) está configurado para cortar el primer flujo de datos de audio (40) en la unidad de acceso predeterminada con el fin de obtener una subsecuencia de paquetes de carga útil del primer flujo de datos de audio dentro de la cual cada paquete de carga útil pertenece a una unidad de acceso respectiva de una serie de unidades de acceso del primer flujo de datos de audio, que incluyen la unidad de acceso predeterminada, y empalmar la subsecuencia de paquetes de carga útil del primer flujo de 106 datos de audio y la secuencia de paquetes de carga útil del segundo flujo de datos de audio de una manera tal que las mismas sean inmediatamente consecutivas una con respecto a la otra y hagan tope entre sí en la unidad de acceso predeterminada, en el que la porción de extremo de la trama de audio con la que la unidad de acceso predeterminada está asociada es una porción de extremo trasero (44) en el caso de la subsecuencia de paquetes de carga útil del primer flujo de datos de audio que precede a la secuencia de paquetes de carga útil del segundo flujo de datos de audio y una porción de extremo delantero (56) en el caso de la subsecuencia de paquetes de carga útil del primer flujo de datos de audio que sucede a la secuencia de paquetes de carga útil del segundo flujo de datos de audio.
- 14El empalmador de flujo de acuerdo con la reivindicación 13, en donde la subsecuencia de paquetes de carga útil del primer flujo de datos de audio precede la secuencia de paquetes de carga útil del segundo flujo de datos de audio y la porción de extremo de la trama de audio con la que la unidad de acceso predeterminada está asociada es una porción de extremo trasero (44).
- 15El empalmador de flujo de acuerdo con la reivindicación 13, en el que el regulador del punto de empalme está configurado para establecer una longitud temporal de la porción de extremo de una manera tal que coincida con un reloj 107 externo.
- 16El empalmador de flujo de acuerdo con la reivindicación 14, en el que el segundo flujo de datos de audio tiene, o el regulador del punto de empalme (106) provoca por medio de inserción, un paquete de unidad de truncamiento adicional (114) insertado en el segundo flujo de datos de audio (110) y ajustable con el fin de indicar una porción de extremo de una trama de audio adicional con la que una unidad de acceso de terminación del segundo flujo de datos de audio (110) está asociada, como para ser descartados durante la emisión, y el primer flujo de datos de audio además comprende un paquete de unidad de truncamiento incluso adicional (58) insertado en el primer flujo de datos de audio (40) y ajustable con el fin de indicar una porción de extremo de una trama de audio incluso adicional con la que la unidad de acceso predeterminada incluso adicional está asociada, como para ser descartados durante la emisión, en el que una distancia temporal entre la trama de audio de la unidad de acceso predeterminada y la trama de audio incluso adicional de la unidad de acceso predeterminada incluso adicional coincide con una longitud temporal de la segunda señal de audio entre una unidad de acceso delantera del mismo sucede, después del empalme, a la unidad de acceso predeterminada y la unidad de acceso trasera, en el que el regulador del punto de empalme (106) está configurado para establecer el paquete de unidad de truncamiento adicional (114) 108 de una manera tal que el mismo indique una porción de extremo trasero (44) de la trama de audio adicional como para ser descartados durante la emisión, y el paquete de unidad de truncamiento incluso adicional (58) de una manera tal que el mismo indique una porción de extremo delantero de la trama de audio incluso adicional como para ser descartados durante la emisión, en el que el multiplexor de empalme (108) está configurado para adaptar la información de marca de tiempo (24) compuesta por el segundo flujo de datos de audio (110) y que indica para cada unidad de acceso una marca de tiempo correspondiente a la que la trama de audio con la que la unidad de acceso respectiva está asociada, se va a emitir, de una manera tal que una marca de tiempo de una trama de audio delantera con la que la unidad de acceso delantera del segundo flujo de datos de audio (110) está asociada coincide con la marca de tiempo de la trama de audio con la que la unidad de acceso predeterminada está asociada más la longitud temporal de la trama de audio con la que la unidad de acceso predeterminada está asociada menos la longitud temporal de la porción de extremo trasero de la trama de audio con la que la unidad de acceso predeterminada está asociada y el regulador del punto de empalme (106) está configurado para establecer el paquete de unidad de truncamiento adicional (114) y el paquete de unidad de truncamiento incluso adicional (58) de una manera tal que una marca de tiempo de la trama de audio incluso 109 adicional sea igual a la marca de tiempo de la trama de audio adicional más una longitud temporal de la trama de audio adicional menos la suma de una longitud temporal de la porción de extremo trasero de la trama de audio adicional y la porción de extremo delantero de la trama de audio incluso adicional.
- 17El empalmador de flujo de acuerdo con la reivindicación 14, en el que el segundo flujo de datos de audio (110) tiene, o el regulador del punto de empalme (106) provoca por medio de inserción, un paquete de unidad de truncamiento adicional (112) insertado en el segundo flujo de datos de audio el cual es ajusfadle con el fin de indicar una porción de extremo de una trama de audio adicional con la que una unidad de acceso delantera del segundo flujo de datos de audio está asociada, como para ser descartados durante la emisión, en el que el regulador del punto de empalme (106) está configurado para establecer el paquete de unidad de truncamiento adicional (112) de una manera tal que el mismo indique una porción de extremo delantero de la trama de audio adicional como para ser descartados durante la emisión, en el que la información de marca de tiempo (24) compuesta por el primer y el segundo flujo de datos de audio y que indica para cada unidad de acceso una marca de tiempo correspondiente a la que la trama de audio con la que la unidad de acceso respectiva del primer y el segundo flujo de datos de audio está asociada, se va a emitir, están alineados temporalmente y el regulador del punto de empalme 110 (106) está configurado para establecer el paquete de unidad de truncamiento adicional de una manera tal que una marca de tiempo de la trama de audio adicional menos una longitud temporal de la trama de audio con la que la unidad de acceso predeterminada está asociada más una longitud temporal de la porción de extremo delantero sea igual a la marca de tiempo de la trama de audio con la que la unidad de acceso predeterminada está asociada-más una longitud temporal de la trama de audio con la que la unidad de acceso predeterminada está asociada menos la longitud temporal de la porción de extremo trasero.
- 18Decodificador de audio que comprende:un núcleo de decodificación de audio (162) configurado para reconstruir una señal de audio (12) , en unidades de tramas de audio (14) de la señal de audio, a partir de una secuencia de paquetes de carga útil (16) de un flujo de datos de audio (120) , en donde cada uno de los paquetes de carga útil pertenece a una respectiva de una secuencia de unidades de acceso (18) en la que el flujo de datos de audio está particionado, en el que cada unidad de acceso está asociada con una respectiva de las tramas de audio;y un truncador de audio (164) configurado para ser sensible a un paquete de unidad de truncamiento (42;58;114) insertado en el flujo de datos de audio con el fin de truncar una trama de audio asociada con una unidad de acceso predeterminada con el fin de descartar, durante la emisión de 111 la señal de audio, una porción de extremo de la misma indicada para ser descartada durante la emisión por el paquete de unidad de truncamiento.
- 19El decodificador de audio de acuerdo con la reivindicación 18, en donde el paquete de unidad de truncamiento (42) comprende un elemento de sintaxis de truncamiento de extremo delantero/trasero (54) y un elemento de longitud de truncamiento (48), en donde el decodificador utiliza el elemento de sintaxis de truncamiento de extremo delantero/trasero como una indicación si la porción de extremo es una porción de extremo trasero (44) o una porción de extremo delantero (56) y el elemento de longitud de truncamiento como una indicación de una longitud (Át) de la porción de extremo de la trama de audio.
- 20Codificador de audio que comprende:un núcleo de codificación de audio (72) configurado para codificar una señal de audio (12), en unidades de tramas de audio (14) de la señal de audio, en paquetes de carga útil (16) de un flujo de datos de audio (40) de una manera tal que cada paquete de carga útil pertenezca a una respectiva de las unidades de acceso (18) en las que el flujo de datos de audio está particionado, cada unidad de acceso está asociada con una respectiva de las tramas de audio, y 112 un insertador de paquetes de truncamiento (74) configurado para insertar en el flujo de datos de audio un paquete de unidad de truncamiento (44;58) que es ajustable con el fin de indicar una porción de extremo de una trama de audio con la que una unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión.
- 21El decodificador de audio de acuerdo con la reivindicación 20, configurado para realizar un control de tasa de modo que una tasa de bits del flujo de datos de audio varia alrededor, y obedece a, una tasa de bits media predeterminada de modo que una desviación de la tasa de bits integrado de la tasa de bits media predeterminada asume, en la unidad de acceso predeterminada, un valor dentro de un intervalo predeterminado que es menor que la de ancho que un intervalo de la desviación de la tasa de bits integrada mientras varia sobre el flujo de datos de audio empalmable completo.
- 22El decodificador de audio de acuerdo con la reivindicación 20 o 21, configurado para realizar un control de tasa de modo que una tasa de bits del flujo de datos de audio varia alrededor, y obedece a, una tasa de bits media predeterminada de modo que una desviación de la tasa de bits integrado de la tasa de bits media predeterminada asume, en la unidad de acceso predeterminada, un valor fijo menor que H de un máximo de la desviación de la tasa de bits integrada mientras varia sobre el flujo de datos de audio empalmable completo. 113
- 23El decodificador de audio de acuerdo con cualquiera de las reivindicaciones 20 a 22, configurado para realizar un control de tasa de modo que una tasa de bits del flujo de datos de audio varia alrededor, y obedece a, una tasa de bits media predeterminada de modo que una desviación de la tasa de bits integrado de la tasa de bits media predeterminada asume, en la unidad de acceso predeterminada así como otras unidades de acceso para las cuales paquetes de unidad de truncamiento son insertadas en el flujo de datos de audio, un valor predeterminado.
- 24El decodificador de audio de acuerdo con cualquiera de las reivindicaciones 20 a 23, configurado para realizar un control de tasa al registrar un nivel de llenado de un búfer de audio codificado, de modo que un nivel de llenado registrado asume, en la unidad de acceso predeterminada, un valor predeterminado.
- 25El decodificador de audio de acuerdo con la reivindicación 24, en donde el valor predeterminado es común entre unidades de acceso para las cuales paquetes de unidad de audio que comprende un primer flujo de datos de audio (40) que 114 comprende una secuencia de paquetes de carga útil (16), cada uno de los cuales pertenece a una respectiva de una secuencia de unidades de acceso (18) en la que el primer flujo de datos de audio está particionado, cada unidad de acceso del primer flujo de datos de audio está asociada con una respectiva de las tramas de audio (14) de una primera señal de audio (12) que está codificada en el primer flujo de datos de audio en unidades de tramas de audio de la primera señal de audio;y un segundo flujo de datos de audio (110) que comprende una secuencia de paquetes de carga útil, cada uno de los cuales pertenece a una respectiva de una secuencia de unidades de acceso en la que el segundo flujo de datos de audio está particionado, cada unidad de acceso del segundo flujo de datos de audio está asociada con una respectiva de las tramas de audio de una segunda señal de audio que está codificada en el segundo flujo de datos de audio en unidades de tramas de audio de la segunda señal de audio;en donde el primer flujo de datos de audio además comprende un paquete de unidad de truncamiento (42;58) insertado en el primer flujo de datos de audio y ajustable con el fin de indicar una unidad de acceso predeterminada, una porción de extremo (44;56) de una trama de audio con la cual la unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión, y el método comprende el establecimiento del 115 paquete de unidad de truncamiento (42;58) de una manera tal que el paquete de unidad de truncamiento indique una porción de extremo (44;56) de la trama de audio con la que la unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión, o el método comprende la inserción de un paquete de unidad de truncamiento (42;58) en el primer flujo de datos de audio y establece el mismo con el fin de indicar una unidad de acceso predeterminada, una porción de extremo (44;56) de una trama de audio con la que una unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión y el establecimiento del paquete de unidad de truncamiento (42;58) de una manera tal que el paquete de unidad de truncamiento indique una porción de extremo (44;56) de la trama de audio con la que la unidad de acceso predeterminada está asociada, como para ser descartados durante la emisión;y el método además comprende el corte del primer flujo de datos de audio (40) en la unidad de acceso predeterminada con el fin de obtener una subsecuencia de paquetes de carga útil del primer flujo de datos de audio dentro de la cual cada paquete de carga útil pertenece a una unidad de acceso respectiva de una serie de unidades de acceso del primer flujo de datos de audio, que incluyen la unidad de acceso predeterminada, y el empalme de la subsecuencia de paquetes de carga útil del primer flujo de datos de audio y la secuencia 116 de paquetes de carga útil del segundo flujo de datos de audio de una manera tal que las mismas sean inmediatamente consecutivas una con respecto a la otra y hagan tope entre sí en la unidad de acceso predeterminada, en el que la porción de extremo de la trama de audio con la que la unidad de acceso predeterminada está asociada es una porción de extremo trasero (44) en el caso de la subsecuencia de paquetes de carga útil del primer flujo de datos de audio que precede a la secuencia de paquetes de carga útil del segundo flujo de datos de audio y una porción de extremo delantero (56) en el caso de la subsecuencia de paquetes de carga útil del primer flujo de datos de audio sucede a la secuencia de paquetes de carga útil del segundo flujo de datos de audio.
- 2628. Método de decodificación de audio que comprende:la reconstrucción de una señal de audio (12), en unidades de tramas de audio (14) de la señal de audio, a partir de una secuencia de paquetes de carga útil (16) de un flujo de datos de audio (120), en el que cada uno de los paquetes de carga útil pertenece a una respectiva de una secuencia de unidades de acceso (18) en la que el flujo de datos de audio está particionado, en el que cada unidad de acceso está asociada con una respectiva de las tramas de audio;y siendo sensible a un paquete de unidad de truncamiento (42;58;114) insertado en el flujo de datos de audio, el truncamiento de una trama de audio asociada con una unidad de 117 acceso predeterminada con el fin de descartar, durante la emisión de la señal de audio, una porción de extremo de la misma indicada para ser descartada durante la emisión por el paquete de unidad de truncamiento.
- 2729. Método de codificación de audio que comprende:la codificación de una señal de audio (12), en unidades de tramas de audio (14) de la señal de audio, en paquetes de carga útil (16) de un flujo de datos de audio (40) de una manera tal que cada paquete de carga útil pertenezca a una respectiva de las unidades de acceso (18) en el que el flujo de datos de audio está particionado, cada unidad de acceso está asociada con una respectiva de las tramas de audio, y la inserción en el flujo de datos de audio de un paquete de unidad de truncamiento (44;58) que es ajustable con el fin de indicar una porción de extremo de una trama de audio con la que una unidad de acceso predeterminada está asociada, como para ser descartada durante la emisión.
- 2830. Medio de almacenamiento digital legible por computadora para el empalme de flujos de datos de audio que comprende el método de la reivindicación 27.
- 2931. Medio de almacenamiento digital legible por computadora para decodificación de audio que comprende el método de la reivindicación 28. *
- 3032. Medio de almacenamiento digital legible por computadora para codificación de audio que comprende el método de la reivindicación 29.
Independent claims30
228 paragraphs in 1 section, as filed
CONCEPT OF AUDIO EMPALME
Field of the Invention
This application refers to the audio splice.
Background of the Invention
The encoded audio usually comes in pieces of samples, often 1024, 2048 or 4096 samples in number per piece. Such pieces are called frames herein. In the context of MPEG, audio codees such as AAC or MPEG-H 3D Audio, these chunks / frames are called granules, scrambled pieces / frames are called access units (AU) and decoded chunks they are called composition units (CU). In the systems of transport of the audio signal it is only accessible and addressable in granularity of these coded pieces (access units). However, it would be favorable to be able to deal with audio data in some final granularity, especially for purposes such as splicing of streams or changes in the configuration of audio data, encoded, synchronous and aligned with another stream such as a stream of video, for example.
What is known so far is the discarding of some samples of a coding unit. The MPEG-4 file format, for example, has the so-called edit lists that can be used in order to discard audio samples at the beginning and end of an encoded audio file / bit stream [3]. Disadvantageously, this method editing list only works with the MPEG-4 file format, that is, it is the specific file format and does not work with flow formats such as MPEG-2 transport streams. Beyond that, the editing lists are deeply rooted in the MPEG4 file format and consequently cannot be easily modified on the fly by flow splicing devices. In AAC [1], the truncation information can be inserted into the data stream in the form of extension_payload. However, such a payload extension in an encoded AAC access unit has the disadvantage that the truncation information is deeply embedded in the AAC AU and cannot be easily modified on the fly by flow splicing devices.
Accordingly, an objective of the present invention is to provide a concept for audio splicing that is more efficient in terms of, for example, the procedural complexity of the splicing process in flow splicers, and / or audio decoders.
This objective is achieved by means of the subject matter of the independent claims appended hereto.
Brief Description of the Invention
The invention of the present application is inspired by the idea that audio splicing can be made more efficient by the use of one or more truncation unit packets inserted in the audio data stream in order to indicate a decoder of audio, for a predetermined access unit, an end portion of an audio frame with which the predetermined access unit is associated, such as to be discarded during broadcast.
In accordance with one aspect of the present application, an audio data stream is initially provided with such a truncation unit packet in order to process the audio data stream provided thereby spliced more easily in the access unit. default to a temporal granularity thinner than the length of the audio frame. The one or more truncation unit packets are, therefore, directed to the audio decoder and flow splicer, respectively. According to some embodiments, a flow splicer simply searches for such a truncation unit package in order to locate a possible splice point. The flow splicer sets the truncation unit packet accordingly in order to indicate an end portion of the audio frame with which the predetermined access unit is associated, such as to be
<img file="MX366276B_D0001.tif" />
discarded during broadcast, cut off the first audio data stream in the default access unit and splice the audio data stream with another audio data stream to support each other in the default access unit. Since the truncation unit packet is already provided within the splicing audio data stream, there is no additional data to be inserted by the splicing process and, consequently, the bit rate consumption remains unchanged to this extent. .
Alternatively, a truncation unit package can be inserted at the time of splicing. Regardless of initially providing an audio data stream with a truncation unit packet or the proportion thereof with a truncation unit packet at the time of splicing, a spliced audio data stream has such a truncation unit packet inserted with its interior the end portion being a rear end portion in case the predetermined access unit is part of the audio data stream that goes in front of the point of splicing and a leading end portion in the event that the predetermined access unit is part of the audio data stream the splice point happens.
Brief Description of the Figures
The advantageous aspects of the implementations of the
MX / a / 2017/002815 This application is subject to the dependent claims. In particular, preferred embodiments of the present application are described below with respect to the figures, among which:
Figure 1 schematically shows from top to bottom an audio signal, the audio data stream that has the encoded audio signal inside in audio frame units of the audio signal, a video consisting of a sequence of frames and other audio data stream and its encoded audio signal therein which are to potentially replace the initial audio signal of a particular video frame in the future;
Figure 2 shows a schematic diagram of a splicing audio data stream, that is, an audio data stream provided with TU packets in order to alleviate splicing actions, in accordance with an embodiment of the present application. ;
Figure 3 shows a schematic diagram illustrating a TU package according to one embodiment;
Figure 4 schematically shows a TU package, according to an alternative embodiment according to which the TU package is capable of indicating a front end portion and a rear end portion, respectively;
Figure 5 shows a block diagram of an audio encoder according to one embodiment;
Figure 6 shows a schematic diagram illustrating an activation source for moments of input splice and output splice according to an embodiment where they depend on a raster of audio frames;
Figure 7 shows a schematic block diagram of a flow splicer according to an embodiment with the figure showing, in addition, the flow splicer during the reception of the audio data flow of Figure 2 and the output of a flow spliced audio data based on it;
Figure 8 shows a flow chart of the operating mode of the flow splicer of Figure 7 during the splicing of the lower to upper audio data stream, in accordance with one embodiment;
Figure 9 shows a flow chart of the operation mode of the flow splicer during the splicing of the lower audio data stream back to the upper one, in accordance with one embodiment;
Figure 10 shows a block diagram of an audio decoder according to an embodiment, which further illustrates the audio decoder during reception of the spliced audio data stream shown in Figure 7;
Figure 11 shows a flow chart of an operation mode of the audio decoder of Figure 10 in order to illustrate the different access unit handling as a function of which are IPF access units and / or access units which comprise TU packages;
Figure 12 shows an example of a TU packet syntax;
Figures 13A to 13C show different examples of how to splice from one audio data stream to the other, with the moment of splicing time determined by a video, herein a video at 50 frames per second and an audio signal encoded in the audio data streams at 48 kHz with 1024 grains of the entire sample or audio frames and with a time stamp time base of 90 kHz in such a way that a video frame duration is equal to 1,800 time base signals while an audio frame or audio granule is equal to 1920 time base signals;
Figure 14 shows a schematic diagram illustrating another representative case of the splicing of two audio data streams in an instant of time determined by a raster splicing of audio frames by the use of the representative plot and the sample rates of the Figures . 13A to C;
Figure 15A-15B shows a schematic diagram illustrating an action of the encoder in the splicing of two audio data streams of different encoding configurations in accordance with one embodiment;
Figure 16A-16B shows different cases of splice use according to one embodiment; Y
Figure 17 shows a block diagram of an audio encoder that supports different encoding configurations in accordance with one embodiment.
Detailed description of the invention
Figure 1 shows a representative portion of an audio data stream in order to illustrate the problems that occur when it comes to splicing the respective audio data stream with another audio data stream. To this extent, the audio data stream of Figure 1 forms a kind of basis for the audio data streams shown in the following figures. Consequently, the description presented below with the audio data stream of Figure 1 is also valid for the audio data streams described below.
The audio data stream of Figure 1 is generally indicated by the reference sign 10. The audio data stream has an audio signal 12 encoded therein. In particular, the audio signal 12 is encoded in the audio data stream in audio frame units 14, that is, temporary portions of the audio signal 12 that can, as illustrated in Figure 1, be do not overlap and butt each other temporarily, or, alternatively, overlap each other. The way in which the audio signal 12 is encoded, in units of the audio frames 14, and the audio data stream 10 can be chosen differently: transformation coding can be used in order to encode the audio signal. audio in the units of the audio frames 14 in the data stream 10. In that case, one or more spectral decomposition transformations may be applied to the audio signal of the audio frame 14, with one or more spectral decomposition transformations temporarily covering the audio frame 14 and extending beyond its front end and rear. The spectral decomposition transformation coefficients are contained within the data flow in such a way that the decoder is capable of reconstructing the respective frame by means of inverse transformation. Portions of overlapping transformations mutually and even beyond the limits of the audio frame in units of which the audio signal is spectrally decomposed are placed in windows with the so-called window functions on the side of the encoder and / or the decoder in such a way that a process called overlap-summing on the decoder side according to which the spectral composition transformations indicated and transformed conversely they overlap each other and are added, it reveals the reconstruction of the audio signal 12.
Alternatively, for example, the audio data stream 10 has the audio signal 12 encoded therein in units of the audio frames 14 by the use of linear prediction, according to which the audio frames are encoded by the use of linear prediction coefficients and the coded representation of residual prediction by the use of, in turn, long-term prediction coefficients (LTP) such as LTP gain and LTP delay, the indices of the code book and / or an excitation transformation coding (residual signal). Even here, the reconstruction of an audio frame 14 on the decoding side may depend on an encoding of a preceding frame or, for example, the temporal predictions of one audio frame to another or the overlapping of windows of transformation to transform the coding of the excitation signal or the like. The circumstance is mentioned herein, as it plays a role in the following description.
For the purposes of transmission and management of the
<td>network, the</td><td>luj o</td><td>from</td><td>data</td><td>from</td><td>Audio</td><td> 10</td><td>is</td><td>composed by</td><td>a</td>
<td>sequence</td><td>. from</td><td colspan="2">packages</td><td>from</td><td>load</td><td>Useful</td><td> 16.</td><td>Each one of</td><td>the</td>
<td>packages</td><td colspan="2">loading</td><td>Useful</td><td> 16</td><td colspan="2">belongs</td><td>to one</td><td>respective of</td><td>the</td>
sequence of access units 18 in which the audio data stream 10 is partitioned along the order of the stream 20. Each of the access units 18 is associated with a respective one of the audio frames 14, in accordance with as indicated by double pointed arrows 22 in Figure 1. As illustrated in Figure 1, the temporal order of the audio frames 14 may coincide with the order of the associated audio frames 18 in the data stream 10: an audio frame 14 immediately following another frame may be associated with an access unit in the data stream 10 immediately after the access unit of the other audio frame in the data stream 10.
That is, according to what is shown in Figure 1, each access unit 18 may have one or more payload packets 16. The one or more payload packets 16 of a certain access unit 18 has / n encoded in its inside the coding parameters mentioned above that describe the associated frame 14 such as the spectral decomposition transformation coefficients, LPC, and / or an excitation signal coding.
The audio data stream 10 may also comprise time stamp information 24 indicating for each access unit 18 of the data stream 10 this timestamp ti in which the audio frame i with which the respective access unit 18 AU ± is associated, it will be issued. The timestamp information 24 may be inserted, as illustrated in Figure 1, into one of the one or more packets 16 of each access unit 18 to indicate the timestamp of the associated audio frame, but Different solutions are also feasible, such as the insertion of the time stamp information t ± of an audio frame i in each of the one or more packets of the associated access unit AU ±.
Due to the packetization, the access unit partition and the time stamp information 24, the audio data stream 10 is especially suitable for transmission between encoder and decoder. That is, the audio data stream 10 of Figure 1 is an audio stream of the stream format. The audio data stream of Figure 1 may, for example, be an audio data stream according to MPEG-H 3D Audio or MHAS [2].
In order to facilitate transport / network handling, packets 16 may have aligned byte sizes and packets 16 of different types can be distinguished. For example, some packages 16 may be related to a first audio channel or a first set of audio channels and have a first type of package associated therewith, while packages that have another type of package associated therewith have another audio channel encoded inside or another set of audio signal audio channels 12 encoded inside. Even additional packages may be of a type of package that rarely carries changing data such as configuration data, valid coding parameters, or being used by, the sequence of access units. Even other packages 16 may be of a type of package that carries valid encoding parameters for the access unit to which they belong, while other payload packages carry the encodings of the sample values, transformation coefficients, LPC coefficients, or similar. Accordingly, each packet 16 may have a packet type indicator therein that is easily accessible by intermediate network entities and the decoder, respectively. The TU packages described below may be distinguishable from payload packages by package type.
While the audio data stream 10 is transmitted as is, no problem occurs. However, it is imagined that the audio signal 12 is going to be emitted on the decoding side to some extent at the time indicated representatively by τ in Figure 1, only. Figure 1 illustrates, for example, that this point in time τ can be determined by some external clock, such as a video frame clock. Figure 1, for example, illustrates in 26 a video composed of a sequence of frames 28 in a time-aligned manner with respect to the audio signal 12, one on top of the other. For example, the timestamp T<sub>frame</sub> it could be the timestamp of the first image of a new scene, a new program or the like, and consequently it could be desired that the audio signal 12 be cut at that time τ = T<sub>frame</sub> and be replaced by another audio signal 12 from that moment forward, which represents, for example, the tone signal of the new scene or program. Figure 1, for example, illustrates an existing audio data stream 30 constructed in the same manner as the audio stream 10, that is, by the use of access units 18 composed of one or more load packets. useful 16 in which the audio signal 32 that accompanies or describes the sequence of frame images 28 from the time stamp T<sub>frame</sub> in the audio frames 14 in such a way that the first audio frame 14 has its front end which coincides with the time stamp T<sub>frarae</sub>, that is, the audio signal 32 is to be output with the leading end of the frame 14 registered for the issuance of the time stamp T<sub>frame</sub>.
Disadvantageously, however, the frame rate of the frame 14 of the audio data stream 10 is completely independent of the frame rate of the video 26. Therefore, it is completely random where within a certain frame 14 of the audio signal 12 τ = T<sub>frame</sub> falls into. That is, without any additional measures, it would be more than possible to completely leave out the access unit AUj associated with the audio frame 14, j, within which τ lies, and add to the predecessor access unit AUj- ± of the audio data flow 10 the sequence of access units 18 of the audio data stream 30, thereby causing a mute in the leading end portion 34 of the audio frame j of the audio signal 12.
The various embodiments described hereinafter overcome the deficiency described above and allow handling such splicing problems.
Figure 2 shows an audio data stream according to an embodiment of the present application. The audio data stream of Figure 2 is generally indicated by the reference sign 40. First, the construction of the audio signal 40 coincides with that explained above with respect to the audio data stream 10, that is, the audio data stream 40 comprises a sequence of payload packets, that is, one or more for each access unit 18 in which the data stream 40 is partitioned. Each access unit 18 is associated with a certain of the audio frames of the audio signal that is encoded in the data stream 40 in the units of the audio frames 14. Beyond this, however, the flow of Audio data 40 has been prepared to be spliced within an audio frame with which any predetermined access unit is associated. In this case, it is representative of the AU access unit<sub>±</sub> and the access unit AUj. First it refers to the access unit AU ±. In particular, the audio data stream 40 becomes spliced by having a truncation unit package 42 inserted therein, the truncation unit package 42 is adjustable to indicate, for the access unit AUj., A portion of end of the associated audio frame i to be discarded during broadcast. The advantages and effects of the truncation unit package 42 will be discussed hereafter. Some preliminary notes, however, will be made with respect to the placement of the truncation unit package 42 and the contents thereof. For example, while Figure 2 shows the truncation unit package 42 as being located within the access unit AUj., That is, the end portion of which indicates the truncation unit package 42, the unit package of truncation 42 may alternatively be located in any access unit preceding the access unit AUi. Similarly, although the truncation unit packet 42 is within the access unit AUJ., The access unit 42 is not required to be the first packet in the respective access unit AUi as illustrated in the form representative in Figure 2.
According to an embodiment illustrated in Figure 3, the end portion indicated by the truncation unit package 42 is a rear end portion 44, that is, a portion of the frame 14 extending from some point in time. time you<sub>nne</sub>r within the audio frame 14 for the rear end of frame 14. In other words, according to the embodiment of Figure 3, there is no syntax element that indicates whether the end portion indicated by the unit package of truncation 42 should be a front end portion or a rear end portion. However, the truncation unit package 42 of Figure 3 comprises a package type index 4 6 indicating that the package 42 is a truncation unit package, and a truncation length element 48 indicating a length of truncation, that is, the temporal length At of the rear end portion 44. The truncation length 48 can measure the length of the portion 44 in units of individual audio samples, or in n-tuples of consecutive audio samples with n being greater than one and being, for example, smaller than the samples N being N the number of samples in the plot.
It is described below that the truncation unit package 42 may optionally comprise one or more markers 50 and 52. For example, the marker 50 could be an output splice marker indicating that the access unit AU ± for the that the truncation unit package 42 indicates the end portion 44, is ready to be used as an exit junction point. Marker 52 could be a marker dedicated to the decoder to indicate whether the current access unit AU ± has actually been used as an exit junction point or not. However, markers 50 and 52 are, as outlined above, merely optional. For example, the presence of the TU 42 packet itself could be a signal to transmit palicers and decoders that the access unit to which the truncation unit 42 belongs is a unit of said access suitable for the output splice, and a Truncation length setting 48 to zero could be an indication to the decoder that a truncation is not to be carried out and without splicing out, accordingly.
The above notes regarding the TU 42 package are valid for any TU package such as the TU 58 package.
In accordance with what will be described below, the indication of a leading end portion of an access unit may also be necessary. In that case, a truncation unit package, such as the TU 58 package, can be adjustable in order to indicate a rear end portion as shown in Figure 3. Said TU package 58 can be distinguished from the truncation unit packages of the front end portion such as 42 by means of the type index of the truncation unit package 46. In other words, the different types of packages could be associated with TU 42 packages indicating rear end portions and TU packages that are to indicate front end portions, respectively.
For the sake of completeness, Figure 4 illustrates a possibility according to which the truncation unit package 42 comprises, in addition to the syntax elements shown in Figure 3, a front / rear indicator 54 indicating whether the truncation length 48 is measured from the front end or the rear end of the audio frame i into the audio frame i, that is, if the end portion, the length of which is indicated by the truncation length 48 is a rear end portion 44 or a front end portion 56. The package type of the TU packages would then be the same.
As described in more detail below, the truncation unit package 42 makes the access unit AU ± suitable for an output splice, since it is feasible for the flow splicers that are described in addition to then adjust the rear end portion 44 in such a way that from the externally defined splice time τ (compare with Figure 1) in the broadcast of the audio frame i stops. Thereafter, the audio frames of the audio data spliced at the input can be broadcast.
However, Figure 2 also illustrates an additional truncation unit package 58 that is inserted into the audio data stream 40, this additional truncation unit package 58 is adjustable in a manner such that it is indicated for the access unit AUj, with j> i, from which an end portion thereof is to be discarded during the broadcast. This time, however, the access unit AUj, that is, the access unit AUj<sub>+</sub>±, its associated audio frame j is encoded inside it in a manner independent of the immediately predecessor access unit AUj_i, namely, that no prediction references or records of the internal decoder have to be set depending on the unit of predecessor access AUj_i, or in which no overlapping process makes a reconstruction of the access unit AUj_i as a requirement to correctly rebuild and issue the access unit AUj. In order to distinguish the access unit AUj, which is an immediate emission access unit, from the other access units that suffer from the access unit interdependencies described above, such as, among others, AU<sub>±</sub>, the access unit AUj is highlighted by the use of hatching.
Figure 2 illustrates the fact that the other access units shown in Figure 2 have their associated encoded audio frame in such a way that their reconstruction is dependent on the immediately predecessor access unit in the sense of that the reconstruction and correct emission of the respective audio frame on the basis of the associated access unit is simply feasible in the case of having the access unit immediately predecessor, as illustrated by the small arrows 60 pointing from the predecessor access unit of the respective access unit. In the case of the access unit AUj, the arrow pointing from the immediately predecessor access unit, namely AUj_<sub>x</sub>, the AU access unit<sub>3</sub> it is crossed out in order to indicate the immediate emission capacity of the access unit AUj. For example, in order to provide this immediate emission capability, the access unit AUj has additional data encoded therein, such as initialization information to initialize the decoder's internal registers, the data that allow an estimate of the cancellation information. of aliases generally provided by the temporary overlap portion of the inverse transformations of the immediately predecessor access unit or the like.
The capacities of the access units AU ± and AUj are different from each other: the access unit AU ± is, as indicated below, suitable as an exit junction point due to the presence of the package of truncation unit 42. In other words, a flow splicer is capable of cutting off the audio data stream 40 in the access unit AU<sub>±</sub> as well as adding access units of another audio data stream, that is, a spliced audio data stream at the input.
This is feasible in the access unit AUj also, with the proviso that the TU packet 58 is capable of indicating a rear end portion 44. Additionally or alternatively, the truncation unit packet 58 is set to indicate a front end portion, and in that case the access unit AUj is suitable to serve as an input splice occasion (again). That is, the truncation unit pack 58 may indicate a front end portion of the audio frame j that is not output and up to that point in time, that is, up to the rear end of this rear end portion, the Audio signal from the audio data stream (preliminary) spliced at the input can be output.
For example, the truncation unit package 42 could have set the output splice marker 50 to zero, while the output splice marker 50 of the truncation unit package 58 may be set to zero or set to 1 Some explicit examples will be described further below as with respect to Figure 16A-16B.
It should be noted that there is no need for the existence of an AUj access unit capable of splicing at the entrance. For example, the audio data stream to be spliced at the input could aim to replace the broadcast of the audio data stream 40 completely from the instant of time τ onwards, that is, without a splicing taking place. input (again) for the audio data stream 40. However, if the audio data stream to be spliced at the input is to replace the audio signal of the audio data stream 40 simply preliminary, then in an input splice again for the audio data stream. 40 it is necessary, and in that case, for any TU 42 packet of the output splice there must be an input splice in the TU 58 packet that follows in the order of the data stream 20.
Figure 5 shows an audio encoder 70 for generating the audio data stream 40 of Figure 2. The audio encoder 70 comprises an audio coding core 72 and the truncation packet inserter 74.
The audio coding core 72 is configured to encode the audio signal 12 entering the audio coding core 72 in units of the audio frames of the audio signal, in the payload packets of the data stream of audio 40 in a manner described above with respect to Figure 1, for example. That is, the audio coding core 72 may be a transformation encoder that encodes the audio signal 12 by the use of an overlapping transform, for example, such as an MDCT, and then the coding of the transformation coefficients, in that the windows of the overlapping transform can, in accordance with the above described, cross the frame boundaries between consecutive audio frames, which thus results in an interdependence of immediately consecutive audio frames and their associated access units. Alternatively, the core of the audio encoder 72 may use coding based on linear prediction to encode the audio signal 12 in the data stream 40. For example, the audio coding core 72 encodes the linear prediction coefficients that describe the spectral envelope of the audio signal 12 or some prefiltered version thereof on a basis at least one frame per frame, with, in addition, the coding of the excitation signal. Continuous updates of predictive coding or loop transformation issues related to the excitation coding signal can lead to interdependencies between immediately consecutive audio frames and their associated access units. However, other coding principles are imaginable too.
The truncation unit packet inserter 74 inserts in the audio data stream 40 the truncation unit packets such as 42 and 58 in Figure 2. As shown in Figure 5, the TU packet inserter 74 may, for this purpose, be sensitive to an activator of the splice position 76. For example, the splice position trigger 76 may be informed of changes in the scene or program or other changes in a video, that is, within the frame sequence, and may accordingly signal the unit packet inserter. of truncation 74 any first frame of such a new scene or program. The audio signal 12, for example, continuously represents the audio accompaniment of the video in case, for example, none of the individual scenes or programs in the video are replaced by other plot sequences or the like. For example, imagine that a video represents a live football match and that audio signal 12 is the tone signal related to it.
Then, the splicing position activator 76 can be operated manually or automatically in order to identify temporary portions of the football game video that are subject to potential substitution by aggregates, that is, video aggregates and, in consequently, the trigger 76 would signal the beginnings of such portions to the TU insert packet 74 in such a way that it can, in response to the signal, insert a TU packet 42 in such a position, that is, in relation to the access unit associated with the audio frame in which the first video frame of the potential lies to be replaced as the portion of the video that is started. In addition, the trigger 76 informs the TU 74 packet inserter at the rear end of such portions potentially to be replaced, such as to insert a TU packet 58 to a respective access unit associated with an audio frame into which the end of such a portion. As regards such TU packets 58, the audio coding core 72 is also sensitive to the trigger 76 in order to encode differently or, exceptionally, the respective audio frame in such an access unit AUj (compare with Figure 2) in a manner that allows immediate issuance in accordance with what was described above. In the middle, that is, within such portions potentially to be replaced from the video, the trigger 76 may intermittently insert TU 58 packages intermittently in order to serve as an entry junction point or exit junction point. According to a specific example, the trigger 76 informs, for example, the audio encoder 7 0 of the timestamps of the first or starting frame of a portion to be potentially replaced, and the timestamp of the frame last or end of such a portion, in which the encoder 70 identifies the audio frames and associated access units with respect to which the insertion of TU packets and, potentially, Immediate broadcast coding will be carried out by identifying the audio frames in which the time stamps received from activator 76 fall.
In order to illustrate this, reference is made to Figure 6 which shows the fixed raster frame in which the audio coding core 72 operates, namely, at 80, together with the fixed raster raster 82 of a video to which audio signal 12 belongs. A portion 84 outside of video 86 is indicated by the use of a bracket. This portion 84 is for example determined manually by an operator or partially or fully automatically by means of scene detection. The first and last frames 88 and 90 have associated with them the timestamps T<sub>b</sub> and T<sup>and</sup>, which are located within the audio frames i and j of the raster of frames 80. Consequently, these audio frames
14, ie yj, the packages of
YOU for
MX / a / 2017/002815 middle of the packet inserter
TU 74, in which audio coding 72 uses the immediate broadcast mode in order to generate the access unit corresponding to the audio frame j.
It should be noted that the TU packet inserter may be configured to insert the TU and 58 packages with default values. For example, the truncation length syntax element 48 can be set to zero. As regards the input splice marker 50, which is optional, it is fixed by the TU 74 packet inserter in the manner indicated above with respect to Figs. two to 4, namely, which indicates the possibility of output splicing for TU 42 packages and for all TU 58 packages in addition to those recorded in the final frame or video image 86. The active splice marker 52 is would be set to zero, since no splicing has been applied so far.
It is observed with respect to the audio encoder of Figure 6, that the way to control the insertion of TU packets, that is, the way to select the access units by which the insertion is carried out, in accordance with explained with respect to Figs. 5 and 6 is only illustrative, other ways of determining those access units for which the insertion is carried out are also feasible. For example, each access unit, each Nth (N> 2) access unit or each IPF access unit could alternatively be provided with a corresponding TU packet.
It has not been explicitly mentioned before, but preferably TU packets are encoded in an uncompressed form in such a way that a bit consumption (encoding bit rate) of a respective TU packet is independent of the actual setting of the TU package. Having said that, it is worth noting that the encoder may, optionally, comprise a rate control (not shown in Figure 5), configured to record a fill level of an encoded audio buffer in order to obtain Surely an audio buffer encoded on the decoder side where the data stream 40 is received or sub-overflows, resulting in stops, or overflows which lead to packet loss 12. The encoder can, for example, control / vary a quantization step size to obey the filling level limitation with the optimization of some measure the rate / distortion. In particular, the rate control can estimate the fill level of the encoded audio buffer of the decoder assuming a transmission capacity / bit rate that can be constant or almost constant and, for example, be preset by an external entity such as a default transmission network. The coding rate of the TU packets of the data stream 40 is taken into account by the rate control. Therefore, in the form shown in Figure 2, that is, in the version generated by the encoder 70, the data stream 40 maintains the preset bit rate with different, however, around it in order to compensate for the variation in coding complexity if the audio signal 12 in terms of its ratio of the rate / distortion or with an overload of the encoded audio fill level of the decoder (which would result in an overflow), nor with reduction of the power of the same (which would give rise to an underflow). However, as described above briefly, and each access unit AUj will be described in more detail below. Output splicing, according to preferred embodiments, is supposed to contribute to the broadcast of the decoder side simply for a time duration smaller than the time length of its audio frame i. As will be clearer from the description presented below, the (forward) access unit of an input splicing audio data stream spliced with the data stream 40 in the respective output splice AU such as AU<sub>±</sub> as a splice interface, it will displace the respective successor AUs of the output splice AU. Therefore, from that moment on, the bit rate control carried out within the encoder 70 is obsolete. Beyond that, said forward AU is preferably coded autonomously in order to allow immediate broadcast, thereby consuming the most coded bit rate compared to the non-IPF AU. Thus, according to one embodiment, the encoder 70 plans or programs the rate control in such a way that the filling level recorded at the respective end of the output splice AU, that is, at its border with the immediate successor AU assumes, for example, a predetermined value such as, for example, <sup>í</sup>/ io a value between <sup>3/</sup>4 and 1/8 of the maximum fill level. Through this measure, Other encoders that prepare the audio data streams that are supposed to be spliced in the data stream 40 at the output splice AUs of the data stream 40 may rely on the fact that the buffer level of encoded audio of the decoder at the time of beginning to receive its own AU (from now on they are sometimes distinguished from the originals by means of an apostrophe) is in the default value in such a way that these other encoders They can further develop rate control accordingly. The description presented so far focused on the output splice AUs of the data stream 40, but adherence to a predetermined estimated / connected fill level can also be achieved by controlling the rate for the input splice AUs. (again) such as AUj even if it does not play a double role as an entry junction point and an exit junction. In this way, said other encoders can also control their rate control in such a way that the estimated or connected fill level assumes a predetermined fill level at a rear AU of the AU sequence of their data stream. . It may be the same as that mentioned for encoder 70 with respect to the output splice AUs. Such rear AUs can be assumed from the return splice AUs that are assumed from a splice point with the input splice AUs of the data stream 40 such as AUj. Therefore, if the encoder rate control 70 has the encoded bit rate planned / programmed in such a way that the estimated / connected fill level assumes the predetermined fill level at (or better after) the AUj, then this bit rate control remains valid even if the splicing has been carried out after encoding and exiting the data stream 40. The default fill level just mentioned could be known to the encoders by default, that is, agreed between them. Alternatively, the respective AUs could be provided with an explicit signaling of the estimated / connected filling level in accordance with what is assumed just after the respective AU of the inlet or outlet splice. For example, the value could be transmitted in the TU packet of the respective input splice AU or output splice. This costs additional general lateral information, but the encoder rate control could be provided with greater freedom in the development of the estimated fill level / connected to the input splice AU or output splice: for example, it may be sufficient then that the fill / connect level after the respective input splice or output splice AU is below or
a certain threshold, such as the maximum fill level, that is, the maximum guaranteed capacity of the decoded audio buffer of the decoder.
With respect to data flow 40, this means that it is of controlled rate to vary around a predetermined average bit rate, that is, it has an average bit rate. The actual bit rate of the splicing audio data stream varies through the sequence of packets, that is, temporarily. The (current) deviation from the predetermined average bit rate may be temporarily integrated. This integrated deviation assumes, in the input splice and output splice units, a value within a predetermined range that may be less than <sup>1/</sup>two of width of a range (max.-min.) of the deviation of the integrated bit rate, or it can assume a fixed value, for example, an equal value for all input splicing and output splicing AUs, which can be less than or ^ 4 of a maximum deviation from the integrated bit rate. As described previously, this value may be pre-configured by default. Alternatively, the value is not fixed and is not the same for all input splicing and output splicing AUs, but may, by means of signaling in the data stream.
Figure 7 shows a flow splicer for splicing audio data streams in accordance with one embodiment. The flow splicer is indicated by the use of reference 100 and comprises a first audio input interface 102, a second audio input interface 104, a splice point regulator 106 and a splice multiplexer 108.
At interface 102, the flow splicer waits for the reception of a splicing audio data stream, that is, an audio data stream provided with one or more TU packets. Figure 7 illustrates by way of example that the audio data stream 40 of Figure 2 enters the flow splice 100 at the interface 102.
Another audio data stream 110 is expected to be received at the interface 104. Depending on the implementation of the flow splicer 100, the audio data stream 110 that entering the interface 104 may be an unprepared audio data stream. as explained and described with respect to Figure 1, or one prepared in accordance with the provisions set forth below.
The junction point regulator 106 is configured to establish the truncation unit packet included in the data stream entering the interface 102, that is, TU 42 and 58 packets of the data stream 40 in the case of Figure 7, and if the truncation unit packets of the other data stream 110 entering the interface 104 are present, in which two such TU packets are shown by way of example in Figure 7, namely, a packet of TU 112 in a front access unit or first AU'i of the audio data stream 110, and a packet of TU 114 in a last or subsequent access unit AU ' <sub>K</sub> of the audio data stream 110. In particular, the apostrophe is used in Figure 7 in order to distinguish between the access units of the audio data stream 110 from the access units of the audio data stream 40. In addition , in the example described with respect to Figure 7, the audio data stream 110 is assumed to be precoded and of fixed length, that is, in this case of
K access units, corresponding to K audio frames that together temporarily cover a time interval in which the audio signal that has been encoded in the data stream 40 is to be replaced. In Figure 7, it is assumed by way of example that this time interval to be
<td>replaced</td><td>I know</td><td colspan="2">extends</td><td>since</td><td>the</td><td>audio plot</td><td>what</td>
<td>corresponds</td><td>to</td><td>Unit</td><td>from</td><td>access</td><td>AUi</td><td>to the audio plot</td><td>what</td>
<td>corresponds</td><td>to</td><td>Unit</td><td>from</td><td>access</td><td>AUJ</td><td></td><td></td>
In particular, the splice point regulator 106 is, in a manner described in more detail below, configured to establish the truncation unit packets in such a way that it becomes clear that a truncation is actually carried out. For example, while the truncation length 48 within the truncation units of the data streams entering the interfaces 102 and 104 can be set to zero, the splice point regulator 106 can change the length adjustment of transformation 48 of TU packets to a non-zero value. How the value is determined is the object of the explanation that follows.
The splice multiplexer 108 is configured to cut off the audio data stream 40 entering the interface 102 in an access unit with a TU package such as AUi access unit with the TU 42 package, in order to obtain a sequence of payload packets of this audio data stream 40, that is, in the present specification in Figure 7 in a representative manner the corresponding payload packet sequence to access the preceding units and which includes the access unit AU ±, and then by splicing this sub-sequence with a sequence of payload packets from the other audio data stream 110 entering the interface 104 in such a way that they are immediately consecutive with respect to each other and stop enter each other in the default access unit. For example, splice multiplexer 108 cuts off the audio data stream 40 in the access unit AUi such that it is sufficient to include the payload package belonging to the access unit AU ± then the access units AU ' added from the audio data stream 110 from the access unit AU'i in such a way that the access units AUi and AU'i butt each other. As shown in Figure 7, splice multiplexer 108 acts similarly in the case of the access unit AUj comprising the TU package 58: this time, splice multiplexer 108 adds data flow 40 , from the payload packets belonging to the access unit AUj, until the end of the audio data stream 110 in a manner such that the access unit AU '<sub>K</sub> stops with the access unit AUj.
Consequently, the splice point regulator
106 It establishes the TU 42 packet of the AU ± access unit as well as to indicate that the end portion to be discarded during the broadcast is a rear end portion since the audio signal of the audio data stream 40 is going to be replaced, preliminary, by the audio signal encoded in the audio data stream 110 from that time forward. In the case of the truncation unit 58, the situation is different: in this case, the splice point regulator 106 establishes the TU 58 package in order to indicate that the end portion to be discarded during the emission is a front end portion of the audio frame with which the access unit AUj is associated. It should be remembered, however, that the fact that the TU 42 package belongs to a rear end portion, while the TU 58 packages belong to a front end portion is already derivable from the input audio data stream. 40 by the use of, for example, different package identifiers of TU 46 for the package of TU 42 on the one hand and the package of TU 58 on the other hand.
The flow splicer 100 outputs the spliced audio data stream thus obtained to an output interface 116, in which the spliced audio data stream is indicated by the use of the reference sign 120.
It should be noted that the order in which the splice multiplexer 108 and the splice point regulator 106 operate in the access units need not be in accordance with that shown in Figure 7. That is, while Figure 7 suggests that splice multiplexer 108 has its input connected to interfaces 102 and 104, respectively, with the output thereof connected to output interface 116 through splice point regulator 106, the Order between splice multiplexer 108 and splice point regulator 106 can be switched.
During operation, the flow splicer 100 may be configured to inspect the input splice syntax element 50 composed of truncation unit packets 52 and 58 within the audio data stream 40 in order to carry out the operation. of splicing with the condition of whether or not the indicative input splice syntax element in the respective truncation unit packet is related to an input splice access unit. This means the following: the splicing process illustrated so far and described in more detail below may have been activated by the TU 42 package, the input splice marker 50 is set to one, as described with with respect to Figure 2. Consequently, the adjustment of this marker to one is detected by the flow splicer 100, whereby the inlet splicing operation is described, which is described in more detail below, but already outlined above.
As indicated above, the splice point regulator 106 may not have to change the configuration within the truncation unit packets as regards discrimination between the incoming splice TU packets such as packets. of TU 42 and TU splice packets such as TU 58 packets. However, the splice point regulator 106 sets the time length of the respective end portion to be discarded during the emission. For this purpose, the splice point regulator 106 may be configured to adjust a time length of the end portion to which the TU packages 42, 58, 112 and 114 refer, according to an external clock. This external clock 122 is derived, for example, from a video frame clock. For example, if you imagine that the audio signal encoded in the audio data stream 40 represents a tone signal that accompanies a video and that this video is video 86 of Figure 6. If you additionally imagine that in addition there is the frame 88, that is, the frame that begins a temporary portion 84 in which an aggregate is to be inserted. The junction point regulator 106 has already detected that the corresponding AUi access unit comprises the TU 42 package, but the external clock 122 informs the junction point regulator 106 the exact time T<sub>b</sub> in which the original tone signal of this video will end and be replaced by the audio signal encoded in the data stream 110. For example, this time point of the splice point may be the time instant corresponding to the first image or frame to be replaced by the added video which in turn is accompanied by a tone signal encoded in the data stream 110.
In order to illustrate the operation mode of the flow splicer 100 of Figure 7 in more detail, reference is made to Figure 8, which shows the sequence of steps carried out by the flow splicer 100. The process begins with a weighting loop 130. That is, the flow splicer 100, such as splice multiplexer 108 and / or splice point regulator 106, verifies that the audio data stream 40 for an output splice point, that is, for a unit of access to which a truncation unit packet 42 belongs. In the case of Figure 7, the access unit i is the first step control of the access unit 132 with itself, until then the control 132 returns again Likewise. As soon as the access unit of the input splice point AU ± has been detected, the TU packet thereof, that is, 42, is established in such a way that it registers the rear end portion of the access unit of the inlet junction point (its front end thereof) with the instant of time derived from external clock 122. After this adjustment 134 by the splice point regulator 106, the splice multiplexer 108 switches to the other data stream, that is, the audio data stream 110, in a manner such that after the access unit AUi of the input splice, the current access units of the data stream 110 are placed in the output interface 116, instead of the subsequent access units of the audio data stream 40. Assuming that the audio signal that is to replace the audio signal of the audio data stream 40 from the moment of input splicing onwards, is encoded in the audio data stream 110 in such a way that this signal audio has been registered, that is, it begins immediately, with the beginning of the first audio frame that is associated with a first AU'i access unit, the flow splicer 100 simply adapts the time stamp information composed of the audio data stream 110 in such a way that a time stamp of the frame that is associated with a first access unit AU'i, for example, matches the moment of input splice time, that is, the time of AU ± plus the time length of the audio frame associated with AUi minus the time length of the rear end portion in accordance with the provisions of step 134. That is, after switching of multiplexer 136, The adaptation 138 is a task carried out continuously by the access unit AU 'of the data stream 110. However, during this time the outgoing splicing routine is also described as described below.
In particular, the output splicing routine carried out by the flow splicer 100 begins with a waiting loop according to which the access units of the audio data stream 110 are continuously checked during the same provided with a packet. TU 114 or for being the last access unit of the audio data stream 110. This check 142 is carried out continuously during the sequence of AU 'access units. As soon as the output splice access unit has been found, namely AU '<sub>K</sub> in the case of Figure 7, then the splice point regulator 106 establishes the TU 114 package of this output splice access unit in order to register the rear end portion to be discarded during the emission, the audio frame corresponding to this AU access unit<sub>K</sub> with an instant of time obtained from the external clock, as a timestamp of a video frame, that is, the first after the addition to which the tone signal encoded in the audio data stream 110 belongs. After this setting 144, the splice multiplexer 108 changes from its input in which the data stream 110 is input, to its other input. In particular, switching 146 is carried out in a manner such that in the spliced audio data stream 12 0, the access unit AUj immediately follows the access unit AU '<sub>K</sub>. In particular, the access unit AUj is the access unit of the data stream 40, the audio frame from which it is temporarily distanced from the audio frame associated with the input splice AUi access unit by a temporary amount corresponding to the temporary length of the audio signal encoded in the data stream 110 or deviated thereof for less than a predetermined amount such as a length or half of a length of the audio frames of the access units of the audio data stream 40.
Thereafter, the splice point regulator 106 establishes in step 148 the package of TU 58 of the access unit AUj to register the front end portion thereof to be discarded during the emission, with the instant of time with which the rear end portion of the audio frame of the access unit AU ' <sub>K</sub> had been registered in step 144. By this measure, the time frame of the audio frame of the access unit AUj matches the time mark of the audio frame of the access unit AU '<sub>K</sub> plus a temporary length of the audio frame of the AU 'access unit<sub>K</sub> minus the sum of the ground end portion of the audio frame of the access unit AU'k and the front end portion of the audio frame of the access unit AUj. This fact will be clearer when looking at the examples provided below.
This splicing routine is also started after switching 146. Similar to ping-pong, the flow splicer 100 switches between the continuous audio data stream 40 on the one hand and the audio data streams of predetermined length with in order to replace the predetermined portions, that is, those between which the access units with TU packages on one side and the TU 58 packages on the other side, and returns back to audio stream 40.
The change of interface 102 to 104 is carried out by the input splice routine, while the output splice routine leads from interface 104 to 102.
However, it is emphasized, once again that the example provided with respect to Figure 7 has only been chosen for illustrative purposes. That is, the flow splicer 100 of Figure 7 is not limited to the bridge portions to be replaced of an audio data stream 40 by the audio data streams 110 having audio signals of appropriate length encoded therein with the first access unit having the first audio frame encoded therein registered for the beginning of the audio signal to be inserted in the temporary portion that is replaced. Rather, the flow splicer can be, for example, to carry out a one-time splicing process only. On the other hand, the audio data stream 110 is not restricted to having its first audio frame registered at the beginning of the audio signal to be spliced at the input. Rather, the audio data stream 110 itself can come from a source that has its own audio frame clock that operates independently of the audio frame clock underlying the audio data stream 40. In that case, the change of the audio data stream 40 to the audio data stream 110 would also comprise, in addition to the steps shown in Figure 8, the adjustment step corresponding to step 148: the adjustment of the TU packet of the audio data stream 110.
It should be noted that the above description of the operation of the flow splicer can be varied with respect to the AU timestamp of the spliced audio data stream 12 0 in such a way that a TU packet indicates an end portion forward to be discarded during the broadcast. Instead of leaving the original time stamp of the AU, the flow multiplexer 108 could be configured to modify the original time stamp thereof by adding the time length of the leading end portion of the time stamp. original time, thus pointing to the rear end of the front end portion and, therefore, at the moment when the audio frame fragment of the AU will actually be broadcast. This alternative is illustrated by the timestamp examples in Figure 16 discussed below.
Fig. 10 shows an audio decoder 160 according to an embodiment of the present application. By way of example, audio decoder 160 is shown during reception of spliced audio data stream 120 generated by stream splicer 100. However, similar to the statement regarding the flow splicer, the audio decoder 160 of Figure 10 is not limited to receiving spliced audio data streams 120 of the type explained with respect to Figures 7 to 9, where a base Audio data stream is preliminarily replaced by other audio data streams that have the corresponding audio signal length encoded inside.
The audio decoder 160 comprises an audio decoder core 162 that receives the spliced audio data stream and an audio trunker 164. The audio decoding core 162 performs the reconstruction of the audio signal in audio frame units of the audio signal of the payload sequence of the input audio data stream 120, in which, as explained above, the payload packets are individually associated with a respective one of the sequence of access units in which the spliced audio data stream 120 is partitioned. Since each access unit 120 is associated with a respective one of the audio frames, the audio decoding core 162 outputs the audio samples reconstructed by the audio frame and the associated access unit, respectively. As described above, decoding may involve a reverse spectral transformation and due to an overlap / aggregate process or, optionally, the concepts of predictive coding, audio decoding core 162 may reconstruct the audio frame. from a respective access unit, while, in addition, by the use of, that is, depending on, a predecessor access unit. However, each time an immediate broadcast access unit arrives, such as the AUj access unit, the audio decoding core 162 is able to use the additional data in order to allow immediate broadcast without needing or waiting for any data from a previous access unit. In addition, as explained above, the audio decoding core 162 may operate by the use of linear predictive decoding. That is, the audio decoding core 162 can use linear prediction coefficients contained in the respective access unit in order to form a synthesis filter and can decode an excitation signal from the access unit that involves, for example, transform decoding, that is, transform in reverse, table searches for the use of indexes contained in the respective access unit and / or predictive coding or internal status updates, then subjecting the excitation signal obtained in this way to the synthesis filter or, alternatively, the signal shaping of excitation in the spectral domain, by the use of a transfer function formed in a manner that corresponds to the transfer function of the synthesis filter. The audio trunker 164 is sensitive to the truncation unit packets inserted in the audio data stream 120 and truncates an audio frame associated with an access unit having such TU packets in order to discard the end portion thereof, which is indicated to be discarded during the issuance of the TU package.
Figure 11 shows an operation mode of the audio decoder 160 of Figure 10. After detecting 170 a new access unit, the audio decoder controls whether this access unit is one encoded by the use of the immediate broadcast mode. If the current access unit is an immediate broadcast frame access unit the audio decoding core 162 treats this access unit as a stand-alone source of information to reconstruct the audio frame associated with this current access unit. That is, in accordance with what has been explained above, the audio decoding core 162 can pre-fill the internal registers for the reconstruction of the audio frame associated with a current access unit based on the data encoded in this unit. access. Additionally or alternatively, audio decoding core 162 refrains from using the prediction of any predecessor access unit as in non-IPF mode. Additionally or alternatively, the audio decoding core 162 does not perform any overlap-sum process with any predecessor access unit or its predecessor audio frame associated with the alias cancellation asset at the temporarily forward end of the audio frame of the current access unit. Rather, for example, audio decoding core 162 derives the temporary alias cancellation information from the current access unit itself. Therefore, if the check 172 reveals that the current access unit is an IPF access unit, then the IPF decoding mode 174 is carried out by the audio decoding core 162, thereby obtaining the reconstruction of the current audio frame. Fit
<img file="MX366276B_D0002.tif" />
Alternatively, if check 172 reveals that the current access unit is not IPF, then audio decoding core 162 is applied as a usual non-IPF decoding mode over the current access unit. That is, the internal registers of the audio decoding core 162 may be adopted, since they are after processing the previous access unit. Alternatively or additionally, an overlap-sum process can be used in order to aid in the reconstruction of the temporary output end of the audio frame of the current access unit. Alternatively or additionally, the prediction of the predecessor access unit can be used. Non-IPF decoding 17 6 also ends in a reconstruction of the audio frame of the current access unit. An upcoming check 178 checks if any truncation is to be carried out. Check 178 is carried out by audio truncator 164. In particular, audio truncator 164 checks if the current access unit has a TU packet and if the TU packet indicates an end portion to be discarded during issue. For example, audio trunker 164 checks if a TU packet is contained in the data flow for the current access unit and if the active splice marker 52 is configured and / or if the truncation length 48 is
MX / a / 2017/002815 non-zero. If a truncation is not performed, the reconstructed audio frame in accordance with what was reconstructed from any of steps 174 or 176 is going to be fully broadcast in step 180. However, if the truncation is to be carried out out, the audio truncator 164 performs the truncation and is limited to the remaining part to be broadcast in step 182. In the case of the end portion indicated by the TU packet being a rear end portion, the rest of the reconstructed audio frame is to be emitted starting with the time stamp associated with that audio frame. In the case of the end portion indicated to be discarded during the broadcast by the TU packet being a front end portion, the rest of the audio frame is output in the time stamp of this audio frame plus the time length of the front end portion. That is, the broadcast of the rest of the current audio frame is differed by the temporal length of the front end portion. The process is then processed additionally with the following access unit.
See the example in Figure 10: Audio decoding core 162 performs normal IPF 176 decoding on AUi ~ i and AU ± access units. However, the latter has the TU 42 package. This TU 42 package indicates a rear end portion to be discarded during the broadcast, and consequently the audio trunker 164 prevents a rear end 184 of the associated audio frame 14 with the AU access unit<sub>±</sub> of being emitted, that is, of participating in the formation of the output audio signal 186. From then on, the access unit AU'i arrives. The same is for an immediate broadcast frame access unit and is treated by audio decoding core 162 in step 174 accordingly. It should be noted that the audio decoding core 162
<td colspan="2">can for example</td><td>, understand the</td><td>capacity</td><td>from</td><td>open more</td><td>from</td>
<td colspan="2">an instantiation</td><td>of himself. AND</td><td>s say,</td><td>each</td><td>Once</td><td>I know</td>
<td>performed</td><td>a</td><td>decoding</td><td>IPF,</td><td>this</td><td>it implies</td><td>the</td>
<td>opening of</td><td>a</td><td>instantiation</td><td>further</td><td>of the</td><td>core</td><td>from</td>
<td>decoding</td><td>from</td><td>audio 162. In</td><td>any</td><td colspan="2">case like</td><td>the</td>
AU'i access unit an IPF access unit, it does not matter that your audio signal is in fact related to a new audio scene compared to its predecessors AUi_i and AUi. The audio decoding core 162 does not care about that. Rather, the access unit AU'i is taken as an autonomous access unit and the audio frame is reconstructed therefrom. Since the length of the rear end portion of the audio frame of the predecessor access unit AUi has probably been set by the flow splicer 100, the beginning of the audio frame of the access unit AU'i stops immediately with the rear end of the rest of the audio frame of the AU ± access unit. That is, they stop at the transition time Ti somewhere in the middle of the audio frame of the access unit AU ±. After meeting the AU 'access unit<sub>Ki</sub> audio decoding core 162 decodes this access unit in step 176 in order to reveal or reconstruct this audio frame, whereby this audio frame is truncated at its rear end, due to the indication of the portion of rear end for your TU 114 package. Therefore, simply the rest of the audio frame of the AU 'access unit<sub>K</sub> until the rear end portion is emitted. Then, the access unit AUj is decoded by the audio decoding core 162 in the IPF decoding 174, that is, independently of the access unit AU '<sub>K</sub> autonomously and the audio frame obtained therefrom is truncated at its front end since its truncation unit package 58 indicates a front end portion. The remains of the audio frames of AU 'access units<sub>K</sub> and AUj butt each other in a snapshot of transition time T2.
The embodiments described above basically use a signaling that describes whether and the number of audio samples of a particular audio frame should be discarded after decoding of the associated access unit. The embodiments described above may for example be applied to extend an audio codec such as MPEG-H 3D Audio. The MEPG-H 3D Audio standard defines an autonomous stream format to transform MPEG-H 3D audio data called MHAS [2]. In line with the embodiments described above, the truncation data of the truncation unit packets described above could be reported at the MHAS level. There, it can be easily detected and easily modified on the fly by means of flow splicing devices such as flow splicer 100 of Figure 7. This new type of MHAS package could be labeled with PACTYP_CUTRUNCATION, for example . The payload of this type of package could have the syntax shown in Figure 12. In order to facilitate concordance between the specific syntax example of Figure 12 and the description carried out previously with respect to Figures 3 and 4, for example, the reference signs of Figures 3 and 4 have been reused in order to identify the corresponding syntax elements in Figure 12. The semantics could be as follows:
isActive: If 1 truncation message is active, if it is 0 the decoder must ignore the message.
canSplice: Tells a splice device that a splice can start or continue at that time. (Note: This is basically an add-start marker but the
MX / a / 2017/002815 splicing device can reset it to 0 since it does not carry any information for the decoder.) TruncRight: If it is 0 the samples are truncated from the end of the AU, if it is 1 the samples are truncated from the beginning of AU.
nTruncSamples: Number of samples to truncate.
It should be noted that the MHAS flow ensures that an MHAS payload packet is always aligned by bytes in such a way that the truncation information is easily accessible on the fly and can be easily inserted, deleted or modified by ex. , by means of a flow splicing device. An MPEG-H 3D Audio stream could contain one type of MHAS packet with PACTYP_CUTRUNCATION pactype for each AU or for a suitable subset of the AUs with isActive set to 0. Next, a flow splice device can modify this MHAS packet. according to your need. Otherwise, a flow splicing device can easily insert such an MHAS package without adding significant bit rate overhead as described hereinafter. The size of the largest granules of MPEG-H 3D Audio is 4096 samples, so that 13 bits for nTruncSamples are sufficient to indicate all significant truncation values. nTruncSamples and the 3 one-bit markers together occupy 16 bits or 2 bytes in a way that no additional byte alignment is needed.
Figures 13A to 13C illustrate how the truncation CU method can be used to implement a precise flow splice sample.
Figure 13A shows a video sequence and an audio stream. In the video frame number 5 the program is connected to a different source. The alignment of video and audio in the new source is different than in the old source. To enable the precise switching sample of PCM audio decoded samples at the end of the last CU of the old stream and at the beginning of the new stream they have to be removed. A short period of cross fading in the decoded PCM domain may be necessary to avoid interference in the output PCM signal. Figure 13A shows an example with concrete values. If for some reason the overlapping of the AU / CU is not desired, there are two possible solutions represented in Figure 13B) and Figure 13C). The first AU of the new flow has to carry the configuration data for the new flow and all the previous rolls that are needed to start the decoder with the new configuration. This can be done by means of an Immediate Broadcasting Frame (IPF) defined in the MPEG-H 3D Audio standard.
Another application of the CU truncation method is changing the configuration of an MPEG-H 3D Audio stream. Different streams of MPEG-H 3D Audio can have very different settings. For example, a stereo program can be followed by a program with 11.1 channels and additional audio objects. The settings will usually change in a video frame limit that is not aligned with the granules of the audio stream. The CU truncation method can be used to implement the precise audio configuration change of the sample, as illustrated in Figure 14.
Figure 14 shows a video stream and an audio stream. In the video frame number 5 the program is changed to a different configuration. The first CU with the new audio configuration is aligned with the video frame in which the configuration change occurred. In order to allow the configuration change the PCM samples of precise audio samples at the end of the last CU with the old configuration have to be removed. The first AU with the new configuration has to carry the new configuration data and all the previous rolls that are needed to start the decoder with the new configuration. This can be accomplished through an Immediate Broadcasting (IPF) Frame defined in the MPEG-H 3D Audio standard. An encoder can use PCM audio samples of the old configuration to encode the previous rolls for the new configuration for the channels that are present in both configurations. Example: If the configuration change is from stereo to 11.1, then the right channels of the new 11.1 configuration can use the previous left and right roll data form of the old stereo configuration. The other channels of the new configuration 11.1 use zeros from previous rolls. Fig. 15A-15B illustrates the operation of the encoder and the generation of the bit stream for this example.
Fig. 16A-16B shows additional examples for spliced or spliced audio data streams. See Fig. 16A, for example. Fig. 16A shows a part of a splicing audio data stream that representatively comprises seven consecutive access units AUi to AU7. The second and sixth access unit are provided with a TU package, respectively. Both are not used, that is, they are not active, by setting marker 52 to zero. The TU packet of the AUg access unit is made up of an IPF type access unit, that is, it allows a splice back into the data stream. In B, Figure 16A shows the audio data stream of A after insertion of an aggregate. The aggregate is encoded in a data stream of the access units AU'1 to AU'4. In C and D, Figure 16A shows a modified case compared to A and B. In particular, in this case the audio encoder of the audio data stream of access units AUi. . ., has decided to change the encoding settings somewhere within the audio frame of the AUg access unit. Accordingly, the original audio data stream of C already comprises two 6.0 timestamp access units, namely AUg and AU'i with a rear end portion and a respective front end portion indicated as to be discarded during the issue, respectively. Here, truncation activation is already predetermined by the audio decoder. However, the AU'i access unit can still be used as an inbound return splice access unit, and this possibility is illustrated in D.
An example of how to change the coding configuration at the output junction point is illustrated in E and F. Finally, in G and H the example of A and B in Fig. 16A is extended through another TU package provided in the AU5 access unit, which can serve as an entry or continuation junction point.
According to the aforementioned, although the pre-arrangement of the access units of an audio data stream with TU packets may be favorable in terms of the ability to take into account the bit rate consumption of These TU packages at a very early stage in the generation of the access unit, this is not mandatory. For example, the flow splicer explained above with respect to Figures 7 to 9 can be modified in that the flow splice identifies points of entry or exit splice by other means than the occurrence of a TU packet in the Incoming audio data stream at the first interface 102. For example, the flow splicer could react for the external clock 122 also with respect to the detection of the input splice and output splice points. According to this alternative, the splice point regulator 106 would not only fix the TU packet but also insert them into the data stream. However, it should be borne in mind that the audio encoder is not released from any preparation task: the audio encoder would still have to choose the IPF encoding mode for the access units that will serve as junction points back from entry.
Finally, Figure 17 shows that the favorable splicing technique can also be used within an audio encoder that is capable of switching between different encoding configurations. The audio encoder 70 in Figure 17 is constructed in the same manner as in Figure 5, but this time the audio encoder 70 is sensitive to a change in the activation configuration 200. That is, see for example case C in Figure 16A: the audio coding core 72 continuously encodes the audio signal in 12 AU access units<sub>±</sub> to AUg. Somewhere within the audio frame of the AUg access unit, the instantaneous configuration change time is indicated by the trigger 200. Accordingly, the audio coding core 72, by the use of the same raster of frames of audio also encodes the current audio frame of the AUg access unit by the use of a new configuration such as an audio coding mode that involves more audio channels, encoded or the like. The audio coding core 72 encodes the audio frame again by the use of the new configuration with, in addition, by the use of the IPF coding mode. This ends at the access unit AU'i, which immediately follows an order of the access unit. Both access units, that is, the access unit AUg and the access unit AU '± are provided with the TU packets by means of the TU 74 packet inserter, the former having a rear end portion indicated for the purpose if discarded during the broadcast and the second one has a front end portion indicated as to be discarded during the broadcast. The latter can also serve, since it is an IPF access unit, as a point of entry inbound return.
For all the embodiments described above, it should be noted that, possibly, cross-fade in the decoder is performed between the reconstructed audio signal from the AU sequence of the spliced audio data stream, up to one AU of output splice (for example, AU ±), which is actually supposed to end at the front end of the rear end portion of the audio frame of this output splice AU on the one hand and the reconstructed audio signal from the AU stream of the data stream of audio from the spliced AU immediately after the output splice AU (such as AU'i) which can be assumed to start immediately from the front end of the audio frame of the successor AU, or at the rear end of the front end portion of the audio frame of this successor AU: That is to say, within a time interval that surrounds and crosses the instant of time where the portions of the immediately consecutive AU, which will be emitted, stop at each other, The actually played audio signal emitted from the audio data stream spliced by the decoder could be formed by a combination of the audio frames of both AUs that immediately stop with a combined contribution of the audio frame of the successor AU temporarily increasing within this time interval and the combined contribution of the audio frame of the output splice AU temporarily decreasing in the time interval. Similarly, a cross fade could be carried out between the input splice AUs such as AUj and its immediate predecessor AU (such as AU '<sub>K</sub>), namely, through the formation of the audio signal actually emitted by a combination of the audio frame of the input splice AU and the audio frame of the predecessor AU within a time interval surrounding and crosses the time in which the front end portion of the audio frame of the input splice AU and the rear end portion of the audio frame of the predecessor AU butt each other.
By the use of other terms, the above embodiments revealed, among other things, a possibility to exploit the available bandwidth by the transport stream, and the available MHz of the decoder: a kind of Audio Splice Point Message is sent along with the audio plot that would replace. Both the outgoing audio and the incoming audio around the splice point are decoded and cross-fade between them can be performed. The Audio Splice Point Message merely tells the decoders where to cross-fade. This is, in essence, a perfect splice because the splice is correctly registered in the PCM domain.
Therefore, the above description reveals, between
<img file="MX366276B_D0003.tif" />
others, the following aspects:
Al. Audio data flow empalmadle 40, comprising:
a sequence of payload packets 16, each of the payload packets belong to a respective one of a sequence of access units 18 in which the flow of audio data spliced is partitioned, each access unit is associated with a respective of the audio frames 14 of an audio signal 12 that is encoded in the audio data stream spliced into units of the audio frames; and a truncation unit package 42; 58 inserted into the audio data stream spliced and adjusted in order to indicate, for a predetermined access unit, an end portion 44; 56 of an audio frame with which the default access unit is associated, such as to be discarded during broadcast.
A2. Audio data stream spliced according to the Al aspect, where the end portion of the audio frame is a rear end portion 44.
A3. Audio data splicing according to the Al or A2 aspect, in which the audio data splicing also includes:
an additional truncation unit packet 58 inserted in the audio data stream spliced and
MX / a / 2017/002815 adjustable in order to indicate an additional predetermined access unit, an end portion 44; 56 of an additional audio frame with which the additional predetermined access unit is associated, such as to be discarded during broadcast.
A4. Audio data flow spliced according to aspect A3, where the end portion of the additional audio frame is a front end portion 56.
TO 5. Audio data flow spliced according to aspect A3 or A4, in which the truncation unit packet 42 and the additional truncation unit packet 58 comprise an output splice syntax element 50, respectively, which indicates whether the respective of the truncation unit package or the additional truncation unit package refers to an outgoing splice access unit or not.
A6 Flow of audio data spliced in accordance with any of the aspects A3 to A5, in which the predetermined access unit, such as AUi, has encoded inside the respective associated audio frame in such a way that a reconstruction thereof on the decoding side it is dependent on an access unit immediately before the default access unit, and a majority of the access units have the respective associated audio frame encoded therein in a
<img file="MX366276B_D0004.tif" />
such that the reconstruction thereof on the decoding side is dependent on the respective immediately preceding access unit, and the additional predetermined access unit AUj has the associated associated audio frame inside it in such a way that the reconstruction of it on the decoding side is independent of the immediately previous access unit
MX / a / 2017/002815 to the additional default access unit, thereby allowing immediate issuance.
A7 Audio data flow spliced in accordance with aspect A6, in which the truncation unit packet 42 and the additional truncation unit packet 58 comprise an output splice syntax element 50, respectively, indicating whether the respective of the truncation unit package or the additional truncation unit package refers to an output splice access unit or not, wherein the output splice syntax element 50 composed of the truncation unit package indicates that the truncation unit package refers to an output splice access unit and the syntax element composed of the unit package Additional truncation indicates that the additional truncation unit packet does not refer to an outbound splice access unit.
A8 Splicing audio data stream according to aspect A6, in which the truncation unit packet 42 and the additional truncation unit packet 58 comprise an output splice syntax element, respectively, indicating whether the respective one of the Truncation unit package or additional truncation unit package refers to an output splice access unit or not, wherein the syntax element 50 composed of the truncation unit package indicates that the truncation unit package refers to an output splice access unit and the output splice syntax element comprised by the unit package Additional truncation indicates that the additional truncation unit package refers to an outbound splice access unit, too, wherein the additional truncation unit package comprises a front / rear end truncation syntax element 54 and a truncation length element 48, wherein the front / rear end truncation syntax element is to indicate whether the end portion of the additional audio frame is a rear end portion 44 or a front end portion 56 and the truncation length element is to indicate an At length of the end portion of the additional audio frame.
A9 Flow of splicing audio data according to any of aspects Al to A8, which is controlled rate to vary around, and is due to a predetermined average bit rate, in such a way that a deviation from the integrated bit rate of the default average bit rate assumes, in the default access unit, a value within a predetermined range that is less than that of width than an interval of the deviation of the integrated bit rate by varying the entire audio splicing data stream.
A10 Audio data flow spliced in accordance with any of aspects Al to A8, which is controlled rate to vary around, and obeys, a predetermined average bit rate in such a way that a deviation from the integrated bit rate of the predetermined average bit rate assumes, in the predetermined access unit, a fixed value less than% of a maximum of the deviation of the integrated bit rate by varying the entire splicing audio data stream.
There. Splicing audio data stream according to any of aspects Al to A8, which is controlled rate to vary around, and is due to a predetermined average bit rate, in such a way that a deviation from the integrated bit rate of the default average bit rate assume, in the default access unit, as well as other access units for which the truncation unit packets are present in the splicing audio data stream, a default value
B1. Spliced audio data flow, comprising: a sequence of payload packets 16, each of the payload packets belong to a respective one of a sequence of access units 18 in which the spliced audio data stream is partitioned, each access unit is associated with a respective one of the audio frames 14;
a truncation unit package 42; 58; 114 inserted into the audio data stream and spliced indicating an end portion 44; 56 of an audio frame with which a predetermined access unit is associated, such as to be discarded during the broadcast, in which in a first sub-sequence of payload packets of the payload packet sequence, each payload packet it belongs to an AU # access unit of a first audio data stream that has a first audio signal encoded within it in audio frame units of the first audio signal, and the access units of the first audio data stream, which include the default access unit, and in a second sequence of payload packets of the payload packet sequence, each payload packet belongs to access units AU '# of a second audio data stream that has a second audio signal encoded within it in audio frame units of the second audio data stream, wherein the first and second subsequence of payload packets are immediately consecutive with respect to each other and abut each other in the predetermined access unit and the end portion is a rear end portion 44 in the case of the first sub-sequence that precedes the second sub-sequence and a front end portion 56 in the case of the second sub-sequence that precedes the first sub-sequence.
B2. Flow of spliced audio data according to aspect Bl, in which the first sub-sequence precedes the second sub-sequence and the end portion as a rear-end portion 44.
B3 Spliced audio data stream according to the Bl or B2 aspect, wherein the spliced audio data stream further comprises an additional truncation unit packet 58 inserted in the spliced audio data stream and indicating a portion of front end 58 of an additional audio frame with which an additional predetermined access unit AUj is associated, such as to be discarded during broadcast, in which in a third sub-sequence of payload packets of the sequence of payload packets, each payload packet belongs to the access units AU '' # of a third audio data stream that is encoded therein a third audio signal, or for the AU # access units of the first audio data stream, after the access units of the first audio data stream to which the payload packets of the first sub-sequence belong, wherein the access units of the second audio data stream include the additional predetermined access unit.
B4 Spliced audio data stream according to aspect B3, in which a majority of the spliced audio data stream access units that include the predetermined access unit have the respective associated audio frame encoded therein such that a reconstruction thereof on the decoding side is dependent on a respective immediately preceding access unit, in which the access unit such as AUi<sub>+</sub>i, immediately after the predetermined access unit and which forms a start of the access units of the second audio data stream has the respective associated audio frame encoded therein in such a way that the reconstruction thereof is independent of the default access unit, such as AU #, thereby allowing immediate issuance, and the additional default access unit AU<sub>D</sub> the additional audio frame is encoded therein in a manner such that the reconstruction thereof is independent of the access unit immediately prior to the additional predetermined access unit, thereby allowing immediate broadcasting, respectively.
B5 Spliced audio data stream according to aspect B3 or B4, wherein the spliced audio data stream further comprises an even additional truncation unit packet 114 inserted in the spliced audio data stream and indicating a portion rear end 44 of an even additional audio frame with which the access unit such as AU '<sub>b</sub> immediately before the most predetermined access unit such as AUj is associated, such as to be discarded during the broadcast, in which the spliced audio data stream comprises time stamp information 24 indicating for each access unit of the stream of spliced audio data a corresponding time stamp in which the audio frame with which the respective access unit is associated, is to be broadcast, wherein a time stamp of the additional predetermined access unit is equal to the time stamp of the access unit immediately preceding the additional predetermined access unit plus a time length of the audio frame with which the unit of immediately previous access to the additional default access unit is associated, minus the sum of a temporary length of the front end portion of the additional audio frame and the rear end portion of the audio frame even additional or equal to the time stamp of the access unit immediately prior to the unit additional default access plus a temporary length of the audio frame with which the access unit immediately prior to the additional default access unit is associated, minus the temporary length of the rear end portion of the audio frame even additional.
B6 Spliced audio data stream according to aspect B2, wherein the spliced audio data stream further comprises an even additional truncation unit packet 58 inserted in the spliced audio data stream and indicates a portion front end 56 of an even additional audio frame with which the access unit such as AUj immediately succeeds the predetermined access unit, such as AU'k is associated, as to be discarded during the broadcast, in which the spliced audio data stream comprises time stamp information 24 indicating for each access unit of the spliced audio data stream a corresponding time stamp to which the frame of audio with which the respective access unit is associated, will be broadcast, wherein a time stamp of the access unit immediately after the default access unit is equal to the time stamp of the predetermined access unit plus a time length of the audio frame with which the default access unit less the sum of a temporary length of the rear end portion of the audio frame with which the predetermined access unit is associated and the front end portion of the access unit is associated even additional or is equal to the time stamp of the predetermined access unit plus a temporary length of the audio frame with which the predetermined access unit is associated minus the time length of the rear end portion of the audio frame with which the default access unit is associated.
B7 Spliced audio data stream according to aspect B6, in which a majority of the access units of the spliced audio data stream have the respective associated audio frame encoded therein in such a way that a reconstruction of the themselves on the decoding side is dependent on a respective access unit immediately before, wherein the access unit immediately after the predetermined access unit and which forms a start of the access units of the second audio data stream has the associated associated audio frame inside it in such a way that the reconstruction of it on the decoding side is independent of the predetermined access unit, thereby allowing immediate broadcast.
B8 Spliced audio data stream in accordance with aspect B7, in which the first and second stream of audio data are encoded by the use of different encoding configurations, in which the access unit immediately after the unit of default access and which forms a start of the access units of the second audio data stream has the CFG configuration data encoded for the configuration of a decoder again.
B9 Spliced audio data stream in accordance with aspect B4, wherein the spliced audio data stream further comprises an even additional truncation unit packet 112 inserted in the spliced audio data stream and indicates a portion front end of an even additional audio frame with which the access unit immediately after the predetermined access unit is associated, such as to be discarded during broadcast, wherein the spliced audio data stream comprises time stamp information 24 indicating for each access unit a corresponding time stamp to which the audio frame with which the respective access unit is associated, it will be issued in which a time stamp of the access unit immediately after the predetermined access unit is equal to the time stamp of the predetermined access unit plus a temporary length of the audio frame associated with the unit default access minus the sum of a temporary length of the front end portion of the even additional audio frame and a temporary length of the rear end portion of the audio frame associated with the default access unit or equal to the time stamp of the default access unit plus a temporary length of the audio frame associated with the default access unit minus the temporary length of the temporary length of the rear end portion of the frame of audio associated with the default access unit.
B10 Flow of spliced audio data according to aspect B4, B5 or B9, in which a time stamp of the access unit immediately after the predetermined access unit is equal to the time stamp of the access unit default plus a temporary length of the audio frame with which the default access unit is associated, minus a temporary length of the rear end portion of the audio frame with which the default access unit is associated.
C1. Flow splicer for splicing audio data streams, comprising:
a first audio input interface 102 for receiving a first audio data stream 40 comprising a sequence of payload packets 16, each of which belongs to a respective one of a sequence of access units 18 in the which the first audio data stream is partitioned, each access unit of the first audio data stream is associated with a respective audio frame 14 of a first audio signal 12 that is encoded in the first audio data stream in audio frame units of the first signal audio;
a second audio input interface 104 for receiving a second audio data stream 110 comprising a sequence of payload packets, each of which belongs to a respective one of a sequence of access units in which the second audio data stream is partitioned, Each access unit of the second audio data stream is associated with a respective audio frame of a second audio signal that is encoded in the second audio data stream in audio frame units of the second audio signal. ;
a splice point regulator; and a splice multiplexer, in which the first audio data stream further comprises a truncation unit packet 42; 58 inserted into the first audio data stream and adjustable in order to indicate a predetermined access unit, an end portion 44; 56 of an audio frame with which a predetermined access unit is associated, such as to be discarded during the broadcast, and the splice point regulator 106 is configured to set the truncation unit payout 42; 58 in a manner such that the truncation unit package indicates an end portion 44; 56 of the audio frame with which the predetermined access unit is associated, such as to be discarded during broadcast or the splice point regulator 106 is configured to insert a truncation unit packet 42; 58 in the first data stream of
<td>audio and set the</td><td>same</td><td>with</td><td>the end</td><td>from</td><td>indicate</td><td>a</td><td>Unit</td>
<td colspan="2">default access,</td><td>a</td><td>portion</td><td>from</td><td>extreme</td><td> 44;</td><td>5 6 of</td>
<td>an audio plot</td><td>with</td><td>the</td><td colspan="2">that one</td><td>Unit</td><td>from</td><td>access</td>
<td>default is</td><td>associated</td><td>ada</td><td>how</td><td>pair</td><td>To be</td><td colspan="2">discarded</td>
during the broadcast, adjust the truncation unit package 42; 58 in a manner such that the truncation unit package indicates an end portion 44; 56
<td>of the</td><td>plot</td><td>audio</td><td>with which</td><td>the</td><td>Unit</td><td>from</td><td>access</td>
<td colspan="2">default</td><td colspan="2">is associated, as</td><td colspan="2">to be</td><td colspan="2">discarded</td>
<td>during</td><td colspan="2">the issue; Y in which the</td><td>multiplexor</td><td>from</td><td>splice</td><td> 108</td><td>is</td>
configured to cut the first audio data stream 40 in the predetermined access unit in order to obtain a sub-sequence of payload packets of the first audio data stream within which each payload packet belongs to a unit of respective access of a series of access units of the first audio data stream, including the default access unit, and splicing the sequence of payload packets of the first audio data stream and the sequence of payload packets of the second audio data stream in such a way that they are immediately consecutive with respect to each other and stop each other in the default access unit, wherein the end portion of the audio frame with which the predetermined access unit is associated is a rear end portion 44 in the case of the payload subsection of the first audio data stream that precedes the sequence of payload packets of the second audio data stream and a leading end portion 56 in the case of the payload subsection of the first audio data stream follows the packet sequence of payload of the second audio data stream.
C2 Flow splicer according to the Cl aspect, in which the sub-sequence of payload packets of the first audio data stream precedes the second sub-sequence of the payload packet sequence of the second audio data stream and the portion end of the audio frame with which the default access unit is associated is a rear end portion 44.
C3 Flow splicer according to aspect C2, in which the flow splicer is configured to inspect an output splice syntax element 50 formed by the truncation unit package and to perform the cutting and splicing of a condition if the output splice syntax element 50 indicates that the truncation unit packet is in relation to an output splice access unit.
C4 Flow splicer according to any aspect of C1 to C3, wherein the splice point regulator is configured to establish a temporary length of the end portion in a manner that matches an external clock.
C5 Flow splicer according to aspect C4, in which the external clock is a video frame clock.
C6 Spliced audio data stream according to aspect C2, in which the second stream of audio data has, or the splice point regulator 106 causes by insertion, an additional truncation unit packet 114 inserted into the second audio data stream 110 and adjustable in order to indicate an end portion of an additional audio frame with which a terminating access unit such as AU '<sub>K</sub> of the second audio data stream 110 is associated, as to be discarded during broadcast, and the first audio data stream further comprises an even additional truncation unit packet 58 inserted in the first audio data stream 40 and adjustable in a manner that indicates an end portion of an even additional audio frame with which the even additional predetermined access unit, such as AUj is associated, as to be discarded during the broadcast, in which a temporary distance between the audio frame of the predetermined access unit, such as AU ± and the even additional audio frame of the even additional predetermined access unit, such as AUj matches with a temporary length of the second audio signal between a front access unit such as AU'i thereof which occurs, after splicing, to the predetermined access unit, such as AU¿ and the rear access unit, such as AU '<sub>K</sub>, in which the splice point regulator 106 is configured to set the additional truncation unit package 114 in such a way that it indicates a rear end portion 44 of the additional audio frame to be discarded during the broadcast , and even additional truncation unit pack 58 in a manner such that it indicates a front end portion of the audio frame even additional to be discarded during the broadcast, wherein the splice multiplexer 108 is configured to adapt the time stamp information 24 composed of the second audio data stream 110 and indicating for each access unit a corresponding time stamp to which the audio frame with the one that the respective access unit is associated with, will be issued, in such a way that a timestamp of a front audio frame to which the front access unit of the second audio data stream 110 is associated coincides with the timestamp of the audio frame with which the audio unit default access is associated plus the temporary length of the audio frame with which the default access unit is associated minus the temporary length of the rear end portion of the audio frame with which the access unit default is associated and the splice point regulator 106 is configured to set the additional truncation unit package 114 and even additional truncation unit package 58 in such a way that an even additional time frame of the audio frame is equal to the time stamp of the additional audio frame plus a temporary length of the additional audio frame minus the sum of a temporary length of the rear end portion of the frame of additional audio and the front end portion of the audio frame even additional.
C7 Spliced audio data flow according to aspect C2, in which the second audio data stream 110 has, or the splice point regulator 106 causes by insertion, an additional truncation unit packet 112 inserted in the second audio data stream and adjustable in order to indicate an end portion of an audio frame additionally with which a front access unit such as AU'i of the second audio data stream is associated, as to be discarded during the broadcast, in which the splice point regulator 106 is configured to set the additional truncation unit packet 112 in such a way that it indicates a leading end portion of the additional audio frame as to be discarded during the broadcast, wherein the timestamp information 24 composed of the first and second stream of audio data and indicating for each access unit a corresponding timestamp to which the audio frame with which the respective access unit of the first and second audio data stream is associated, it will be broadcast, is temporarily aligned and the splice point regulator 106 is configured to set the additional truncation unit packet 112 in such a way that a time frame of the additional audio frame minus a time length of the audio frame with which the default access unit, just as AUi is associated plus a time length of the leading end portion is equal to the time frame of the audio frame with which the predetermined access unit is associated plus a time length of the audio frame with which the Default access unit is associated minus the temporary length of the rear end portion.
GAVE. Audio decoder comprising:
an audio decoding core 162 configured to reconstruct an audio signal 12, in audio frame units 14 of the audio signal, from a sequence
<td>of packages</td><td>payload</td><td colspan="2">16 of a flow of</td><td colspan="2">Data of</td><td>Audio</td>
<td>120, in the</td><td>that each one</td><td>of the</td><td>packages</td><td>from</td><td>load</td><td>Useful</td>
<td>belonging to</td><td>a respective</td><td>of a</td><td>sequence</td><td>from</td><td colspan="2">units of</td>
<td>access 18</td><td>in which</td><td>luj o</td><td>of data</td><td>from</td><td>Audio</td><td>is</td>
partitioned, in which each access unit is associated with a respective one of the audio frames; and an audio truncator 164 configured to be sensitive to a truncation unit packet 42; 58; 114 inserted into the audio data stream in order to truncate an audio frame associated with a predetermined access unit in order to discard, during the broadcast of the audio signal, an end portion thereof indicated to be discarded during the broadcast by the truncation unit package.
D2 Audio decoder according to the DI aspect, wherein the end portion 44 is a front end portion or a rear end portion 56.
D3 Audio decoder according to the DI or D2 aspect, in which a majority of the access units of the audio data stream have the respective associated audio frame encoded therein in such a way that the reconstruction thereof is dependent on a respective immediately preceding access unit, and the audio decoding core 162 is configured to reconstruct the audio frame with which each of the majority of the access units is associated based on the respective immediately preceding access unit.
D4 Audio decoder according to aspect D3, in which the predetermined access unit has encoded inside the respective associated audio frame in such a way that the reconstruction thereof is independent of an access unit immediately prior to the default access unit, wherein the audio decoding unit 162 is configured to reconstruct the audio frame with which the predetermined access unit is associated independent of the access unit immediately preceding the predetermined access unit.
D5 Audio decoder according to appearance
D3 or D4, in which the predetermined access unit has configuration data encoded therein and the audio decoding unit 162 is configured to use the configuration data to configure the decoding options according to the configuration data and apply the configuration options. decoding to reconstruct the audio frames with which the default access unit and a series of access units immediately after the access unit Default is associated.
D6 Audio decoder according to any of the aspects DI to D5, wherein the audio data stream comprises time stamp information 24 indicating for each access unit of the audio data stream of a time stamp corresponding to the one that the audio frame with which the respective access unit is associated, is going to be broadcast, wherein the audio decoder is configured to broadcast audio frames with front ends temporarily aligned with the audio frames according to the time stamp information and leaving out the end portion of the audio frame with which The default access unit is associated.
D7 Audio decoder according to any of the aspects DI to D6, configured to perform cross-fade in a joint of the end portion and a remaining portion of the audio frame.
The audio encoder comprising:
an audio coding core 72 configured to encode an audio signal 12, in audio frame units 14 of the audio signal, in payload packets 16 of an audio data stream 40 in a manner such that each packet payload belongs to a respective of the access units 18 in which the audio data stream is partitioned, each access unit is associated with a respective of the audio frames, and a truncation packet inserter 74 configured to insert into audio stream of a truncation unit packet 44; 58 being adjustable in order to indicate an end portion of an audio frame with which a predetermined access unit is associated, such as to be discarded during broadcast.
E2 Audio encoder according to the or El aspect , in which the audio encoder is configured to generate a spliced audio data stream according to any of aspects Al to A9.
E3 Audio encoder according to the El or E2 aspects, in which the audio encoder is configured to select the default access unit among the access units based on an external clock.
E4 Audio encoder according to aspect E3, in which the external clock is a video frame clock.
E5. Audio encoder according to any of aspects El to E5, configured to carry out a rate control in such a way that a bit rate of the audio data stream varies around, and obeys, a bit rate predetermined average, in such a way that a deviation from the integrated bit rate of the predetermined average bit rate assumes, in the predetermined access unit, a value within a predetermined range that is less than that of width than a range of the deviation of the integrated bit rate by varying the entire splicing audio data stream.
E6 Audio encoder according to any of aspects El to E5, configured to carry out a rate control in such a way that a bit rate of the audio data stream varies around, and obeys, a bit rate predetermined average, in such a way that a deviation from the integrated bit rate of the predetermined average bit rate assumes, in the predetermined access unit, a fixed value less than<sup>3</sup>^<sub>4</sub> of a maximum of the deviation of the integrated bit rate by varying the entire splicing audio data stream.
E7 Audio encoder according to any of aspects El to E5, configured to carry out a rate control in such a way that a bit rate of the audio data stream varies around, and obeys, a bit rate predetermined average, in such a way that a deviation from the integrated bit rate of the predetermined average bit rate assumes, in the default access unit as well as other access units for which the truncation unit packets are inserted into the audio data stream, a predetermined value.
E8 Audio encoder according to any of the aspects E to E7, configured to carry out a rate control by entering an encoded audio decoder buffer to fill state in such a way that a connected fill state assumes, in the default access unit, a default value.
E9. Audio encoder according to aspect E8, in which the default value is common among the access units by which the truncation unit packets are inserted into the audio data stream.
E10 Audio encoder according to aspect E8, configured to indicate the default value within the audio data stream.
While some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a process step or a function of a process step. Similarly, the aspects described in the context of a process step also represent a description of a corresponding block or an element or characteristic of a corresponding apparatus. Some or all of the process steps can be executed by (or by means of) a hardware device, such as a microprocessor, a computer program or an electronic circuit. In some embodiments, some one or more of the most important process steps may be executed by said apparatus.
The inventive spliced or spliced audio data streams can be stored in a digital storage medium or they can be transmitted in a transmission medium such as a wireless transmission medium or a cable transmission medium, such as the Internet.
Depending on certain implementation requirements, embodiments of the invention may be implemented in hardware or software. The implementation can be carried out by the use of a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH, which has electronically readable control signals stored therein, which cooperate (or are capable of cooperating) with a computer system program in a manner such that the respective method is carried out. Therefore, the digital storage medium can be computer readable.
Some embodiments according to the invention comprise a data carrier that has electronically readable control signals, which are capable of cooperating with a programmable computer system, in such a way that one of the methods described herein is carried out. cape.
Generally, the embodiments of the present invention can be implemented as a computer program product with a program code, the program code is operative to carry out one of the methods, when the computer program product is executed in a computer. The program code, for example, can be stored on a machine-readable media.
Other embodiments comprise the computer program for carrying out one of the methods described herein, stored on a machine-readable media.
In other words, an embodiment of the method of the invention is, therefore, a computer program that has a program code for carrying out one of the methods described herein, when the computer program is executed in a computer.
A further embodiment of the methods of the invention is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recording the same, the computer program for carrying out one of the methods described herein. The data carrier, the digital storage medium or the registered medium are in tangible and / or non-transient type.
A further embodiment of the method of the invention is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the signal sequence can, for example, be configured to be transferred through a data communication connection, for example over the Internet.
An embodiment further comprises a processing means, for example a computer, or a programmable logic device, configured for or adapted to perform one of the methods described herein.
An embodiment further comprises a computer that has the same computer program installed to carry out one of the methods described herein.
A further embodiment according to the invention comprises an apparatus or system configured to transfer (eg, electronically or optically) a computer program to carry out one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may comprise, for example, a file server to transfer the computer program to the receiver.
In some embodiments, a programmable logic device (eg, a field-programmable door array) can be used to perform all or some of the functionalities of the methods described herein. In some embodiments, a field-programmable door array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably carried out by means of any hardware device.
The apparatus described herein may be implemented by the use of a hardware apparatus, or by the use of a computer, or by the use of a combination of a hardware apparatus and a computer.
<td></td><td>The methods described</td><td>in</td><td>the</td><td colspan="2">Present</td><td>memory</td><td>I know</td>
<td>they can</td><td>carry out by use</td><td>from</td><td>a</td><td>apparatus</td><td>from</td><td>hardware,</td><td>or</td>
<td>the use</td><td>of a computer, for</td><td>the</td><td>use</td><td>of a</td><td colspan="2">combination</td><td>from</td>
a hardware device and a computer.
The embodiments described above are merely illustrative of the principles of the present invention. It is understood that the modifications and variations of the provisions and details described herein will be apparent to those skilled in the art. It is the intention, therefore, to be limited only by the scope of the impending patent claims.
<td>and not for</td><td>the details</td><td>specific presented</td><td>to</td><td>mode of</td>
<td>description</td><td>and explanation</td><td>of the realizations of</td><td>the</td><td>Present</td>
<td>memory.</td><td></td><td></td><td></td><td></td>
<td>References</td><td></td><td></td><td></td><td></td>
<td> [1]</td><td colspan="2">Method And Encoder And Decoder</td><td>For</td><td>S ample-</td>
Accurate Representation Of An Audio Signal, IISlb-10 F51302 WO-ID, FH110401PID [2] ISO / IEC 23008-3, Information Technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio [3] ISO / IEC DTR 14496-24: Information technology - Coding of audio-visual objects - Part 24: Audio and systems interaction
21 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21
51 members in 17 offices
Members51
| Document | Office | Kind | |
|---|---|---|---|
| EP2996269A1 | European Patent Office (EPO) | A1 | |
| CA2960114A1 | Canada | A1 | |
| WO2016038034A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201626803A | Taiwan Province of China | A | |
| AR101783A1 | Argentina | A1 | |
| SG11201701516TA | Singapore | A | |
| AU2015314286A1 | Australia | A1 | |
| KR20170049592A | Republic of Korea | A | |
| MX2017002815A | Mexico | A | |
| EP3192195A1 | European Patent Office (EPO) | A1 | |
| US2017230693A1 | United States of America | A1 | |
| CN107079174A | China | A | |
| JP2017534898A | Japan | A | |
| BR112017003288A2 | Brazil | A2 | |
| TWI625963B | Taiwan Province of China | B | |
| RU2017111578A | Russian Federation | A | |
| RU2017111578A3 | Russian Federation | A3 | |
| AU2015314286B2 | Australia | B2 | |
| MX366276BThis record | Mexico | B | |
| KR101997058B1 | Republic of Korea | B1 | |
| RU2696602C2 | Russian Federation | C2 | |
| CA2960114C | Canada | C | |
| JP6605025B2 | Japan | B2 | |
| US10511865B2 | United States of America | B2 | |
| JP2020008864A | Japan | A | |
| AU2015314286C1 | Australia | C1 | |
| US2020195985A1 | United States of America | A1 | |
| CN107079174B | China | B | |
| US11025968B2 | United States of America | B2 | |
| CN113038172A | China | A | |
| JP6920383B2 | Japan | B2 | |
| US2021352342A1 | United States of America | A1 | |
| MY189151A | Malaysia | A | |
| US11477497B2 | United States of America | B2 | |
| US2023074155A1 | United States of America | A1 | |
| CN113038172B | China | B | |
| EP3192195B1 | European Patent Office (EPO) | B1 | |
| EP3192195C0 | European Patent Office (EPO) | C0 | |
| EP4307686A2 | European Patent Office (EPO) | A2 | |
| US11882323B2 | United States of America | B2 | |
| EP4307686A3 | European Patent Office (EPO) | A3 | |
| US2024129560A1 | United States of America | A1 | |
| ES2969748T3 | Spain | T3 | |
| PL3192195T3 | Poland | T3 | |
| EP4307686B1 | European Patent Office (EPO) | B1 | |
| EP4307686C0 | European Patent Office (EPO) | C0 | |
| EP4546794A2 | European Patent Office (EPO) | A2 | |
| ES3030539T3 | Spain | T3 | |
| PL4307686T3 | Poland | T3 | |
| EP4546794A3 | European Patent Office (EPO) | A3 | |
| US12495170B2 | United States of America | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 366276
- Application
- 2815
Titles2
- Spanish
- CONCEPTO DE EMPALME DE AUDIO
- English
- CONCEPT OF AUDIO EMPALME
Classification
- CPC, 8
- H04N21/233
- H04N21/439
- H04H20/103
- H04N21/4302
- H04N21/44004
- H04N21/23424
- H04L65/70
- H04L47/34
- IPC, 6
- H04N21 234
- H04H20 10
- H04L47 43
- H04L47 30
- H04N21 233
- H04N21 439
