Audio encoder and decoder with program loudness and boundary metadata.
Abstract
Apparatus and methods for generating an encoded audio bitstream, including by including program loudness metadata and audio data in the bitstream, and optionally also program boundary metadata in at least one segment (e.g., frame) of the bitstream. Other aspects are apparatus and methods for decoding such a bitstream, e.g., including by performing adaptive loudness processing of the audio data of an audio program indicated by the bitstream, or authentication and/or validation of metadata and/or audio data of such an audio program. Another aspect is an audio processing unit (e.g., an encoder, decoder, or post-processor) configured (e.g., programmed) to perform any embodiment of the method or which includes a buffer memory which stores at least one frame of an audio bitstream generated in accordance with any embodiment of the method.

Term
7.3 yearsleft in the term
Expires 15 January 2034.
- Priority
- Filed
- Granted
- Today
- Expires
8 claims: 4 independent, 4 dependent
- 1REIVINDICACIONES 1. Un aparato - procesador de audio para decodificar un flujo de bits de audio codificado, comprendiendo el aparato procesador de audio:una memoria intermedia de entrada que almacena al menos una porción del flujo de bits de audio codificado, incluyendo el flujo de bits de audio codificado datos de audio y un contenedor de metadatos;un analizador del flujo de bits que analiza los datos de audio;y un decodificador que decodifica los datos de audio, en donde el flujo de bits de audio codificado se segmenta en uno o más cuadros, incluyendo cada cuadro: una sección de información de sincronización que incluye una palabra de sincronización de cuadro, una sección de información del flujo de bits que sigue a la sección de información de sincronización, incluyendo la información del flujo de bits, metadatos de audio, una sección de información del flujo de bits adicional ubicada en un extremo de la sección de información del flujo de bits, hasta seis bloques de datos de audio siguiendo la sección de información del flujo de bits, 130 una sección de información auxiliar* que-rs jguo 1 loo hasta seis bloques de datos de audio, una palabra de corrección de errores siguiendo la sección de información auxiliar, y uno o más campos de omisión que contienen el espacio no utilizado que queda en el cuadro, en donde al menos uno de los uno o más cuadros incluye el contenedor de metadatos, estando el contenedor de metadatos ubicado en un espacio de datos reservado seleccionado del grupo que consiste en uno o más campos de omisión, la sección de información del flujo de bits adicional, la sección de información auxiliar o combinación de los mismos, en donde el contenedor de metadatos incluye: un encabezado que identifica el inicio del contenedor de metadatos, incluyendo el encabezado una palabra de sincronización seguida por un campo de longitud que especifica una longitud del contenedor de metadatos, un campo de versión de formato después del encabezado, especificando el campo de versión de formato una versión de formato del contenedor de metadatos, una o más cargas útiles de metadatos después del campo de versión de formato, incluyendo cada carga útil de metadatos un identificador que identifica de manera única la carga útil de metadatos seguida de los metadatos de la carga IMPI INSTITUTO MEXICANO Ϊ-, DE LA PROPIEDAD INDUSTRIAL útil de metadatos, y 131 datos de protección que siguen la carga útil de uno o más metadatos, siendo los datos de protección para autenticar o validar el contenedor de metadatos o la una o más cargas útiles de metadatos dentro del contenedor de metadatos, y en donde la una o más cargas útiles de metadatos incluye una carga útil de sonoridad de programa, y la carga útil de sonoridad incluye un campo del tipo regulación de sonoridad, consistiendo el campo del tipo regulación de sonoridad en un campo de 4 bits que indica que la norma de regulación de sonoridad se utilizó para calcular la sonoridad del programa asociada con los datos de audio.
- 2El aparato procesador de audio de la reivindicación 1, en donde la palabra de sincronización es un campo de 16 bits que tiene un valor de 0x5838.
- 3El aparato procesador de audio de la reivindicación 1, en donde la una o más cargas útiles de metadatos incluye una carga útil de sonoridad de programa, y la carga útil de sonoridad incluye un campo de canal de diálogo, consistiendo el campo de canal de diálogo en un campo de 3 bits que indica si, el canal izquierdo, derecho o central de los datos de audio contiene un diálogo hablado
- 4El aparato procesador de audio de la reivindicación 1 en donde la una o más cargas útiles de 132 metadatos incluye una carga útil de sonoí'ld'STTTO' prugiáilld, éh donde la carga útil de sonoridad de programa incluye un tipo de corrección de sonoridad, consistiendo el tipo de corrección de sonoridad de un campo de 1 bit que indica si los datos de audio se corrigieron con un proceso de corrección de sonoridad infinita anticipada o basada en archivos.
- 5El aparato procesador de audio de la reivindicación 1, en donde el flujo de bits de audio codificado es un flujo de bits AC-3 o un flujo de bits E-AC3.
- 6Un método para decodificar un flujo de bits de audio codificado, comprendiendo el método:recibir al menos una porción del flujo de bits de audio codificado, incluyendo el flujo de bits de audio codificado datos de audio y un contenedor de metadatos;analizar los datos de audio;y decodificar los datos de audio, en donde el flujo de bits de audio codificado se segmenta en uno o más cuadros, incluyendo cada cuadro: una sección de información de sincronización que incluye una palabra de sincronización de cuadro, una sección de información del flujo de bits que sigue a la sección de información de sincronización, incluyendo la información del flujo de bits, metadatos de éiMti 133 IMPI - 1 - J 0 INSTITUTO MEXICANO UE LA PROPIEDAD INDUSTRIAL audio, — - i, r , , una sección de información del flujo de bits adicional ubicada en un extremo de la sección de información del flujo de bits, hasta seis bloques de datos de audio que siguen la sección de información del flujo de bits, una sección de información auxiliar que sigue los hasta seis bloques de datos de audio, una palabra de corrección de errores que sigue la sección de información auxiliar, y uno o más campos de omisión que contienen el espacio no utilizado que queda en el cuadro, en donde al menos uno de los uno o más cuadros incluye el contenedor de metadatos, estando el contenedor de metadatos ubicado en un espacio de datos reservado seleccionado del grupo que consiste del uno o más campos de omisión, la sección de información del flujo de bits adicional, la sección de información auxiliar o una combinación de los mismos, en donde el contenedor de metadatos incluye: un encabezado que identifica un inicio del contenedor de metadatos, incluyendo el encabezado una palabra de sincronización seguida por un campo de longitud que especifica una longitud del contenedor de metadatos, un campo de versión de formato que sigue al 134 encabezado, especificando el campo de versión de formato una versión de formato del contenedor de metadatos, una o más cargas útiles de metadatos que sigue al campo de versión de formato, incluyendo cada carga útil de 5 metadatos un identificador que identifica de manera única la carga útil de metadatos seguida por los metadatos de la carga útil de metadatos, y datos de protección que siguen la una o más cargas útiles de metadatos, siendo los datos de protección para
- 710 autenticar o validar el contenedor de metadatos o la una o más cargas útiles de metadatos dentro del contenedor de metadatos, y en donde la una o más cargas útiles de metadatos incluye una carga útil de sonoridad de programa, y la carga
- 815 útil de sonoridad incluye un campo del tipo regulación de sonoridad, consistiendo el campo tipo regulación de sonoridad de un campo de 4 bits que indica que la norma de regulación de sonoridad se utilizó para calcular la sonoridad del programa asociada con los datos de audio. 135 IMPI
Independent claims8
786 paragraphs in 99 sections, as filed
(54) Title: AUDIO ENCODER AND DECODER WITH LIMIT METADATA AND PROGRAM LOUDNESS. (54) Title: AUDIO ENCODER AND DECODER WITH PROGRAM LOUDNESS AND BOUNDARY METADATA.
(57) Summary
Apparatus and methods for generating an encoded audio bitstream, including including program loudness metadata and audio data in the bitstream, and optionally also program boundary metadata in at least one segment (eg, frame) of the bit stream. Other aspects are apparatus and methods for decoding said bitstream, for example, including by performing adaptive loudness processing of the audio data of an audio program indicated by the bitstream, or authentication and / or validation of metadata and / or audio data of said audio program. Another aspect is an audio processing unit (for example, an encoder, decoder, or post processor) configured (for example, programmed) to perform any mode of the method or which includes a buffer memory that stores at least one frame of a stream. audio bits generated according to any modality of the method.
(57) Abstract
Apparatus and methods for generating an encoded audio bitstream, including by including program loudness metadata and audio data in the bitstream, and optionally also program boundary metadata in at least one segment (eg, trame) of the bitstream. Other aspects are apparatus and methods for decoding such a bitstream, eg, including by performing adaptive loudness Processing of the audio data of an audio program indicated by the bitstream, or authentication and / or validation of metadata and / or audio data of such an audio program. Another aspect is an audio Processing unit (eg, an encoder, decoder, or post-processor) configured (eg, programmed) to perform any embodiment of the method or which ineludes a buffer memory which stores at least one trame of an audio bitstream generated in accordance with any embodiment of the method.
^ Wrimoim fr
PATENT TITLE No. 356196
Headlines):
Home:
Denomination:
Classification:
Inventors):
DOLBY LABORATORIES LICENSING CORPORATION
100 Potrero Avenue, San Francisco, California, 94103-4813, USA.
AUDIO ENCODER AND DECODER WITH LIMIT METADATA AND PROGRAM LOUDNESS.
CIP:
CPC:
G10L19 / Q02; Gp Ο | 1 φ | Β; Hp% p9 / 0 (L H03G9 / 02 G1QL1S / 00> Gt0tlá / 06; GieLÍ9 /, 1 | 7;, íAO £ G9 / 005; H03G9 / 025
MICHAEUX3RÁÑT; SCOTT MICHAHLWARD
Number:
MX / a / 2016/004804
<img file="MX356196B_D0001.tif" />
NORCROSf; JEFFREY RIEDMILLER;
OR, <sup>1</sup> <
<sup>5</sup>*'’ ·
Intentional:
PaTss US US *
Validity: Vefriló years'
Date of VStgeilnientae 15 dtr eperc> óe 2034 m Date of Exp ^ dtc ^ ón: mlíyoMb ^ OéiÉk ^ * »
The patent of reference with fundamental ^ e ^ nte ^ Αίοβ-Ί<sup>0</sup>, ^ Irácotófc
In accordance with artfo¿íftjptó> the Leysjf Jae / opiedfáJnáfttf8M ^ (^ nte patera from the date of filing ^ Áyteda | estictra request
Who subscribes the present title hijeeAfun ^ ajfleníweeJo di ^ 'uest ^^^^ ticulj β * ^ seciáfSS (Official Gazette of the Federation (β3Λ) | 2 // 06/1991 25/01/2006, 05/06/2009, 01/06/2010,
Regulations of the Mexican Institute of) aWÍ¡) lR £ ad articles 1st, 3rd, 4th, 5th fraction V Subsection a), apr »
12/27/1999, reformed 10/10/2002, 07/29/2006!
Deputy Generals, Coordinator, DI Directors<sup>1 </sup>Departmental and other subordinates of the Mext Institute 08/04/2004 and 09/13/2007)
Number:
61 / 754,882 61/824, P10 fa feytié the PrO ^ e ^ á Industrial.
íh «Wiawiger» 4> de veiqte «A ^ ipromogables, told nO * fBN * mMener vigenws ^ fe rights.
Industrial Property Law I, Í / 1ffl4MA 12/26/ ^ 85, ^ 8 ^) 6 * 1999, 01/26/2004, 06/16/2005, '20Vla «* ft®Wl', Wra ^ wCl ^ Bo a), 4 ° and 12 “sections I and III of the |» TWorm ^ ¿el 01Ih «SSÍ # ÍÍwD7 / 2004, 07/28/2004 and 7/09/2007); OrqagicS jgCRytnauMexicano de la Propiedad Industrial (DOF Agreement that delegates powers to the Divisional Deputy Directors, Coordinators 12/15/1999, amended on 02/04/2000, 07/29/2004,
994 istrial. '
This letter is signed with an advanced electronic signature (FIEL), based on articles 7 BIS 2 of the Industrial Property Law; 3 of its Regulations, and 1 fraction III, 2 fraction V, 26 BIS and 26 TER of the Agreement establishing the guidelines for the use of the Electronic Payment and Services Portal (PASE) of the Mexican Institute of Industrial Property, in the procedures indicated.
THE DIVISIONAL DIRECTOR OF PATENTS
NAHANNY CANAL REYES
<img file="MX356196B_D0002.tif" />
Original string:
NAHANNY MARISOL CANAL REYES | 00001000000403252793 | Administration Service
Tr¡butarla | 1695 || MX / 2018/43039 | MX / a / 2016/004804 | Normal patent title with divisional PCT | 1220 | RRGO | Pág (s) | tyKV6xQoB0EnRRuP42N8jf7aKMA =
Digital stamp:
LGDYNotxWKnf5OL + VaDfySD¡ZMtRdWAmltK¡sEh + Qe79iVxLII + jpvJxl7M8V6E / wclbdt1aUNeUUofpzxmAjtd54f
AWLK2KGJGZ06qhbzVQhcp9r8Q08WgXBXRXxYN0qQRWkXDuZ5lbulLxtFyonw3CouOJv / K8ZEZbtjV29SLK7d + TsHrr
27qAPPXOzZGdRJNHFFw655NQ1aDVH3lddZjZIN3wSbTwrJUUG2gyQcu7tOKfTua6eOQb2trTglcqpruLpNqvKTNp8s
Y5NvLvlN7CCfnuD42gQtp3kzRTYZf7NCgkpkMvYVhlAc0NfulvMNBvVHBcqYBIr6m1uy228g ==
Arenal No 550, Piso 1, Pueblo Santa María Tepepan, Xochlinilco, 16020. Mexico City, (55) 53340700 www.gob.mx/lmpi
<img file="MX356196B_D0003.tif" />
ΙΜ
MEXICAN INSTITUTE
<img file="MX356196B_D0004.tif" />
AuSl <ep¿ © 0N ENCODER AND DECODER
PROQRAMft SOUND LIMIT AND SOUND DATA
CROSS REFERENCE Ά RELATED REQUESTS
This application claims priority of U.S. Provisional Patent Application No. 61 / 754,882, filed on January 21, 2013, and U.S. Provisional Patent Application No. 61 / 824,010, filed May 16, 2013.
2013, each of which is incorporated herein in its entirety.
TECHNICAL FIELD
The invention relates to audio signal processing, and more particularly, to encoding and decoding audio data bit streams with metadata indicating the loudness processing status of audio content and the location of program boundaries. of audio indicated by bit streams. Some embodiments of the invention generate or decode audio data in one of the formats known as AC-3, Enhanced AC-3, or E-AC-3, or Dolby E.
BACKGROUND OF THE INVENTION
Dolby, Dolby Digital, Dolby Digital Plus and Dolby E are trademarks of Dolby. Laboratories Licensing Corporation. Dolby Laboratories provides proprietary implementations of AC-3 and E-AC-3 known as Dolby Digital and Dolby Digital.
IMPI
INDUSTRIAL MEXICAN INSTITUTE OF THE PRÜRIEDM)
<img file="MX356196B_D0005.tif" />
Plus, respectively. . -Data processing units operate ΊΓΓpTc only blindly and do not pay attention to the audio data processing history that occurs before data is received. This can work in a processing framework where a single entity performs all of the audio data processing and encodes a variety of target media playback devices, while one target media playback device performs all of the decoding and playback. encoded audio data. However, this blind processing does not work well (or directly does not work) in situations where multiple audio processing units are scattered across a diverse network or located in tandem (i.e., chain) and are expected optimally perform their respective types of audio processing. For example, some audio data may be encoded by high-performance media systems and may have to be converted to a reduced form suitable for a mobile device in conjunction with a media processing chain. Accordingly, an audio processing unit may unnecessarily perform a type of processing on the audio data that has already been performed. For example, a unit of
IMPI
INSTITUTO MEXICANO M IA INDUSTRIAL PROPERTY volume leveling can perform processing in
<img file="MX356196B_D0006.tif" />
an input audio clip, regardless of whether the same volume leveling or similar volume leveling has previously been performed on the input audio clip.
As a result, the volume leveling unit can perform leveling even when it is not necessary. This unnecessary processing can also cause degradation and / or deletion of specific features while playing the audio data.
A typical audio data stream includes both audio content (for example, one or more channels of audio content) and metadata indicating at least one characteristic of the audio content. For example, in an AC-3 bit stream there are several audio metadata parameters that are specifically intended to change the sound of the supplied program to a listening environment. One of the metadata parameters in the DIALNORM parameter, which is intended to indicate the average level of dialogue that occurs in an audio program, and is used to determine the level of the audio playback signal.
During playback of a bit stream comprising a sequence of different audio program segments (each with a different DIALNORM parameter), a decoder ιΜΡΙ
<img file="MX356196B_D0007.tif" />
AC-3 uses the DIALNORM parameter of each segment to perform a type of loudness processing in which it modifies the level of reproduction or loudness so that the perceived loudness of the segment sequence dialog is at a consistent level. Each encoded audio segment (item) in a sequence of encoded audio items would (in general) have a different DIALNORM parameter, and the decoder will increase the level of each of the items so that the level of reproduction or loudness of the dialogue to each item is the same or very similar, although this may require applying different amounts of gain to different items during playback.
DIALNORM is typically set by a user, and is not automatically generated, although there is a default DIALNORM value if no value is set by the user. For example, a content creator can make loudness measurements with an external device to an AC-3 encoder, and then transfer the result (indicating the loudness of spoken dialogue from an audio program) to the encoder to set the DIALNORM value. ' Therefore, the content creator is trusted to correctly set the parameter
DIALNORM.
IMPI
<img file="MX356196B_D0008.tif" />
There are several different reasons why the DIALNORM parameter in an AC-3 bit stream may be incorrect. First, each AC-3 encoder has a default DIALNORM value that is used during bitstream generation if a DIALNORM value is not set by the content creator. This default value can be considerably different from the actual dialogue loudness level of the audio.
Second, even if a content creator measures loudness and sets the DIALNORM value accordingly, a loudness meter or algorithm may have been used that does not comply with the recommended AC-3 loudness measurement method , which results in an incorrect DIALNORM value. Third, even if a bit stream
AC-3 has been created with the measured DIALNORM value and was set correctly by the content creator, it may have been changed to an incorrect value during transmission and / or storage of the bitstream. For example, in television broadcast applications it is common for AC-3 bit streams to be decoded, modified, and recoded using incorrect DIALNORM metadata information. Therefore, a DIALNORM value included in an AC-3 bit stream may be incorrect or inaccurate and therefore may have a negative impact on the quality of the
IMPI
Mexican Institute of Industrial Property
<img file="MX356196B_D0009.tif" />
listening experience.
Also, the DIALNORM parameter does not indicate the loudness processing status of the corresponding audio data (for example, what types of loudness processing have been performed on the audio data). Until the present invention, a bitstream did not include metadata, which indicates the loudness processing status (eg, types of loudness processing applied to) the audio content of the bitstream or the loudness processing status and the loudness of the audio content of the bitstream, in a format of a type described in the present disclosure. Loudness processing status metadata in such a format is useful to facilitate adaptive loudness processing of an audio bitstream and / or validity check of loudness processing status and loudness of audio content, of a particularly efficient way. Although the present invention is not limited to use with an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream, for convenience it will be described in embodiments where it generates, decodes, or otherwise it processes said bitstream which includes metadata in loudness processing state.
<img file="MX356196B_D0010.tif" />
Τ Τ Τ · 1 1 <sup>1</sup> ι - ι, -ν | _ | I ί »tx w« 5 I Ι \ ΕΛΙ »
An encoded AC-3 bit stream comprises metadata and each other
IMPI six channels of audio content. Audio content is audio data that was compressed using perceptual audio encoding. Metadata includes various audio metadata parameters that are intended to change the sound of a supplied program to a listening environment.
The details of the AC-3 (also known as Dolby Digital) encoding are well known and are set out in many published references including the following:
ATSC Standard A52 / A: Digital Audio Compression Standard (A-3 / A: Digital Audio Compression Standard] (AC-3), Revision A, Advanced Television Systems
Committee [Advanced Television Systems Committee], August 20, 2001; and US Patents 5,583,962; 5,632,005; 5, 633,981;
5,727,119; and 6,021,386, which are incorporated herein in their entirety by this reference.
The details of Dolby Digital Plus (E-AC3) encoding are set out in Introduction to Dolby Digital Plus, an Enhancement to the Dolby Digital Coding System, AES Convention Paper 6196, 117.<sup>to</sup> AES Convention [Society of Audio Engineers], October 28, 2004.
Dolby E encoding details are set
<img file="MX356196B_D0011.tif" />
IMPI in Efficient Bit Allocation, Quantization, and Coding in an Audio Distribution System, AES Preprint 5068, 107.<sup>to</sup>
AES Conference, August 1999 and in Professional Audio Coder Optimized for Use with Video, AES Preprint 5033, 107.<sup>to</sup> Conference of the, August 1999.
Each frame of an AC-3 encoded audio bitstream contains audio content and metadata for 1536 digital audio samples. For a 48 kHz sampling rate, this represents 32 milliseconds of digital audio or a rate of 31.25 frames per second of audio.
Each frame of an E-AC-3 encoded audio bitstream contains audio content and metadata for 256, 512, 768 or
1536 digital audio samples, depending on whether the frame contains one, two, three or six blocks of audio data respectively. For a 48 kHz sampling rate, this represents 5,333, 10,667, 16, or 32 milliseconds of digital audio, respectively, or a rate of 189.9, 93.75, 62.5, or 31.25 frames per second of audio, respectively.
As indicated in Figure 4, each AC-3 frame is divided into sections (segments), including: a section of
Synchronization Information (SI) containing (as shown in Figure 5) a synchronization word (SW) and the first of
<img file="MX356196B_D0012.tif" />
two error correction words (CRC1); a Bit Flow Information (BSI) section that contains most of the metadata; six Audio Blocks (ABO a, AB5) that contain compressed audio data content (and may also contain metadata); leftover bit segments (W) that contain any unused leftover bits after you compress the audio content; an information section
Auxiliary (AUX) that can contain more metadata; and the second of two error correction words (CRC2). The leftover bit segment (W) can also be called a jump field.
As indicated in Figure 7, each E-AC-3 frame is divided into sections (segments), including: a section of
Synchronization Information (SI) containing (as shown in Figure 5) a synchronization word (SW);
a Bit Flow Information (BSI) section that contains most of the metadata; between one and six Audio Blocks (ABO to .ΆΒ5) that contain compressed audio data content (and can also contain metadata); bit segments before (W) containing any unused bit left over after the audio content is compressed (although only one bit segment left over is displayed,
<img file="MX356196B_D0013.tif" />
a different leftover bit segment will typically follow each audio block); an Auxiliary Information (AUX) section that may contain more metadata; and an error correction word (CRC). The leftover bit segment (W) can also be called a jump field.
In an AC-3 (or E-AC-3) bit stream, there are several audio metadata parameters that are specifically intended to change the sound of the supplied program to a listening environment. One of the metadata parameters is the parameter
DIALNORM, which is included in the BSI segment.
As shown in Figure 6, the BSI segment of an AC-3 frame includes a five-bit parameter (DIALNORM) that indicates the DIALNORM value for the program. A five-bit parameter (DIALNORM2) is included that indicates the DIALNORM value for a second audio program in the same AC-3 frame if the audio encoding mode (acmod) of the AC-3 frame is 0, which indicates that a dual-mono or 1 + 1 channel configuration is being used.
The BSI segment also includes an indicator (addbsie) indicating the presence (or absence) of additional bitstream information following the addbsie bit, a parameter (addbsil) indicating the length of any additional bitstream information to continuation of
IMPI
<img file="MX356196B_D0014.tif" />
flow information of the addbsil value.
metadata values that don't
Figure 6.
invention is a unit of e a buffer, an addbsil value, and up to 64 bits extra bits (addbsi) after
The BSI segment includes others specifically shown in the
Brief description of the invention
In a modality class, audio processing that included an audio decoder and a parser. The buffer stores at least one frame of an encoded audio bitstream. The encoded audio bitstream includes audio data and a metadata container. The metadata container includes a header, one or more metadata payloads, and protection data. The header includes a sync word that identifies the start of the container. The one or more metadata payloads describes an audio program associated with the audio data. Protection data is located after the one or more metadata payloads. Protection data can also be used to verify the integrity of the metadata container and the one or more payloads within the metadata container. The audio decoder is coupled to the buffer memory and is capable of decoding the audio data. The parser is
<img file="MX356196B_D0015.tif" />
IMPI coupled or integrated to the audio decoder and capable * of parsing the metadata container?
In typical embodiments, the method includes receiving an encoded audio bitstream where the encoded audio bitstream is segmented into one or more frames. The audio data is extracted from the encoded audio bitstream, along with a metadata container. The metadata container includes a header following one or more metadata payloads following protection data.
Finally, the integrity of the container and the one or more metadata payloads is verified through the use of protection data. The one or more metadata payloads may include a program loudness payload that contains data indicating the measured loudness of an audio program associated with the audio data.
A program loudness metadata payload, referred to as loudness processing state metadata (LPSM), embedded in an audio bitstream in accordance with typical embodiments of the invention can be authenticated and validated for example to allow loudness regulators to verify if a particular program loudness is already within a specified range and that the data
<img file="MX356196B_D0016.tif" />
Corresponding audio IMPIs have not been modified thereby ensuring compliance with applicable regulations.) A loudness value included in a data block comprising the loudness processing state metadata can be read to verify that, rather than computing the loudness again. In response to the LPSM, a regulatory agency may determine that the corresponding audio content complies (as indicated by the LPSM) with the regulatory and / or legal requirements for loudness (for example, the regulations promulgated by the
Loudness Mitigation in Commercial Advertising, also known as the CALM Act) without the need to compute the loudness of the audio content.
Loudness measurements necessary to meet some regulatory and / or legal requirements (for example, the regulations promulgated by the CALM Act) are based on the loudness of the integrated program. The built-in program loudness requires that a loudness measurement be made, either of the dialogue level or full mix level in a full audio program. Therefore, to perform program loudness measurements (for example, at various stages in the transmission chain) to verify compliance with typical legal requirements, it is essential
<img file="MX356196B_D0017.tif" />
IMPI η 1 II, INLOJiniñL I —— that measurements be made with the knowledge of which audio data (and metadata) determine an entire audio program, and typically this requires knowledge of the location of the start and end of the program (for example, during bitstream processing that indicates a sequence of audio programs).
In accordance with typical embodiments of the present invention, an encoded audio bitstream indicates at least one audio program (eg, a sequence of audio programs), and the program limit metadata and LPSM included in the stream Bits allow you to reinitialize the program loudness measurements at the end of a program and thereby provide an automatic way to measure the built-in program loudness. Typical embodiments of the invention include program boundary metadata in an efficiently encoded audio bitstream, which enables the accurate and robust determination of at least one boundary between consecutive audio programs indicated by the bitstream. Typical modalities allow for the robust and exact determination of a program limit in the sense that they allow for the determination of the program limit accurately even in cases where bit streams indicating different programs are spliced together (to generate the
<img file="MX356196B_D0018.tif" />
bitstream of the invention) so that it truncates one or both of the spliced bit streams (and thus discards the program boundary metadata that has been included in at least one of the pre-splice bit streams).
In typical embodiments, the loudness limit metadata in a frame of the bitstream of the invention is a program limit indicator that indicates a frame count. Typically, the indicator indicates the number of frames between the current frame (the frame that includes the indicator) and a program limit (the start or end of the current audio program). In some preferred embodiments, program limit flags are efficiently and symmetrically inserted at the beginning and end of each bitstream segment that indicates a single program (i.e., in frames that occur within some number of predetermined frames after the start of the segment and in frames that occur within some predetermined number of frames before the end of the segment), so that when two such bitstream segments are concatenated (to indicate a sequence of two programs), the program boundary metadata may be present (eg symmetrically) on both sides of the boundary between two
<img file="MX356196B_D0019.tif" />
INSTITUTO MEXICaN'- ȣ programs).
To limit the increase in data rate that results from including program boundary metadata in an encoded audio bitstream (which may indicate an audio program or sequence of audio programs), in typical modes, the Program boundaries are inserted only into a subset of the bitstream frames. Typically, the limit flag insertion rate is a non-increasing function of increasing separation of each of the bitstream frames (where a flag is inserted) from the program limit that is closest to each of these plots, where the insertion rate of limit indicators denotes the average ratio of the number of frames (indicating a program) that a program limit indicator includes to the number of frames (indicating the program) that 'does not include a program limit indicator, where the average is a moving average of a number (eg, relatively small amount) of consecutive frames of the encoded audio bitstream. In a modality class', the limit indicator insertion rate is a function that logarithmically decreases in increasing distance (from each indicator insertion location) from the nearest program limit, and for
<img file="MX356196B_D0020.tif" />
For each frame containing flags that includes one of the flags, the size of the flag in that frame containing flags is equal to or greater than the size of each flag in a frame located closest to the closest program boundary that is the frame that contains indicators (i.e. the size of the program limit flag in each frame containing flags is a non-decreasing function of increasing gap of that frame containing flags from the nearest program limit).
Another aspect of the invention is an audio processing unit (APU) configured to perform any embodiment of the method of the invention. In another class of embodiments, the invention is an APU that includes a buffer that stores (eg, non-transiently) at least one frame of an audio bitstream. encoded that has been generated by any embodiment of the method of the invention. Examples of APUs include, but are not limited to, encoders (eg, transcoders), decoders, codeos, preprocessing systems (preprocessors), postprocessing systems (postprocessors), audio bitstream processing systems, and combinations thereof. elements.
IMP
<img file="MX356196B_D0021.tif" />
iWTtWtO MEXICANO w
t) E THE INDUSTRIAL MÍWlMD
In another class of embodiments, the invention is an audio processing unit (APU) configured to generate an encoded audio bitstream comprising audio data segments and metadata segments, where the audio data segments indicate the data Audio and each of at least some of the metadata segments includes loudness processing state (LPSM) metadata and optionally also program boundary metadata.
Typically, at least one of said metadata segments in a bitstream frame includes at least one segment of
LPSM indicating whether a first loudness type has been made in the frame audio data (i.e. the audio data in at least one segment of frame audio data) and at least one other LPSM segment indicating the loudness of at least some of the frame audio data (for example, the loudness of the dialog for at least some of the frame audio data indicating the dialog). In one embodiment in this class, the APU is an encoder configured to encode input audio to generate encoded audio, and the audio data segments include the encoded audio. In typical modalities in this class, each of the metadata segments has a
IMPI Mexican Institute M THE PROPERTY
JZ. j; . _J,. ,. INDUSTRIAL, preferred format to be described herein.
In some modes, each of the encoded bitstream metadata segments (an AC-3 bitstream or an E-AC-3 bitstream in some modes) that includes LPSM (for example, LPSM and boundary metadata program bits) are included in a skip field segment leftover bit of a bitstream frame (for example, a leftover bit segment W of the type shown in Figure 4 or Figure
7). In other embodiments, each of the metadata segments of the encoded stream bits (a bit stream
AC-3 or an E-AC-3 bitstream in some modes) that includes LPSM (for example, LPSM and program limit metadata) is included as additional bitstream information in the addbsi field of the Information segment Bitstream (BSI) of a bitstream frame or an auxiliary data field (for example, an AUX segment of the type shown in Figure 4 or Figure 7) at the end of a bitstream frame . Each metadata segment that LPSM includes may be in the format specified herein with reference to Tables 1 and 2 below (i.e. includes the nuclear elements specified in
Table 1 or a variation of these, followed by a payload ID (which identifies the metadata as LPSM) and the values
IMPI
<img file="MX356196B_D0022.tif" />
payload size, followed by payload (LPSM data having the format indicated in Table 2, or the format as indicated in a variation in Table 2 described herein). In some embodiments, a frame may include one or two · metadata segments, each of which includes LPSM, and if the frame includes two metadata segments, one may be present in the addbsi field of the frame and the other in the AUX field of the frame.
In one class of embodiments, the invention is a method that includes the steps of encoding audio data to generate an AC-3 or E-AC-3 encoded audio bitstream, including including in a metadata segment (from to minus one frame of the bit stream) LPSM and program limit metadata and optionally also other metadata for the audio program to which the frame belongs. In some embodiments, each such metadata segment is included in either an addbsi field in the frame, or an auxiliary data field in the frame. In other embodiments, each of the metadata segments is included in a segment of bits left over from the frame. In some embodiments, each LPSM-containing metadata segment and the program boundary metadata contain a nuclear header (and optionally additional nuclear elements as well) and
<img file="MX356196B_D0023.tif" />
INDUSTRIAL after the nuclear header (or the nuclear header and other nuclear elements) an LPSM payload segment (or container) that has the following format:
a header, typically including at least one identification value (eg, format version, length, period, count, and LPSM subflow association values, as indicated in Table 2 herein), and after the header, LPSM, and program limit metadata. Program limit metadata can include a program limit frame count and a code value (for example, an offset_exist value) that indicates whether the frame includes only a program limit frame count or both a frame count program limit as an offset value), and (in some cases) an offset value.
LPSMs can include:
at least one dialogue indication value indicating whether the corresponding audio data indicates dialogue or does not indicate dialogue (eg, which channels of corresponding audio data indicate dialogue). Dialog indication values can indicate whether dialogue is present on any combination or all channels of the corresponding audio data; .
at least one loudness regulation compliance value
IMPI
<img file="MX356196B_D0024.tif" />
indicating whether the corresponding audio data complies with an indicated set of loudness regulations; at least one loudness processing value indicating a type of loudness processing that has been performed on the corresponding audio data; and at least one loudness value indicating at least one loudness characteristic (eg, peak or average loudness) of the corresponding audio data.
In other embodiments, the encoded bit stream is a 10 bit stream that is not an AC-3 bit stream or an EAC-3 bit stream, and each of the metadata segments it includes
LPSM (and optionally also program limit metadata) is included in a segment (or field or slot) of the bitstream reserved for storing additional data. Each metadata segment that includes LPSM may have a format similar or identical to that specified herein with reference to Tables 1 and 2 below (i.e., includes nuclear elements similar or identical to those specified in Table 1, followed by a Payload ID (which identifies metadata as LPSM) and payload size values, followed by payload (data from
LPSM that have the format similar or identical to the format indicated in Table 2, or a variation in Table 2
<img file="MX356196B_D0025.tif" />
described herein).
In some embodiments, the encoded bit stream comprises a sequence of frames, each frame including a Bit Stream Information (BSI) segment that includes an addbsi field (sometimes referred to as a segment or slot) and a field or auxiliary data slot (for example, the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream), and comprises audio data segments (for example, the segments AB0-AB5 of the frame shown in Figure 4) and metadata segments, where the audio data segments indicate audio data, and each of at least some of the metadata segments includes metadata in the state of loudness processing (LPSM) and optionally also program boundary metadata. LPSMs are present in the bitstream in the following format. Each of the metadata segments that includes
LPSM is included in an addbsi field of the
BSI of a bit stream frame, or in an auxiliary data field of a bit stream frame, or in a segment of bits left over from a bit stream frame. Each metadata segment that LPSM includes includes an LPSM payload segment (or container) that has the following format:
IMPI
<img file="MX356196B_D0026.tif" />
a header (typically including at least one identification value, for example, the format version, length, period, count, and LPSM subflow association values indicated in Table 2 below); and after the header, the LPSMs and optionally also the program limit metadata. Program limit metadata can include a program limit frame count and a code value (for example, an offset_exist value) that indicates whether the frame includes only a program limit frame count or both a frame count program limit and an offset value), and (in some cases) an offset value. LPSMs can include:
at least one dialog indication value (eg, Dialog Channels parameters in Table 2) indicating whether the corresponding audio data indicates dialogue or does not indicate dialogue (eg, which corresponding audio data channels indicate dialogue). Dialog indication values can indicate whether dialogue is present on any combination or all channels of the corresponding audio data;
at least one loudness regulation compliance value (for example, parameter 'Type of Loudness Regulation of Table 2) that indicates whether the corresponding audio data
IMP
MEXICAN INSTITUTE OF THE INDUSTRIAL REOPÍEÍJAO comply with a set indicated dn eleven- ~ + ic regulation;
at least one loudness processing value (for example, one or more of the Correction Indicator parameters of
Loudness with Dialogue Door Function, Type of
Loudness Correction from Table 2) indicating at least one type of loudness processing that has been performed on the corresponding audio data; and at least one loudness value (for example, one or more of the Loudness with Gate Function parameters relative to ITU,
Loudness with ITU Speech Gate Function, ITU Short-Term 3s Loudness (EBU 3341) and True Peak of the
Table 2) indicating at least one loudness characteristic (eg peak or average loudness) of the corresponding audio data. .
In any embodiment of the invention that contemplates uses or generates at least one loudness value that indicates the corresponding audio data, the loudness values can indicate at least one loudness measurement characteristic used to process the loudness and / or dynamic range of audio data.
In some implementations, each of the metadata segments in an addbsi field, or an auxiliary data field, or a segment of bits left over from a stream frame
IMPI
INSTITUTO MEXICANO □ E THE PROPERTY IN & USTRIAL bits has the following format:
a nuclear header (typically including a sync word that identifies the start of the metadata segment followed by identification values, for example, nuclear element version, length, and period, extended element count, and subflow association values indicated in Table 1 below); and after the nuclear header, at least one protection value (for example, an HMAC digest and audio fingerprint values, where the HMAC digest can be a 256-bit HMAC digest (using the SHA-2 algorithm) computerized over the audio data, the core element, and all the expanded elements, of an entire frame, as indicated in Table 1) useful for at least one of decryption, authentication, or validation of at least one of metadata in loudness processing state or the corresponding audio data); and also after the nuclear header, if the metadata segment includes LPSM, LPSM payload identification (ID), and LPSM payload size values that identify the following metadata as an LPSM payload and indicate the size of the LPSM payload. The
<img file="MX356196B_D0027.tif" />
LPSM payload segment (preferably having the
IMPI ^ format specified above) follows the LPSM payload ID and LPSM payload size values.
In some modalities of the type described in the previous paragraph, each of the metadata segments in the auxiliary data field (or addbsi field or leftover bit segment) of the frame has three levels of structure:
a high-level structure, which includes an indicator that indicates whether the auxiliary data field (or addbsi) includes metadata, at least one ID value that indicates what type of metadata is present, and typically also a value that indicates how many bits of metadata (for example, of each type) are present (if metadata exists). One type of metadata that could be present is LPSM, another type of metadata that could be present is program boundary metadata, and another type of metadata that could be present is media research metadata;
an intermediate tier structure, comprising a core element for each identified metadata type (eg, nuclear header, payload ID and protection values, and payload size values, eg of the type mentioned above, for each identified type of metadata); and
<img file="MX356196B_D0028.tif" />
ΙΜΡΙ a low-level structure, comprising each payload for a nuclear element (for example, an LPSM payload, if the nuclear-element identifies that one is present, and / or a metadata payload of another type, if the nuclear element identifies that one is present).
Data values can be nested in this three-level structure. For example, protection values for an LPSM payload and / or other metadata payload identified by a nuclear element may be included after each payload identified by the nuclear element (and therefore ', after the nuclear header of the nuclear element). In one example, a nuclear header could identify an LPSM payload and another metadata payload, payload ID, and payload size values for the first payload (for example, the LPSM payload) could follow the nuclear heading, the first payload could follow the size and ID values, the payload ID and the payload size value for the second payload could follow the first payload, the second payload could follow these size and ID values, and protection values for one or both of the payloads (or for nuclear element values and one or both of the payloads) could follow the last payload.
<img file="MX356196B_D0029.tif" />
IMPI
MEXICAN INSTITUTE OF LA MOHEDAL ·
INDUSTRIAL
In some embodiments, the nuclear element of a metadata segment in an auxiliary data field (or addbsi field or leftover bit segment) of a frame comprises a nuclear header (typically including identification values, eg, nuclear element version ) and after the nuclear heading: values that indicate whether fingerprint data is included for metadata in the metadata segment, values that indicate whether there is external data (related to audio data that corresponds to metadata in the metadata segment), the payload ID, and the payload size values for each type of metadata (eg LPSM, and / or metadata of a type other than LPSM) identified by the nuclear element, and protection values for at least one type of metadata identified by the nuclear element. The metadata payloads of the metadata segments follow the nuclear heading, and (in some cases) nest within the values of the nuclear element.
In another preferred format, the encoded bitstream is a Dolby bitstream. E, and each of the metadata segments that LPSM includes (and optionally also program boundary metadata) is included in the first N sample locations of the Dolby E guardband interval.
IMPI
INDI ISTUIAL
In another class of embodiments, the industrial invention is an APU (eg, a decoder) coupled and configured to receive an encoded audio bitstream comprising audio data segments and metadata segments, where the audio data segments indicate the audio data, and each of at least some of the metadata segments includes loudness processing state (LPSM) metadata and optionally also program limit metadata, and to extract the LPSMs from the bit stream, to generate decoded audio data in response to the audio data, and to perform at least one adaptive loudness processing operation on the audio data using the
LPSM. Some modalities in this class also include an APU-coupled post processor, where the post processor is coupled and configured to perform at least one adaptive loudness processing operation on the audio data using the LPSM.
In another class of embodiments, the invention is an audio processing unit (APU) that includes a buffer memory and a buffer coupled processing subsystem, where the APU is coupled to receive an encoded audio bitstream that It comprises audio data segments and metadata segments, where the segments of
<img file="MX356196B_D0030.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL Property audio data indicates audio data, and each of at least some of the metadata segments includes loudness processing state (LPSM) metadata and optionally also program limit metadata, the buffer stores ( for example, non-transiently) at least one frame of the encoded audio bitstream, and the processing subsystem is configured to extract the LPSMs from the bit stream and to perform at least one adaptive loudness processing operation on the audio data using the LPSMs. In typical embodiments in this class, the APU is one of an encoder, a decoder, and a post processor.
In some implementations of the method of the invention, the generated audio bitstream is one of a bitstream
AC-3, an E-AC-3 bitstream, or a Dolby E bitstream, which includes loudness processing state metadata, as well as other metadata (for example, a DIALNORM metadata parameter, metadata parameters dynamic range control, and other metadata parameters).
In some other implementations of the method, the generated audio bitstream is a coded bitstream of another type.
<img file="MX356196B_D0031.tif" />
IMPI
Aspects of the invention include a system or device configured (eg, programmed) to perform any embodiment of the method of the invention, and a computer-readable medium (eg, a disk) that stores codes (eg, not transient) to implement any embodiment of the method of the invention or stages thereof. For example, the system of the invention may or may include a general-purpose programmable processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform any of a variety of data operations , including an embodiment of the method of the invention or stages thereof. Said general use processor may be or may include a computer system that includes a data input device, a memory, and processing circuits programmed (and / or configured in another way) to carry out an embodiment of the method of the invention (or stages of this) in response to the data stated in this.
Brief description of the 'figures
Figure 1 is a block diagram of an embodiment of a system that can be configured to perform an embodiment of the method of the invention.
<img file="MX356196B_D0032.tif" />
Figure 2 is a block diagram of an encoder that is an embodiment of the audio processing unit of the invention.
FIG. 3 is a block diagram of a decoder that is one embodiment of the audio processing unit of the invention, and a post processor attached to it that is another embodiment of the audio processing unit of the invention.
Figure 4 is a diagram of an AC-3 frame, including the segments into which it is divided.
Figure 5 is a diagram of the Information segment of
Synchronization (SI) of an AC-3 frame, which includes the segments into which it is divided.
Figure 6 is a diagram of the Bit Stream Information (BSI) segment of an AC-3 frame, including the segments into which it is divided.
Figure 7 is a diagonal diagram of an E-AC-3 frame, which includes the segments into which it is divided.
Figure 8 is a frame diagram of an encoded audio bitstream including program limit metadata in a format according to an embodiment of the invention.
Figure 9 is a diagram of other frames of the coded audio bitstream of Figure 9. Some of these frames
<img file="MX356196B_D0033.tif" />
! Mexican mifΜι, include program limit metadata with fc®Wm, according to an embodiment of the invention.
Figure 10 is a diagram of two encoded bit streams: one bit stream (IEB) in which one program boundary (labeled Boundary) is aligned with a transition between two frames of the bit stream, and another bit stream ( TB) in which a program limit (labeled
True limit) is deviated by 512 samples from a transition between two frames of the bitstream.
Figure 11 is a set of diagrams showing four encoded audio bit streams. The bitstream at the top of Figure 11 (labeled Scenario 1) indicates a first audio program (Pl) that includes program limit metadata followed by a second audio program (P2) that also includes Program; the second bit stream (labeled Scenario 2) indicates a first audio program (Pl) that includes program limit metadata followed by a second audio program (P2) that does not include program limit metadata; the third bit stream (labeled Scenario 3) indicates a first truncated audio program (Pl) that includes program limit metadata, and that has been spliced with a second integer audio program (P2) that includes limit metadata
IMPI
MEXICAN INSTITUTE OF PROPERTY,. industrial program; and the fourth bit stream (labeled Scenario
4) indicates a first truncated audio program (Pl) that includes program limit metadata, and a second truncated audio program (P2) that includes program limit metadata, and that it is spliced with a part of the first audio program .
Notation and Nomenclature
Throughout the present description including in the claims, the expression performing an operation on a signal or data (for example, filtering, increasing, transforming or applying gain to the signal or data) is used broadly to indicate the performance of an operation directly on the signal or data, or on a processed version of the signal or data (for example, in a version of the signal that has undergone preliminary filtering or preprocessing before performing the operation on these).
Throughout the present description including in the claims, the term system is used in a broad sense to indicate a device, system or subsystem. For example, a subsystem that implements a decoder can refer to a decoder system, and a system that includes that subsystem (for example, a system that generates output signals X in response to multiple inputs
IMPIí ^ tHjTtWft) MtXlCANO
FROM LA MOPILUAU
... _ <sub>n</sub> industrial data, in which the subsystem generates M from the data inputs and the other data inputs X-M are received from an external source) can also refer to a decoder system.
Throughout the present description including in the claims, the term processor is used in a broad sense to indicate a 'system or device programmable or otherwise configurable (for example, with software or firmware) to perform operations on the data (for example, audio or video data or other image data). Examples of processors include a programmable field array of gates (or other configurable IC or chipset), a digital signal processor programmed and / or otherwise configured to perform conduit processing on audio or other data. sound, a general-purpose programmable processor or computer, and a microprocessor chip or programmable chipset.
Throughout the present description including in the claims, the expressions audio processor and
<td>unit of</td><td>processing</td><td>audio are used</td><td>of</td><td>way</td>
<td colspan="4">interchangeable, and in a broad sense, to indicate a</td><td>system</td>
<td>configured</td><td>to process</td><td>audio data. The</td><td colspan="2">examples of</td>
<td>units of</td><td>processing</td><td>audio include,</td><td>of</td><td>no way</td>
<img file="MX356196B_D0034.tif" />
IMPI
UUTITUTO MEXICANO Say LA INDUSTRIAL PROPERTY, encoders (eg, transcoders), decoders, codees, preprocessing systems, postprocessing systems and bitstream processing systems (sometimes called bitstream processing tools).
Throughout the present description including in the claims, the term "processing state metadata" (eg, as the term "loudness processing state metadata") refers to separate and different data from corresponding audio data (the content audio from an audio data stream that also includes metadata in the processing state). Processing metadata is associated with audio data, indicates the loudness processing status of the corresponding audio data (for example, what type or types of processing have already been performed on the audio data), and typically it also indicates at least one trait or characteristic of the audio data. The association of the metadata in the processing state with the audio data is synchronous.
Therefore, current processing state metadata (most recently received or updated) indicates that the corresponding audio data comprises,
IMPI
MEXICAN INSTITUTE t> AND INDUSTRIAL PROPERTY
INDUS 'ina.
<img file="MX356196B_D0035.tif" />
contemporary, the results of the indicated type or types of audio data processing. In some cases, metadata in the processing state may include processing history and / or some or all of the parameters that are used in and / or derived from the indicated types of processing. Additionally, metadata in the processing state may include at least one feature or characteristic of the corresponding audio data, which has been computed or extracted from the audio data. Processing metadata may also include other metadata that is not related to or derived from any of the corresponding audio data. For example, third party data, tracking information, identifiers, specific or standard information, user annotation data, user preference data, etc. can be added. using a particular audio processing unit to switch to other audio processing units.
Throughout the present description including in the claims, the term loudness processing state metadata (or LPSM) means processing state metadata indicating the loudness processing status of the corresponding audio data (for example, what type or types of
<img file="MX356196B_D0036.tif" />
example, loudness) of the corresponding audio data. Loudness processing metadata can include data (for example, other metadata). they are not (that is, when considered alone) metadata in loudness processing state.
Throughout the present description including in the claims, the term channel (or audio channel) means a monophonic audio signal.
Throughout the present description including in the claims, the term audio program means a set of one or more to audio channels and optionally also associated metadata (eg, metadata describing a desired spatial audio presentation, and / or
LPSM, and / or program limit metadata).
Throughout the present description including in the claims, the expression program boundary metadata indicates metadata of an encoded audio bitstream, where the encoded audio bitstream indicates at least one audio program (eg two or more audio programs), and the program limit metadata indicates the location in the bitstream of at least one
INSTITUTO MEXICANO limit (start and / or end) of at least one of cfiWQ & iAprb ^ 3Sffl ^ audio. For example, meta-data-of-program (from a stream of coded audio bits indicating an audio program) may include metadata indicating the location (for example, the beginning of the N frame of the data stream). bits, or the location of sample M of frame N of the bitstream) of the beginning of the program, and additional metadata indicating the location (for example, the beginning of frame J of the bitstream, or the location of sample K of frame J of the bit stream) at the end of the program.
Throughout the present description including in the claims, the term coupled or coupled is used to indicate a connection either direct or indirect. Therefore, if a first device is coupled to a second device, that connection can be through a direct connection or through an indirect connection through other devices and connections.
Detailed description of the embodiments of the invention
In accordance with typical embodiments of the invention, a payload of program loudness metadata, called loudness processing state metadata (LPSM) and optionally also program boundary metadata are
IMP
MEXICAN INSTITUTE OF PROPERTY
<img file="MX356196B_D0037.tif" />
embedded in one or more reserved fields (<sup>1</sup>S<sup>) UST</sup>f ^ nu metadata segments - from a d / ά 'stream of SUdifi, also include audio data in other segments (audio data segments). Typically, at least one segment of each leg of the bitstream includes LPSM, and at least one other segment of the leg · includes the corresponding audio data (i.e., audio data whose loudness and loudness processing status is indicated by the LPSM ). In some modalities, the volume of data of the
LPSM can be small enough to be carried out without affecting the assigned bit rate to carry the audio data.
Metadata communication in loudness processing state in an audio processing chain is particularly useful when two or more audio processing units need to work in tandem with each other through the processing chain (or content life cycle) . Without the inclusion of loudness processing state metadata in an audio bitstream, serious processing problems such as quality, level and spatial degradations can occur, for example, when two or more audio codes are used in the string and asymmetric volume leveling is applied more than once during the bitstream path
IMPI
<img file="MX356196B_D0038.tif" />
MEXICAN INSTITUTE A
t) t THE PROPERTY
INDUSTRIAL pzSr 'to a media consuming device (or a playback point of the bitstream audio content).
Figure 1 is a block diagram of an example of an audio processing chain (an audio data processing system), in which one or more elements of the system can be configured according to an embodiment of the present invention. The system includes the following elements, coupled together as shown: A preprocessing unit, an encoder, a signal analysis and metadata correction unit, a transcoder, a decoder and a preprocessing unit. In variations on the system shown, one or more of the elements is omitted, or additional audio data processing unit is included.
In some implementations, the preprocessing unit of Figure 1 is configured to accept PCM (time-domain) samples that comprise audio content as input, and to generate processed PCM samples. The encoder can be configured to accept PCM samples as input and to generate an encoded (eg compressed) audio bitstream indicating the audio content. The. bitstream data indicating
IMPí'é >>
MEXICAN INSTITUTE OF PROPERTY the audio content is sometimes referred to in<sup>1N</sup>^ S<sup>R1A</sup>^ res ^ TTte audio data. If the encoder is contiguous 3 <3Ó “~ 3β ~ according to a typical embodiment of the present invention, the encoder audio bitstream output includes loudness processing state metadata (and typically other metadata as well, which. optionally include program limit metadata) as well as audio data.
The signal analysis and metadata correction unit of Figure 1 can accept one or more encoded audio bit streams as input and determine (eg validate) whether the metadata is in processing state in each encoded audio bit stream they are correct, when performing signal analysis (for example, using program limit metadata in an encoded audio bitstream). If the signal analysis and metadata correction unit discovers that the included metadata is invalid, it typically replaces the incorrect value (s) with the correct value (s) obtained from the signal analysis. Therefore, each encoded audio bitstream output from the signal analysis and metadata correction unit may include corrected (or uncorrected) processing state metadata as well as encoded audio data.
MEXICAN INSTITUTE. '' Ά 'hi ι * property VV I.
iNStmiT
Γ1Ε THE PROPERTY Vj .....
The transcoder in Figure 1 can accept £ ^<sup>R</sup>^ lujOis ~ clé audio bits encoded as input, and '' ^ enefár<sup>ri</sup>It cracks ^ '' changed audio bits (e.g., encoded differently) in response (e.g., by decoding an input stream and re-decoding the decoded stream in a different encoding format). If the transcoder is configured in accordance with a typical embodiment of the present invention, the encoder audio bitstream output includes loudness processing state metadata (and typically also other metadata) as well as encoded audio data. Metadata may have been included in the bit stream.
The decoder of Figure 1 can accept encoded (eg compressed) audio bitstreams as input, and outgoing (in response) streams of decoded PCM audio samples. If the decoder is configured in accordance with a typical embodiment of the present invention, the output of a decoder in normal operation is or includes any of the following:
an audio sample stream, and a corresponding loudness processing state metadata stream (and typically also other metadata) extracted from an input encoded bit stream; or
IMPI
Yz * MEXICAN INSTITUTE
GIVE THE SP.OPISUAU V>. -,
INDUSTRIAL an audio sample stream, and a corresponding stream of control bits determined from loudness processing state metadata (and typically also other metadata) extracted from an input encoded bitstream; or a stream of audio samples, without a corresponding stream of processing status metadata or control bits determined from processing status metadata. In the latter case, the decoder can extract loudness processing (and / or other metadata) · metadata from the input encoded bitstream and perform at least one operation on the extracted metadata (eg validation), including if it does not generate the extracted metadata or control bits determined from it.
By configuring the post-processing unit of Figure 1 according to a typical embodiment of the present invention, the post-processing unit is configured to accept a stream of decoded PCM audio samples, and to post-process these (for example, volume leveling of audio content) using loudness processing state metadata (and typically other metadata as well) received with
IMPI Mexican Institute Dt THE PROPERTY
<img file="MX356196B_D0039.tif" />
samples, or control bits (determined by the metadata in sonic processing state <T ~ and 'typically also other metadata) received with the samples. The post-processing unit is also typically configured to play the post-processed audio content for playback by one or more speakers.
Typical embodiments of the present invention provide an improved audio processing chain where the audio processing units (eg, encoders, decoders, transcoders, and pre and post processing units) adapt their respective processing to be applied to data data according to a simultaneous state of the media data as indicated by the loudness processing state metadata respectively received by the processing units of Audio.
The input of audio data to any audio processing unit of the system of Figure 1 (for example, the encoder or transcoder of Figure 1) can include loudness processing state metadata (and optionally also other metadata) as well as also audio data (eg encoded audio data). These metadata could have been included in the
IMPI
MEXICAN INSTITUTE
<img file="MX356196B_D0040.tif" />
L> E FROPIEDAD audio input using another element of the s'ÍSféfrta
Figure 1 (or other source, not shown in Figure according to an embodiment of the present invention. The processing unit that receives the input audio (with metadata) can be configured to perform at least one operation on the metadata (for example, validation) or in response to metadata (eg. adaptive input audio processing), and typically also to include the metadata in the output audio, a processed version of the metadata, or control bits determined from the metadata.
A typical embodiment of the invention's audio processing unit (or audio processor) is configured to perform adaptive processing of audio data based on the state of the audio data as indicated by metadata in the state of audio processing. loudness corresponding to audio data. In some embodiments, adaptive processing is (or includes) loudness processing (if metadata indicates that loudness processing, or processing similar to this, has not yet been performed on the audio data, but is not (and is not includes) loudness processing (if metadata indicates that loudness processing, or
<img file="MX356196B_D0041.tif" />
IMPI<sup>4</sup>*
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL (processing similar to this, has already been performed on the audio data). In some embodiments, adaptive processing is or includes metadata validation (for example, performed on a metadata validation subunit) to ensure that the audio processing unit performs other adaptive processing of audio data based on the state of audio data as indicated by loudness processing metadata. In some embodiments, validation determines the reliability of loudness processing state metadata associated with (for example, included in a bitstream with) the audio data. For example, if the metadata is validated to be reliable, then the results of a previously performed type of processing can be reused and the same type of audio processing can be prevented from being performed. On the other hand, if the metadata is found to have been altered (or otherwise unreliable), then the type of media processing allegedly previously performed (as indicated by the unreliable data) may be repeated by from the audio processing unit, and / or other processing by the audio processing unit may be performed on the metadata and / or the audio data. The unit of
<img file="MX356196B_D0042.tif" />
Mexican INSTITUTE t) F. THE PROPERTY
t) F. THE TIP
<img file="MX356196B_D0043.tif" />
ur. the ·<sup>r</sup> '<sup>, t</sup> 'audio processing can also be configured to send signals to other processing units'<sup>ro</sup>subsequent audio in an enhanced media processing string indicating that loudness processing state metadata (for example, present in a media bit stream) is valid, if the unit determines that the processing state metadata is valid ( for example, based on a match of a extracted cryptographic value and a reference cryptographic value).
Figure 2 is a block diagram of an encoder (100) that is an embodiment of the audio processing unit of the invention. Any of the components or elements of encoder 100 can be implemented as one or more processes and / or one or more circuits (eg, ASIC, FPGA, or other integrated circuits), in hardware, software, or a combination of hardware and software. Encoder 100 comprises framebuffer 110, lll parser, decoder 101, audio status validator 102, loudness processing stage 103, audio stream selection stage 104, encoder 105, fill / format stage
107, metadata generation step 106, dialogue loudness measurement subsystem 108 and framebuffer 109, connected as shown. Typically also, the
IMPI
MEXICAN INSTITUTE OF ERUITEDAD
<img file="MX356196B_D0044.tif" />
encoder 100 includes other prodflSW & fere elements shown). -<sup>11</sup> ·
Encoder 100 (which is a transcoder) is configured to convert an input audio bitstream (which, for example, can be one of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream) into an encoded output audio bitstream (which, for example, may be another of an AC-3 bitstream, an E-AC-3 bit stream or Dolby E) bit stream including performing adaptive and automated loudness processing using the loudness processing state metadata included in the input bitstream. For example, encoder 100 can be configured to convert an input Dolby E bitstream (a format typically used in production and broadcast facilities but not in consumer devices that receive audio programs that have been broadcast on it) into an encoded output audio bitstream (suitable for transmission to consumer devices) in AC-3 format or
E-AC-3.
The system of Figure 2 also includes an encoded audio supply subsystem 150 (which stores and / or supplies the output of encoded bit streams from the
IMPI
<img file="MX356196B_D0045.tif" />
<sub>η</sub> . . <sub>Ί</sub> ,,. MEXICAN INSTITUTE, encoder 100) and decoder 152. An encoded audio f-bit of a 100 pnprip ς encoder<sub>Ρ</sub>γ stored by subsystem 150 (for example, in the form of a DVD or Blu ray disc), or transmitted by subsystem 150 5 (which can implement a transmission link or network), or can be both stored and transmitted via the subsystem
150. Decoder 152 is configured to decode an encoded audio bitstream (generated by encoder 100) that is received by subsystem 150, including by extracting loudness processing state (LPSM) metadata from each bitstream frame (and optionally also bitstream program boundary metadata extraction), and generate decoded audio data. Typically, the decoder 152 is configured to perform adaptive loudness processing on the decoded audio data using the LPSMs (and optionally also program limit metadata), and / or to send the decoded audio data and the LPSMs to a post processor configured to perform adaptive loudness processing on decoded audio data using LPSMs (and optionally also program limit metadata). Typically, decoder 152 includes a
IMPI
MEXICAN INSTITUTE
-1 rr «rr <sub>Ί</sub> , 'PROPERTY τι buffer that stores (for example, no'NWSfi-siT the encoded audio bitstream tree-i-ferreters — from .....' it '* subsystem 150.
Various implementations of encoder 100 and decoder 152 are configured to perform different embodiments of the method of the invention.
Framebuffer 110 is a buffer memory attached to receive an encoded input audio bitstream. In operation, buffer 110 stores (eg, non-transiently) at least one frame of the encoded audio bitstream, and a sequence of the frames of the encoded audio bitstream is committed from the buffer
110 to the lll parser.
The lll parser is docked and configured to extract Loudness Processing State (LPSM) metadata, and optionally also Program Limit metadata (and / or other metadata) from each frame of the encoded input audio in which they are included said metadata, to confirm the LPSM (and optionally also program limit metadata and / or other metadata) in the audio state validator 102, loudness processing step 103, step 106 and subsystem 108, to extract audio data from the encoded input audio, and to confirm the data from
IMPI
MEXICAN INSTITUTE OF PROPERTY decodif i d & 'dS'i *<sup>1</sup>·
<img file="MX356196B_D0046.tif" />
decoded, v audio face on decoder 101. Encoder 100 is configured for audio to generate audio data confirming decoded audio data in loudness processing step 103, audio stream selection step 104, subsystem 108, and typically also for status validator 102.
The state validator 102 is configured to authenticate and validate the LPSMs (and optionally other metadata) committed to it. In some embodiments, LPSMs are (or are included in) a block of data that has been included in the input bitstream (eg, in accordance with one embodiment of the present invention). The block may comprise a cryptographic hash (a hash-based message authentication code or HMAC) to process the LPSMs (and optionally other metadata) and / or the underlying audio data (provided from the decoder 101 up to validator 102). The data block can be digitally signed in these modalities, so that a subsequent audio processing unit can relatively easily authenticate and validate metadata in the processing state.
IMPI
For example, the HMAC is used to generate a
MEXICAN INSTITUTE OF PROPERTY
<img file="MX356196B_D0047.tif" />
the protection values included in e 1 f lújó (Té bita dS '"may also include the digest. The digest can be generated as follows for an AC-3 frame:
one. After the AC-3 and LPSM data are encoded, the data bytes of the frame (frame_d<sup>tie</sup> # 1 and frame_data # 2 concatenated) and the LPSM data bytes are used as input for the HMAC hash function. Other data, which may be present within an auxiliary data field, are not taken into account to calculate the digest. Other data may be bytes that do not belong to AC-3 data or data
LSPSM. The protection bits included in the LPSM may not be taken into account to calculate the HMAC digest.
2. After the digest is calculated, it is written to the bitstream in a field reserved for protection bits.
3. The last step in generating the complete AC-3 frame is the calculation of the CRC check. This is written to the far end of the frame and all data belonging to this frame is taken into account, including the bits of
LPSM.
Other cryptographic methods, including but not limited to any one or more cryptographic methods not
IMPI
MEXICAN INSTITUTE 'Zií' & TSSé & ffi
FROM THE PROPERTY VfeateífiJrffJi »'
INDUSTRIAL ^ * Λ-ϊ £ «ϊ · *
HMAC can be used for validation of. LPSMs (eg, in validator 102) to ensure secure transmission and reception of LPSMs and / or underlying audio data. For example, validation (using said cryptographic method) can be performed on each audio processing unit that receives an embodiment of the audio bitstream of the invention to determine whether the loudness processing state metadata and the corresponding audio data included in the bitstream has undergone (and / or been the result of) specific loudness processing (as indicated by the metadata) and has not been modified after the completion of said specific loudness processing.
The state validator 102 confirms the control data in the audio stream selection step 104, metadata generator 106, and dialogue loudness measurement subsystem 108, to indicate the results of the validation operation. In response to the control data, step 104 can select (and pass through to encoder 105) either:
the adaptively processed output of loudness processing step 103 (for example, when the LPSMs indicate that the audio data output from the decoder
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX356196B_D0048.tif" />
101 it has not undergone a specific type of loudness processing, and the validator 102 control bits indicate that the LPSMs are valid); or the audio data output from decoder 101 (for example, when LPSMs indicate that the audio data output from decoder 101 has already undergone the specific type of loudness processing that would be performed by step 103, and the bits Check Values 102 indicate that the LPSMs are valid.)
Step 103 of encoder 100 is configured to perform adaptive loudness processing on the decoded audio data output from decoder 101, based on one or more audio data characteristics indicated by LPSMs extracted by decoder 101. Step 103 may be an adaptive real-time loudness transformation-domain processor and dynamic range control. Step 103 may receive user data input (eg user target dynamic range / loudness values or dialnorm values), or other metadata input (eg one or more third party data types, trace information , identifiers, specific or standard information, user annotation data, user preference data, etc.) and / or other
<img file="MX356196B_D0049.tif" />
IMPI
INSTITUTO MEXICANO DE IX PROPIEDAD INDUSTRIAL data input (for example, from a fingerprint process), and use said data input to process the output of decoded audio data from the decoder
101. Step 103 can perform adaptive loudness processing on the decoded audio data output (output from decoder 101) indicating a single audio program (as indicated by the program limit metadata extracted by the lll parser. ), and it may reinitialize loudness processing in response to receiving the decoded audio data (output from decoder 101) indicating a different audio program as indicated by the program limit metadata extracted by the lll parser.
Dialogue loudness measurement subsystem 108 can operate to determine the loudness of decoded audio segments (from decoder 101) indicating dialogue (or other vocal broadcasts), for example, using LPSM (and / or other metadata) fetched by decoder 101, when the validator 102 control bits indicate that the LPSMs are invalid. The operation of the dialogue loudness measurement subsystem 108 can be disabled when the LPSMs indicate previously determined dialogue loudness segments (or other vocal emissions) of the audio
IMPI
<img file="MX356196B_D0050.tif" />
decoded (from decoder 101) when the _check bits of validator 102 indicate that the LPSMs are valid.
Subsystem 108 can perform a loudness measurement on the decoded audio data indicating a single audio program (as indicated by the program limit metadata extracted by parser 111), and can reset the measurement in response to the receiving the decoded audio data indicating a different audio program as indicated by said program limit metadata.
There are useful tools (for example, the Dolby LM100 Loudness Meter) to measure the level of dialogue in audio content easily and conveniently. Some embodiments of the APU of the invention (eg, step 108 of encoder 100) are implemented to include (or to perform the functions of) such a tool to measure the average dialogue loudness of the audio content of a stream of audio bits (eg a confirmed decoded AC-3 bit stream for step 108 from decoder 101 of encoder 100).
If step 108 is implemented to measure the true average loudness of the audio data, the measurement may include a segment isolation step of the
IMPIr
INSTITUTO MRXICAN ,, '»*
<img file="MX356196B_D0051.tif" />
GIVE PROPERTY andio content that mainly contain<sup>1</sup>*<sup>1</sup> Vocal surveys. Audio segments that are primarily vocal broadcasts are then processed according to a loudness measurement algorithm. For audio data decoded from an AC-3 bitstream, this algorithm can be a standard measure of K-weighted loudness (according to the international standard ITU-R
BS.1770). Alternatively, other loudness measures can be used (for example, those based on psychoacoustic loudness models).
Isolation of vocal broadcast segments is not essential to measure the average dialogue loudness of audio data. However, it improves measurement precision and typically provides more satisfactory results from a listener's perspective. Since not all audio content contains dialogue (vocal broadcasts), the loudness measurement of the entire audio content can provide a sufficient approximation of the dialogue level of the audio, if the vocal emissions are present.
Metadata generator 106 generates (and / or passes through to step 107) metadata that will be included in the bitstream encoded by step 107 for encoder 100 to generate as output. The generator
IMPI
<img file="MX356196B_D0052.tif" />
metadata 106 can pass, through step 107, the LPSMs (and optionally also program limit metadata and / or other metadata) extracted by encoder 101 and / or parser lll (for example, when control bits Validator 102 indicate that the LPSM and / or other metadata is valid), or generate new LPSM (and optionally also program limit metadata and / or other metadata) and confirm the new metadata in step
107 (for example, when the validator control bits
102 indicate that the LPSM and / or other metadata extracted by decoder 101 is invalid, or you can confirm in step 107 a combination of metadata extracted by decoder 101 and / or parser lll and newly generated metadata. Metadata generator 106 may include loudness data generated by subsystem 108, and at least one value indicating the type of loudness processing performed by subsystem 108, in LPSMs this confirms in step 107 for inclusion in the stream of coded bits for encoder 100 to generate as output.
Metadata generator 106 can generate protection bits (which may consist of or include a hash-based message authentication code or HMAC) useful for
IMPI
<img file="MX356196B_D0053.tif" />
at least one of decryption, authentication, or validation of the LPSMs (and optionally also other metadata) to be included in the encoded bitstream and / or the underlying audio data to be included in the encoded bitstream. Metadata generator 106 may provide such protection bits in step 107 for inclusion in the encoded bit stream.
In normal operation, the dialogue loudness measurement subsystem 108 processes the audio data output from decoder 101 to generate, in response to this, loudness values (eg, dialog loudness values with gate function or without gate function) and dynamic range values. In response to other values, the metadata generator 106 can generate loudness processing state (LPSM) metadata for inclusion (by padding / formatting 107) in the encoded bit stream for encoder 100 to generate as output.
Additionally, optionally or alternatively, the 106 and / or 108 subsystems of encoder 100 may perform further analysis of audio data to generate metadata indicating at least one characteristic of the audio data for inclusion in the 'stream of bits encoded so that the
Μ
<img file="MX356196B_D0054.tif" />
I saw stage 107 generate them as output.
Encoder 105 encodes (eg, by compressing these) the audio data output from selection stage 104, and commits the encoded audio in step 107 for inclusion in the encoded bitstream for the stage 107 generate them as output.
Step 107 multiplexes the encoded audio from encoder 105 and the metadata (including LPSMs) from generator 106 to generate the encoded bit stream for output by step 107, preferably so that the encoded bit stream has the format specified by a preferred embodiment of the present invention.
Framebuffer 109 is a buffer memory that stores (for example, non-transiently) at least one frame of the output of the encoded audio .bit stream from the stage
107, and then a sequence of the encoded audio bitstream frames from buffer 109 is reconfirmed as data output from encoder 100 to delivery system 150.
The LPSMs generated by the metadata generator 106 and included in the bitstream encoded by step 107 indicates the loudness processing status of the data
<img file="MX356196B_D0055.tif" />
Corresponding audio IMPIs (eg, what type or types of loudness processing has been performed on the audio data) and loudness (eg, measured dialogue loudness, loudness with gate function and / or without gate function, and / or dynamic range) of the corresponding audio data.
Herein, loudness gate function and / or level measurements made on the audio data refers to a specific loudness level or threshold where the computed value or values that exceed the threshold are included in the final measurement (for example , ignoring short-term loudness values below -60 dBFS in the final measured values). Gate function at an absolute value refers to a fixed level or loudness, while gate function at a relative value refers to a value that depends on a current measurement value without gate function.
In some implementations of encoder 100, the encoded bitstream stored in buffer 109 (and generated as output to supply system 150) is either an AC-3 bitstream or an E-AC-3 bitstream, and it comprises audio data segments (for example, segments AB0AB5 of the frame shown in Figure 4) and metadata segments, where the audio data segments indicate
<img file="MX356196B_D0056.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<td>Data of</td><td>Audio,</td><td>and every</td><td>one of</td><td>at least</td><td>some of</td><td>the</td>
<td>segments</td><td colspan="2">metadata</td><td>It includes</td><td>metadata</td><td>in state</td><td>of</td>
<td colspan="2">processing of</td><td colspan="2">loudness (LPSM)</td><td>. The phase</td><td>107 inserts</td><td>the</td>
<td>LPSM (and</td><td colspan="2">optionally</td><td>too</td><td>metadata</td><td>limit</td><td>of</td>
<td>Program)</td><td>at</td><td>flow of</td><td>bits in</td><td colspan="2">the following format.</td><td>Every</td>
one of the metadata segments including LPSM (and optionally also program boundary metadata) is included in a segment of excess bits in the bit stream (eg segment of excess bits W as shown in Figure 4 or Figure 7), or an addbsi field of the bitstream information segment (BSI) of a bitstream frame, or in an auxiliary data field (for example, the AUX segment shown in Figure 4 or Figure 7) at the end of a bitstream frame. A frame may include one or two metadata segments, each of which includes LPSM, and if the frame includes two metadata segments, one may be present in the frame's addbsi field and the other in the frame's AUX field. . In some embodiments, each metadata segment that LPSM includes includes an LPSM payload segment (or container) that has the following format:
a header (typically including a sync word that identifies the start of the payload of
<img file="MX356196B_D0057.tif" />
IMPI
INéTIfU'rO MEXICANO
OF THE PROPERTY .
LPSM followed by at least one identi'FTS'S ^ ió ^ '^ T ^ r example value, format version, length,' pLflÓÜ'b ”corifeo, and LPSM subflow association values indicated in the
Table 2 below); · and after the header, at least one dialog prompt value (eg Dialog Channels parameter in Table 2) indicating whether the corresponding audio data indicates dialog or does not indicate dialog (eg , which corresponding audio data channels indicate dialogue);
at least one loudness regulation compliance value (eg parameter 'Loudness Regulation Type in Table 2) indicating whether the corresponding audio data complies with an indicated set of loudness regulations;
at least one loudness processing value (for example, one or more of the Loudness Correction Indicator parameters with Dialog Door Function, Type of
Loudness Correction from Table 2) indicating at least one type of loudness processing that has been performed on the corresponding audio data; and at least one loudness value (for example, one or more of the Loudness with Gate Function parameters relative to ITU,
<img file="MX356196B_D0058.tif" />
IMPI
Loudness with ITU Vocal Emissions Gate Function, ITU Short-Term 3s Loudness (EBU 3341) and True Peak from Table 2) indicating at least one loudness characteristic (e.g. peak or average loudness) of the corresponding audio.
In some embodiments, each LPSM-containing metadata segment and the program boundary metadata contain a nuclear header (and optionally additional nuclear elements as well) and, after the nuclear header (or the nuclear header and other nuclear elements) a load segment LPSM (or container) tool that has the following format:
a header, typically including at least one identification value (eg, format version, length, period, count, and LPSM subflow association values, as indicated in Table 2 herein), and after the header, .lpsm and program limit metadata. Program limit metadata can include a program limit frame count and a code value (for example, an offset_exist value) that indicates whether the frame includes alone, a program limit frame count, or both a count of program limit frame and an offset value), and (in some cases) an offset value.
<img file="MX356196B_D0059.tif" />
MEXICAN INSTITUTE OF THE f-KOriWAD
In some implementations, each of the metadata inserted by step 107 into a leftover agmar or an addbsi field or ancillary data field of a bitstream frame has the following format: a nuclear header (typically including a sync word that identifies the start of the metadata segment followed by identification values, for example, version, length, and period, extended element count, and subflow association values of the nuclear element indicated in Table 1 below); and after the nuclear header, at least one protection value (for example, HMAC digest and audio fingerprint values from Table 1) useful for at least one of decryption, authentication, or validation of at least one metadata in the state of loudness processing or the corresponding audio data); and and also after the nuclear header, if the metadata segment includes LPSM, LPSM payload identification (ID), and LPSM payload size values that identify the following metadata as a payload of
LPSM and indicate the size of the LPSM payload.
The LPSM payload (or container) segment (preferably having the format specified above)
ΙΜΡΙ6Β »
MEXICAN INSTITUTE • <sub>η T</sub> _,,,. PROPERTY C follows the LPSM payload ID and the 'Nvawres ^ -I know LPSM payload size. ........... —................
In some modes, each of the metadata segments in the auxiliary data field (or addbsi field) of a frame has three levels of structure:
a high-level structure, which includes an indicator indicating whether the auxiliary data field (or addbsi) includes metadata, at least one value, of ID indicating what type of metadata is present, and typically also a value indicating how many bits metadata (for example, of each type) are present (if metadata is present). One type of metadata that could be present is LPSM, another type of metadata that could be present is program boundary metadata, and another type of metadata that could be present is media research metadata (for example, Media Research metadata. Nielsen);
an intermediate level structure, comprising a core element for each type of metadata identified (for example, nuclear header, protection values, and LPSM payload ID and LPSM size values, as mentioned above, for each type metadata identifier); and
MEXICAN INSTITUTE a low-level structure, comprising dá ^^^ ggá ^^ á ^ for a nuclear element (for example, a r.arga úr-ii LPSM, if the nuclear element identifies that one is present, and / or a metadata payload of another type, if the nuclear element identifies that one is present).
Data values can be nested in this three-level structure. For example, protection values for a payload and / or other metadata payload identified by a nuclear element may be 'included after each payload identified by the nuclear element (and therefore after the nuclear heading of the element nuclear). In one example, a nuclear header could identify an LPSM payload and another metadata payload, payload ID, and payload size values for the first payload (for example, the LPSM payload) could follow the nuclear header, the first payload could follow the size and ID values, the payload ID and the payload size value for the second payload could follow the first payload, the second payload could follow these IDs and size values, and protection bits for both payloads (or for nuclear element values and both payloads) could follow the last payload.
ΙΜΡΙ »
MEXICAN INSTITUTE OF PROPERTY
In some modes, if the lOT decoder<sup>UST</sup>VeciBte ^ generated audio bitstream according to the invention of the invention with cryptographic hash, the decoder is configured to parse and retrieve the cryptographic hash of a determined data block from the bit streams, said block comprising loudness processing state metadata (LPSM) and optionally also program limit metadata. Validator 102 can use the cryptographic hash to validate the received bitstream and / or associated metadata. For example, if validator 102 finds that LPSMs are valid based on a value match between a reference cryptographic hash and the cryptographic hash retrieved from the data block, then you can disable processor operation
103 in the corresponding audio data and causing selection step 104 to pass through (unchanged) the audio data. Additionally, optionally or alternatively, other types of cryptographic techniques can be used instead of a method based on a cryptographic hash.
Encoder 100 of Figure 2 can determine (in response to LPSM, and optionally also program limit metadata, extracted by decoder 101) that a
MEXICAN INSTITUTE
OF PROPERTY | \ ΠΓ.
INDUSTRIAL post / preprocessing unit has performed a kind of loudness processing on the audio data to be encoded (in elements 105, 106 and 107) and therefore can create (in generator 106) metadata in processing state of loudness that includes the specific parameters used in and / or derived from the loudness processing performed previously. In some implementations, the encoder
100 You can create (and include in the output of bitstreams encoded in it) metadata in processing state that indicates the processing history on the audio content as long as the encoder detects the types of processing that have been performed on the audio content .
Figure 3 is a block diagram of a decoder (200) which is an embodiment of the audio processing unit of the invention, and of a post processor (300) coupled thereto. The post processor 300 is also an embodiment of the audio processing unit of the invention.
Any of the components or elements of the decoder
200 and post processor 300 can be implemented as one or more processes and / or one or more circuits (eg, ASIC, FPGA, or other integrated circuits), in hardware, software, or a combination of hardware and software. Decoder 200
IMPI
<img file="MX356196B_D0060.tif" />
it comprises framebuffer 201, parser 205, audio decoder 202, audio status validation stage (validator) 203 and control bit generation stage 204, connected in the manner shown.
Typically in addition, decoder 200 includes other processing elements (not shown).
Framebuffer 201 (a buffer) stores (eg, non-transiently) at least one frame of the encoded audio bitstream received by decoder 200. A sequence of the frames of the encoded audio bitstream is committed from buffer 201 to parser 205.
The parser 205 is coupled and configured to extract loudness processing state (LPSM) metadata and optionally also program boundary metadata, and other metadata from each frame of the encoded input audio, to confirm at least the LPSM (and program limit metadata if any is extracted) to audio status validator 203 and step 204, to confirm the LPSMs (and optionally also program limit metadata) as output (for example, to the post processor 300) to extract audio data from the encoded input audio, and to confirm the extracted audio data to the
ΙΜΡΪ
<img file="MX356196B_D0061.tif" />
decoder 202,
The encoded audio bitstream input 'to decoder 200 may be one of an AC-3 bitstream, an E-AC-3 bitstream, or a Dolby E bitstream.
The system in Figure 3 also includes a post processor
300. Postprocessor '300 comprises a framebuffer 301 and other processing elements (not shown) that include at least one processing element coupled to buffer 301. Framebuffer 301 stores (eg, non-transiently) at least one frame of the decoded audio bitstream received by post processor 300 of decoder 200. The processing elements of post processor 300 are coupled and configured to adaptively receive and process a sequence of frames from the decoded audio bitstream output of the buffer.
301, using metadata output (including LPSM values) from decoder 202 and / or control bit output from step 204 of decoder 200.
Postprocessor 300 adaptive loudness processing is configured on the decoded audio data using the LPSM values and optionally also program limit metadata (for example, depending on loudness processing status, and / or one or
Typically the to perform
IMPI
MEXICAN INSTITUTE
FROM PROPERTY ____ plus audio data characteristics, indicated jJ '^ r'TÉ'SM for audio data indicating a Tt5' "'single audio program;
Various implementations of decoder 200 and post processor 300 are configured to perform different embodiments of the method of the invention.
The audio decoder 202 of the decoder 200 is configured to decode the audio data extracted by the parser 205 to generate decoded audio data and to confirm the decoded audio data as output (eg to the post processor)
<img file="MX356196B_D0062.tif" />
300)
The state validator 203 is configured to authenticate and validate the LPSMs (and optionally other metadata) committed to it. In some embodiments, LPSMs are (or are included in) a data block that has been included in the input bitstream (eg, in accordance with one embodiment of the present invention). The block may comprise a cryptographic hash (a hash-based message authentication code or HMAC) to process the LPSMs (and optionally also other metadata) and / or the underlying audio data (provided from the 205 parser and / or decoder 202 to validator 203).
The data block can be digitally signed in these
IMPIOS
MEXICAN INSTITUTE,
OE THE PROPERTY Otejr -.- iU.sp.r,
INDUSTRIAL modalities, so that a post audio processing unit can relatively easily authenticate and validate metadata in the processing state.
Other cryptographic methods, including but not limited to any one or more non-HMAC cryptographic methods, may be used for LPSM validation (eg, in validator 203) to ensure the secure transmission and reception of LPSMs and / or underlying audio data. For example, validation (using said cryptographic method) can be performed on each audio processing unit that receives an embodiment of the audio bitstream of the invention to determine whether the loudness processing state metadata and the corresponding audio data included in the bitstream has undergone (and / or been the result of) specific loudness processing (as indicated by the metadata) and has not been modified after the completion of said specific loudness processing.
Status validator 203 confirms the control data for control bit generator 204, and / or confirms the control data as output (eg, to post processor 300) to indicate the results of the validation operation. In response to control data (and optionally also
<img file="MX356196B_D0063.tif" />
IMPI other metadata extracted from input bitstream), step 204 can generate (and commit to post processor 300) if:
The control bits indicating that the decoded audio data output from decoder 202 has undergone a specific type of loudness processing (when the
LPSMs indicate that the audio data output from decoder 202 has undergone the specific type of loudness processing, and the control bits of validator 203 indicate that the LPSMs are valid); or the control bits indicating that the decoded audio data output from decoder 202 must undergo a specific type of loudness processing (for example, when LPSMs indicate that the audio data output from decoder 202 has not experienced the type specific loudness processing, or when the LPSMs indicate that the audio data output from decoder 202 has undergone the specific type of loudness processing but the check bits of validator 203 indicate that the LPSMs are invalid).
Alternatively, decoder 200 commits the metadata extracted by decoder 202 from the input bitstream, and LPSMs (and optionally also
I Ml pi
MEXICAN INSTITUTE> / ·
OE PROPERTY '--¿.-Λ' program limit metadata) extracted by the syntactic Tcffi ^ ÍLizl't 205 from the postprocessor 300 input bitstream, and the postprocessor 300 performs loudness processing on the decoded audio data using the LPSM (and optionally also program limit metadata) or perform validation of the
LPSM then performs loudness processing on the decoded audio data using the LPSMs (and optionally also program limit metadata) if the validation indicates that the LPSMs are valid.
In some embodiments, if encoder 200 receives an audio bitstream generated in accordance with an embodiment of the invention with cryptographic hashing, the encoder is configured to parse and retrieve the cryptographic hash of a given data block from the bit streams, said block comprising loudness processing state metadata (LPSM). Validator 203 can use the cryptographic hash to validate the received bitstream and / or associated metadata. For example, if validator 203 finds that the LPSMs are valid based on a value match between a reference cryptographic hash and the cryptographic hash retrieved from the data block, then it can send a signal to a data unit.
<img file="MX356196B_D0064.tif" />
IMPI
MEXICAN INSTITUTE
OE THE PROP1EUAD O * s «g *» »» Pp, INDUSTRIAL · * - 'post audio processing (eg postprpcets.adnjs.— 300, which can be or include a volume leveling unit) to go to through (unchanged) the bitstream audio data. Additionally, optionally or alternatively, other types of cryptographic techniques can be used instead of a cryptographic hash-based method.
In some implementations of decoder 200, the encoded (and stored in memory 201) bitstream is either an AC-3 bitstream or an E-AC-3 bitstream, and comprises audio data segments (eg, the segments AB0AB5 of the frame shown in Figure 4) and metadata segments, where the audio data segments indicate audio data, and each of at least some of the metadata segments includes loudness processing state (LPSM) metadata and optionally also program boundary metadata. The decoder stage
202 (and / or parser 205) is configured to extract from bitstream LPSMs (and optionally also program limit metadata) that have the following format. Each of the metadata segments that include
LPSM (and optionally also program limit metadata) is included in a segment of leftover bits of
IMPI * ® *
INSTITUTO MEXICANO a bitstream frame, or an addbs field<sup>D</sup>í '^ Ss ^ sf of bitstream information (BSI ^ ... do.-wna bitstream frame, or in an auxiliary data field (for example, the AUX segment shown in Figure 4) when end of a bitstream frame A bitstream frame can include one or more metadata segments, each of which can include LPSM, and if the frame includes two metadata segments, one can be present in the field frame addbsi and the other in the frame's AUX field. In some modalities, each metadata segment that includes
LPSM includes an LPSM payload segment (or container) that has the following format:
a header (typically including a sync word that identifies the start of the payload of
LPSM followed by identification values, eg format version, length, period, count, and LPSM subflow association values indicated in Table 2 below); and after the header, at least one dialogue indication value (for example, Dialogue Channels parameter in Table 2) that indicates whether the corresponding audio data indicates dialogue or does not indicate dialogue (for example, which data channels of Audio
<img file="MX356196B_D0065.tif" />
IMPI
MEXICAN INSTITUTE,. ,,, OF THE r »OflEBAD corresponding indicate dialogue); industrial at least one SUnó-ETflacT regulation compliance value (for example, Loudness Regulation Type parameter in Table 2) indicating whether the corresponding audio data complies with an indicated set of loudness regulations;
at least one loudness processing value (for example, one or more of the Correction Indicator parameters of
Loudness with Dialogue Door Function, Type of
Loudness Correction from Table 2) indicating at least one type of loudness processing that has been performed on the corresponding audio data; and at least one loudness value (for example, one or more of the Loudness with Gate Function parameters relative to ITU,
Loudness with ITU Speech Gate Function, ITU Short-Term 3s Loudness (EBU 3341) and True Peak of the
Table 2) indicating at least one loudness characteristic (eg peak or average loudness) of the corresponding audio data.
In some embodiments, each LPSM-containing metadata segment and the program boundary metadata contain a nuclear header (and optionally additional nuclear elements as well) and after the nuclear header (or the
<img file="MX356196B_D0066.tif" />
IMPÍ "nuclear header and other nuclear elements" - an LPSM payload segment (or container) that has the following format:
a header, typically including at least one identification value (for example, format version, length, period, count, and LPSM subflow association values, as indicated in Table 2 below), and after the header , 'LPSMs and program limit metadata. Program limit metadata can include a program limit frame count and a code value (for example, an offset_exist value) that indicates whether the frame includes only 'a program limit frame count or both a count of program limit frame as an offset value), and (in some cases) an offset value.
In some implementations, parser 205 (and / or decoder stage 202) is configured to extract, from an excess bit segment or an addbsi field or ancillary data field from a bitstream frame, each metadata segment it has the following format:
a nuclear header (typically including a sync word that identifies the start of the metadata segment followed by at least one identification value,
IMPI
MEXICAN INSTITUTE OF PROPERTY, <sub>(z</sub> INDUSTRIAL (eg, version, length, and period, extended element count, and subflow association values of the nuclear element indicated in Table 1 below); and after the nuclear header, at least one protection value (for example, HMAC digest and audio fingerprint values from Table 1) useful for at least one of decryption, authentication, or validation of at least one metadata in the state of loudness processing or the corresponding audio data); and and also after the nuclear header, if the metadata segment includes LPSM, LPSM payload identification (ID), and LPSM payload size values that identify the following metadata as a payload of
LPSM and indicate the size of the LPSM payload.
The LPSM payload segment (or container) (preferably in the format specified above) follows the LPSM payload ID and the LPSM payload size values.
More generally, the encoded audio bitstream generated by preferred embodiments of the invention has a structure that provides a mechanism for labeling metadata elements and sub-elements as nuclear expanded (optional) or (optional) elements. This
<img file="MX356196B_D0067.tif" />
IMPI
MEXICAN INSTITUTE DB INDUSTRIAL PROPERTY
<img file="MX356196B_D0068.tif" />
It allows the data rate of the bitstream (including its metadata) to increase across various applications. The core (mandatory) elements of the preferred bitstream syntax should be able to point out that the expanded (optional) elements associated with the audio content are present (in-band) and / or at a remote location (out-of-band) ).
Nuclear elements are required to be present in each frame of the bit stream. Some sub-elements of nuclear elements are optional and can be present in any combination. Expanded elements are not required to be present in every frame (to limit bitrate overload). Therefore, expanded elements may be present in some frames and not others. Some subelements of an expanded element are optional and may be present in any combination, while some subelements of an expanded element may be required (that is, if the expanded element is present in a frame of the bitstream).
In one class of modalities, an encoded audio bitstream comprising a sequence of audio data segments and metadata segments is generated (eg, by an audio processing unit that translates the
IMPI
MEXICAN INSTITUTE OF PROPERTY invention). The audio data segments indicate * the audio data, each of at least some ^ two metadata segments includes loudness processing state (LPSM) metadata and optionally also program boundary metadata, and the Audio is transmitted simultaneously with a time division with the metadata segments. In preferred embodiments in this class, each of the metadata segments has a preferred format that will be described herein.
In a preferred format, the encoded bitstream is either an AC-3 bitstream or an E-AC-3 bitstream, and each of the metadata segments that LPSM includes is included (eg, by step 107 of a preferred implementation of encoder 100) as additional bitstream information in the addbsi field (shown in Figure 6) of the Bitstream Information (BSI) segment of a bitstream frame, or in an auxiliary data field of a bitstream frame or in a segment of leftover bits of a bitstream frame.
In the preferred format, each of the frames includes a nuclear element that has the format shown in Table 1 below, in the addbsi (or bit segment left over) field of the frame:
<img file="MX356196B_D0069.tif" />
IMPI
MEXICAN INSTITUTE OF LA RROHEDAD
INDUSTRIAL
Table 1
<td colspan="3">Parameter</td><td colspan="2">Description</td><td>Obliged River (M) / Option onal (0)</td>
<td colspan="3"></td><td colspan="2"></td><td></td>
<td>SINC [ID]</td><td></td><td></td><td>The word of</td><td>synchronization</td><td>M</td>
<td></td><td></td><td></td><td colspan="2">can be a 16 bit value</td><td></td>
<td></td><td></td><td></td><td>fixed at the value of</td><td>0x5838</td><td></td>
<td>Version</td><td></td><td>of</td><td></td><td></td><td>M</td>
<td colspan="2">nuclear element</td><td></td><td></td><td></td><td></td>
<td>Length</td><td></td><td>of</td><td></td><td></td><td>M</td>
<td colspan="2">nuclear element</td><td></td><td></td><td></td><td></td>
<td>Period</td><td></td><td>of</td><td></td><td></td><td>M</td>
<td>element</td><td colspan="2">nuclear</td><td></td><td></td><td></td>
<td>(xxx)</td><td></td><td></td><td></td><td></td><td></td>
<td>Count</td><td colspan="2">element</td><td colspan="2">Indicates the number of items</td><td>M</td>
<td>extended</td><td></td><td></td><td>metadata</td><td>extended</td><td></td>
<td></td><td></td><td></td><td>associated with</td><td>the element</td><td></td>
<td></td><td></td><td></td><td>nuclear. East</td><td>value can</td><td></td>
<td></td><td></td><td></td><td>increase decrease</td><td>as</td><td></td>
<td></td><td></td><td></td><td>bit stream</td><td>passes from the</td><td></td>
<td></td><td></td><td></td><td colspan="2">production until</td><td></td>
<td></td><td></td><td></td><td colspan="2">distribution and final emission.</td><td></td>
<img file="MX356196B_D0070.tif" />
<td>Association</td><td>of</td><td>Describe what subflows it is with</td><td>M</td>
<td>subflow</td><td></td><td>associated the nuclear element.</td><td></td>
<td>Signature HMAC)</td><td>(digest</td><td>256-bit HMAC digest (using SHA-2 algorithm) calculated on the data of audio, the nuclear element and all elements expanded of the entire plot.</td><td>M</td>
<td>Count</td><td>regressive</td><td>The field only appears for</td><td> 0</td>
<td>limit</td><td>from PGM</td><td>some frame numbers in the beginning or end of a flow bits / program file Audio. Therefore, a change version element core can be used to signal the inclusion of this parameter.</td><td></td>
<td colspan="2">Fingerprint of Audio</td><td>Audio fingerprint taken in some of the samples of PCM audio represented by the period field the nuclear element.</td><td> 0</td>
<img file="MX356196B_D0071.tif" />
Video fingerprint
<td>Paw print</td><td>digital de</td><td colspan="2">video taken</td>
<td colspan="2">in some of the</td><td>samples</td><td>of</td>
<td>video</td><td>tablets</td><td>(yes</td><td>the</td>
<td>would have</td><td colspan="2">represented by</td><td>the</td>
field of the period of the nuclear element.
URL / UUID
<td colspan="2">This field</td><td>this</td><td>defined for</td>
<td>wear</td><td>a</td><td>Url</td><td>and / or a UUID</td>
<td>(can</td><td>to be</td><td colspan="2">redundant for the</td>
<td>paw print</td><td colspan="2">digital)</td><td>which refers to</td>
an external location for additional program content (essence) and / or metadata associated with the bitstream.
In the preferred format, each of the addbsi (or auxiliary data) fields or leftover bit segments they contain
LPSMs contain a nuclear heading (and optionally additional nuclear elements as well) and after the nuclear heading (or the nuclear heading and other nuclear elements), the following LPSM values (parameters):
a payload ID (which identifies the metadata as LPSM) after the values -of nuclear elements (for example, as specified in Table 1);
<img file="MX356196B_D0072.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY a payload size (indicating the size of the LPSM payload) after the payload ID; and LPSM data (after the payload ID and payload size value) that is formatted as indicated in the following table (Table 2):
Table 2
<td>Parameter from LPSM [Sonority intelligent ]</td><td>Description</td><td>Number of state unique</td><td>Obligator io (M) / Option nal (0)</td><td>Cup of inserted n (Period of update action of the parameter or)</td>
<td>Version of LPSM</td><td></td><td></td><td>M</td><td></td>
<td>Period of LPSM (xxx)</td><td>Applicable to xxx fields only</td><td></td><td>M</td><td></td>
<td>Count LPSM</td><td></td><td></td><td>M</td><td></td>
<td>Association subflow from LPSM</td><td></td><td></td><td>M</td><td></td>
IMPI
MEXICAN INSTITUTE
FROM THE PROPERTY 'Q-Tr. . “'T ·? INDUSTRIAL
Channels of dialogue
Loudness regulation type
Indicates which combination of L, C, and R audio channels contain vocal broadcasts in the previous 0.5 seconds. When there are no speech broadcasts in any L, C or R combination, then this parameter should indicate without dialogue.
Indicates that the associated audio data stream complies with a specific set of regulations (for example, ATSC A / 85 or EBU
<td></td><td>R128)</td>
<td>Indicator</td><td>Indicates whether the flow of</td>
<td>of</td><td>associated audio has been</td>
<td>Correction</td><td>corrected based on</td>
<td>of</td><td>gate function</td>
<td>Sonority</td><td>dialogue</td>
<td>with Function</td><td></td>
<td>Door</td><td> •</td>
<td>Dialogue</td><td></td>
<img file="MX356196B_D0073.tif" />
Plot
O (only present if
Loudness_
Regulatio n_Type indicates that the corresponding audio is UNCORRECTED)
Plot
MEXICAN INSTITUTE '¿ΖΖ + ϊϊ 1'
OF INDUSTRIAL PROPERTY
<td rowspan="2">Kind of correction of sonority</td><td rowspan="2">Indicates whether the flow of associated audio has been corrected with a controller 'range dynamic and loud in real time (RT) or infinite prospective (file based).</td><td> 2</td><td>0 (only</td><td>Plot</td>
<td></td><td>this Present yes Loudness_ Regulatio n_Type indicates that he Audio correspond tooth is WITHOUT TO CORRECT)</td><td></td>
<td>Sonority with function door relative ITU (INF)</td><td>Indicates loudness Integrated ITU-R BS.1770- 3 of the audio streams partners without metadata applied (eg 7 bits: -58</td><td> 128</td><td>OR</td><td>1 s</td>
<td></td><td>-> +5.5 LKFS 0.5 LKFS stages)</td><td></td><td></td><td></td>
<td>Sonority with function door with emissions ITU vowels (INF)</td><td>Indicates loudness Integrated ITU-R BS.1770- 1/3 of the emissions vowels / dialysis of audio streams partners without metadata applied (eg 7 bits: -58</td><td> 128</td><td> 0</td><td>1 s</td>
<td></td><td>-> +5.5 'LKFS 0.5 LKFS stages)</td><td></td><td></td><td></td>
<img file="MX356196B_D0074.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<td>Sonority</td><td></td><td>Indicates the loudness of</td><td rowspan="2"> 256</td><td rowspan="2"> 0 _</td><td>Π, 1 s -</td>
<td colspan="2">from 3s to Short term ITU (EBU 3341)</td><td>ITU without function 3 second door (ITU-BS.1771-1) from associated audio stream no metadata applied (sliding window) a ~ 10Hz rate insertion (eg 8 bits: 116 -> +11.5 LKFS 0.5 LKFS stages)</td><td></td>
<td>Value peak true</td><td>of</td><td>Indicates the Peak value True Annex 2 of ITU-R BS.1770-3 (dB TP) audio stream associate without metadata applied. (That is to say, largest value in the frame period marked in the field of item period)</td><td> 256</td><td> 0</td><td>0.5 s</td>
<td></td><td></td><td>116 -> +11.5 LKFS 0.5 LKFS stages</td><td></td><td></td><td></td>
<td>Offset mixture</td><td>of</td><td>Indicates the offset of mixing loudness</td><td></td><td></td><td></td>
Program limit
IMPI sross® _INDUSTRIA '.
Indicates, in frames, when a 'program limit will occur or has occurred. When the program limit is not at the frame limit, the optional sample offset will indicate how far in the frame the actual program limit occurs
In another preferred format of an encoded bitstream generated in accordance with the invention, the bitstream is either an AC-3 bitstream or an E-AC-3 bitstream, and each of the metadata segments is included. including LPSM (and optionally also program limit metadata) (eg, by step 107 of a preferred implementation of encoder 100) in any of: a bit segment left over from a bitstream frame; or an addbsi field (shown in Figure 6) from the segment of
Bitstream Information (BSI) of a bitstream frame; or an auxiliary data field (for example, the AUX segment shown in Figure 4) at the end of a bitstream frame. A frame can include one or more metadata segments, each of which includes
<img file="MX356196B_D0075.tif" />
LPSM, and if the frame includes two Β ^ τηοηί-ης metadatoo -; - u-t can be present in the addbsi field of the frame and the other in the AUX field of the frame. Each metadata segment that includes LPSM has the format specified above with reference to Tables 1 and 2 below (i.e. includes the core elements specified in Table 1, followed by the Payload ID (which identifies the metadata as LPSM ) and the payload size values specified above, followed by payload (LPSM data in the format shown in Table 2).
In another preferred format, the encoded bitstream is a Dolby E bitstream, and each of the metadata segments that LPSM includes (and optionally also program boundary metadata) is in the first N sample locations in the range Dolby E's Guardian Band A Dolby E bitstream that includes such a metadata segment that includes LPSM preferably includes a value indicating the LPSM payload length noted in the Pd word of the SMPTE 337M preamble (the word repetition rate
SMPTE 337M Pa preferably remains identical to the associated video frame rate).
In an additional format, where the encoded bitstream is an E-AC-3 bitstream, each of the segments is included
1ΜΡΙ6 ^ |
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL ¿aTJgg ·· metadata that includes LPSM (and optionally also program limit metadata) (for example, by step 107 of a preferred implementation of encoder 100) as additional bitstream information in a segment of leftover bits or in the addbsi field of the Bit Stream Information (BSI) segment of a bit stream frame. Additional aspects of encoding an E-AC-3 bit stream with LPSM in this preferred format are described below:
one. During the generation of an E-AC-3 bit stream, while the E-AC-3 encoder (which inserts the LPSM values into the 'bit stream) is active, for each frame (synchronized frame) generated, the Bitstream must include a metadata block (which includes LPSM) carried in the addbsi (or segment of excess bits) field of the frame. The bits required to call the metadata block must not increase the encoder bit rate (frame length);
2. All metadata blocks (containing LPSM) must contain the following information:
loudness_correction_type'_f lag: where '1' indicates that the loudness of the corresponding audio data was corrected before the encoder, and '0' indicates that the loudness was corrected with a loudness corrector embedded in the
<img file="MX356196B_D0076.tif" />
encoder (for example, loudness processor 103 del · ι ii | ii ijh.<sub>¿</sub> bu ,,,,,, encoder 100 of Figure 2);
speech_channel - Indicates which source channels contain speech broadcasts (within the previous 0.5s). If no vocal broadcasts are detected, this should be indicated;
speech_loudness: indicates the loudness of the integrated vocal emissions of each corresponding audio channel that contains vocal emissions (in the previous 0.5 s);
ITU_loudness: indicates the built-in ITU BS.1770-3 loudness of each corresponding audio channel; and gain: the loudness compound gains for investing in a decoder (to demonstrate investment capacity);
3. While the E-AC-3 encoder (which inserts the LPSM values into the bitstream) is active and receives an AC-3 frame with a 'trust' flag, the loudness handler in the encoder should be avoided (eg loudness processor 103 of encoder 100 of Figure 2). The 'trusted' source dialnorm and DRC values should be passed through (eg, encoder 100 generator 106) to encoder component E-AC-3 (eg, encoder step 100 107). LPSM block generation continues and
<img file="MX356196B_D0077.tif" />
IMPI
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL loudness_correction_type_flag is set to '1'. The loudness controller offset sequence should be synchronized at the start of the decoded AC-3 frame where the 'trust' flag appears. The loudness controller deviation sequence should be implemented as follows: the leveler_amount control is decreased from a value of 9 to a value of 0 over 10 audio block periods (i.e. 53.3 ms) and the leveler_back_end_meter control is placed in offset mode (this operation should result in a seamless transition). The term reliable leveler deviation implies that the dialnorm value of the source bitstream is also reused in the encoder output, (for example, if the trusted source bitstream has a dialnorm value of -30 then the output encoder must use -30 for the output dialnorm value);
Four. While the E-AC-3 encoder (which inserts the LPSM values into the bitstream) is active and receives an AC-3 frame without the 'trust' flag, the loudness driver inserted into the encoder (by For example, loudness processor 103 of encoder 100 of Figure 2) should be active. LPSM block generation continues and loudness_correction_type_flag is set to
IΛ4 ΡI
INSTITUTE ΜΕΧιγαμγ.
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY '0'. The loudness driver activation sequence should be synchronized at the start of the AC-3 frame decode where the 'trust' flag disappears. The loudness driver activation sequence must be implemented as
<td>The next</td><td>way:</td><td>the</td><td>control of</td><td colspan="2">leveler amount se</td>
<td>increases from</td><td>a value</td><td>of</td><td>0 to a value</td><td>from 9 ι</td><td>over 1 period</td>
<td>block</td><td>Audio.</td><td>(is</td><td>say 5.3</td><td>ms)</td><td>and control</td>
<td colspan="2">leveler back end meter</td><td>I know</td><td>place in the</td><td>mode</td><td>'active' (this</td>
<td>operation should</td><td>result</td><td>in</td><td colspan="2">a transition without</td><td>interruptions</td>
<td>and include</td><td>a</td><td colspan="2">reinitialization</td><td>of</td><td>integration</td>
back_end_meter); and
5. During encoding, a graphical user interface (GUI) must prompt a user for the following parameters: Input audio program:
[Reliable / Untrustworthy] - the status of this parameter is based on the presence of the confidence indicator within the input signal; and Time Loudness Correction
Actual: [Enabled / Disabled] - the status of this parameter is based on whether this loudness driver embedded in the encoder is active.
When decoding an AC-3 or E-AC-3 bit stream that has LPSM (in the preferred format) included in a segment of leftover bits, or the addbsi field of the segment
<img file="MX356196B_D0078.tif" />
Bitstream Information (BSI), of each of the —-Iij II I1HI .ai'TiT — ΙΊΠΠΓΤΠ.
IMPI frames of the bitstream, the decoder must parse the LPSM block data (in the leftover bit segment or addbsi field) and pass all the extracted LPSM values · to a graphical user interface (GUI). The set of extracted LPSM values are updated in each frame.
In another preferred format of an encoded bitstream generated in accordance with the invention, the encoded bitstream is either an AC-3 bitstream or an E-AC3 bitstream, and each of the metadata segments is included which include LPSM (eg, by step 107 of a preferred implementation of encoder 100) in a segment of leftover bits, or in an Aux segment, or as additional bitstream information in the addbsi field (shown in Figure 6) of the Bitstream Information (BSI) segment of a bitstream frame. In this format (which is a variation on the format described above with reference to Tables 1 and 2), each of the addbsi (or Aux or leftover bit) fields that LPSM contains contains the following LPSM values:
the nuclear elements specified in Table 1, followed by the payload ID (which identifies the metadata as
ΙΜΡΙ
<img file="MX356196B_D0079.tif" />
LPSM) and payload size values, followed by payload (LPSM data) that has the following format (similar to the mandatory elements indicated in Table 2 above):
LPSM payload version: a 2-bit field indicating the version of the LPSM payload;
dialchan: a 3-bit field that indicates whether the channels
Left, Right and / or Center of the corresponding audio data contains spoken dialogs. The location of the bits in the dialchan field can be as follows: bit 0, which indicates the presence of dialogue in the left channel, is stored in the most significant bit of the dialchan field; and bit 2, which indicates the presence of dialogue on the center channel, is stored in the least significant bit of the dialchan field.
Each bit of the dialchan field is set to '1' if the corresponding channel contains spoken dialogue during the previous 0.5 seconds of the program;
loudregtyp: a 4-bit field that indicates which loudness regulation standard the loudness of the program meets.
Setting the loudregtyp field to '000' indicates that LPSMs do not indicate loudness regulation compliance. For example, a value in this field (for example, 0000) can
100
IMPI
<img file="MX356196B_D0080.tif" />
indicate that compliance with a norm ^ J ^ j ^ gg ^^ g ^^ n ^ of ,, loudness is not indicated, another value in this field (for example, 0001) may indicate that the program's audio data complies with the ATSC A / 85 standard, and another value in this field (for example, 0010) may indicate that the program's audio data complies with the EBU R128 standard. In the example, if the field is set to any value other than '0000', the loudcorrdialgat and loudcorrtyp fields should still be in the payload;
loudcorrdialgat: a one-bit field that indicates whether loudness correction with a dialog gate function has been applied. If the loudness of the program has been corrected using the dialog gate function, the value of the loudcorrdialgat field is set to '1'. Otherwise, it is set to '0';
loudcorrtyp: a one-bit field indicating the type of loudness correction applied to the program. If the loudness of the program has been corrected with an infinite (file-based) prospective loudness correction process, the loudcorrtyp value is set to '0'. If the loudness of the program has been corrected using a combination of real-time loudness measurement and dynamic range control, the value of this field is set to '1';
101
<img file="MX356196B_D0081.tif" />
IMPI loudrelgate: a one-bit field that indicates whether the relative gate function loudness data (ITU) exists. If the loudrelgate field is set to '1', a 7-bit ituloudrelgat field must remain in the payload;
loudrelgat: a 7-bit field indicating the loudness of the program with relative gate function (ITU). This field indicates the built-in loudness of the audio program, measured in accordance with ITU-R BS.1770-3 without any gain adjustment because dynamic range and dialnorm compression is applied. Values from 0 to 127 are interpreted as -58 LKFS to +5.5 LKFS, in 0.5 LKFS stages;
loudspchgate - A one-bit field that indicates whether loudness data with Voice Emission Gate (ITU) function exists. If the loudspchgate field is set to '1', a 7-bit loudspchgat field must remain in the payload;
loudspchgat: a 7-bit field indicating the loudness of the program with the Vocal Gate function. This field indicates the built-in loudness of the corresponding integer audio program, measured in accordance with formula (2) of ITU-R BS.1770-3 and without any gain adjustment because dynamic range compression and dialnorm are applied. Values from 0 to 127 are interpreted as -58 to +5.5 LKFS, in
0.5 LKFS stages;
102
<img file="MX356196B_D0082.tif" />
IMPI loudstrm3se: a one-bit field that indicates whether loudness data exists in the short term (3 seconds). If the field is set to '1', a 7-bit loudstrm3s field must remain in the payload;
loudstrm3s: a 7-bit field indicating the loudness without gate function of the previous 3 seconds of the corresponding audio program, measured according to ITU-R
BS.1771-1 and without any gain adjustment because dynamic range compression and dialnorm are applied. Values from 0 to 256 are interpreted as -116 LKFS to +11.5
LKFS, in 0.5 LKFS stages;
truepke: a one-bit field that indicates whether true peak loudness data exists. If the truepke field is set to '1', an · 8 bit truepk field must remain in the payload;
y truepk: an 8-bit field that indicates the program's true peak sample value, measured in accordance with Annex 2 of ITU-R BS.1770-3 and without any gain adjustment because compression of dynamic range and dialnorm. Values from 0 to 256 are interpreted as -116 LKFS to +11.5
LKFS, in 0.5 LKFS stages;
In some embodiments, 'the nuclear element of a metadata segment in a segment of leftover bits or in a field of
<img file="MX356196B_D0083.tif" />
103
1viIP auxiliary data (or addbsi field) of a frame of an AC-3 bit stream or E-AC-3 bit stream comprises a nuclear header (typically including identification values, eg, nuclear element version) and after the nuclear heading: values that indicate whether fingerprint data (or other protection data) is included for metadata in the metadata segment, values that indicate whether there is external data (related to audio data that corresponds to metadata in the metadata segment) , the payload ID and payload size values for each type of metadata (eg LPSM, and / or metadata of a type other than LPSM) identified by the nuclear element, and protection values for at least one type of metadata identified by the nuclear element. The metadata payloads of the metadata segments follow the nuclear heading, and (in some cases) nest within the values of the nuclear element.
Typical modalities of. The invention includes program boundary metadata in an efficiently encoded audio bitstream, which enables the robust and accurate determination of at least one boundary between consecutive audio programs indicated by the bitstream. The modalities
- tW? ' * · \ ^ ·> 'Al ·
MEXICAN INSTITUTE
IX THE PROPERTY ΐ
INDUSTRIAL
104 Typical allow solid determination and. _exact ^ of a program limit in the sense that they allow the determination of the program limit exactly, even in cases where the bit streams indicating different programs are combined (to generate the bit stream of the invention) in a way that one or both of the spliced bit streams is truncated (and thus discard the program boundary metadata that has been included in at least one of the pre-splice bit streams).
In typical embodiments, the program limit metadata in a frame of the bitstream of the invention is a program limit indicator indicating the frame count. Typically, the indicator indicates the number of frames between the current frame (the frame that includes the indicator) and a program limit (the start or end of the current audio program). In some preferred embodiments, program limit flags are efficiently and symmetrically inserted at the beginning and end of each bitstream segment that indicates a single program (i.e., in frames that occur within some number of predetermined frames after the start of the segment and on frames that occur within some predetermined number of frames before the end of the segment) so that
105
IMPI
<img file="MX356196B_D0084.tif" />
when two such bitstream segments are concatenated (to indicate a sequence of two programs), the program boundary metadata may be present (eg symmetrically) or both sides of the boundary between two programs).
Maximum robustness can be achieved by inserting a program limit flag into each frame of a bit stream indicating a program, but this would not be practical typically due to the associated increase in data rate.
In typical embodiments, program limit flags are inserted only into a subset of the frames of an encoded audio bitstream (which may indicate an audio program or sequence of audio programs) and the rate of flag insertion Boundary is a non-increasing function of increasing separation of each of the bitstream frames (where a flag is inserted) from the program limit that is closest to each of those frames, where the limit indicator insertion rate denotes the average ratio of the number of frames (indicating a program) that includes a program limit indicator to the number of frames (indicating the program) that does not include a limit indicator program where the average is a moving average of a
106
IMPI
<img file="MX356196B_D0085.tif" />
amount (eg, relatively small amount) of consecutive frames in the encoded audio bitstream.
Increasing the rate of insertion of limit flags (for example, at locations in the bitstream near a program limit) increases the rate of data required to supply the bitstream. To compensate for this, the size (number of bits) 'of each inserted indicator preferably decreases as the insertion rate of the limit indicator increases (for example, so that the size of the program limit indicator in frame N of the bitstream, where N is an integer, is a function that does not increase the distance (number of frames) between frame N and the nearest program boundary). In a class of modalities, the limit indicator insertion rate is a function that logarithmically decreases in increasing distance (from each indicator insertion location) from the nearest program limit, and for each frame containing indicators that includes one of the indicators, the size of the indicator in said indicator-containing frame is equal to or greater than the size of each indicator in a frame located closer to the nearest program boundary than said indicator-containing frame. Typically, the size of each indicator is determined by an increasing function of the
107
<img file="MX356196B_D0086.tif" />
ΙΜΡΙ number of frames from the location of the insertion of the indicator to the nearest program limit.
For example, the modalities of Figures 8 and 9 are considered, where each column identified by a frame number (in the top row) indicates a frame of an encoded audio bitstream. The bitstream indicates an audio program that has a first program limit (indicating the start of the program) that occurs immediately to the left of the column identified by frame number 17 on the left side of Figure 9, and a second program limit (indicating the end of the program) that occurs immediately to the right of the column identified by frame number 1 on the right side of Figure 8. The program limit indicators included in the frames shown in Figure 8 count backwards the number of frames between the current frame and the second program limit.
The program limit indicators included in the frames shown in Figure 9 count forward the number of frames between the current frame and the first program limit.
In the embodiment of Figures 8 and 9, a program limit indicator is inserted only in each of frames 2<sup>n</sup> of the first X frames of the encoded bitstream after the start of the audio program indicated by the
108
-i -i ϊ Ji .. ij
MEXICAN INSTITUTE OF PROPERTY <sub>n</sub> . INDUSTRIAL _ bitstream, and in each of frames 2<sup>N</sup> (of the last X-frames in the bitstream) nearest''toΡ'Ρη'ΈθΤ program indicated by the bitstream, where the program comprises Y-frames, X is an integer less than or equal to Y / 2, and N is a positive integer on an interval from 1 to log<sub>2</sub>(X). Therefore, (as indicated in Figures 8 and 9), a program limit flag is inserted into the second frame (N = 1) of the bit stream (the frame containing flag closest to the start of the program ), in the fourth frame (N = 2), in the eighth frame (N = 3), and so on, and in the eighth frame since the end of the bitstream, in the fourth frame from the end of the bit stream and in the second frame from the end of the bit stream (the frame containing pointer closest to the end of the program). In this example, the program limit indicator in frame 2<sup>K</sup> from the start (or end) of the program it comprises log binary bits<sub>2</sub>(2<sup>N + 2</sup>), as indicated in Figures 8 and 9. Therefore, the program limit indicator in the second frame (N = 1) from the start (or end) of the program comprises log<sub>2</sub>(2<sup>N + 2</sup>) = log2 (2<sup>3</sup>) = 3 binary bits, and the flag in the fourth frame (N = 2) since the start (or end) of the program
<img file="MX356196B_D0087.tif" />
understand log<sub>2</sub>(2<sup>N + 2</sup>) = log<sub>2</sub>(2<sup>4</sup>) = 4 binary bits, and so on.
109
IMPI
MEXICAN INSTITUTE OF PROPERTY the format'o ^ cfe
<img file="MX356196B_D0088.tif" />
Ca ^ a ^ indicator '' of front, one or one or more bits
In the example of Figures 8 and 9, the program flag is the following program limit consisting of a sequence bit of bits 0 (either any consecutive bits) after the leading bit, and a two-bit rear code. The back code is 11 for flags in the last X frames of the bit stream (the frames closest to the end of the program), as indicated in Figure 8. The back code is 10 for the flags in the first X frames of the bitstream (the frames closest to the start of the program), as indicated in Figure 9. Therefore, to read (decode) each flag, counts the number of zeros between the leading bit 1 and the trailing code. If the rear code is identified as 11, the indicator indicates that there is (2<sup>Z + 1</sup> - 1) frames between the current frame (the frame that includes the indicator) and the end of the program, where Z is the number of zeros between the leading bit 1 and the back code of the indicator. The decoder can be effectively implemented to ignore the first and last bits of each such flag, to determine the inverse of the sequence of the other (intermediate) bits of the flag (for example, if the intermediate bit sequence is 0001 where bit 1 is the last bit
110
<img file="MX356196B_D0089.tif" />
IMPI in the sequence, the inverted intermediate bit sequence is 1000 where bit 1 is the first inverted sequence bíV'Vñ ia), and to identify the binary value of the inverted intermediate bit sequence as the index of the current frame (the plot in which the indicator is included) with respect to the end of the program. For example, if inverted sequence of intermediate bits is 1000, this inverted sequence has the binary value 2<sup>4</sup> = 16, and the plot is identified as 16.<sup>to</sup> frame before the end of the program (as indicated in the column in Figure 8 that describes frame 0).
If the rear code is identified as 10, the indicator indicates that there is (2<sup>Z + 1</sup> - 1) frames between the start of the program and the current frame (the frame that includes the indicator), where Z is the number of zeros between the leading bit 1 and the back code of the indicator. The decoder can be efficiently implemented to ignore the first and last bit of each indicator, to determine the inverse of the sequence of the intermediate bits of the indicator (for example, if the intermediate bit sequence is 0001 where the bit is the last bit in the sequence, the inverted intermediate bit sequence is 1000 where bit 1 is the first bit in the inverted sequence), and to identify the value
111
<img file="MX356196B_D0090.tif" />
binary of the inverted sequence of intermediate bits as the index of the current frame (the frame in which the flag is included) with respect to the start of the program. For example, if inverted sequence of intermediate bits is
1000, this inverted sequence has the binary value 2<sup>4</sup> =
16, and the plot is identified as 16.<sup>to</sup> plot after program start (as indicated in the column of the
Figure 9 describing frame 32).
In the example of Figures 8 and 9, a program limit indicator is present only in each of frames 2<sup>N</sup> of the first X frames of an encoded bitstream after the start of an audio program indicated by the stream of
<td>bits, and in each</td><td>a</td><td>of the</td><td>frames 2<sup>N</sup> (of</td><td colspan="2">the latest frames</td>
<td>X in the flow</td><td>of</td><td>bits)</td><td>closest to</td><td>end of</td><td>Program</td>
<td>indicated by the</td><td colspan="2">flow of</td><td>bits where the</td><td>Program</td><td>understands</td>
<td>frames Y, X is</td><td>a</td><td>whole</td><td>less than or equal</td><td>to Y / 2, and</td><td>N is a</td>
positive integer on an interval from 1 to log2 (X). The inclusion of program limit flags adds only an average bit rate of 1,875 bits / frame to the bit rate required to transmit the bitstream without flags.
In a typical implementation of the embodiment of Figures 8 and 9, where the bit stream is an audio bit stream
112
χ. Ρ ί
ΝΧΤΤΤΙΙ'ΐΤ) MtXlCAN '/ UbArtCV · A' ΙΜ LA Ρ «η» Κ · ί) ΛΙ, encoded AC-3, each frame contains content <sup>M</sup>cé "<sup>IA</sup>áucfirry metadata for 1536 digital audio samples'. '' For 'nail ~ ”48 kHz sampling rate, this represents 32 milliseconds of digital audio or a rate of 31.25 frames per second of audio. Therefore, in such mode, a program limit indicator in a frame separated by some number of frames (X frames) from a program limit indicates that the limit occurs 32X milliseconds after the end of the frame containing the indicator ( or 32X milliseconds before the start of the frame containing the flag).
In a typical implementation of the embodiment of Figures 8 and 9, where the bit stream is an audio bit stream
<td colspan="2">E-coded</td><td>AC-3,</td><td>each frame of the stream</td><td>bit</td><td>contains</td>
<td>content</td><td>of</td><td>Audio</td><td>and metadata for 256,</td><td colspan="2">, 512, 768 or 1536</td>
<td>samples</td><td>of</td><td>Audio</td><td>digital, depending</td><td>of if</td><td>the plot</td>
<td>contains</td><td>one,</td><td>two,</td><td>three or six blocks</td><td>of data</td><td>audio</td>
respectively. For a 48 kHz sampling rate, this represents 5,333, 10,667, 16, or 32 milliseconds of digital audio, respectively, or a rate of 189.9, 93.75, 62.5, or 31.25 frames per second of audio, respectively. Therefore, in such mode, (assuming that each frame indicates 32 milliseconds of digital audio), a program limit indicator in a frame separated by some number of frames
113
<img file="MX356196B_D0091.tif" />
(X frames) of a program limit indicates that the limit occurs 32X milliseconds after the end of the frame containing the flag (or 32X milliseconds before the start of
IMPI the frame containing the flag).
In some modes where a program boundary may occur within a frame of an audio bitstream (that is, not aligned with the start or end of a frame), the program boundary metadata included in one frame of the stream Bit rate includes a count of program limit frames (that is, metadata indicating the number of integer frames between the start or end of the frame containing the frame count and a program limit) and an offset value.
The offset value indicates an offset (typically a number of samples) between the start or end of a frame containing the program limit, and the actual location of the program limit within the frame containing the program limit. An encoded audio bitstream may indicate a sequence of programs (audio tracks) from a corresponding sequence of video programs, and the boundaries of such audio programs tend to occur at the edges of the frames rather than at the edges of the audio frames. Also, some audio encodings (eg E-AC-3 encodings) use audio frame sizes that are not
114
<img file="MX356196B_D0092.tif" />
Plνι Pl
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY are aligned with the frames. Also, in some cases, an initially encoded audio bitstream undergoes transcoding to generate a transcoded bitstream, and the initially encoded bitstream has a different frame size than the transcoded bitstream so that it is not ensures that a program limit (determined by the initially encoded bitstream) occurs at a frame limit of the transcoded bitstream. For example, if the initially encoded bitstream (eg, IEB bitstream in Figure 10) has a frame size of 1536 samples per frame, and the transcoded bitstream (eg, TB bitstream of the Figure 10) has a frame size of 1024 samples per frame, the transcoding process may cause the actual program limit to occur not at a frame limit of the transcoded bit stream but somewhere in a frame of this (for example , 512 samples in a frame of the transcoded bitstream, as indicated in Figure 10), due to differentiated frame sizes of the different codees.
The embodiments of the present invention in which the program boundary metadata included in a frame of an encoded audio bitstream include an offset value
115
<img file="MX356196B_D0093.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY as well as a program limit plot count are useful in the three cases highlighted in this paragraph (as well as in other cases).
The modality described above with reference to
Figures 8 and 9 do not include an offset value (eg, an offset field) in any of the frames of the encoded bitstream. In variations of this mode, an offset value is included in each frame of an encoded audio bitstream that includes a program limit indicator (for example, in frames that correspond to frames numbered 0, 8, 12, and 14 in Figure 8, and frames numbered 18, 20, 24 and 32 in Figure 9).
In one class of modalities, a data structure (in each frame of an encoded. Bitstream containing the program boundary metadata of the invention) includes a code value indicating whether the frame includes only a frame count of program limit, or both, a program limit frame count and an offset value. For example, a code value may be the value of a single-bit field (which will be called the offset_exist field here), the offset_exist value = 0 may indicate that offset values are not included in the frame, and the offset value exists = 1 can indicate that both a frame count
116
IMPI
<img file="MX356196B_D0094.tif" />
Program limit values as an offset value are included in the frame.
In some embodiments, at least one frame of an AC-3 or E-AC-3 encoded audio bitstream includes a metadata segment that includes the LPSM and program boundary metadata (and optionally other metadata) for a program audio determined by the bit stream. Each of said metadata segments (which may be included in an addbsi field, or an auxiliary data field, or a segment of bits left over from the bit stream) contains a nuclear header, (and optionally also additional nuclear elements), and after the nuclear header (or the nuclear header and other nuclear elements) an LPSM payload segment (or container) that has the following format:
a header (typically including at least one identification value, for example, format version, length, period, count, and LPSM subflow association values), and
after the header, the program limit metadata (which can include a count of program limit frames, a code value (for example, an offset_exist value) that indicates whether the frame includes only one count
117
<img file="MX356196B_D0095.tif" />
IMPI
INSTITUTO MEXICANO ΠΕ the INDUSTRIAL property of the program limit frame or both a program limit frame count and an offset value (and in some cases an offset value) and the LPSMs. LPSMs can include: at least one dialogue indication value indicating whether the corresponding audio data indicates dialogue or does not indicate dialogue (for example, which channels of corresponding audio data indicate dialogue). Dialog indication values can indicate whether dialogue is present on any combination or all channels of the corresponding audio data;
at least one loudness regulation compliance value indicating whether the corresponding audio data complies with an indicated set of loudness regulations;
at least one loudness processing value indicating a type of loudness processing that has been performed on the corresponding audio data; and at least one loudness value indicating at least one loudness characteristic (eg, peak or average loudness) of the corresponding audio data.
In some embodiments, the LPSM payload segment includes a code value (an offset_exist value) that indicates whether the frame includes only a program limit frame count or both, a limit limit frame count of
118
MEXICAN INSTITUTE
INDUSTRIAL program and an offset value. For example, in one of those modes, when that code value indicates (for example, when offset_exist = 1) that the frame includes a program limit frame count and an offset value, the LPSM payload segment may include an offset value that is an 11-bit unsigned integer (that is, having a value from 0 to 2048) and indicating the number of additional audio samples between the indicated frame limit (the frame limit that includes the program limit) and the actual program limit '. If the program limit frame count indicates the number of frames (with the current frame rate) up to the frame containing program limits, the exact location (in units of sample numbers) of the program limit (relative to at the beginning or end of the frame that includes the payload segment of
LPSM) would be calculated as:
S = (frame_counter * frame size) + offset, where S is the number of 'samples up to the program limit (from the beginning or end of the frame that includes the LPSM payload segment), frame_counter is the count of frames indicated by the program limit frame count, frame size is the number of samples per frame, and offset is the number of samples indicated by the value
119
IMPI
MEXICAN INSTITUTE OE LA EROPIEDAf) INDUSTRIAL
<img file="MX356196B_D0096.tif" />
offset.
Some modes where the rate of inserting program limit flags increases near the actual program limit implements a rule that states that an offset value is never included in a frame if the frame is less than or equal to some number (AND ) of frames from the frame that includes the program limit. Typically Y = 32. For an E-AC-3 encoder that implements this rule (with Y = 32), the encoder never inserts an offset value at the end second of an audio program. In this case, the receiving device is responsible for maintaining a timer and therefore performing its own offset calculation (in response to program limit metadata, which includes an offset value, in a frame of the encoded bitstream which is more than frames AND from the frame containing program limit). For programs whose audio programs are known to be frame-aligned with corresponding video program frames (for example, typical contribution feed with Dolby E-coded audio), it would be superfluous to include offset values in the coded bitstreams indicating programs audio. Therefore, offset values will not typically be included in such encoded bit streams.
120
IMPIfB *
MEXICAN INSTITUTE OF!> PROPERTY
INDUSTRIAL ^ TuJ ± L
Referring to Figure 11, we consider below cases where encoded audio bitstreams are spliced together to generate one embodiment of the audio bitstream of the invention.
The bit stream at the top of Figure 11 (labeled Scenario 1) indicates an entire first audio program (Pl) that includes program limit metadata (program limit flags, F) followed by a second audio program (P2) An integer that also includes program limit metadata (program limit flags, F).
The program limit indicators at the end of the first program (some of which are shown in the
Figure 11) are identical or similar to those described with reference to Figure 8, and determine the location of the boundary between the two programs (i.e., the boundary at the start of the second program). The program limit indicators in the initial part of the second program (some of which are shown in Figure 11) are identical or similar to those described with reference to Figure 9, and also determine the location of the limit. In typical modes, an encoder or decoder implements a timer (calibrated by the indicators in the first program) that counts down to the limit of the
121
<img file="MX356196B_D0097.tif" />
program, and the same timer (c indicators in the second program) cu
<img file="MX356196B_D0098.tif" />
from the very limit of the program. As indicated by the limit timer graph in Scenario 1 of the
Figure 11, said timer countdown (calibrated by the indicators in the first program) reaches zero at the limit, and the forward countdown of the timer (calibrated by the indicators in the second program) indicates the same location of the limit.
The second bitstream at the top of Figure 11 (labeled Scenario 2) indicates an entire first audio program (Pl) that includes program limit metadata (program limit flags, F) followed by a second program of audio (P2) integer that does not include program limit metadata. The program limit indicators at the end of the first program (some of which are shown in Figure 11) are identical or similar to those described with reference to Figure 8, and determine the location of the limit between the two programs ( that is, the limit at the start of the second program), the same as in the
Scenario 1. In typical modes, an encoder or decoder implements a timer (calibrated by the flags in the first program) that counts down
122
IMPI
<img file="MX356196B_D0099.tif" />
up to the program limit, and the same timer (without additional calibration) continues to count forward from the program limit '(as indicated by the limit timer graph in Scenario 2 of Figure 11).
The third bitstream at the top of Figure 11 (labeled Scenario 3) indicates a truncated first audio program (Pl) that includes program limit metadata (program limit flags, F) and has been spliced with a second whole audio program (P2) that also includes program limit metadata (program limit indicators, F). The splice has removed the last N frames from the first program. The program limit indicators in the initial part of the second program (some of which are shown in Figure 11) are identical or similar to those described with reference to Figure 9, and determine the location of the limit (splice) between the first truncated program and the second whole program. In typical modes, an encoder or decoder implements a timer (calibrated by the flags in the first program) that counts down to the end of the first untruncated program, and the same timer (calibrated by the flags in the second program) counts toward forward from the start of the second program. The
123
IMPI
<img file="MX356196B_D0100.tif" />
The start of the second program is the program limit in Scenario 3. As indicated by the limit timer graph in Scenario 3 in Figure 11, said timer countdown (calibrated by the indicators in the first program) is reset. (in response to the program limit metadata in the second program) before it reaches zero (in responses to the program limit metadata in the first program). Therefore, although truncation of the first program (via splicing) prevents the timer from identifying the program boundary between the first truncated program and the start of the second program in response to (i.e. calibrated by) program boundary metadata only in the first program, the program metadata in the second program reset the timer, so that the reset timer correctly indicates (as the location that corresponds to the reset timer zero count) the location of the program boundary between the first program and the start of the second program.
The fourth bitstream · (labeled Scenario 4) indicates a first truncated audio program (Pl) that includes program limit metadata (program limit flags, F), and a second truncated audio program (P2) that includes
124
IMPI
<img file="MX356196B_D0101.tif" />
program limit metadata (program limit indicators, F) and that it has been spliced with a part (the untruncated part) of the first audio program. The program limit indicators in the initial part of the entire second program (pre-truncation) (some of which are shown in Figure 11) are identical or similar to those described with reference to Figure 9, and the program limit indicators at the end of the first whole program (pre-truncation) (some of which are shown in the
Figure 11) are identical or similar to those described with reference to Figure 8. The splice has removed the last N frames of the first program (and therefore some of the program limit flags that had been included in these prior to splicing ) and the first M frames of the second program (and therefore some of the program limit flags that had been included in these before splicing). In typical modes, an encoder or decoder implements a timer (calibrated by the flags in the first truncated program) that counts back to the end of the first untruncated program, and the same timer (calibrated by the flags in the second truncated program) counts forward from the start of the second non-truncated program. As indicated
125 the limit timer graph in Scenario 4 of the
Figure 11, said timer countdown (calibrated by the program limit metadata in the first program) is reset (in response to the program limit metadata in the second program) before it reaches zero (in responses to the program limit metadata in the first program). Truncation of the first program (via splicing) prevents the timer from identifying the program limit between the first truncated program and the start of the second program in response to (i.e. calibrated by) program limit metadata only in the first program . However, the reset timer does not correctly indicate the location of the program boundary between the end of the first truncated program and the start of the second truncated program. Therefore, truncation of both spliced bit streams can avoid precise determination between the boundary between them.
The embodiments of the present invention can be implemented in hardware, firmware, or software, or a combination of both (eg, as a programmable logic array). Unless otherwise specified, the algorithms or processes included as part of the invention do not inherently refer to any computer or other
IMPI
<img file="MX356196B_D0102.tif" />
126 particular apparatus. In particular, various general-purpose machines may be used with programs written in accordance with the indications herein, or it may be desirable to construct more specialized apparatus (eg, integrated circuits) to perform the necessary steps of the method.
Therefore, the invention can be implemented in one or more computer programs that run in one or more programmable computer systems (for example, an implementation of any of the elements of Figure 1, or encoder 100).
<td>of</td><td>the</td><td>Figure 2</td><td>(or</td><td>a</td><td>element</td><td>of</td><td>this),</td><td>or decoder</td><td> 200</td>
<td>of</td><td>the</td><td>Figure 3</td><td>(or</td><td>a</td><td>element</td><td>of</td><td>this),</td><td>or post processor</td><td> 300</td>
<td>of</td><td>the</td><td>Figure</td><td>3 i</td><td>[or</td><td colspan="2">an element</td><td colspan="2">of this)) where each</td><td>one</td>
it comprises at least one processor, at least one data storage system (including volatile and nonvolatile memory and / or storage elements), at least one input device or port, and at least one output device or port. The program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known way. Each of these programs can be implemented in any desired computer language (which includes machine, assembly, or high-level procedural languages,
127
IMPI
<img file="MX356196B_D0103.tif" />
logical, or object-oriented programming languages) to communicate with a computer system. In either case, the language can be a compiled or interpreted language.
For example, when implemented by computer software instruction sequences, various functions and steps of the embodiments of the invention can be implemented by multi-threaded software instruction sequences that are executed on suitable digital signal processing hardware, in which case the various devices, steps, and mode functions may correspond to parts of the software instructions.
Each computer program is preferably stored or downloaded to a storage medium or device (for example, memory or solid state medium, or magnetic or optical medium) readable by a programmable computer for general or special use, to configure and operate the computer when the storage medium or device is read by a computer system to perform the procedures described herein. The system of the invention can also be implemented as a computer-readable storage medium, configured with (i.e., storing) a computer program, where the medium
<img file="MX356196B_D0104.tif" />
128
Storage IMPI configured in this way causes a computer system to operate in a specific and predefined manner to perform the functions described herein.
A number of embodiments of the invention have been described.
However, it will be understood that different modifications may be made without departing from the spirit or the scope of the description. Various modifications and variations of the present invention can be made in light of the above indications. It is understood that, within the scope of the appended claims, the invention may be carried out in any way other than as specifically described herein.
129
IMPI
<img file="MX356196B_D0105.tif" />
Contents99
111 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98 Sheet 99 Sheet 100 Sheet 101 Sheet 102 Sheet 103 Sheet 104 Sheet 105 Sheet 106 Sheet 107 Sheet 108 Sheet 109 Sheet 110 Sheet 111
589 members in 28 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 201361754882 | United States of America | P | |
| 61754882210120161824010 | United States of America | – | |
| 201361824010 | United States of America | P | |
| 2014011672 | United States of America | W |
Members589
| Document | Office | Kind | |
|---|---|---|---|
| CA2816889A1 | Canada | A1 | |
| CA2998405A1 | Canada | A1 | |
| CA3216692A1 | Canada | A1 | |
| WO2012075246A2 | World Intellectual Property Organization (WIPO) | A2 | |
| TW201236446A | Taiwan Province of China | A | |
| AR084086A1 | Argentina | A1 | |
| DE202013001075U1 | Germany | U1 | |
| AU2011336566A1 | Australia | A1 | |
| JP3183637U | Japan | U | |
| MX2013005898A | Mexico | A | |
| IL226100A0 | Israel | A0 | |
| IL226100D0 | Israel | D0 | |
| SG190164A1 | Singapore | A1 | |
| DE202013006242U1 | Germany | U1 | |
| CN203134365U | China | U | |
| US2013246077A1 | United States of America | A1 | |
| EP2647006A1 | European Patent Office (EPO) | A1 | |
| JP3186472U | Japan | U | |
| KR20130111601A | Republic of Korea | A | |
| CL2013001571A1 | Chile | A1 | |
| CN103392204A | China | A | |
| TWM467148U | Taiwan Province of China | U | |
| CN203415228U | China | U | |
| JP2014505898A | Japan | A | |
| CN103943112A | China | A | |
| CA2888350A1 | Canada | A1 | |
| WO2014113465A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014113471A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014113478A1 | World Intellectual Property Organization (WIPO) | A1 | |
| FR3001325A3 | France | A3 | |
| UA106163C2 | Ukraine | C2 | |
| WO2014124377A2 | World Intellectual Property Organization (WIPO) | A2 | |
| KR20140106760A | Republic of Korea | A | |
| KR101438386B1 | Republic of Korea | B1 | |
| TWM487509U | Taiwan Province of China | U | |
| TW201442020A | Taiwan Province of China | A | |
| WO2014124377A3 | World Intellectual Property Organization (WIPO) | A3 | |
| CA2898891A1 | Canada | A1 | |
| CN104240709A | China | A | |
| WO2014204783A1 | World Intellectual Property Organization (WIPO) | A1 | |
| FR3007564A3 | France | A3 | |
| KR20140006469U | Republic of Korea | U | |
| RU2013130293A | Russian Federation | A | |
| TW201506911A | Taiwan Province of China | A | |
| SG11201502405RA | Singapore | A | |
| IL237561A0 | Israel | A0 | |
| IL237561D0 | Israel | D0 | |
| KR20150047633A | Republic of Korea | A | |
| AU2014207590A1 | Australia | A1 | |
| HK1198674A | Hong Kong, China | A | |
| HK1198674A1 | Hong Kong, China | A1 | |
| CN104737228A | China | A | |
| FR3001325B3 | France | B3 | |
| MX2015004468A | Mexico | A | |
| AU2014281794A1 | Australia | A1 | |
| EP2901449A1 | European Patent Office (EPO) | A1 | |
| TWI496461B | Taiwan Province of China | B | |
| AU2014207590B2 | Australia | B2 | |
| AU2014281794B2 | Australia | B2 | |
| IN1633MUN2015A | India | A | |
| IN1765MUN2015A | India | A | |
| IN1766MUN2015A | India | A | |
| SG11201505426XA | Singapore | A | |
| IL239687A0 | Israel | A0 | |
| IL239687D0 | Israel | D0 | |
| KR20150099586A | Republic of Korea | A | |
| KR20150099615A | Republic of Korea | A | |
| KR20150099709A | Republic of Korea | A | |
| KR200478147Y1 | Republic of Korea | Y1 | |
| AU2014281794B9 | Australia | B9 | |
| KR20150105955A | Republic of Korea | A | |
| CN104937844A | China | A | |
| CN104995677A | China | A | |
| MX2015010477A | Mexico | A | |
| JP2015531498A | Japan | A | |
| CN105027478A | China | A | |
| HK1204135A | Hong Kong, China | A | |
| HK1204135A1 | Hong Kong, China | A1 | |
| US2015325243A1 | United States of America | A1 | |
| FR3007564B3 | France | B3 | |
| TW201543469A | Taiwan Province of China | A | |
| RU2568372C2 | Russian Federation | C2 | |
| EP2946469A1 | European Patent Office (EPO) | A1 | |
| EP2946495A1 | European Patent Office (EPO) | A1 | |
| US2015348558A1 | United States of America | A1 | |
| EP2954515A1 | European Patent Office (EPO) | A1 | |
| US2015363160A1 | United States of America | A1 | |
| US2015372820A1 | United States of America | A1 | |
| IL239687A | Israel | A | |
| TWI524329B | Taiwan Province of China | B | |
| JP2016507088A | Japan | A | |
| JP5879362B2 | Japan | B2 | |
| JP2016507779A | Japan | A | |
| TW201610984A | Taiwan Province of China | A | |
| KR20160032252A | Republic of Korea | A | |
| JP2016510544A | Japan | A | |
| MX338238B | Mexico | B | |
| CA2888350C | Canada | C | |
| CA2898891C | Canada | C | |
| CN103392204B | China | B |
Numbers
- Publication
- 356196
- Application
- 2016004804
Titles2
- Spanish
- CODIFICADOR Y DECODIFICADOR DE AUDIO CON METADATOS DE LIMITE Y SONORIDAD DE PROGRAMA.
- English
- AUDIO ENCODER AND DECODER WITH PROGRAM LOUDNESS AND BOUNDARY METADATA.
Classification
- CPC, 7
- G10L19/002
- G10L19/167
- H03G9/005
- H03G9/025
- G10L19/06
- G10L19/16
- G10L19/008
- IPC, 4
- G10L19 002
- G10L19 16
- H03G9 00
- H03G9 02