Audio processing apparatus with loudness processing state metadata
Abstract
In one type of embodiment of this creation, this creation is an audio processing device that includes an input register memory for storing at least one audio frame ( frame), a parser coupled to the input register memory to retrieve the audio data, an AC-3 or E-AC-3 decoder, coupled to the parser to generate a decoded audio The data stream and an output register memory are coupled to the decoder to store the decoded audio data.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
8 claims: 8 independent, 0 dependent
- 1An audio processing device comprising:an input register memory for storing at least one frame of an encoded audio bit stream including LPSM and audio data;a parser coupled to the Input register memory to retrieve the audio data;an AC-3 or E-AC-3 decoder coupled to the parser to generate a decoded audio data stream;and an output register memory , Which is coupled to the decoder to store the decoded audio data. 一種音訊處理設備,其包含:一輸入暫存器記憶體,用來儲存一包含LPSM及音訊資料之經過編碼的音訊位元流的至少一音框(frame);一剖析器,其耦合至該輸入暫存器記憶體以擷取該音訊資料;一AC-3或E-AC-3解碼器,其耦合至該剖析器以產生一經過解碼的音訊資料流;及一輸出暫存器記憶體,其耦合至該解碼器以儲存該經過解碼的音訊資料。
- 2For example, the audio processing equipment of the first item of the patent application includes a loudness processor coupled to the AC-3 or E-AC-3 decoder to use the LPSM to implement the decoded audio data stream The adaptive loudness processing. 如申請專利範圍第1項之音訊處理設備,其更包含一響度處理器,其耦合至該AC-3或E-AC-3解碼器,用以使用該LPSM來實施該經過解碼的音訊資料流之可調適的(adaptive)響度處理。
- 3For example, the audio processing equipment of item 2 of the scope of patent application further includes an audio status verifier, which is coupled to the AC-3 or E-AC-3 decoder to identify and/or confirm the LPSM and/or use The LPSM authenticates and/or confirms the decoded audio data stream, and the audio state verifier is further coupled to the loudness processor to control the adaptive loudness processing of the loudness processor. 如申請專利範圍第2項之音訊處理設備,其更包含一音訊狀態驗證器,其耦合至該AC-3或E-AC-3解碼器,用以鑑定及/或確認該LPSM及/或使用該LPSM來鑑定及/或確認該經過解碼的音訊資料流,其中該音訊狀態驗證器更被耦合至該響度處理器以控制該響度處理器的該可調適的響度處理。
- 4For example, the audio processing equipment of the second item of the patent application further includes a post-processor, which is coupled to the AC-3 or E-AC-3 decoder to use the LPSM to implement the decoded audio data stream Its adjustable loudness processing. 如申請專利範圍第2項之音訊處理設備,其更包含一後處理器,其耦合至該AC-3或E-AC-3解碼器,用以使用該LPSM來實施該經過解碼的音訊資料流之可調適的響度處理。
- 5For example, the audio processing equipment of item 4 of the scope of patent application, it further includes an audio status verifier, which is coupled to the AC-3 or E-AC-3 decoder to identify and/or confirm the LPSM and/or use The LPSM authenticates and/or confirms the decoded audio data stream, wherein the audio state verifier is further coupled to the loudness processor and the post-processor to control the adaptability of the loudness processor and the post-processor The loudness processing. 如申請專利範圍第4項之音訊處理設備,其更包含一音訊狀態驗證器,其耦合至該AC-3或E-AC-3解碼器,用以鑑定及/或確認該LPSM及/或使用該LPSM來鑑定及/或確認該經過解碼的音訊資料流,其中該音訊狀態驗證器更被耦合至該響度處理器及該後處理器以控制該響度處理器及該後處理器的該可調適的響度處理。
- 6For example, the audio processing equipment of the first item in the scope of patent application, wherein the LPSM is a container of one or more loudness processing state interpretation data located after the header of the at least one sound frame. 如申請專利範圍第1項之音訊處理設備,其中該LPSM是位在該至少一音框的標頭之後的一或多個響度處理狀態詮釋資料的容器。
- 7For example, in the audio processing equipment of item 1 in the scope of patent application, the LPSM includes a loudness adjustment slot. 如申請專利範圍第1項之音訊處理設備,其中LPSM包含一響度調整欄位(slot)。
- 8For example, the audio processing equipment of item 1 in the scope of patent application, in which the LPSM includes a loudness correction field. 如申請專利範圍第1項之音訊處理設備,其中LPSM包含一響度校正欄位。
Independent claims8
121 paragraphs, as filed
Audio processing equipment with loudness processing state interpretation data
Audio processing apparatus with loudness processing state metadata
Related Application Cases
This case is related to the U.S. Provisional Patent Application No. 61/754,882 filed on January 21, 2013 by Michael Ward and Jeffrey Riedmiller named "Audio Encoder and Decoder with Loudness Processing State Metadata".
This creation is related to audio processing, and more specifically, it is a device that decodes the bit stream with the interpretation data (metadata) representing the loudness processing state of the audio content. Some embodiments of the present creation generate or decode audio in one of Dolby Digital (AC-3), Dolby Digital Plus (Enhanced AC-3 or E-AC-3), or Dolby E formats.
Dolby, Dolby Digital, Dolby Digital Plus, or Dolby E are trademarks of Dolby Laboratories Licensing Corporation. Dolby Laboratories provides services called Dolby Digital and Dolby Digital Plus's exclusive implementation of AC-3 and E-AC-3.
The audio data processing unit typically operates in a blind manner and does not pay attention to the processing history of the audio data before the audio data is received. This completes all the audio data processing and encoding processing architectures for various target media rendering devices while a single entity implements all the decoding and providing operations of the encoded audio data on a target media rendering device It works. However, this kind of blind processing is difficult when multiple audio processing units are spread across a changeable network or are set in a random (eg, chain) manner and are expected to implement their respective types of audio processing. Implemented (or impossible to implement at all). For example, some audio data can be encoded for high-performance media systems and must be converted into a reduced format suitable for mobile devices along a media processing chain. Therefore, an audio processing unit may unnecessarily perform a kind of processing on an audio data, and this kind of processing is a processing that has already been implemented. For example, a volume leveling unit can perform processing on an input audio clip, regardless of whether the same or similar volume leveling processing has been performed on the input audio clip before. Therefore, the volume leveling unit will implement the leveling even when it is not necessary. This unnecessary processing will also cause the degradation and/or elimination of specific features while providing the content of the audio data.
A typical audio data stream includes both audio content (eg, one or more channels of the audio content) and metadata indicative of at least one feature of the audio content. For example, there are several audio interpretation data parameters in an AC-3 bit stream, which are especially used when changing the sound of a program that is sent to a listening ring mirror. Interpretative resources One of the material parameters is the DIALNORM parameter, which is used to indicate the average level of the dialogue occurring in an audio program, and is used to determine the audio playback signal level.
During the playback of a bit stream containing a series of different audio program segments (each segment has a different DIALNORM parameter), an AC-3 decoder uses the DIALNORM parameter of each segment to perform a loudness process. During the process, the playback level or loudness is changed, so that the perceived loudness of the dialogue of the series of segments is at the same level. Each encoded audio segment (item) in a string of encoded audio items (item) will (substantially) have a different DIALNORM parameter, and the decoder will adjust the position of each item of these items This makes the dialogs playback level or loudness for each item the same or very close, although this may require different amounts of gain to be applied to different items during playback.
The DIALNORM is typically set by the user and is not automatically generated, but if no value is set by the user, there will be a default DIALNORM value. For example, a content creator will use a device outside the AC-3 encoder to perform loudness measurement, and then send the result (which is an indication of the loudness of an audio programs spoken dialogue) to the encoder. Set the DIALNORM value. Therefore, the correct setting of the DIALNORM parameters depends on the content creator.
The DIALNORM parameter in the AC-3 bitstream may be incorrect for several different reasons. First, if a DIALNORM value is not If generated by the content creator, each AC-3 encoder has a default DIALNORM value, which is used during the generation of the bitstream. This default value may be substantially different from the actual dialog loudness level of the audio. Second, even if a content creator measures the loudness and sets the DIALNORM value accordingly, the loudness measurement algorithm or measurer used may not match the memorized AC-3 loudness measurement method, resulting in an incorrect DIALNORM value. Third, even though the AC-3 bitstream has been generated with the measured DIALNORM value and correctly set by the content creator, it may be changed to an incorrect one during the transmission and/or storage of the bitstream. Numerical value. For example, it is very common in TV broadcast applications where AC-3 bitstreams will use incorrect DIALNORM interpretation data to be decoded, modified, and then re-encoded. Therefore, the DIALNORM value included in the AC-3 bit stream may be incorrect or inaccurate, which will have a negative impact on the quality of the listening experience.
In addition, the DIALNORM parameter does not indicate the loudness processing status of the corresponding audio data (for example, which loudness processing has been implemented on the audio data). Before this creation, the audio bitstream did not include interpretation data, which is the loudness processing state of the audio content of the audio bitstream (for example, the type of loudness processing applied to the audio content) or the audio of the audio bitstream The loudness processing state of the content and the indication of the loudness, the format of which is a form described in this article. The loudness processing state interpretation data in this format is useful for promoting the adaptive loudness processing of an audio bit stream in a particularly efficient manner and/or the loudness processing state of the audio content and the validity of the loudness. Verification is helpful.
PCT International Patent Application Publication No. WO 2012/075246 A2 (which filed an international application on December 1, 2011 and was assigned to the applicant in this case) discloses the use of interpretative data (which is A method and system for the audio bitstream of the processing state (such as loudness processing state) and characteristics (such as the representation of loudness) of audio content. This reference also describes the adaptive processing of the audio content of the bit stream using the interpretation data, and the verification of the loudness processing state and the correctness of the audio content of the bit stream using the interpretation data. However, the reference does not describe that an audio bitstream contains interpretation data (LPSM), which is the loudness processing of the audio data and the representation of the loudness in the format described in this article. As mentioned, this format of LPSM is very effective in facilitating the adaptive loudness processing of the bit stream and/or the loudness processing status of the audio content and the validity of the loudness. Verification is helpful.
Although this creation is not limited to use with AC-3 bitstream, E-AC-3 bitstream, or Dolby E bitstream, for convenience, in the embodiment of this creation, this creation will be described Generate, decode, or otherwise process the bit stream containing the interpretation data of the loudness processing state.
An AC-3 encoded bit stream contains interpretation data and audio content of 1 to 6 channels. The audio content is audio data that has been compressed using perceptual audio coding. The interpretation data includes several audio interpretation data parameters, which are sent after changing To a parameter used when listening to the sound of a program in the environment.
The details of AC-3 (also known as Dolby Digital) encoding are well-known and proposed in many references, including:<i>ATSC Standard A52/A: Digital Audio Compression Standard (AC-3), Revision A</i>, Advanced Television Systems Committee, 20 Aug. 2001; and US Patent Nos. 5,583,962; 5,632,005; 5,633,981; 5,727,119; and 6,021,386.
The details of Dolby Digital Plus (E-AC-3) are described in "Introduction to Dolby Digital Plus, an Enhancement to the Dolby Digital Coding System," AES Convention Paper 6196,117 published on October 28, 2004.<sup>th</sup>AES Convention.
The details of Dolby E encoding are described in "Efficient Bit Allocation, Quantization, and Coding in an Audio Distribution System", AES Preprint 5068, 107th AES Conference in August 1999 and "Professional Audio Coder Optimized for Use with Video", AES Preprint 5033, 107th AES Conference.
Each frame of an AC-3 encoded bit stream contains audio content and interpretation data for 1536 digital audio samples. For a 48kHz sampling rate, this means 32 milliseconds of digital audio or 31.25 frames of audio per second.
Each frame of an E-AC-3 encoded bit stream contains The audio content and interpretation data for 256, 512, 768 or 1536 digital audio samples are related to the audio frame containing 1, 2, 3 or 6 digital data blocks. For a 48kHz sampling rate, this means 5.333, 10.667, 16 or 32 milliseconds of digital audio or 189.9, 93.75, 62.5 or 31.25 frames per second.
As shown in Figure 4, each AC-3 sound frame is divided into sections (segments), which include: a synchronization information (SI) section, which includes a synchronization character (SW) and two error correction characters. The first error correction character (CRC1) (as shown in Figure 5); the bit stream information (BSI) section, which contains most of the interpretation data; 6 audio blocks (AB0 to AB5), which Contains compressed audio content data (and may also include interpretation data); waste bits (W), which include any unused bits left over after the audio bits are compressed; auxiliary (AUX) Information section, which can contain more interpretation data; and the second error correction character (CRC2) of the two error correction characters.
As shown in FIG. 7, each E-AC-3 sound frame is divided into sections (segments), which include: a synchronization information (SI) section which contains a synchronization character (SW) (as shown in FIG. 5) ); Bitstream Information (BSI) section, which contains most of the interpretation data; 1 to 6 audio blocks (AB0 to AB5), which contain the compressed audio content data (and may also include interpretation Data); useless bits (W), which include any unused bits left after the audio bit is compressed; auxiliary (AUX) information section, which can contain more interpretation data; and error correction Characters (CRC).
There are several audios in an AC-3 (or E-AC-3) bit stream Interpretation data parameters are especially used when changing the sound of a program sent to a listening environment. One of the interpretive data parameters is the DIALNORM parameter, which is included in the BSI segment.
As shown in Figure 6, a BSI segment of an AC-3 frame includes a 5-bit parameter ("DIALNORM") that displays the DIALNORM value for the program. A 5-bit parameter ("DIALNORM2") showing the DIALNORM value of the second audio program carried in the same AC-3 frame is included if the audio coding mode of the AC-3 frame ("acmod") If it is "0", it means using double 1 or "1+1" channel configuration.
The BSI fragment also includes a flag ("addbsie"), which is displayed after the "addbsie" bit with (or no) additional bitstream information, and a parameter ("addbsil"), which is displayed in the "addbsie" bit. The length of any additional bitstream information after the "" value, and up to 64 bits of additional bitstream information ("addbsi") after the "addbsil" value.
The BSI segment includes other interpretation data values that are not specifically shown in FIG. 6.
In one type of embodiment of the present creation, the present creation is an audio processing device that includes an input register memory for storing at least one audio frame of an encoded audio bitstream including LPSM and audio data, A parser, which is coupled to the input register memory to retrieve the audio data, and an AC-3 or E-AC-3 decoder, which is coupled to the parser To generate a decoded audio data stream and an output register memory, which is coupled to the decoder to store the decoded audio data.
<p>100Encoder</p><p>101Decoder</p><p>102Audio Status Verifier</p><p>103Loudness processing stage</p><p>104Audio stream selection stage</p><p>105Encoder</p><p>106Interpretation data generation stage</p><p>107Stuffer/Formatter stage</p><p>108Dialogue loudness measurement subsystem</p><p>109Sound frame register (memory)</p><p>150Encoded audio transmission subsystem</p><p>152Decoder</p><p>110Sound frame register</p><p>111Analyzer</p><p>200Decoder</p><p>300Post processor</p><p>201Sound frame register</p><p>202Audio Decoder</p><p>203Audio Status Verification Phase (Verifier)</p><p>204Control bit generation stage</p><p>205Parser</p><p>301Sound frame register</p>
Figure 1 is a block diagram of an embodiment of a system that can be constructed to implement an embodiment of the method of authoring.
Fig. 2 is a block diagram of an encoder of an embodiment of the audio processing unit of the present creation.
FIG. 3 is a block diagram of a decoder of an embodiment of the audio processing unit of the present creation and a post-processor coupled to the decoder. The post-processor is another embodiment of the audio processing unit of the present creation.
Fig. 4 is a diagram including an AC-3 sound frame divided into segments.
FIG. 5 is a diagram of a synchronization information (SI) segment including AC-3 sound frames divided into segments.
FIG. 6 is a diagram of a bit stream information (BSI) segment including AC-3 sound frames divided into segments.
Fig. 7 is a diagram including an E-AC-3 sound frame divided into segments.
Signs and terms
In this disclosure including the scope of the patent application, an operation (eg, filtering, scaling, etc.) is performed on a signal or data. (scaling, deforming, or applying gain to the signal or data) This description is used in a broad sense to indicate that the signal or data is directly or a processed version of the signal or data (e.g., The signal has received an early filtering or pre-processing version before performing the operation) to perform the operation.
In this disclosure including the scope of the patent application, the term "system" is used in a broad sense to mean a device, system, or subsystem. For example, a subsystem that implements a decoder can be referred to as a decoder system, and a system that includes this subsystem (e.g., a system that generates X output signals in response to multiple outputs, in which the The subsystem generates M inputs and the other XM inputs are received from an external source) can also be referred to as a decoder system.
In this disclosure including the scope of the patent application, the term "processor" is used in a broad way to mean a programmable or configurable (configurable) (eg, using Software or firmware) systems or devices used to perform operations on data (such as audio, or video, or other image data). Examples of processors include a field programmable gate array (or other configurable integrated circuit or chipset), one that is programmed and/or configured in other ways for audio or other sound data A digital signal processor that implements pipeline processing, a programmable general-purpose processor or computer, and a programmable microprocessor chip or chipset.
In this disclosure including the scope of the patent application, the terms "audio processor" and "audio processing unit" are used interchangeably And used in a broad sense to indicate a system constructed to process audio data. Examples of audio processing units include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems, post-processing systems, and bit stream processing systems (sometimes referred to as bit stream Processing tools).
In this disclosure including the scope of the patent application, the term "processing state interpretation data" (for example, expressed as "loudness processing state interpretation data") refers to the corresponding audio data (one also includes processing state interpretation data). The separate and different data of the audio content of the data stream. The processing status interpretation data is associated with the audio data, shows the loudness processing status of the corresponding audio data (for example, what processing has been performed on the audio data), and typically also shows at least one characteristic or characteristic of the audio data. The association between the processing state interpretation data and the audio data is time synchronization. Therefore, the current (recently received or updated) processing state interpretation data shows that the corresponding audio data also includes the results of the displayed audio data processing type. In some examples, the processing status interpretation data may include processing history and/or part or all of the parameters used in and/or derived from the displayed processing category. In addition, the processing state interpretation data may include at least one characteristic or feature of the corresponding audio data (which has been calculated or extracted from the audio data). The processing status interpretation data also includes other interpretation data that is irrelevant to any processing of the corresponding audio data or not derived from any processing of the corresponding audio data. For example, third-party data, audio track information, identification code, exclusive or standard information, user annotation data, user preference data, etc. can be added with a special audio processing unit Pass it to other audio processing units.
In this disclosure including the scope of the patent application, the term "loudness processing state interpretation data" (or "LPSM") means the corresponding audio datas processing state interpretation data representation of the loudness processing state (e.g., already What kind of processing is performed on the audio data) and typically also refers to at least one characteristic or characteristic (for example, loudness) of the corresponding audio data. The loudness processing state interpretation data (ie, when it is considered separately) may include data that is not the loudness processing state interpretation data (eg, other interpretation data).
In this disclosure including the scope of the patent application, words such as "coupled" or "coupled" are used to indicate direct or indirect connection. Therefore, if a first device is coupled to a second device, the connection can be through a direct connection or through an indirect connection through other devices or other connections.
Detailed description of the embodiment of this creation
According to the exemplary embodiment of the present creation, the loudness processing state interpretation data (LPSM) is embedded in one or more reserved areas (or slots) of the interpretation data segment of an audio bit stream, and the audio bit The stream also includes audio data in other segments (audio data segments). Typically, at least one segment of each sound frame of the bit stream includes LPSM, and at least one other segment of the sound frame includes corresponding audio data (ie, the loudness processing state and the audio data indicated by the LPSM) . In some embodiments, the data volume of the LPSM can be small enough to be carried without affecting To the bit rate allocated to carry the audio data.
When two or more audio processing units must work cooperatively with each other in an audio processing chain (or content life cycle), it is helpful to communicate loudness processing state interpretation data in the audio processing chain. If the loudness processing status interpretation data is not included in the audio bitstream, several media processing issues (such as quality, level, and spatial degradation) will be used for the audio in, for example, two or more audio codecs In the processing chain and single-ended volume leveling occurs when the bit stream travels to the media consuming device (or the rendering point of the audio content of the bit stream).
FIG. 1 is a block diagram of an exemplary audio processing chain (audio data processing system). In this figure, one or more components of the system can be constructed according to embodiments of the present invention. The system includes the following elements, which are coupled together as shown: a preprocessing unit, an encoder, a signal analysis and interpretation data correction unit, a conversion encoder, a decoder, and a preprocessing unit. In the variation of the system shown, one or more of these components are omitted, or additional audio data processing units are added.
In some implementations, the preprocessing unit of FIG. 1 is constructed to accept PCM (time domain) samples containing audio content as input, and output processed PCM samples. The encoder can be constructed to accept the PCM samples as input and output an encoded (eg, compressed) audio bitstream indication of the audio content. The data of the bitstream (which is the representation of the audio content) is sometimes referred to as "audio Data". If the encoder is constructed according to a typical embodiment of this creation, the audio bitstream output from the encoder includes loudness processing state interpretation data (and typically other interpretation data) and audio data.
The signal analysis and interpretation data correction unit of Figure 1 can accept one or more encoded audio bitstreams as input and determine (eg, verify) in each encoded audio bitstream by performing signal analysis The processing status interprets whether the data is correct. If the signal analysis and interpretation data correction unit finds that the included interpretation data is invalid, it will replace the incorrect value with the correct value obtained from the signal analysis. Therefore, each encoded audio bit stream from the signal analysis and interpretation data correction unit may include corrected (or uncorrected) processing state interpretation data and encoded audio data.
The transcoder of Figure 1 accepts an encoded audio bitstream as input, and outputs a modified (eg, differently coded) audio bitstream in response (eg, decodes the input stream and decodes the The decoded stream is re-encoded in a different encoding format). If the transcoder is constructed according to a typical embodiment of this creation, the audio bitstream output from the transcoder includes loudness processing state interpretation data (and typically other interpretation data) and encoded audio material. The interpretation data may already be included in the bitstream.
The decoder of FIG. 1 accepts an encoded (eg, compressed) audio bit stream as input, and (in response) outputs a decoded PCM audio sample stream. If the decoder is constructed according to a typical embodiment of this creation, the output of the decoder in typical operation is Any of the following may include any of the following: an audio sample stream, and a corresponding loudness processing state interpretation data (and typically other interpretation data) stream extracted from an encoded bitstream input ; Or an audio sample stream, and a corresponding control bit stream, which is determined using the loudness processing state interpretation data (and typically other interpretation data) extracted from an encoded bit stream input; or An audio sample stream that does not have a corresponding processing state interpretation data stream or a control bit stream determined by the processing state interpretation data. In this last example, the decoder can extract loudness processing state interpretation data (and/or other interpretation data) from the encoded bitstream input and perform at least one operation on the captured interpretation data (For example, verification), even if the decoder does not output the captured interpretation data or the control bits determined by the interpretation data.
By constructing the post-processing unit of FIG. 1 according to a typical embodiment of the present creation, the post-processing unit is constructed to receive a decoded PCM audio sample stream and use the received loudness processing state along with the samples Interpretation data (and other interpretation data) or (determined by the decoder from the loudness processing state interpretation data and other interpretation data) control bits to perform post-processing on the decoded PCM audio sample stream (for example, the volume of the audio content) Leveling). The post-processing unit is also typically constructed to render the post-processed audio content for playback by one or more speakers.
The typical embodiment of this creation provides an enhanced audio processor In the audio processing chain, the audio processing units (such as encoders, decoders, transcoders, and pre-processing units and post-processing units) adapt their respective processing, which will be based on these The contemporaneous state (contemporaneous state) of the media data displayed in the loudness processing state interpretation data respectively received by the audio processing unit is applied to the audio data.
The audio data input to any audio processing unit of the system in Figure 1 (such as the encoder or transcoder in Figure 1) may include loudness processing state interpretation data (and optionally, other interpretation data) and audio data (e.g., Encoded audio data). According to the embodiment of the present creation, the interpretation data may have been included in the input audio by another element of the system of FIG. 1 (or another source not shown in FIG. 1). The processing unit that receives the audio input (with interpretive data) can be configured to perform at least one operation (such as verification) on the interpretive data or respond to the interpretive data (such as adaptive processing of the input audio). )), and typically include the interpretation data, the processed version of the interpretation data, or the control bits determined by the interpretation data in its output audio.
In a typical embodiment of this creation, the audio processing unit (or audio processor) is constructed to implement the audio data status according to the audio data status displayed by the loudness processing status interpretation data and the audio data corresponding to the audio data. Adjusted treatment. In some embodiments, the adaptive processing is (or includes) loudness processing (if the interpretation data shows that the loudness processing, or similar processing, has not been implemented on the audio data If the adaptive processing is not (and does not include) loudness processing (if the interpretation data shows that the loudness processing, or similar processing, has been implemented on the audio data). In some embodiments, the adaptive processing is or includes interpretation data verification (for example, implemented in an interpretation data verification subunit) to ensure that the audio processing unit displays the interpretation data according to the loudness processing state. The state of the audio data is used to implement other adaptive processing of the audio data. In some embodiments, the verification determines the reliability of the loudness processing state interpretation data associated with the audio data (eg, included with it in a bit stream). For example, if the interpretive data is verified to be reliable, the result from the type of audio processing that was previously implemented can be reused and the implementation of the same type of audio processing can be avoided. On the other hand, if the interpretation data is found to have been tampered with (or unreliable for other reasons), then (as shown by the unreliable interpretation data) the type of media processing that has been allegedly implemented must be The audio processing unit is repeated, and/or other processing can be implemented by the audio processing unit on the interpretation data and/or the audio data. The audio processing unit can also be constructed to send a signal that the loudness processing state interpretation data is valid to other audio processing units downstream in an enhanced media processing chain (e.g., present in a media bit stream), If the unit determines that the loudness processing state interpretation data is valid (for example, based on the result that a retrieved cryptographic value (cryptographic value) matches a reference cryptographic value).
Figure 2 is a block diagram of an encoder (100), which is an embodiment of the audio processing unit of the present creation. Encoder 100 any The components or elements may be implemented as one or more processors and/or one or more circuits (for example, ASIC, FPGA, or other integrated circuits), hardware, software, or a combination of hardware and software. The encoder 100 includes a sound frame register 110, a parser 111, a decoder 101, an audio state verifier 102, a loudness processing stage 103, an audio stream selection stage (MUX) 104, an encoder 105, and a filler/formatting stage 107 The interpretation data generation stage 106, the dialog loudness measurement subsystem 108, and the sound frame register 109 are connected as shown in the figure. Typically, the encoder 100 also includes other processing elements (not shown).
The encoder 100 (which is a conversion encoder) is constructed to implement adaptive and automated loudness processing by using the loudness processing state interpretation data included in the bit stream input to input the audio bit stream (such as It can be one of AC-3 bit stream, E-AC-3 bit stream, or Dolby E bit stream) converted into an encoded audio bit stream output (which can be, for example, AC-3 bit stream , E-AC-3 bitstream, or the other of Dolby E bitstream). For example, the encoder 100 can be constructed to convert a Dolby E bitstream input (a format that is typically used in production and broadcasting equipment rather than in consumer devices that receive audio programs) into an AC-3 or E-AC-3 format encoded audio bitstream output (which is suitable for broadcasting to consumer devices).
The system of FIG. 2 also includes an encoded audio transmission subsystem 150 (which stores and/or transmits the encoded bit stream output from the encoder 100) and a decoder 152. An encoded audio bitstream output from the encoder 100 can be output by the subsystem 150 (for example, in DVD or blue It is stored in the form of an optical disc, or transmitted by the subsystem 150 (which can be embodied as a transmission chain or network), or stored and transmitted by the subsystem 150 at the same time. The decoder 152 is constructed to decode an encoded audio bitstream (generated by the encoder 100) received by the subsystem 150, which includes the loudness processing state interpretation data (LPSM) from the bitstream. Extract and generate decoded audio data from each audio frame. Typically, the decoder 152 is constructed to use the LPSM to perform adaptive loudness processing on the decoded audio data, and/or to send the decoded audio data and LPSM to a post-processor, which is constructed for use The LPSM performs adaptive loudness processing on the decoded audio data. Typically, the decoder 152 includes a register that stores (eg, in a non-transitory manner) the decoded audio data received from the subsystem 150.
The different materializations of the encoder 100 and the decoder 152 are constructed to implement different embodiments of the authoring method.
The sound frame register 110 is a temporary memory coupled to receive a decoded audio bit stream input. In operation, the register 110 stores (eg, in a non-transitory manner) at least one frame of the decoded audio bit stream, and the sequence of the decoded audio bit stream is from the register 110 to the parser 111 is asserted.
The parser 111 is coupled and constructed to extract the loudness processing state interpretation data (LPSM) and other interpretation data from the decoded audio data, declare at least the LPSM to the audio state verifier 102, the loudness processing stage 103, The interpretation data generation stage 106 and the subsystem 108 are used to extract audio data from the decoded audio input And announce the audio data to the decoder 101. The decoder 101 of the encoder 100 is constructed to decode the audio data to generate decoded audio data, and declare the decoded audio data to the loudness processing stage 103, the audio stream selection stage 104, the subsystem 108, And typically it is also announced to the state verifier 102.
The state verifier 102 is constructed to authenticate and confirm the declared LPSM (and optionally, other interpretation data). In some embodiments, the LPSM is a data block that has been included in the bitstream input (or is included in the data block) (eg, according to an embodiment of the present invention). The data block may include a cryptographic hash used to process the LPSM (and optionally other interpretation data) and/or the underlying (provided by the decoder 101 to the verifier 102) audio data. hash) (a hash-based message authentication code or "HMAC"). The data block can be digitally signed in these embodiments, so that a downstream audio processing unit can relatively easily identify and confirm the processing state interpretation data.
For example, the HMAC is used to generate a digest, and the protection value included in the bit stream of the author may include the digest. The abstract can be generated for an AC-3 sound frame as follows:
1. After the AC-3 data LPSM is encoded, the frame data bytes (which are concatenated into frame_data#1 and frame_data#2) and the LPSM data bytes are used as the input of the hash function HMAC. Other data (which can be stored in an auxiliary data field) is used to calculate the abstract It is not considered when necessary. These other data may be bytes that do not belong to the AC-3 data or the LSPSM data. The protection bits included in the LPSM may not be considered when calculating the HMAC digest.
2. After the digest is calculated, it is written into a bit stream in a field reserved for protection bits.
3. The final step in generating the complete AC-3 tone frame is the calculation of the CRC-check. This is written at the end of the frame and all data belonging to this frame are considered, including the LPSM bit.
Other encryption methods (including but not limited to any of one or more non-HMAC encryption methods) can be used for LPSM verification (for example, in the verifier 102) to ensure that the LPSM and/or the underlying Safe transmission and reception of audio data. For example, verification (using this encryption method) can be implemented in each audio processing unit, which receives an embodiment of the created audio bitstream to determine the loudness processing state interpretation data and include Whether the corresponding audio data in the bit stream has been subjected to specific loudness processing (and/or the corresponding audio data has been obtained from the specific loudness processing (indicated by the interpretive data)) and after the specific loudness processing has been implemented Has not been modified.
The state verifier 102 declares control data to the audio stream selection stage 104, the interpretation data generation stage 106, and the dialog loudness measurement subsystem 108 to display the result of the verification operation. When responding to the control data, the stage 104 can select (and send it to the encoder 105) the following: the output of the loudness processing stage 103 that is adaptively processed (For example, when the LPSM indicates that the audio data output from the decoder 101 has not yet received a specific type of loudness processing, and the control bit from the verifier 102 indicates that the LPSM is valid); or from the decoder 101 The audio data output (for example, when the LPSM shows that the audio data output from the decoder 101 has accepted the specific type of loudness processing that will be implemented in stage 103, and the control bit from the verifier 102 shows that the LPSM Is valid).
The stage 103 of the encoder 100 is constructed to perform adaptive loudness processing on the decoded audio data output from the decoder 101 according to one or more audio data characteristics displayed by the LPSM extracted by the decoder 101. Stage 103 may be an adaptable conversion domain real-time loudness and dynamic range control processor. Stage 103 can receive user output (for example, user target loudness/dynamic range value or dialogue normalization value (dialnorm value)), or other interpretation data input (for example, third-party data, audio track information, identification element, exclusive or One or more of standard information, user annotation data, user preference data, etc.) and/or other input (for example, input from authentication processing), and use this input to process the decoded audio data from decoder 101 .
When the control bit from the verifier 102 shows that the LPSM is valid, the dialog loudness measurement subsystem 108 is operable, for example, using the LPSM (and/or other interpretation data) extracted by the decoder 101 to determine The loudness of the segments of the decoded audio (from the decoder 101), which are indicative of the dialogue (or other speech). When the control bit from the validator 102 shows that When the LPSM is valid, the operation of the dialogue loudness measurement subsystem 108 can be disabled when the LPSM displays (from the decoder 101) the previously determined loudness of the dialogue (or other speech) segment of the decoded audio ( disabled).
There are useful tools (eg, Dolby LM100 loudness meter) that can conveniently and easily measure the level of dialogue in the audio content. Some embodiments of the APU (eg, stage 108 of the encoder 100) of the present creation are embodied as including this tool (or the function of implementing this tool) for measuring an audio bit stream (eg, a slave encoder 100). The decoder 101 is announced to the average dialog loudness of the audio content of the decoded AC-3 bitstream of stage 108.
If stage 108 is embodied as measuring the true average dialog loudness of the audio data, the measurement may include the step of isolating segments of the audio content that significantly contain speeches. The audio segments that significantly contain speeches are then processed according to a loudness measurement algorithm. For the audio data decoded from the AC-3 bit stream, this algorithm can be a standard K-weighted loudness measurement (according to the international standard ITU-R BS.1770). Alternatively, other loudness measures may be used (e.g., loudness measurements according to the psychoacoustic model of loudness).
The isolation of the speech fragments is not critical for measuring the average dialog loudness of the audio data. However, from the listener's perspective, this can improve the accuracy of the measurement and typically provide more satisfactory results. Because not all audio content contains dialogue (speech), so such as If there is a speech, the loudness measurement of the entire audio content provides a sufficient approximation of the dialogue level of the audio.
The interpretation data generator 106 generates interpretation data to be included in the encoded bit stream to be output from the encoder 100 by the stage 107. The interpretive data generator 106 can send the LPSM (and/or other interpretive data) extracted by the encoder 101 to the stage 107 (for example, when the control bit from the validator 102 shows that the LPSM and/or When other interpretation data is valid), or generate a new LPSM (and/or other interpretation data) and declare the new interpretation data to stage 107 (for example, when the control bit from the validator 102 shows that the LPSM and / Or when other interpretation data retrieved by the decoder 101 is invalid), or it can declare to the stage 107 a combination of the interpretation data retrieved by the decoder 101 and the newly generated interpretation data. The interpretive data generator 106 may include the loudness data generated by the subsystem 108 and at least one numerical representation of the type of loudness processing performed by the subsystem 108. In the LPSM, it is announced to the stage 107 for inclusion in the From the encoded bit stream output from the encoder 100.
The interpretive data generator 106 can generate protection bits (which can include or include a hash-based message authentication code or "HMAC") for the LPSM to be included in the encoded bit stream ( And optionally, other interpretation data) and/or at least one of the decryption, authentication, or verification of the audio data to be included under the encoded bits is very useful. The interpretive data generator 106 can provide these protection bits to the stage 107 to be included in the encoded bit In the stream.
In a typical operation, the dialog loudness measurement subsystem 108 processes the audio data output from the decoder 101 to generate a loudness value in response to the audio data output (eg, gated or un-gated loudness value) and dynamic range value. In response to these values, the interpretive data generator 106 can generate loudness processing state interpretation data (LPSM) for (by the filler/formatting stage 107) to include in the encoded bit output from the encoder 100 In the stream.
Additionally, optionally, or alternatively, the subsystems 106 and/or 108 of the encoder 100 may perform additional analysis of the audio data to generate the audio data that will be included in the process that will be output from the stage 107 Metadata indicative of at least one feature in the encoded bit stream.
The encoder 105 (by comparing the audio data output from the selection stage 104) encodes the audio data, and declares the encoded audio data to the stage 107 to include the audio data output from the stage 107 The encoded bit stream.
The stage 107 multiplexes the encoded bit stream from the encoder 105 and the interpretation data (including LPSM) from the generator 106 to generate the encoded bit stream to be output from the stage 107 , It is preferable to make the encoded bit stream have the format defined by a preferred embodiment of the present creation.
The frame register 109 is a temporary memory that stores (e.g., in a non-transitory manner) the encoded At least one sound frame output by the bitstream, and a series of sound frames linked by the encoded audio bits are then declared from the register 109 as the output from the encoder 100 to the transmission system 150.
The LPSM generated by the interpretive data generator 106 and included in the encoded bit stream by the stage 107 is the loudness processing state of the corresponding audio data (eg, what loudness processing has been implemented on the audio data ) And the corresponding indication of the loudness of the audio data (eg, the loudness of the measured dialogue, the gated and/or un-gated loudness, and/or the dynamic range).
In this article, the "gating" of the loudness and/or level measurement implemented on audio data refers to a specific level or loudness threshold, and the calculated value that exceeds the threshold is included In the final measurement (for example, ignore the short-term loudness value below -60dBFS in the final measurement value). Absolute value gating refers to a fixed level or loudness, while to a relative value gating refers to a value that is dependent on the current "non-gated" measurement value.
In some implementations of the encoder 100, the encoded bit stream temporarily stored in the memory 109 (and output to the transmission system 150) is an AC-3 bit stream or an E-AC-3 bit stream , And include audio data fragments (for example, the audio frame fragments AB0-AB5 shown in Figure 4) and interpretation data fragments, where these audio data fragments are representations of audio data, and each of at least some of these interpretation data fragments One includes Loudness Processing State Interpretation Data (LPSM). Stage 107 inserts the LPSM into the bit stream of the following format. Each piece of interpretative data including the LPSM is included in In the "addbsi" field of the bit stream information ("BSI") segment of a frame of the bit stream, or in the auxiliary data field of the terminal of a frame of the bit stream (for example, as shown in Figure 4) AUX clip shown). The sound frame of the bit stream can include one or two interpretive data fragments, each interpretive data fragment includes LPSM, and if the sound frame includes two interpretive data fragments, one of them is reproduced in the addbsi column of the sound frame One and the other are in the AUX column of the sound frame. Each interpretive data segment including the LPSM includes an LPSM payload (or container) segment having the following format: a header (which typically includes a syncword used to identify the beginning of the LPSM payload) , Followed by at least one identity value (for example, the LPSM format, length, period, number, and substream associated values described in Table 2 below); and after the header, at least one dialogue indicator value (For example, the "dialog channel" in Table 2 channel)" parameter), which shows whether the corresponding audio data indicates dialogue or does not indicate dialogue (for example, which channels of the corresponding audio data indicate dialogue); at least one loudness adjustment conforming value (for example, "loudness" in Table 2 "Adjustment Type" parameter), which shows whether the corresponding audio data is consistent with a set of pointed loudness adjustments); at least one loudness processing value (for example, "Dialogue gated loudness correction flag (flag)" in Table 2 , One or more of the "loudness positive type" parameters), which show that at least one of the loudness processing has been implemented on the corresponding audio data; and At least one loudness value (e.g., "ITU relative to gated loudness" in Table 2, "ITU speech gated loudness", "ITU (EBU 3341) short-term 3s loudness", and one of the "true peak" parameters Or more), which displays at least one characteristic of loudness (for example, peak or average loudness) of the corresponding audio data.
In some embodiments, each piece of interpretation data inserted by stage 107 into the "addbsi" field of a frame of the bitstream has the following format: a core header (which typically includes a Interpret the sync word at the beginning of the data segment, followed by the identity value (for example, the core component version, length, and period, the number of extension components, and the substreams associated with these values listed in Table 1 below); And after the core header, at least one protection value (such as the HMAC digest and automatic fingerprint value of Table 1) for decryption, authentication, or at least one of the loudness processing state interpretation data or the corresponding audio data At least one of the verifications is useful; and also after the core header, if the interpretive data segment includes LPSM, then the LPSM payload identity ("ID") and the LPSM payload size value will treat the following interpretive data as LPSM The payload also shows the size of the LPSM payload.
The LPSM payload (or container) fragment (which preferably has the above-mentioned format) follows the LPSM payload identity and the LPSM payload size value.
In some embodiments, the auxiliary data in a sound frame (Or "addbsi") each piece of interpretation data in the field has a three-level structure: a high-level structure, which includes a flag indicating whether the auxiliary data (or "addbsi") field includes interpretation data, At least one ID value shows which type of interpretation data appears, and a value, which shows how many bits of the interpretation data (for each type) appear (if any interpretation data appears). One type of interpretive data that can appear is LSPM, and another type of interpretive data that can appear is media research interpretation data (for example, Nielsen Media Research Interpretation Data); an intermediate level structure that contains an interpretive data for each type of identified ( For example, the above-mentioned core components used for the core header, protection value, and LPSM payload ID and LPSM payload size value for each of the identified interpretive data; and the following hierarchical structure, which includes the core components for Each payload of a core component (eg, an LPSM payload, if the core component identifies an LPSM payload, and/or an interpretive data payload, if the core component identifies an interpretive data payload If it exists).
The data values in this three-level structure can be nested. For example, the protection value for an LPSM payload identified by a core component and/or another interpretive data payload can be included after each payload is identified by the core component (and thus included in the core component After the core header). In one example, a core header can identify an LPSM payload and another interpretive data payload, and the payload ID and payload size value used for the first payload (eg, the LPSM payload) can be followed by In the core standard After the head, the first payload itself can follow the ID and the size values, the payload ID and the payload size value used for the second payload can follow the first payload, the second The payload itself can follow these IDs and equivalent values, and the protection value used for these two payloads (or for the core component value and two payloads) can follow the last payload.
In some embodiments, if the decoder 101 accepts an audio bitstream with cryptographic hash generated according to an embodiment of the present creation, then the decoder is constructed to parse and derive data from a data determined by the bitstream. Obtain the cryptographic hash from the block, and the data block contains Loudness Processing State Interpretation Data (LPSM). The verifier 102 can use the cryptographic hash to verify the received bit stream and/or associated interpretation data. For example, the verifier 102 finds the LPSM to be verified based on the match between a reference cryptographic hash and the cryptographic hash obtained from the data block, and then it can use the processor 103 to determine the corresponding audio data. The processing stops and causes the selection stage 104 to transmit (unchanged) the audio data. Additionally, optionally, or alternatively, other types of encryption techniques can be used to replace cryptographic hash-based methods.
The encoder 100 of FIG. 2 can (in response to the LPSM captured by the decoder 101) (in components 105, 106, and 107) determine that a post/pre-processing unit has performed a loudness process on a to-be-encoded audio data. Therefore, it is possible to generate (in the generator 106) loudness processing state interpretation data, which includes specific parameters used in the previously implemented loudness processing and/or derived from the previously implemented loudness processing. In some embodiments, as long as the encoder knows what has been implemented on the audio content According to the processing type, the encoder 100 can generate (and be included in the encoded bit stream output generated by the encoder) the loudness processing state interpretation data representation of the processing history of the audio content.
Fig. 3 is a block diagram of a decoder (200) and a post-processor (300) coupled to the decoder. The decoder is an embodiment of the audio processing unit of the present invention. The post-processor (300) is also an embodiment of the audio processing unit of the present invention. Any component or element of the decoder 200 and the post-processor 300 may be embodied as one or more programs and/or one or more circuits (such as ASIC, FPGA, or other integrated circuits, hardware, software, or hardware). A combination of body and software. The decoder 200 includes a frame register 201, a parser 205, an audio decoder 202, an audio state verification stage (audio state verifier) 203, and a control bit generation stage 204, which are shown in the figure Connected as shown. Typically, the decoder 200 also includes other processing elements (not shown).
The frame register 201 (a temporary memory) stores (eg, in a non-transitory manner) at least one frame of the encoded audio bitstream received by the decoder 200. A series of audio frames of the encoded audio bit stream are asserted from the register 201 to the parser 205.
The parser 205 is coupled and constructed to extract loudness processing state interpretation data (LPSM) from each frame of the encoded audio input to at least declare the LPSM to the audio state verifier 203 and stage 204, for declaring the LPSM as output (sent to the post-processor 300), for extracting audio from the encoded audio input Data, and used to declare the captured audio data to the decoder 202.
The encoded audio bitstream input to the decoder 200 may be one of AC-3 bitstream, E-AC-3 bitstream, or Dolby E bitstream.
The system of FIG. 3 also includes a post-processor 300. The post-processor 300 includes a sound frame register 301 and other processing elements (not shown), which includes at least one processing element coupled to the register 301. The frame register 301 stores (eg, in a non-transitory manner) at least one frame of the decoded audio bit stream received by the post-processor 300 from the decoder 200. The processing elements of the post-processor 300 are coupled and constructed to receive and use the interpretation data (including the LPSM value) output from the decoder 202 and/or the control bits output from the stage 204 of the decoder 200 to adaptively process the slave A series of audio frames of the decoded audio bit stream output by the register 301. Typically, the post-processor 300 is constructed to use the LPSM value (eg, according to the loudness processing status displayed by the LPSM and/or one or more audio data characteristics) to implement an adjustable loudness of the decoded audio data handle.
Various embodiments of the decoder 200 and the post-processor 300 are constructed to implement different embodiments of the method of this authoring.
The audio decoder 202 of the decoder 200 is constructed to decode the audio data extracted by the parser 205 to generate decoded audio data, and declare the decoded audio data as (for example, sent to the post-processor 300 Output.
The state verifier 203 is constructed to authenticate and confirm the declared LPSM (and optionally, other interpretation data). In some embodiments, the LPSM is a data block that has been included in the bitstream input (or is included in the data block) (eg, according to an embodiment of the present invention). The data block may include an underlying (underlying) used to process the LPSM (and optionally other interpretation data) and/or (provided by the parser 205 and/or the decoder 202 to the validator 203) ) The cryptographic hash of the audio data (a hash-based message authentication code or "HMAC"). The data block can be digitally signed in these embodiments, so that a downstream audio processing unit can relatively easily identify and confirm the processing state interpretation data.
Other encryption methods (including but not limited to any of one or more non-HMAC encryption methods) can be used for LPSM verification (eg, in the verifier 203) to ensure the LPSM and the underlying audio data Safe transmission and reception. For example, verification (using this encryption method) can be implemented in each audio processing unit, which receives an embodiment of the created audio bitstream to determine the loudness processing state interpretation data and include Whether the corresponding audio data in the bit stream has been subjected to specific loudness processing (and/or the corresponding audio data has been obtained from the specific loudness processing (indicated by the interpretive data)) and after the specific loudness processing has been implemented Has not been modified.
The state verifier 203 declares control data to the control bit generator 204 and/or declares the control data as an output (for example, sent to the post-processor 300) to display the result of the verification operation. In response When the control data (and optionally other interpretation data extracted from the bit stream input), stage 204 can generate the following list (and declare it to the post-processor 300): control bit, its Shows that the decoded audio data output from the decoder 202 has received a specific type of loudness processing (when the LPSM shows that the decoded audio data output from the decoder 202 has received the specific type of loudness processing, and comes from the The control bit of the verifier 203 indicates that the LPSM is valid); or the control bit indicates that the decoded audio data output from the decoder 202 should receive a specific type of loudness processing (when the LPSM display comes from the When the decoded audio data output of the decoder 202 has not yet received the specific type of loudness processing, or when the LPSM indicates that the decoded audio data output from the decoder 202 has received the specific type of loudness processing, but comes from the verification When the control bit of the device 203 shows that the LPSM is invalid).
Alternatively, the decoder 200 may declare the LPSM (and any other interpretation data) extracted from the bitstream input by the decoder 202 to the post-processor 300, and the post-processor 300 uses the LPSM to decode the decoded data. Perform loudness processing on the audio data of the LPSM, or implement the verification of the LPSM, and then (if the verification shows that the LPSM is valid) use the LPSM to perform loudness processing on the decoded audio data.
In some embodiments, if the decoder 201 accepts an audio bitstream with cryptographic hash generated according to an embodiment of the present creation, the decoder is constructed to parse and derive from a data determined by the bitstream. The cryptographic hash is obtained from the data block, and the data block contains Loudness Processing State Interpretation Data (LPSM). The verifier 203 can use the cryptographic hash to verify the received bit stream and/or associated interpretation data. For example, if the verifier 203 finds the LPSM to be verified based on the match between a reference cryptographic hash and the cryptographic hash obtained from the data block, it can notify a downstream audio processing unit (e.g., The post-processor 300, which may be or may include a volume leveling unit, is used to transmit (unchanged) the audio data. Additionally, optionally, or alternatively, other types of encryption techniques can be used to replace cryptographic hash-based methods.
In some embodiments of the decoder 100, the received (and temporarily stored in the memory 201) encoded audio bitstream is an AC-3 bitstream or an E-AC-3 bitstream , And include audio data fragments (eg, the AB0-AB5 fragments of the sound frame shown in Figure 4) and interpretive data fragments, where the audio data fragments are indicative of the audio data, and at least some of the interpretive data fragments Each piece of interpretation data includes Loudness Processing State Interpretation Data (LPSM). The decoder stage 202 is constructed to extract LPSM from the bit stream with the following format. Each piece of interpretative data including LPSM is included in the "addbsi" field of the bitstream information ("BSI") segment of a frame of the bitstream, or at the end of a frame of the bitstream In the auxiliary resource field of (for example, the AUX segment shown in Figure 4). The sound frame of the bit stream may include one or two interpretive data fragments, each interpretive data fragment includes LPSM, and if the sound frame includes two interpretive data fragments, then One is reproduced in the addbsi field of the sound frame and the other is in the AUX field of the sound frame. Each interpretive data segment including the LPSM includes an LPSM payload (or container) segment having the following format: a header (which typically includes a syncword used to identify the beginning of the LPSM payload) , Followed by at least one identity value (for example, the LPSM format, length, period, number, and substream associated values described in Table 2 below); and after the header, at least one dialogue indicator value (For example, the "dialog channel" parameter in Table 2), which shows whether the corresponding audio data indicates a dialogue or not (for example, which channels of the corresponding audio data indicate a dialogue); at least one loudness Adjust the coincidence value (for example, the "loudness adjustment type" parameter in Table 2), which shows whether the corresponding audio data is consistent with a set of pointed loudness adjustments); at least one loudness processing value (for example, the "dialog One or more of the "Gated Loudness Correction Flag" and "Loudness Correction Type" parameters), which display at least one loudness process that has been implemented on the corresponding audio data; and at least one loudness value (e.g., table 2 "ITU relative to the gated loudness", "ITU speech gated loudness", "ITU (EBU 3341) One or more of short-term 3s loudness" and "true peak" parameters), which displays at least one characteristic of loudness (for example, peak or average loudness) of the corresponding audio data.
In some embodiments, the decoder stage 202 is constructed to retrieve each interpretive data segment with the following format from the "addbsi" field or an auxiliary data field of a frame of the bitstream: Header (which typically includes a synchronization word used to identify the beginning of the interpretation data segment, followed by at least one identity value (for example, the core component version, length, and period, and the number of extension components listed in Table 1 below) , And the sub-streams associated with these values); and after the core header, at least one protection value (for example, the HMAC summary and automatic fingerprint value in Table 1), which interprets the data for the loudness processing state or the corresponding At least one of the decryption, authentication, or verification of at least one of the audio data is useful; and also after the core header, if the interpretive data segment includes LPSM, then the LPSM payload identity ("ID") and The LPSM payload size value treats the following interpretation data as the LPSM payload and shows the size of the LPSM payload.
The LPSM payload (or container) fragment (which preferably has the above-mentioned format) follows the LPSM payload identity and the LPSM payload size value.
More generally, the encoded audio bitstream generated by the preferred embodiment of this creation has a structure that provides a mechanism to mark interpretive data components and subcomponents as core (essential) components or extensions (not Necessary) components. This makes the data rate of the bit stream (including its own interpretation data) suitable for many applications. The core (necessary) components of the better bitstream syntax structure (syntax) should also be able to be signaled The extended (non-essential) component associated with the audio content is the message present (in-band) and/or at a remote location (out of band).
The core element is required to appear in each frame of the bit stream. Some sub-elements of the core elements are unnecessary and can appear in any combination. These expansion elements are not required to appear in every frame (to limit the bitrate overhead). Therefore, the expansion elements can appear in some sound frames but not in other sound frames. Some sub-elements of an expansion element are unnecessary and can appear in any combination, and some sub-elements of an expansion element may be necessary (ie, if the expansion element appears in a frame of the bit stream).
In one type of embodiment, an encoded audio bitstream containing a series of audio data fragments and interpretive data fragments is generated (eg, by an audio processing unit embodying the creation). The audio data fragments are representations of audio data, each of at least some of the interpretive data fragments includes Loudness Processing State Interpretation Data (LPSM), and the audio data fragments are divided by the interpretive data fragments Time-division multiplexed. In this type of preferred embodiment, each interpretive piece of information has a better format that will be described in this document.
In a preferred format, the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each segment of interpretation data including the LPMS is included (e.g., by the code A preferred implementation of device 100 Example stage 107 includes) in the "addbsi" field of the bitstream information ("BSI") segment of a frame of the bitstream (as shown in FIG. 6), or in a segment of the bitstream In the auxiliary data field of the sound frame.
In the preferred format, each frame includes a core element in the addbsi field of the frame, which has the format shown in Table 1 below:<tables><img alt="" file="twm467148u_d0001.tif" he="2221" id="" img-content="drawing" img-format="tif" inline="no" orientation="portrait" wi="2114" /></tables>
In the preferred format, each addbsi (or auxiliary data) field containing LPSM includes a core header (and optionally additional core components), and in the core header (or the core flag) Header and other core components) are followed by the following LPSM values (parameters): Payload ID (which refers to the interpretive data as LPSM), which follows the core components (for example, the core components specified in Table 1) ; Payload size (which indicates the size of the LPSM payload), which follows the payload ID; and LPSM data (which follows the payload ID and the payload size value), which has the following table ( Table 2) specified format:<tables><img alt="" file="twm467148u_d0002.tif" he="3040" id="" img-content="drawing" img-format="tif" inline="no" orientation="portrait" wi="2174" /></tables><tables><img alt="" file="twm467148u_d0003.tif" he="2027" id="" img-content="drawing" img-format="tif" inline="no" orientation="portrait" wi="2143" /></tables>
In another preferred format of the coded bitstream generated according to this creation, the bitstream is an AC-3 bitstream or an E-AC-3 bitstream, and each includes the LPSM Interpretation data fragments are included (eg, included by stage 107 of a preferred embodiment of the encoder 100) in the "addbsi" of the bitstream information ("BSI") fragment in a frame of the bitstream In the field (as shown in Figure 6); or in the auxiliary data field at the end of a frame of the bit stream (for example, the AUX shown in Figure 4) Fragment). A sound frame can include one or two interpretive data fragments, each interpretive data fragment includes LPSM, and if the sound frame includes two interpretive data fragments, then an interpretive data fragment is in the addbsi field of the sound frame and Another piece of interpretation data is in the AUX field of the sound frame. Each piece of interpretative data that includes the LPSM has the format defined in Table 1 and Table 2 above (that is, it includes the core components defined in Table 1, followed by the payload ID (this interpretive data refers to the LPSM) and the payload size value defined above, followed by the payload (the LPSM data has the format listed in Table 2).
In another preferred format, the encoded bitstream is a Dolby E bitstream, and each interpretive data segment including the LPSM is the first N sample positions of the Dolby E guard band interval. A Dolby E bitstream including this interpretive data segment including LPSM preferably includes a numerical representation of the length of the LPSM payload, which is formed as the signal in the Pd character of the SMPTE 337M preamble (the SMPTE The 337M Pa character repetition rate is better to keep the same as the related video frame rate).
In a preferred format where the encoded bitstream is an E-AC-3 bitstream, each interpretive data segment containing the LPSM is included (e.g., by a preferred implementation of the encoder 100 The example stage 107 includes) additional bitstream information in the "addbsi" field as a segment of bitstream information ("BSI") of a frame of the bitstream. Next, we will describe the additional aspects of encoding an E-AC-3 bitstream with this better format of LPSM: 1. During the generation of an E-AC-3 bit stream, when the E-AC-3 encoder (which inserts the LPSM value into the bit stream) is "active", for For each generated frame (synchronous frame), the bit stream should include an interpretation data block (including LPSM) carried in the addbsi field of the frame. The bits required to carry the interpretation data block should not increase the encoder bit rate (frame length); 2. Each interpretation data block (including LPSM) should contain the following information: 1. Loudness_Correction _Type_flag: where '1' indicates that the corresponding audio data's loudness is corrected upstream of the encoder, and '0' indicates that the loudness is embedded in the encoder by a loudness corrector (such as , The loudness processor 103 of the encoder 100 in FIG. 2) calibration; 2. Speech_Channel: It shows which source channels contain speech (in the previous 0.5 second period). If no speech is detected, it should be displayed as such; 3. Speech_loudness: It displays the combined speech loudness of each corresponding audio channel including the speech (in the previous 0.5 second period); 4. ITU_Loudness: It displays the combined ITU of each corresponding audio frequency BS.1770-2 Loudness; 5. Gain: used for inverted loudness composite gain in a decoder (to demonstrate reversibility); 3. In the E-AC-3 encoder (which inserts the LPSM value into the bit Streaming) is "active" and is receiving an AC-3 frame with a'trust' flag, the loudness controller in the encoder (such as The loudness processor 103 of the encoder 100 of FIG. 2 should be bypassed. The'trusted' source dialogue normalization value (dialnorm value) and the DRC value should be transmitted (e.g., the generator 106 of the encoder 100) to the E-AC-3 encoder component (e.g., the stage 107 of the encoder 100). The LPSM block generation continues and the loudness_correction_type_flag is set to '1'. The loudness controller bypass program must be synchronized with the beginning of the decoded AC-3 sound frame where the'trust' flag appears. The loudness controller bypass procedure should be implemented as follows: the leveler_number control is reduced from the value of 9 to the value of 0 during 10 audio blocks (ie, 53.3 milliseconds) and the leveler_backend _Counter control is put into detour mode (this operation should produce a seamless conversion result). The term'trusted' detour of the leveler means that the dialogue normalization value of the source bit stream is also reused in the output of the encoder. (For example, if the'trusted' source bit stream has a dialogue normalization value of -30, the output of the encoder should use -30 as the outbound dialogue normalization value); 4. When the E-AC-3 encoder (which inserts the LPSM value into the bit stream) is "active" and is receiving an AC-3 frame without the'trust' flag, set it The loudness controller in the encoder (for example, the loudness processor 103 of the encoder 100 in FIG. 2) should be active. The LPSM block generation continues and the loudness_correction_type_flag is set to '0'. The loudness controller bypass program must be synchronized with the beginning of the decoded AC-3 sound frame without the'trust' flag. The loudness controller bypass procedure should be implemented as follows: The leveler_ The quantity control is increased from the value of 0 to the value of 9 during 1 audio block (that is, 5.3 milliseconds) and the leveler_backend_counter control is put into the'active' mode (this operation should generate The result of seamless conversion and including a backend_counter merge reset); and 5. During encoding, a graphical user interface (GUI) should display the following parameters to the user: "Input audio program: [trusted/ Untrusted"-the status of this parameter is based on the presence of the "trust" flag in the input signal; and "Real-time loudness correction: [start/stop action]"-the status of this parameter is based on the code embedded in the code Whether the sound controller in the device is active.
When decoding an AC-3 or E-AC-3 bitstream with a (preferred format) LPSM (the LPSM is included in the bitstream information of each frame of the bitstream ("BSI") ) In the "addbsi" field of the fragment), the encoder should analyze (in the addbsi field) the LPSM block data and send all the captured LPSM values to the graphical user interface (GUI). The group of captured LPSM values is updated every frame.
In another preferred format of a coded bitstream generated according to this creation, the coded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, each including the The LPSM interpretation data segment is included (as included in stage 107 of a preferred embodiment of the encoder 100) as a bitstream information ("BSI") segment (or Aux segment) of a frame of the bitstream Additional bitstream information in the "addbsi" field (as shown in Figure 6). In this format (it refers to Table 1 and 2 above) Variations of the description format), each addbsi (or Aux) field that includes the LPSM contains the following LPSM values: the core components defined in Table 1, followed by the payload ID (which identifies the interpretation data as LPSM) And the size of the load, followed by the payload (LPSM data), which has the following format (which is similar to the necessary elements listed in Table 2 above): LPSM payload version: it is a 2-bit column Bit, which shows the version of the LPSM payload; dialchan: it is a 3-bit field, which shows which of the left, right, and/or center channels of the corresponding audio data contains the spoken dialogue. The bit allocation of the dialchan field can be as follows: bit 0 (which represents the dialogue system appears in the left channel) is stored in the most important bit of the dialchan field; and bit 2 (which represents the dialogue It appears in the central channel) is stored in the least important bit of the dialchan column. Each bit in the dialchan field is set to '1', if the corresponding channel contains spoken dialogue during the first 0.5 seconds of the program; loudregtyp: it is a 3-bit field, which displays this Which loudness adjustment standard the program loudness complies with. Setting the "loudregtyp" field to '000' means that the LPSM does not show loudness adjustment compliance. For example, a value in this field (e.g., 000) may indicate that the loudness adjustment standard is not complied with, and another value in this field (e.g., 001) may indicate that the audio data of the program complies with the ATSC A/85 standard. This field Another value of (eg, 010) can represent that the audio data of the program complies with the EBU R128 standard. In this example, if the field is set to something other than '000' If there is any value, in the payload, the loudcorrdialgat and loudcorrtyp fields should follow this field; loudcorrdialgat: a 1-bit field, which shows whether the loudness correction of the dialogue gate has been implemented. If the loudness of the program is corrected by using dialog gating, the value of the loudcorrdialgat field is set to '1'. Otherwise, it is set to '0'; loudcorrtyp: a 1-bit field that displays the type of loudness correction applied to the program. If the loudness of the program has been corrected by infinite look-ahead (file-based) loudness correction processing, the loudcorrtyp field is set to '0'. If the loudness of the program has been calibrated by a combination of real-time loudness measurement and dynamic range control, this field is set to '1'; loudrelgate: a 1-bit field that shows whether there is a relative gate-controlled loudness The data (ITU) exists. If the loudrelgate field is set to '1', then in the payload, a 7-bit ituloudrelgat field should follow this field; loudrelgat: a 7-bit field, which displays the relative gate Controlled and eye loudness (ITU). This field displays the combined loudness of the audio program, which is measured in accordance with ITU-R BS.1770-2, without any gain adjustment, because dialogue normalization and dynamic range compression are applied. The value from 0 to 127 is interpreted as -58LKFS to +5.5LKFS, in the 0.5LKFS step; loudspchgate: a 1-bit field, which shows whether the loudness data (ITU) of the speech gate control exists. If the loudspchgate field is set to '1', then in the payload, a 7-bit loudspchgat The field should follow this field; loudspchgat: a 7-bit field, which displays the gated and loudness of the speech. This field displays the combined loudness of the entire corresponding audio program, which is measured according to the formula (2) of ITU-R BS.1770-3 without any gain adjustment, because dialogue normalization and dynamic range compression are applied. Values from 0 to 127 are interpreted as -58LKFS to +5.5LKFS, in the 0.5LKFS step; loudstrm3se: a 1-bit field, which shows whether short-term (3 seconds) loudness data exists. If this field is set to '1', then in the payload, a 7-bit loudstrm3s field should follow this field; loudstrm3s: a 7-bit field, which displays the corresponding The un-gated loudness of the audio program in the first 3 seconds is measured in accordance with ITU-R BS.1771-1 without any gain adjustment, because dialogue normalization and dynamic range compression are applied. Values from 0 to 256 are interpreted as -116LKFS to +11.5LKFS, in the 0.5LKFS step; truepke: a 1-bit field, which shows whether the true peak loudness data exists. If the truepke field is set to '1', then in the payload, an 8-bit truepk field should follow this field; and truepk: an 8-bit field, which displays the The true peak sample value of the program is measured in accordance with Annex 2 of ITU-R BS.1770-3 without any gain adjustment, because dialogue normalization and dynamic range compression are applied. Values from 0 to 256 are interpreted as -116LKFS to +11.5LKFS, in 0.5LKFS step.
In some embodiments, the core element of an interpretive data segment in the auxdata field (or "addbsi" field) of a frame of an AC-3 bitstream or E-AC-3 bitstream includes a The core header (which typically includes the identity value, such as the core component version), and after the core header: whether the fingerprint data (or other protection value) of the interpretation data for the interpretation data fragment has a value included Indicates whether external data (which is related to the audio data of the interpretive data corresponding to the interpretive data segment) exists, and is used for each type of interpretive data identified by the core component (eg, LPSM, and/or a non- LPSM's interpretation data) payload ID and payload size value, and a protection value for at least one of the interpretation data identified by the core component. The interpretive data payload of the interpretive data segment follows the core header and (in some cases) is nested within the value of the core element.
The embodiments of the present invention can be embodied as hardware, firmware, or software, a combination of the two (eg, programmable logic array). Unless defined differently, the algorithms or processes included as part of this creation are not necessarily related to any particular computer or other equipment. In detail, various general-purpose machines containing programs written in accordance with the teachings of this creation can be used, or this creation can be more convenient to build more specialized equipment (such as integrated circuits) to implement all The required method steps. Therefore, the present creation can be embodied as one or more computer systems that can be programmed with one or more programs (such as the embodiment of any element in FIG. 1, or the encoder 100 (or its elements) in FIG. 2), or The decoder 200 (or its components) of FIG. 3, Or a computer program executed on the post-processor 300 (or its components) in FIG. 3, each of these systems includes at least one processor, at least one data storage system (which includes volatile and non-volatile memory and/or Storage element), at least one input device or input port, and at least one output device or output port. Code is applied to the output data to implement the functions described in this article and generate output information. The output information is applied to one or more output devices in a conventional manner.
Each such program can be implemented in any desired computer language (including machine language, assembly language, or high-level sequential, logical, or object-oriented programming language) to communicate with the computer system. In any case, the language can be a compiled or interpreted language.
For example, when implemented by a computer software instruction program, the various functions and steps of the embodiment of this creation can be implemented with a multithreaded instruction sequence that can be executed on appropriate digital signal processing hardware. In this example The various devices, steps, and functions of the embodiments can correspond to the parts of the software instructions.
Each such computer program is preferably stored in or downloaded to a storage medium or device (e.g., solid-state memory or medium, or magnetic medium or optical medium) that can be read by a general-purpose or special-purpose programmable computer, It is used to configure and operate the computer when the storage medium or device is read by the computer system to implement the procedures described herein. The created system can also be implemented as a computer-readable storage medium, which is configured with (eg, stored) computer programs, wherein the storage medium configured in this way causes the computer system to operate in a specific and predetermined manner Reality Apply the functions described in this article.
Several embodiments of this creation have been described. However, it will be understood that various changes can be achieved without departing from the spirit and scope of this creation. Based on the above teachings, this creation can have many modifications and changes. It should be understood that, within the scope of the following patent applications, this creation can be embodied in a way different from that explicitly described in this article.
3 sheets
Sheet 1 Sheet 2 Sheet 3
510 members in 28 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 201361754882 | United States of America | P | |
| 201361754882 | United States of America | P | |
| 61754882 | United States of America | – | |
| 201361754882P | – | – | – |
| US201361754882P | – | – | – |
Members510
| Document | Office | Kind | |
|---|---|---|---|
| CA2816889A1 | Canada | A1 | |
| CA2998405A1 | Canada | A1 | |
| CA3216692A1 | Canada | A1 | |
| WO2012075246A2 | World Intellectual Property Organization (WIPO) | A2 | |
| TW201236446A | Taiwan Province of China | A | |
| AR084086A1 | Argentina | A1 | |
| DE202013001075U1 | Germany | U1 | |
| AU2011336566A1 | Australia | A1 | |
| JP3183637U | Japan | U | |
| MX2013005898A | Mexico | A | |
| IL226100D0 | Israel | D0 | |
| SG190164A1 | Singapore | A1 | |
| DE202013006242U1 | Germany | U1 | |
| CN203134365U | China | U | |
| US2013246077A1 | United States of America | A1 | |
| EP2647006A1 | European Patent Office (EPO) | A1 | |
| JP3186472U | Japan | U | |
| KR20130111601A | Republic of Korea | A | |
| CL2013001571A1 | Chile | A1 | |
| CN103392204A | China | A | |
| TWM467148UThis record | Taiwan Province of China | U | |
| CN203415228U | China | U | |
| JP2014505898A | Japan | A | |
| CN103943112A | China | A | |
| CA2888350A1 | Canada | A1 | |
| WO2014113465A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014113471A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014113478A1 | World Intellectual Property Organization (WIPO) | A1 | |
| FR3001325A3 | France | A3 | |
| UA106163C2 | Ukraine | C2 | |
| WO2014124377A2 | World Intellectual Property Organization (WIPO) | A2 | |
| KR20140106760A | Republic of Korea | A | |
| KR101438386B1 | Republic of Korea | B1 | |
| TWM487509U | Taiwan Province of China | U | |
| TW201442020A | Taiwan Province of China | A | |
| WO2014124377A3 | World Intellectual Property Organization (WIPO) | A3 | |
| CA2898891A1 | Canada | A1 | |
| CN104240709A | China | A | |
| WO2014204783A1 | World Intellectual Property Organization (WIPO) | A1 | |
| FR3007564A3 | France | A3 | |
| KR20140006469U | Republic of Korea | U | |
| RU2013130293A | Russian Federation | A | |
| TW201506911A | Taiwan Province of China | A | |
| SG11201502405RA | Singapore | A | |
| IL237561D0 | Israel | D0 | |
| KR20150047633A | Republic of Korea | A | |
| AU2014207590A1 | Australia | A1 | |
| HK1198674A1 | Hong Kong, China | A1 | |
| CN104737228A | China | A | |
| FR3001325B3 | France | B3 | |
| MX2015004468A | Mexico | A | |
| AU2014281794A1 | Australia | A1 | |
| EP2901449A1 | European Patent Office (EPO) | A1 | |
| TWI496461B | Taiwan Province of China | B | |
| AU2014207590B2 | Australia | B2 | |
| AU2014281794B2 | Australia | B2 | |
| IN1633MUN2015A | India | A | |
| IN1765MUN2015A | India | A | |
| IN1766MUN2015A | India | A | |
| SG11201505426XA | Singapore | A | |
| IL239687D0 | Israel | D0 | |
| KR20150099586A | Republic of Korea | A | |
| KR20150099615A | Republic of Korea | A | |
| KR20150099709A | Republic of Korea | A | |
| KR200478147Y1 | Republic of Korea | Y1 | |
| AU2014281794B9 | Australia | B9 | |
| KR20150105955A | Republic of Korea | A | |
| CN104937844A | China | A | |
| CN104995677A | China | A | |
| MX2015010477A | Mexico | A | |
| JP2015531498A | Japan | A | |
| CN105027478A | China | A | |
| HK1204135A1 | Hong Kong, China | A1 | |
| US2015325243A1 | United States of America | A1 | |
| FR3007564B3 | France | B3 | |
| TW201543469A | Taiwan Province of China | A | |
| RU2568372C2 | Russian Federation | C2 | |
| EP2946469A1 | European Patent Office (EPO) | A1 | |
| EP2946495A1 | European Patent Office (EPO) | A1 | |
| US2015348558A1 | United States of America | A1 | |
| EP2954515A1 | European Patent Office (EPO) | A1 | |
| US2015363160A1 | United States of America | A1 | |
| US2015372820A1 | United States of America | A1 | |
| IL239687A | Israel | A | |
| TWI524329B | Taiwan Province of China | B | |
| JP2016507088A | Japan | A | |
| JP5879362B2 | Japan | B2 | |
| JP2016507779A | Japan | A | |
| TW201610984A | Taiwan Province of China | A | |
| KR20160032252A | Republic of Korea | A | |
| JP2016510544A | Japan | A | |
| MX338238B | Mexico | B | |
| CA2888350C | Canada | C | |
| CA2898891C | Canada | C | |
| CN103392204B | China | B | |
| MX339611B | Mexico | B | |
| HK1212091A1 | Hong Kong, China | A1 | |
| RU2568372C9 | Russian Federation | C9 | |
| EP2901449A4 | European Patent Office (EPO) | A4 | |
| UA111927C2 | Ukraine | C2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Expiration of patent term of a granted utility modelGrantedMK4K | MK4K |
Numbers
- Publication
- M467148
- Publication, DOCDB
- M467148
- Publication, EPODOC
- TWM467148U
- Application
- 102201885
- Application, DOCDB
- 102201885
- Application, EPODOC
- TW20132201885U
Titles3
- English
- AUDIO PROCESSING APPARATUS WITH LOUDNESS PROCESSING STATE METADATA
- Chinese
- 具響度處理狀態詮釋資料之音訊處理設備
- English
- Audio processing equipment with loudness processing state interpretation data
Classification
- CPC, 4
- H03G5/005
- G10L19/167
- H03G9/005
- H03G9/025
- IPC, 2
- G10L19 00
- G10L19 20