Metadata time marking information for indicating a section of an audio object
19 claims: 5 independent, 14 dependent
- 1A method of encoding time indicator information into audio data, wherein the audio data is a bit stream, the method is a joint bit by encoding the time indicator information as audio metadata in the audio data. The time indicator information comprises a step of forming a stream, the time indicator information indicates a plurality of divisions of an audio object in the audio data, and the time indicator information is a metadata container of the joint bit stream at a plurality of positions of the audio data. Encoded within, the plurality of positions occur at a particular rate of occurrence in the voice data bit stream, whereby the corresponding decoder is in the division of the voice object indicated by the time indicator information. A method capable of starting playback of the audio object from the beginning.
- 9The method comprises encoding the labeling information in the audio data, the labeling information labels the plurality of compartments of the audio object, and the labeling information is a meta of the joint bit stream. Claim 1 encoded as data Or 8 The method described in any one of the above.
- 17A method of decoding time indicator information in a joint bit stream that includes audio data and audio metadata, comprising decoding the time indicator information provided as said audio metadata in the joint bit stream. The indicator information indicates a plurality of divisions of the audio object encoded in the audio data, and the time indicator information is encoded in the metadata container of the joint bit stream at a plurality of positions of the audio data. The plurality of positions occur at a specific rate of occurrence in the audio data bit stream, thereby allowing the audio object to start playing from the beginning of the segment of the audio object indicated by the time indicator information. ,Method.
- 18An encoder configured to encode time indicator information as audio metadata in audio data, the audio data being a bit stream, thereby forming a joint bit stream, the time indicator information The time indicator information is encoded in the metadata container of the joint bit stream at a plurality of positions of the voice data to indicate a plurality of divisions of the voice object encoded in the voice data. The position occurs at a particular rate of occurrence in the audio data bit stream, which allows the corresponding decoder to start playing the audio object from the beginning of the segment of the audio object indicated by the time indicator information. Encoder.
- 19A decoder configured to decode time indicator information provided as audio metadata in a joint bit stream containing audio data, wherein the time indicator information is an audio object encoded in the audio data. The time indicator information is encoded in the metadata container of the joint bit stream at a plurality of positions of the voice data, and the plurality of positions are specific in the voice data bit stream. A decoder that occurs at an rate of occurrence, thereby allowing the decoder to start playing the audio object from the beginning of the segment of the audio object indicated by the time indicator information.
Independent claims5
117 paragraphs, as filed
The present application relates to audio coding, and more specifically to metadata within audio data, which indicates a section of an audio object.
A piece of music can often be recognized by listening to the characteristic parts of the piece of music (such as a refrain chorus). Also, in order to evaluate whether a music consumer likes or dislikes a music, it may be sufficient to listen to the characteristic parts of the music. If a music consumer is looking for a feature of a song stored as digital audio data, the music consumer must manually fast forward in the song to find the feature. This is cumbersome, especially when music consumers are browsing multiple songs in a large music collection to find a particular song.
A first aspect of the invention relates to a method for encoding time marking information in speech data.
Preferably, the encoded audio data containing the time indicator information is stored in a single audio file, such as in an MP3 (MPEG-1 Audio Layer 3) file or an AAC (Advanced Audio Coding) file. Will be done.
According to this method, the time indicator information is encoded as voice metadata in the voice data. The time indicator information indicates at least one segment of the voice object encoded in the voice data. For example, the time indicator information may specify only the start position and end position of the division, or the start position.
The at least one division may be a feature portion of the audio object. Such feature parts often allow the audio object to be instantly recognized by listening to the feature part.
The time indicator information encoded in such voice data makes it possible to instantly browse a specific section of the voice object. Therefore, it is avoided to manually search through the audio object to find a specific segment.
The time indicator information encoded in the voice data enables extraction of a specific division (for example, a feature division (particularly chorus)). The division can be used as a ringtone or an alarm signal. For this purpose, the division can be saved in a new file, or when the ringtone or alarm sound or alarm signal is played, from that particular division using the time indicator in the audio data. Playback can be started.
When the at least one division is a characteristic part (that is, an important part or a representative part) of a voice object, by using this labeled division together with the time indicator information, it is recognized instantly by listening. An audio thumbnail of the audio object that enables it is provided.
Even if the consumer device supports automatic audio data analysis for finding a particular category (eg, a feature category of a piece of music), such analysis for finding that category is not necessary. This is because the time indicator information has already been specified in advance and is included in the voice data.
Audio data can be pure audio data, multiplexed multimedia video / audio data (such as MPEG-4 video / audio bitstream or MPEG-2 video / audio bitstream), or such multiplexing. Note that it may be the audio part of the video / audio data.
The time indicator information may be encoded at the time of generation of the voice data, or the time indicator information may be included in the given voice data.
The audio data output from the encoder or the audio data input to the audio decoder typically forms a bitstream. Therefore, the term "bitstream" may be used in place of the term "audio data" throughout this application. The encoded voice data including the time indicator information is preferably stored in a single file stored on the storage medium.
Nevertheless, the encoded audio data (in other words, the encoded bitstream) has a separate file, one audio file with audio information and one or more time markers. May be generated by multiplexing information from a single metadata file.
Audio data may be used in streaming applications such as Internet radio bitstreams or multimedia bitstreams including video and audio. Alternatively, the audio data may be stored in a consumer-side storage medium (such as flash memory or a hard disk).
Preferably, the audio object is encoded by a perceptual coding scheme (such as MP3, Dolby Digital, or the coding method used in (HE-) AAC). Alternatively, the audio object may be a PCM (Pulse Code Modulation) encoded audio object.
For example, the audio object may be a recording of a piece of music or speech (such as an audiobook).
Preferably, the coding of the time indicator information allows forward compatibility. That is, the time indicator information is encoded in such a way that the decoder that does not correspond to the time indicator information can skip the time indicator information.
Preferably, both backwards and forwards compatibility are achieved. Backward compatibility means that a decoder that supports time-labeled information (eg, a HE-AAC decoder with an extractor and processor for time-marked metadata) does not contain time-marked information in conventional audio data (eg, a HE-AAC decoder). For example, it means that both conventional HE-AAC bitstreams) and audio data with time-marked information (eg, HE-AAC bitstreams with additional time-marked metadata) can be read. Forward compatibility means that a decoder that does not support time indicator information (for example, a conventional HE-AAC decoder) has conventional audio data that does not include time indicator information and conventional audio data that includes time indicator information. It means that both parts of the expression can be read (in this case, the time indicator information is skipped because it is not supported).
According to one embodiment, the time indicator information indicates the position of a feature portion of the speech object. For example, in the case of a musical piece, the time indicator information may indicate a chorus, a refrain, or a part thereof. In other words, the time-labeled metadata points to important or representative parts. As a result, a music player that decodes the audio bitstream can start playing at a critical moment.
The time indicator information may indicate multiple compartments within the audio object (eg, within a piece of music or audiobook). In other words, the time indicator information may include a plurality of time indicators associated with a plurality of sections of the voice object. For example, the time indicator information may indicate the time position of the start point and end point of a plurality of divisions. As a result, it is possible to browse to various sections within the audio object.
The time indicator information may specify various temporal positions related to the temporal musical structure of the music. In other words, the time indicator information may indicate a plurality of divisions in the music, and the plurality of divisions are related to divisions having different temporal and musical structures. For example, the time indicator information may indicate the beginning of one or more of the following categories: That is, the introductory part, the first lyrics, the first refrain or chorus, the second (third) lyrics, the second (third) refrain or chorus, or the bridge.
The time indicator information may also indicate the motive, subject and / or variant of the subject in the song.
In addition, the time indicator information may specify other musical aspects (such as the generation of a singing voice (eg, the first vocal entry)) or the musical composition (the generation of a particular instrument (particularly, particularly)). It may be related to the appearance of a solo of a particular instrument) or a group of instruments (eg, brass section, back vocals) or the loudest part of the song.
The time indicator information may also indicate a section having a particular musical characteristic. Musical characteristics may be, for example, a particular musical style or genre, a particular mood, a particular tempo, a particular tonality, a particular articulation.
The time indicator indicator may be associated with the labeling information used to label the indicator. For example, the labeling information may describe a particular musical characteristic of the segment. Specific musical characteristics include a musical style or genre designation (eg, soft, classic, electronic, etc.), an associated mood designation (eg, happy, sad, aggressive), tempo (eg, per minute). The speed or pace of the audio signal specified by the number of beats or labeled by musical terms such as Allegro, Andante, etc., the tonality of that division of the audio signal (eg, in a major, c minor), or Articulations (eg portato, legato, pichikart).
The labeling information may be contained in another metadata field. The labeling information may include text labels. Alternatively, for labeling purposes, the time indicator may be associated with an index in a table that specifies, for example, a musical structure or musical characteristic as described above. In this case, the index of each label is included in the audio data as labeling information. An example of such a reference table [lookup table] is shown below.<tables num="1"><img id="000002" he="39" wi="123" file="JP5771618B2_D0001.tif" img-format="tif" img-content="drawing" /></tables>
In this example, only the index (here 4 states, thus 2 bits) will be carried within the audio bitstream. The reference table is typically known to the decoder. However, it is also possible to carry the table within the audio bitstream.
The time indicator information and one or more labels associated with the time indicator information (eg, a label encoded in the metadata as a text label or as an index of a reference table that allows label extraction from the table). Together, it allows the user to easily browse through a large database of audio objects (such as a large collection of songs) to find a particular part (eg, a guitar solo).
Time-marked information may also allow loop playback over parts of interest (eg, guitar solos, vocal parts, refrains), which may allow rehearsals and practice of the instrumental or vocal parts of the song. Is facilitated.
The time indicator information may be stored as metadata in an audio file (eg, AAC file or MP3 file) and the time information (eg, the start and end points of a particular segment, or the start and end points of a particular segment). The continuation length) may be encoded in one or more of the following formats. · Seconds (eg 20 seconds) and optionally fractional seconds (eg 0.2 seconds) Sample number (for example, a 28-bit wide sample number field covers a length of more than 1 hour at a sampling rate of 44100 Hz). · Frame numbers (for example, at a sampling rate of 44100 Hz and at 1024 samples / frame, the 18-bit wide frame number field covers lengths greater than 1 hour). -Integer frame number and integer sample number, or Integer frame numbers and fractional frame values (for example, adding a 2-bit wide fractional frame value to an 18-bit wide frame counter results in 5 ms accuracy at a sampling rate of 44100 Hz and 1024 samples / frame. ).
The degree of accuracy of the various formats for coding the above time information varies. The format used typically depends on the requirements of the application. For "chorus finder" applications, time resolution is not so important, so the format does not need to be highly accurate either. However, for "practicing an instrument to a song" application that uses very strict loops, the time resolution requirement may also be high, and therefore a high precision format is preferably used.
The time indicator metadata may be included (eg, once) at the beginning of the audio data (eg, the header of the audio bitstream).
Alternatively, the time indicator information may be encoded in a plurality of sections of the voice data. For example, multiple compartments may occur in a bitstream with a particular incidence (eg, every n seconds or every n audio frames (n 1, eg, n = 1)). In other words, the time indicator information may be encoded at a particular fixed update rate.
When encoding the time indicator information in a plurality of sections, the time indicator information in a given section among the plurality of sections may be specified in relation to the occurrence of the given section in the bitstream. .. In other words, the time designation of the time indicator can be specified in relation to the time when the metadata is inserted. For example, the time indicator may specify the temporal distance between the regularly spaced metadata update positions and the segment of interest (eg, 3 seconds before the chorus of the audio signal begins).
Inclusion of time indicator information at a particular update rate in this way facilitates browsing functionality for streaming applications (eg, broadcasting).
Further embodiments of the coding method are described in the independent claims.
A second aspect of the application relates to a method of decoding the time indicator information provided in the voice data. According to this method, the time indicator information provided as voice metadata is decoded. This decoding is typically done with the decoding of the audio object given in the audio data. The time indicator information indicates at least one segment (eg, the most characteristic part) of the voice objects encoded in the voice data, as described above in connection with the first aspect of the present invention.
The above description relating to the coding method according to the first aspect of the present application also applies to the decoding method according to the second aspect of the present application.
According to one embodiment, after decoding the time-labeled information, reproduction begins at the beginning of the labeled section. The beginning of the labeled section is specified by the time indicator information. To start playback from the beginning of the labeled section, the decoder may start decoding from the labeled section. Playback from the beginning of the labeled section may be initiated by user input. Alternatively, playback may start automatically (for example, in the case of playback of feature portions of a plurality of songs).
Preferably, the regeneration of the section is stopped at the end of the section. The end is indicated by time indicator information. In the loop mode, it is possible to resume playback from the beginning of the division.
Decoding of the time indicator information and reproduction from the beginning of each division may be performed for a plurality of voice objects. This makes it possible to browse through multiple songs (eg, browse the most distinctive parts of multiple songs in a large music collection).
The encoded time indicator information indicating the feature portion of the music also facilitates browsing of various radio channels (eg, various Internet radio channels).
In order to browse various radio channels, the time indicator information in the plurality of audio bitstreams associated with the plurality of radio channels is decoded. Playback begins at the beginning of each of at least one segment indicated by the time indicator information for each bitstream, one for each of the plurality of bitstreams. Therefore, according to this embodiment, the characteristic division (or the characteristic division of a plurality of songs) of the songs on the first radio channel may be reproduced. After that, the characteristic division (or the characteristic division of a plurality of songs) of the songs on the second radio channel (and then on the third radio channel) may be reproduced. This allows radio consumers to get an impression of the type of music being played on a variety of radio channels.
This method may also be used to reproduce a medley of various songs being played on a given radio channel. In order to generate such a medley, the time indicator information of a plurality of audio objects in the bitstream of the radio channel is decoded. Each section of each audio object is played, one for each of the plurality of audio objects. The method may also be performed on multiple radio channels. This makes it possible to play a medley of songs for each of a plurality of radio channels and provide an impression of what kind of music is being played on the various channels.
The concepts described above may be used in connection with both real-time radio and on-demand radio. In the case of real-time radio, the user typically cannot jump to a particular point in the radio program (in real-time radio, the user sometimes jumps to a past point in the radio program depending on the buffer size. It is possible). For on-demand radio, the listener can start and stop at any point in the radio program.
In the case of real-time radio, the playback device preferably has the ability to store a particular amount of music in memory. By decoding the time indicator information, the device captures important parts of each of the last one or more songs on one or more radio channels and stores these important sections in memory for later playback. You may. The playback device may record the continuous audio stream received by the radio channel, optionally delete non-essential parts (to free up memory) later, or the playback device may You may record the important part directly.
The same concept can be used for television over the Internet.
According to certain embodiments, the labeled compartment may be used as a ringtone or alarm signal. For this purpose, the division may be stored in different files used for playing the ringtone or alarm signal, or the time indicator information indicating the division may be used to indicate the ringtone or alarm. For signal reproduction, reproduction may be started from the beginning of the division.
A third aspect of the present application relates to a encoder configured to encode time indicator information as audio metadata in audio data. The time indicator information indicates at least one segment of the voice object encoded in the voice data.
The above description relating to the coding method according to the first aspect of the present application also applies to the encoder according to the third aspect of the present application.
A fourth aspect of the application relates to a decoder configured to decode time indicator information provided as audio metadata in audio data. The time indicator information indicates at least one segment of the voice object encoded in the voice data.
The above description relating to the decoding method according to the first aspect of the present application also applies to the decoder according to the fourth aspect of the present application.
The decoder may be used in a voice player (eg, a music player such as in a portable music player having flash memory and / or a hard disk). The term "portable music player" also covers mobile phones with music player functionality. If the audio decoder allows browsing through the songs by playing back each feature portion of each song, the display displaying the song titles may be omitted. In that case, the size of the music player can be further reduced and the device cost can be reduced.
A fifth aspect of the application relates to audio data (eg, audio bitstreams). The voice data includes time indicator information as voice metadata. The time indicator information indicates at least one segment of the voice object encoded in the voice data. The voice data may be a bitstream (such as (Internet) radio bitstream) that is streamed from the server to the client (ie, the consumer). Alternatively, the audio data may be contained within a file stored on a storage medium (such as flash memory or a hard disk). For example, the audio data may be AAC (Advanced Audio Coding), HE-AAC (High Efficiency AAC), Dolby Pulse, MP3 or Dolby Digital bitstreams. Dolby Pulse is based on HE-AACv2 (HE-AAC version 2), but provides additional metadata. Throughout this application, the term "AAC" includes all extended versions of AAC (such as HE-AAC or Dolby Pulse). The term "HE-AAC" (as well as "HE-AACvl" and "HE-AACv2") also covers Dolby Pulse. The audio data may be multimedia data that includes both audio and video information.
Hereinafter, the present invention will be described with reference to the accompanying drawings by various exemplary examples.
<figref num="1">It is a figure which shows the schematic embodiment of the encoder which encodes the time indicator information.</figref><figref num="2">It is a figure which shows the schematic embodiment of the decoder which decodes the time indicator information.</figref>
The various uses of metadata time information are discussed below. Metadata time markers may indicate different types of divisions and may be used in different applications.
<u style="single">Metadata that indicates the characteristic part of the song (for example, chorus) Time indicator information</u>
Time indicator information may be used to indicate a feature portion of the song (eg, chorus, refrain or part thereof). A song is often easier to recognize by listening to a feature (eg, a chorus) than by reading the song title. By using the metadata time indicator indicating the characteristic part of the song, it is possible to search for the song that is known, and it becomes easy to browse by listening through the song database. Music consumers can instantly recognize and identify songs by listening to the most important parts of each song. In addition, such functionality is available when browsing songs on a portable music player device that has no display at all, or when the device is currently invisible to the user because it is in a pocket or bag. Very convenient.
The time indicator information indicating the characteristic part of the song is also useful when discovering a new song. By listening to the feature part (for example, chorus), the user can easily judge whether he likes or dislikes the song. Thus, based on listening to the most distinctive parts, the user may decide whether he or she wants to listen to the entire song, or whether he or she wants to pay to buy the song. it can. This functionality is useful, for example, in the applications of music stores and music discovery services.
<u style="single">Metadata related to the temporal and musical structure of a song Time indicator information</u>
The time indicator information specifies various temporal positions related to the temporal musical structure of the song (for example, to indicate the position of an intro, lyrics, refrain, bridge, another refrain, another lyrics, etc.). May be used for.
This allows the user to easily browse different parts of the song during the song. For example, the user can easily browse the part of the song that the user likes.
Metadata time indicator information related to musical structure is also useful for practicing musical instruments or singing. Such time indicator information provides the possibility of navigating through different parts of the song, thereby accessing the section of interest and once while practicing the instrument or singing. It can be played only or in a loop.
<u style="single">Metadata related to the occurrence of a particular instrument or singing voice Time indicator information</u>
The time indicator information may also be used to specify the generation of a particular instrument or the generation of a singing voice (and optionally a pitch range). Such time indicator information is useful, for example, in musical instrument or singing practice. When the user is learning to play an instrument (eg, a guitar), the user can easily find the part of the song that he or she wants to play (eg, a guitar solo). For singers, it is useful to find the part of the desired pitch range in the song.
<u style="single">Metadata time indicator information that indicates a division with a specific musical characteristic</u>
To find a section with a musical description of a particular musical characteristic, such as an articulation (eg, legato, pizzicato), style (eg, Allegro, Andante) or tempo (eg, beats per minute). , Time indicator information may be used. This may help, for example, in practicing a musical instrument. This is because the user can easily find the relevant and interesting parts of the song to practice. Reproduction may be looped across such specific segments.
<u style="single">Metadata time indicator information that indicates a division with a specific mood or tempo</u>
Metadata time indicator information may indicate a section with a particular mood (eg, energetic, aggressive, or mild) or tempo (eg, beats per minute). Such metadata helps to find the part of the song that corresponds to the mood. The user can search for song categories in a particular mood. This also makes it possible to create a medley with a division from multiple songs or all available songs according to a particular mood.
Such metadata may be used to find suitable music for exercise (eg, running, spinning, home trainer, or aerobics). Metadata may also make it easier to adapt the music to the training intensity level when training at different levels of intensity. Therefore, using such metadata helps the user align a particular planned workout with the appropriate music. For example, in the case of interval training (alternately short high-intensity workouts followed by breaks), energetic, aggressive or fast divisions are regenerated during the high-intensity periods, while During the break period, the gentle or slow division is regenerated.
In the various uses of the metadata time information as described above, the time indicator information is preferably integrated into the audio file (eg, in the header of the song file, instead of file-based usage. The metadata time indicator information may also be used within the context of a streaming application (eg, a radio streaming application (eg, via the Internet)), eg, a feature portion of a song (eg, a chorus or one of them). If there is metadata time indicator information indicating a part), such metadata can be used in the context of browsing various radio channels. Such metadata can be used for multiple radio stations (eg, internet radio). ), And facilitates browsing various radio channels on devices capable of storing a certain amount of music in memory (eg, on a hard disk or flash memory). An important part of a song. By signaling the position of (eg, chorus), the device has the last few songs (eg, for the last n songs; n 1) for more than one of those channels. For example, n = 5) each important part can be determined. The device may capture these important parts and keep these compartments in memory (and to free the memory). You may remove the rest of the last few songs). You listen to each channel through this collection of choruses, what kind of music is being broadcast from that channel, and whether you like it or not. You can easily get an overview of the data.
<u style="single">Metadata time indicator information that indicates a particular segment of a voice object</u>
Time indicator information may be used to indicate a particular division of audio objects (eg, audiobooks, audio podcasts, educational materials) that include speech and optional music and optional sounds. These compartments can be related to the content of the audio object (for example, specifying chapters or theater scenes in an audiobook, specifying some segments that give a summary of the entire audio object, and so on). These divisions can also be related to the characteristics of the audiobook (eg, in an audiobook that is a collection of stories, indicating whether a division is cheerful or not). In the case of audio materials for education, the time indicator information may indicate various parts of the audio object with respect to the difficulty of the material. In addition, the time indicator information in the educational material may indicate a category that requires the active participation of the learner (for example, comprehension problem in the language course, pronunciation exercise).
Metadata After discussing the various exemplary uses of time-marked information, we discuss the exemplary sources of time-marked information. The time indicators written in the metadata may come from one or more of the following sources, for example:
-Automatic extraction (eg, by a Music Information Retrieval (MIR) algorithm or service on the consumer side (ie, client side) or music provider side (ie, server side)). Examples of automatic extraction algorithms are discussed below. "A Chorus-Section Detection Method for Musical Audio Signals and Its Application to a Music Listening Station" (Masataka Goto, IEEE Transactions on Audio, Speech and Language Processing, Vol.14, No.5, pp.1783-1794, 2006 September), and "To Catch a Chorus: Using Chroma-Based Representations for Audio Thumbnailing" (MA Bartsch, MA) and GH Wakefield, IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, 2001). These documents are incorporated herein by reference.
-Transfer from an external database. For example, the audio library may be synchronized with an external database. Data may be fetched remotely because the external database hosting the metadata may be accessible, for example, through a computer network or cellular network (Gracenote's Compact Disc Database). (Same as for CDs that get artist / track information from (CDDB))
· Manually entered in the editor on the client side (ie by the consumer).
In the following, we discuss a variety of exemplary metadata containers for carrying metadata time indicator information. Carrying metadata over voice or multimedia bitstreams can be done in a number of ways. It may be desirable to include such data in a forward-compatible manner (ie, non-destructively for decoders that do not support the extraction of time-labeled metadata). One of the following commonly used metadata embedding methods may be used to embed the metadata in the audio data.
<u style="single">ID3 container</u>
ID3 tags (ID3-"Identify an MP3") are metadata containers often used with MP3 (MPEG-1 / 2 Layer III) audio files. The embedding is rather simple, as it basically inserts the ID3 tag at the beginning of the file (for ID3v2) or appends it at the end (for ID3v1). In particular, ID3 tags have become the de facto standard for MP3 players, so forward compatibility is usually achieved. Unused data fields in the ID3 tag may be used (or data fields for different uses may be diverted from their intended use) or ID3 to carry the time indicator. Tags may be extended by one or more data fields for time-marked transport.
<u style="single">MPEG-1 / 2 auxiliary data</u>
The MPEG-1 or MPEG-2 layer I / II / III audio bitstream provides an auxiliary data container that may be used for time-labeled metadata. These auxiliary data containers are described in the standardized literature ISO / IEC 11172-3 and ISO / IEC 13818-3. These are incorporated herein by reference. Such ancillary data containers are signaled in a fully forward compatible manner by the "AncDataElement ()" bitstream element, which allows variable size data containers. If the decoder does not support time indicator information, the decoder typically ignores this additional data. This data container mechanism makes it possible to transmit metadata at any frame of the bitstream.
<u style="single">Extended payload in MPEG-2 / 4 AAC bitstream</u>
MPEG-2 or MPEG-4 For AAC (advanced audio coding) audio bitstreams, use AAC's "extension_payload ()" mechanism as described in the standardized documents ISO / IEC 13818-7 and ISO / IEC 14496-3. The time indicator information may be stored in the data container. These documents are incorporated herein by reference. This approach can be used not only in basic AAC, but also in extended versions of AAC (such as HE-AACv1 (high efficiency AAC version 1), HE-AACv2 (high efficiency AAC version 2) and Dolby Pulse). It is possible. This "extension_payload ()" mechanism is signaled in a fully forward compatible way that allows variable size data containers. If the decoder does not correspond to the time indicator information encoded by the "extension_payload ()" mechanism, the decoder typically ignores this additional data. This data container mechanism makes it possible to transmit metadata at any frame of the bitstream. Therefore, the metadata may be updated continuously (eg, for each frame). A detailed example of the integration of time indicator information into an AAC bitstream will be described later in this application.
<u style="single">ISO-based media file format (MPEG-4 Part 12)</u>
Alternatively, an ISO-based media file format (MPEG-4 Part 12), as specified in ISO / IEC 14496-12, may be used. This container standard already has a hierarchical substructure for metadata. Metadata can include, for example:
-iTunes metadata, -The "extension_payload ()" element as part of the MPEG-4 AAC audio bitstream as discussed above, or -Customized metadata partition.
This ISO-based media file format may be used to include such time-labeled metadata in the context of Dolby Digital audio data or Dolby Pulse audio data or other audio data formats. For example, time-labeled metadata may be added to the Dolby pulse audio bitstream, which further differentiates Dolby pulse from traditional HE-AAC.
The hierarchical structure specified in ISO / IEC 14496-12 can be used to include metadata specific to, for example, Dolby Pulse or Dolby Media Generators. This metadata is carried in mp4 files within the "moov" atom. The "moov" atom includes the user data atom "udta". The user data atom "udta" is a unique ID (universal unique identifier). identifier)-By using "uuid"), identify what you are carrying. The box contains several metaatoms, each of which can carry different types of metadata. The type of metadata is specified by the handler "hdlr". Existing types may carry information about, for example, song titles, artists, genres, and so on. For example, it may be possible to specify a new type of extended markup language (XML) structure that contains the required information. The exact format is determined based on the information you want to send. The example below shows the structure in which the time indicator metadata is part of an atom named "xml_data".<tables num="2"><img id="000003" he="41" wi="123" file="JP5771618B2_D0001.tif" img-format="tif" img-content="drawing" /></tables>
The time indicator metadata atom "xml_data" coded in XML format can be structured as shown in the example below.<tables num="3"><img id="000004" he="69" wi="123" file="JP5771618B2_D0001.tif" img-format="tif" img-content="drawing" /></tables>
Such atoms can contain information about their size. That is, the parser that does not recognize the type can skip the division and continue the analysis of the subsequent data. Therefore, forward compatibility is achieved.
<u style="single">Other formats for metadata</u>
Other multimedia container formats that support metadata and may be used to transport time-labeled metadata are widely used industry standards (MPEG-4 Part 14 (also known as MP4, standardized literature ISO / IEC 14496). (Specified in -14) and 3GP format etc.).
Below are two examples of integrating time-labeled metadata into bitstream syntax.
<u style="single">First example of audio thumbprint bitstream syntax</u>
Some metadata container formats define the use of text strings (for example, in an extended markup language (XML) framework), while other metadata container formats are simply common for binary data chunks. It is a container. Table 1 below shows an example of a binary format bitstream specified by the pseudo-C syntax (which is a common practice in ISO / IEC standards). Bitstream elements larger than one bit are usually written / read as an unsigned-integer-most-significant-bit-first (uimsbf) with the most significant bit first.<tables num="4"><img id="000005" he="121" wi="136" file="JP5771618B2_D0001.tif" img-format="tif" img-content="drawing" /></tables>
These bitstream elements have the following meanings:
The integer element "BS_SECTION_ID" is, for example, 2 bits in length and describes the type of content of the labeled section (eg 0 = chorus, 1 = lyrics, 2 = solo, 3 = vocals).
The integer element "BS_NUM_CHAR" has, for example, 8 bits in length, and the length of the text string "BS_ARTIST_STRING" is described in bytes. In this example, the integer element "BS_NUM_CHAR" and the text string "BS_ARTIST_STRING" are used only in special cases (ie, when the integer element "BS_SECTION_ID" indicates vocal inclusion). See the statement "if (BS_SECTION_ID == 3)" in the pseudo-C syntax.
The text string element "BS_ARTIST_STRING" contains the name of the vocal artist in the labeled section. The text string may be coded in, for example, 8-bit ASCII (eg, UTF-8 as specified in ISO / IEC 10646: 2003). In this case, the bit length of the text string is 8 × BS_NUM_CHAR.
The integer element "BS_START" indicates the start frame number of the labeled section.
The integer element "BS_LENGTH" indicates the length of the labeled partition (here, represented by the number of frames).
An example of a bitstream based on the above pseudo-C syntax is "11 00001101 01000001 01110010 01110100 00100000 01000111 01100001 01110010 01100110 01110101 01101110 01101011 01100101 01101100 001010111111001000 01100001101010".
The above exemplary bitstream specifies:
The VOCAL_ENTRY section with the text tag "Art Garfunkel" starts at frame number 45000 and has a continuation length of 6250 frames (thus this section stops at frame 51250).
<u style="single">Second example of voice thumbprint bitstream syntax</u> The second example is based on the first example and uses the extension_payload () mechanism from ISO / IEC 14496-3. The syntax of the extension_payload () mechanism is described in Table 4.51 (Subordinate section 4.4.2.7, ISO / IEC14496-3: 2001 / FDAM: 2003 (E)). This is incorporated herein by reference.
Compared with the syntax of the extension_payload () mechanism in Table 4.51 (Subordinate Section 4.4.2.7, ISO / IEC14496-3: 2001 / FDAM: 2003 (E)), in the second example, as shown in Table 2. Adds an additional extension_type (ie, an extension_type called "EXT_AUDIO_THUMBNAIL") to the syntax of extension_payload (). If the decoder does not support this additional extension_type, this information is typically skipped. In Table 2, additional bitstream elements for audio thumbprints are underlined. The extension type "EXT_AUDIO_THUMBNAIL" is associated with the metadata "AudioThumbprintData ()", and Table 3 shows an example of the syntax of "AudioThumbprintData ()". The syntax of "Audio ThumbprintData ()" in Table 3 is similar to the syntax in Table 1. The rules for the bitstream elements "BS_SECTION_ID", "BS_NUM_CHAR", "BS_ARTIST_STRING", "BS_START" and "BS_LENGTH" are the same as those discussed in connection with Table 1. The variable "numAuThBits" counts the number of additional bits associated with AudioThumbprintData ().
The variable "numAlignBits" corresponds to the required number of fill bits, including the total number of bits of extension_payload (variable "cnt" (unit: bytes)), voice thumbprint (variable "numAuThBits"), and variable "extension type" (this). Is determined as the difference from the number of bits used in extension_payload () to identify the extension type). In this given example, "numAlignBits" is equal to 4, "AudioThumbprintData ()" returns the total number of bytes read.<tables num="5"><img id="000006" he="192" wi="139" file="JP5771618B2_D0001.tif" img-format="tif" img-content="drawing" /></tables><tables num="6"><img id="000007" he="151" wi="137" file="JP5771618B2_D0001.tif" img-format="tif" img-content="drawing" /></tables>
FIG. 1 shows an exemplary embodiment of an encoder 1 for coding time indicator information. The encoder receives the audio signal 2. The audio signal 2 may be a PCM (Pulse Code Modulation) encoded audio signal 2 or a perceptually encoded audio bitstream (MP3 bitstream, Dolby Digital bitstream, conventional HE-AAC bitstream or Dolby). It may be (such as a pulse bitstream). The audio signal 2 is the aforementioned audio bit extended by a multimedia transport format (eg, such as "MP4" (such as MPEG-4 Part 14) or a metadata container (such as "ID3")). It may be in any of the stream formats. The audio signal 2 includes an audio object (eg, a musical piece). The encoder 1 further receives the time indicator data 7. The time indicator data 7 indicates one or more divisions (such as the most characteristic part) in the speech object. The time indicator data 7 may be automatically identified, for example, by a music information retrieval (MIR) algorithm, or may be manually entered. The encoder 1 may further receive labeling information 8 for labeling one or more labeled compartments.
Based on signals 2 and 7 and optionally signal 8, the encoder 1 contains a voice object and a bitstream 3 containing time indicator information for marking one or more compartments in the voice object. To generate. The bitstream 3 may be an MP3 bitstream, a Dolby Digital bitstream, a HE-AAC bitstream or a Dolby pulse bitstream, where the bitstream optionally associates the time indicator information with the labeling information. It differs from conventional bitstreams in that it includes. Bitstream 3 is either a multimedia transport format (eg, "MP4" (MPEG-4 Part 14)) or the aforementioned audio bitstream format extended by a metadata container (eg, "ID3"). Bitstream 3 may be stored as audio files in a storage medium (not shown) (such as flash memory or hard disk) for later playback, or a streaming application (such as a flash memory or hard disk). It may be streamed on (such as internet radio).
Bitstream 3 may include header section 4. The header section 4 may include a time indicator metadata section 5. The time indicator metadata section 5 has encoded time indicator information and associated labeling information. The time indicator information may include a start point and a stop point for one or more labeled sections, or each start point and each continuation length for one or more labeled sections. The time-marked metadata section 5 may be included in the metadata container as described above. Bitstream 3 further includes audio object 6. Thus, time information for one or more segments is included in the bitstream metadata, which allows navigation to important parts of the audio object, for example.
FIG. 2 shows an exemplary embodiment of the decoder 10. The decoder 10 is configured to decode the bitstream 3 generated by the encoder 1. The decoder 10 generates an audio signal 11 based on a bitstream 3 (such as a PCM audio signal 11). The decoder 10 is typically part of a consumer device for audio playback (especially music playback). Consumer devices are such as portable music players without mobile phone functionality, mobile phones with music player functionality, notebooks, set-top boxes, or DVD players. Consumer devices for audio reproduction may be utilized for combined audio / video reproduction. The decoder 10 further receives the selection signal 13. Depending on the selection signal 13, the decoder 10 jumps to the labeled section of the audio object to decode the labeled section, or ends the normal decoding of the audio object from the beginning of the audio object. Do up to. If the decoder jumps to the labeled section of the audio object, the consumer device starts playing from the labeled section.
The decoder 10 may optionally further output the decoded labeling information 12. The decrypted labeling information 12 may be input to the display driver (not shown) so that it appears on the display of the device.
In the present specification, a method and a system for encoding time indicator information as metadata in voice data are described. This time indicator information allows music consumers to quickly identify feature parts of audio files.
The methods and systems described herein may be implemented as software, firmware and / or hardware. Certain components may be implemented as software running on, for example, a digital signal processor or microprocessor. Other components may be implemented, for example, as hardware or as a purpose-built integrated circuit. The signals that emerge in the described methods and systems may be stored on media (such as random access memory or optical storage media). These may be transferred over a network (such as a radio network, satellite network, wireless network, or wired network (eg, the Internet)). Typical devices that utilize the methods and systems described herein are portable electronic devices or other consumer devices used for storing and / or rendering audio signals. These methods and systems may be used on a computer system (eg, an internet web server) that stores and provides audio signals (eg, music signals) for download.
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both waysCites: the store holds 5 of 6
| Document | Relation | Office |
|---|---|---|
| JP2006163063A | Cites | Japan |
| WO2009101703A1 | Cites | World Intellectual Property Organization (WIPO) |
| JP2007248895A | Cites | Japan |
| JP2007520727A | Cites | Japan |
| JP2000206973A | Cites | Japan |
| Eoin Brazil,Cue Point Processing: An Introduction,Proc. COST G-6 Conference on Digital Audio Effects(DAFX-01),IE,2001年12月 6日 | Non-patent | – |
9 members in 5 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 25278809 | United States of America | P | |
| 61252788 | United States of America | – | |
| 2010065463 | European Patent Office (EPO) | W | |
| 61252788 | – | – | – |
| EP2010065463 | – | – | – |
| US20090252788P | – | – | – |
| WO2010EP65463 | – | – | – |
Members9
| Document | Office | Kind | |
|---|---|---|---|
| WO2011048010A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2012197650A1 | United States of America | A1 | |
| EP2491560A1 | European Patent Office (EPO) | A1 | |
| CN102754159A | China | A | |
| JP2013509601A | Japan | A | |
| US9105300B2 | United States of America | B2 | |
| JP5771618B2This record | Japan | B2 | |
| CN102754159B | China | B | |
| EP2491560B1 | European Patent Office (EPO) | B1 |
12 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5771618
- Publication, DOCDB
- 5771618
- Publication, EPODOC
- JP5771618B
- Application
- 2012533640
- Application, DOCDB
- 2012533640
- Application, EPODOC
- JP20120533640
Titles2
- Japanese
- 音声オブジェクトの区分を示すメタデータ時間標識情報
- English
- Metadata that indicates the classification of the audio object Time indicator information
Classification
- CPC, 3
- G11B27/034
- G10L21/04
- G10L25/00
- IPC, 1
- G10L19 00
