Metadata preparing device, preparing method therefor and retrieving device
Abstract
The present invention relates to a metadata generating device having a content reproduction unit (1) for reproducing and outputting content, a monitor (3) for monitoring the content reproduced by the content reproduction unit, a sound input unit (4), and a recognition sound input unit The voice recognition section (5) of the input voice signal, the metadata generation section (6) which converts the information recognized by the voice recognition section into metadata, and the identification information addition section (7), the identification information addition section (7) from The reproduced content supplied by the content reproducing unit acquires identification information for identifying each part in the content and assigns the metadata; the metadata generating device makes the generated metadata and each part in the content Partially establish associations.

Term
Term ended
Projected expiry passed 23 June 2023, 3.3 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
21 claims: 3 independent, 18 dependent
- 1一种元数据生成装置,其特征在于,具有再现内容并输出的内容再现部、声音输入部、识别所述声音输入部输入的声音信号的声音识别部、将所述声音识别部识别的信息转换成元数据的元数据生成部,以及识别信息附加部,该识别信息附加部从所述内容再现部供给的再现后的内容,获取用于识别所述内容中的各部分的识别信息,并将其附加到所述元数据上;该元数据生成装置使生成的所述元数据与所述内容内的各部分相关联。
- 2如权利要求1所述的元数据生成装置,其特征在于,还具有与所述内容相关的词典,在通过所述声音识别部识别由所述声音输入部输入的声音信号时,与所述词典相关联地进行识别。
- 3如权利要求2所述的元数据生成装置,其特征在于,所述声音识别部以单词为单位,与所述词典相关联地识别所述声音信号。
- 4如权利要求1或3所述的元数据生成装置,其特征在于,还具有包含键盘的信息处理部,可以通过从所述键盘输入,由所述信息处理部修正元数据。
- 5如权利要求1~5中任一项所述的元数据生成装置,其特征在于,使用附在所述内容上的时间代码信息,作为所述识别信息。
- 6如权利要求1~6任一项所述的元数据生成装置,其特征在于,使用附在所述内容上的内容的地址、编号或帧号码,作为所述识别信息。
- 7如权利要求1所述的元数据生成装置,其特征在于,所述内容为静态图像内容,使用所述静态图像内容的各个地址,作为所述识别信息。
- 8如权利要求1所述的元数据生成装置,其特征在于,所述内容再现部由内容数据库构成;所述声音输入部按照与所述内容数据库供给的同步信号同步的时钟,将所输入的关键词的声音信号转化为数据并供给所述声音识别部;所述声音识别部根据在所述声音输入部转化为数据的声音信号数据,识别所述关键词;所述元数据生成部构成文件处理部,该文件处理部将表示所述内容中包含的图像信号的时间位置的时间代码用作所述识别信息,并将所述声音识别部输出的关键词与所述时间代码相结合,来生成元数据文件。
- 9如权利要求8所述的元数据生成装置,其特征在于,还具有记录部,将所述内容数据库供给的内容与所述元数据文件一道作为内容文件进行记录。
- 10如权利要求9所述的元数据生成装置,其特征在于,还具有生成控制文件的内容信息文件处理部,所述控制文件用于管理应记录所述内容文件的记录位置和所述元数据文件之间的关系,在所述记录部中,与所述内容文件和所述元数据文件一道记录所述控制文件。
- 11如权利要求8所述的元数据生成装置,其特征在于,还具有词典数据库,所述声音识别部可以从多种类型词典中选择适合于所述内容的种类词典。
- 12如权利要求11所述的元数据生成装置,其特征在于,可以将与内容相关的关键词供给所述声音识别部,所述声音识别部优先识别所述关键词。
- 13一种元数据生成方法,其特征在于,对与内容相关的信息进行声音输入,由声音识别装置对输入的声音信号进行声音识别,将所述声音识别后的信息转换成元数据,将用于识别所述内容内的各部分的、附加在所述内容上的识别信息附加到所述元数据上,使生成的所述元数据与所述内容内的各部分相关联。
- 14如权利要求13所述的元数据生成方法,其特征在于,再现所述内容并在监视器上显示,同时对与内容相关的信息进行声音输入。
- 15如权利要求13所述的元数据生成方法,其特征在于,利用与所述内容相关的词典,并由所述声音识别装置与所述词典相关联地识别所述输入的声音信号。
- 16如权利要求13所述的元数据生成方法,其特征在于,使用附在所述内容上的时间代码信息,作为所述识别信息。
- 17如权利要求13所述的元数据生成方法,其特征在于,使用静态图像内容作为所述内容,并使用所述静态图像内容的各个地址作为所述识别信息。
- 18一种元数据检索装置,其特征在于,包括:再现内容并输出的内容数据库;声音输入部,按照与再现的所述内容的同步信号同步的时钟,将所输入的关键词的声音信号转化为数据;声音识别部,根据由所述声音输入部转化为数据的声音信号数据识别关键词;文件处理部,将所述声音识别部输出的关键词与表示所述内容中包含的图像信号的时间位置的时间代码相结合,来生成元数据文件;内容信息文件处理部,生成控制文件,所述控制文件用于对内容文件的记录位置和所述元数据文件之间的关系进行管理;记录部,记录所述内容文件、所述元数据文件和所述控制文件;检索部,确定包含所输入的检索关键词的所述元数据文件,并参照所述控制文件提取与所述内容文件的所述关键词对应的记录位置;所述内容文件的记录位置是所述记录部中的记录位置。
- 19如权利要求18所述的元数据检索装置,其特征在于,所述内容信息文件处理部输出的控制文件是记载与所述内容的记录时间对应的、所述记录部中所述内容的记录位置的表,可以根据所述时间代码检索所述内容的记录位置。
- 20如权利要求18所述的元数据检索装置,其特征在于,还具有词典数据库和向所述声音识别部供给与内容相关的关键词的关键词供给部,所述声音识别部可以从多种类型的词典中选择适合于所述内容的种类词典,并且优先识别所述关键词。
- 21如权利要求18所述的元数据检索装置,其特征在于,还具有词典数据库,所述声音识别部可以从多种类型的词典中选择适合于所述内容的种类的词典,所述检索部利用从所述声音识别部使用的通用词典中选定的关键词进行检索。
Independent claims21
93 paragraphs, as filed
Metadata generating device, its generating method and retrieval device
Technical field
The present invention relates to a metadata generating device and a metadata generating method for generating metadata associated with content such as created images and sounds. It also relates to a retrieval device that uses the generated metadata to retrieve content.
Background technique
In recent years, people have added metadata associated with these content to the created image and sound content.
However, the existing metadata addition operation generally adopts the method of reproducing the created image and sound content based on the script or commentary manuscript of the created image and sound content and confirming the information that should be used as the metadata. Generated by manual computer input. Therefore, considerable labor is required in the process of generating metadata.
Japanese Patent Laid-Open No. 09-130736 describes a system that uses voice recognition to attach a mark when a camera takes a picture. However, this system is used at the same time of photography and cannot be adapted to attach metadata to already produced content.
Summary of the invention
The present invention solves the above-mentioned problems, and its object is to provide a metadata generation device and a metadata generation method that can easily generate metadata for a prepared content through voice input.
Another object of the present invention is to provide a retrieval device that can easily retrieve content using the generated metadata.
The metadata generating device of the present invention includes a content reproduction unit that reproduces and outputs content, a voice input unit, a voice recognition unit that recognizes a voice signal input by the voice input unit, and converts information recognized by the voice recognition unit into metadata The metadata generating unit of the, and the identification information adding unit that acquires the identification information used to identify each part in the content from the reproduced content supplied by the content reproduction unit, and adds it to On the metadata; the metadata generating device associates the generated metadata with each part in the content.
The metadata generation method of the present invention is to perform voice input for content-related information, voice recognition is performed on the input voice signal by a voice recognition device, and the voice recognition information is converted into metadata, which will be used to identify the information. The identification information of each part in the content that is added to the content is added to the metadata, and the generated metadata is associated with each part in the content.
The metadata retrieval device of the present invention includes: a content database for reproducing and outputting content; a sound input unit, which converts the inputted keyword sound signal into data according to a clock synchronized with the synchronization signal of the reproduced content; The recognition unit recognizes keywords based on the voice signal data converted into data by the voice input unit; the file processing unit combines the keywords output by the voice recognition unit with the time indicating the time position of the image signal contained in the content The code is combined to generate a metadata file; the content information file processing unit generates a control file for managing the relationship between the recording location of the content file and the metadata file; the recording unit, the recording location The content file, the metadata file, and the control file; the search unit determines the metadata file that contains the input search keyword, and refers to the control file to extract the keyword with the content file Corresponding recording location; the recording location of the content file is the recording location in the recording section.
Description of the drawings
Fig. 1 is a structural block diagram of a metadata generating device according to Embodiment 1 of the present invention.
Fig. 2 shows an example of metadata with time code in the first embodiment of the present invention.
Fig. 3 is a structural block diagram of a metadata generating device according to Embodiment 2 of the present invention.
Fig. 4 shows an example of a still image content/metadata display unit in the same device.
Fig. 5 is a block diagram showing another structure of the metadata generating device according to the second embodiment of the present invention.
Fig. 6 is a structural block diagram of a metadata generating apparatus according to Embodiment 3 of the present invention.
FIG. 7 is a structural diagram showing an example of the dictionary DB in the device of the same embodiment.
Fig. 8 shows a recipe as an example of a content script used in the device of the same embodiment.
FIG. 9 is a data diagram in TEXT format showing an example of a metadata file generated by the apparatus of the same embodiment.
Fig. 10 is a structural block diagram of a metadata generating apparatus according to Embodiment 4 of the present invention.
Fig. 11 is a structural diagram of an example of an information file generated by the apparatus of the same embodiment.
Fig. 12 is a structural block diagram of a metadata retrieval device according to Embodiment 5 of the present invention.
Fig. 13 is a structural block diagram of a metadata generating apparatus according to Embodiment 6 of the present invention.
detailed description
According to the metadata generating device of the present invention, when generating content-related metadata or adding tags, voice recognition and voice input are used to generate metadata or tags, and the metadata or tags are matched with the time or scene of the content. Wait to establish associations. Therefore, it is possible to automatically generate metadata generated by keyboard input in the past through voice input. The so-called metadata refers to a collection of tags, and the term "metadata in the present invention also includes the tag itself. In addition, the so-called content refers to all objects that are generally called content, such as created images, audio content, static image content, database-based images, and audio content.
Preferably, the metadata generating device of the present invention further includes a dictionary related to the content, and when the voice signal input by the voice input unit is recognized by the voice recognition unit, it is associated with the dictionary for recognition. According to this structure, the keywords extracted in advance from the script of the created content are input as the voice signal, and the dictionary field is set and the keyword priority is given according to the script, so that the voice recognition unit can be efficiently and accurately used to generate elements. data.
The voice recognition unit may recognize the voice signal in association with the dictionary in units of words. Furthermore, the metadata generating device preferably further has an information processing unit including a keyboard, and the metadata can be corrected by the information processing unit by input from the keyboard. The time code information attached to the content can be used as the identification information. Alternatively, the address, number, or frame number of the content attached to the content may be used as the identification information. In addition, the content may be static image content, and each address of the static image content is used as identification information.
As an application example of the present invention, the following metadata generating device can be constructed. That is, the content reproduction unit is composed of a content database. The voice input unit converts the input voice signal of the keyword into data according to a clock synchronized with the synchronization signal supplied from the content database and supplies it to the voice recognition unit. The voice recognition unit recognizes the keyword based on the voice signal data converted into data by the voice input unit. The metadata generating unit constitutes a file processing unit that uses a time code indicating the time position of the image signal contained in the content as the identification information, and combines the keywords output by the voice recognition unit with The time codes are combined to generate a metadata file.
According to this structure, metadata can be added efficiently even in an interval of several seconds. Therefore, metadata can be generated in a short time interval that is difficult to rely on conventional key input.
Preferably, this structure further has a recording unit that records the content supplied by the content database together with the metadata file as a content file. Furthermore, it is preferable to further have a content information file processing unit that generates a control file for managing the relationship between the recording position where the content file should be recorded and the metadata file, and in the recording unit, The control file is recorded together with the content file and the metadata file. In addition, it is preferable to have a dictionary database, and the voice recognition unit can select a dictionary suitable for the type of content from a plurality of types of dictionaries. Preferably, keywords related to the content may be supplied to the voice recognition unit, and the voice recognition unit preferentially recognizes the keywords.
Preferably, the metadata generating method of the present invention reproduces the content and displays it on a monitor, and at the same time performs sound input of information related to the content. Furthermore, it is preferable that a dictionary related to the content is used, and the voice recognition device recognizes the input voice signal in association with the dictionary. In addition, it is preferable to use time code information attached to the content as the identification information. In addition, static image content may be used as the content, and each address of the static image content may be used as the identification information.
According to the metadata retrieval device of the present invention, by using the control file indicating the recording location of the content and the metadata file indicating the metadata, time code, etc., a desired part of the content can be quickly retrieved based on the metadata.
In the metadata search device of the present invention, it is preferable that the control file output by the content information file processing unit is a table describing the recording position of the content in the recording unit corresponding to the recording time of the content, and The recording position of the content is retrieved according to the time code.
Furthermore, it is preferable to further have a dictionary database and a keyword supply unit that supplies keywords related to the content to the voice recognition unit, and the voice recognition unit can select a category dictionary suitable for the content from a plurality of types of dictionaries. , And prioritize the identification of the keywords.
In addition, it is preferable to further have a dictionary database, the voice recognition unit can select a dictionary suitable for the type of content from a plurality of types of dictionaries, and the search unit uses a general dictionary used by the voice recognition unit to select a dictionary. Search for the specified keywords.
Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
(Embodiment 1) Fig. 1 is a block diagram showing the structure of a metadata generating device in Embodiment 1 of the present invention. The content reproduction unit 1 is a component for confirming the produced content at the time of metadata generation. The output of the content reproduction section 1 is supplied to the image monitor 2, the sound monitor 3, and the time code addition section 7. The microphone 4 is provided as a voice input unit for generating metadata. The voice input from the microphone 4 is input to the voice recognition unit 5. A dictionary 8 for voice recognition is connected to the voice recognition unit 5, and data therein can be referenced. The recognition output of the voice recognition unit 5 is supplied to the metadata generating unit 6, and the generated metadata is supplied to the time code adding unit 7, and then can be output from the time code adding unit 7 to the outside.
Video and sound signal reproduction devices such as VTRs, hard disk devices, and optical disk devices can be used, video and sound signal reproduction devices that use storage units such as semiconductor memories as recording media, or video and sound signal reproduction devices that are supplied by transmission or broadcasting. An image and sound reproduction device, etc., serve as the content reproduction unit 1.
The reproduced image signal is supplied to the image monitor 2 from the image signal output terminal 1a of the content reproduction section 1. The reproduced sound signal is supplied to the sound monitor 3 from the sound signal output terminal 1b. The reproduced time code is supplied to the time code adding section 7 from the time code output terminal 1c. In addition, the image monitor 2 and the sound monitor 3 are not essential as components of the metadata generating device, and can be connected and used as needed.
When generating metadata, the operator confirms one or both of the image monitor 2 and the audio monitor 3, and at the same time refers to the script or commentary manuscript according to the situation, reads out the metadata to be input, and inputs it through the microphone 4. The voice signal output from the microphone 4 is supplied to the voice recognition unit 5. In addition, the voice recognition unit 5 refers to the data of the voice recognition dictionary 8 as necessary. The voice data recognized by the voice recognition unit 5 is supplied to the metadata generation unit 6 and converted into metadata.
To the metadata generated as described above, in order to add information corresponding to the time or scene relationship of each part of the content, the time code adding unit 7 adds the time code information obtained from the reproduced content and supplied from the content reproduction unit 1.
In order to explain the above actions more specifically, assume a scene in the case where the content is cooking instructions as an example. In this case, the operator confirms the display screen of the image monitor 2 and at the same time makes the sound of "1 scoop of salt" through the microphone 4. The voice recognition unit 5 refers to the dictionary 8 and recognizes "salt" and "1 scoop", and then metadata The generating unit 6 converts them into marks of "salt" and "1 scoop", respectively. In addition, the structure of the voice recognition unit 5 is not particularly limited, and various voice recognition units commonly used can be used for voice recognition, as long as it is recognized as data such as "salt" and "1 scoop". In general, metadata refers to the collection of these tags. The result of the voice recognition is shown in FIG. 2, the metadata 9 a is output from the metadata generating unit 6 and then supplied to the time code adding unit 7.
The time code adding unit 7 generates packet data composed of time code-added metadata 10 with time code based on the time code signal 9 b supplied from the content reproduction unit 1. The generated metadata can be directly output or stored in a recording medium such as a hard disk.
This example shows the case of generating metadata in the form of a packet, but it is not limited to this.
(Embodiment 2) Fig. 3 is a structural block diagram of a metadata generating device according to Embodiment 2 of the present invention. In this embodiment, the content of a static image is used as a metadata generation object. In order to identify the content of the still image, the generated metadata is associated with the content of the still image by using the content address corresponding to the time code in the case of the moving image.
In FIG. 3, the camera 11 is an element for producing still image content. The still image content recording unit 12 adds address information to the output of the camera 11 and records it. In order to generate metadata, the static image content and address information recorded here are supplied to the static image content·metadata recording unit 13. The address information is also supplied to the metadata address adding section 19.
The microphone 16 is used for voice input of information related to the still image, and its output is input to the voice recognition unit 17. The voice recognition unit 17 is connected to the voice recognition dictionary 20 and can refer to its data. The recognition output of the voice recognition unit 17 is supplied to the metadata generating unit 18, and the generated metadata is supplied to the metadata address adding unit 19.
The still image content and metadata recorded in the still image content/metadata recording unit 13 are reproduced by the still image content/metadata reproducing unit 14 and then displayed on the still image content/metadata display unit 15.
The operation of the metadata generating device with the above-mentioned structure will be described in detail below.
The still image content captured by the camera 11 is recorded on a recording medium (not shown) by the still image content recording unit 12, and address information is added and the address information is also recorded on the recording medium. The recording medium is generally composed of a semiconductor memory, but is not limited to a semiconductor memory, and various recording media such as a magnetic memory, an optical recording medium, and a magneto-optical recording medium can also be used. The recorded static image content is supplied to the static image content/metadata recording unit 13 through the output terminal 12a and the input terminal 13a, and the address information is also supplied to the static image content/metadata recording unit 13 through the address output terminal 12b and the input terminal 13b. The address information is also supplied to the metadata address adding section 19 through the output terminal 12b and the input terminal 19b.
On the other hand, information related to the still image captured by the camera 11 is input to the voice recognition unit 17 through the microphone 16. Information related to the still image includes the title, shooting time, photographer, shooting location (where), subject (who), and subject (what), etc. In addition, the data of the voice recognition dictionary 20 may be supplied to the voice recognition unit 17 as necessary.
The voice data recognized by the voice recognition unit 17 is supplied to the metadata generation unit 18 and converted into metadata or tags. In general, metadata refers to a collection of information related to content, such as title, shooting time, photographer, shooting location (where), subject (who), and subject (what). In order to add information corresponding to the relationship of the content or scene of the still image content, the metadata or mark generated as described above is supplied to the metadata address adding unit 19. The metadata address adding section 19 adds the address information supplied from the still image content recording section 12 to the metadata. The address-added metadata to which the address information is added as described above is supplied to the static image content/metadata recording unit 13 through the output terminal 19c and the input terminal 13c. In the static image content/metadata recording unit 13, the static image content of the same address and the metadata of the same address are associated and recorded.
In order to more specifically explain the metadata with the address, FIG. 4 shows that the static image content and metadata reproducing section 14 reproduces the static image content and metadata recorded by the static image content and metadata recording section 13, and the static image content The image content/metadata display unit 15 displays an example of the result.
The screen of the static image content/metadata display unit 15 in FIG. 4 is an example, and is composed of a static image content display unit 21, an address display unit 22, and a metadata display area 23. The metadata display area 23 is composed of, for example, 1) a title description section 23a, 2) a date and time description section 23b, 3) a photographer description section 23c, 4) a shooting location description section 23d, and the like. These metadata are generated from the voice data recognized by the voice recognition unit 17 described above.
In the above-mentioned operation, metadata is generated when the content of the captured still image does not need to be confirmed before the shooting of the still image content, almost simultaneously with the shooting, or after the shooting.
In the following, with reference to FIG. 5, in the case where metadata is generated after the static image content is created, a case where the static image content is reproduced and metadata is generated for the monitored static image content will be described. The same components as those in FIG. 3 are assigned the same reference numerals, and descriptions of their functions and the like are omitted. In this case, a still image content/address reproduction unit 24 is provided between the still image content recording unit 12 and the still image content/metadata recording unit 13. In addition, a monitor 25 for supplying the output of the still image content/address reproducing unit 24 is provided.
The still image content photographed by the camera 11 and supplied to the still image content recording unit 12 is recorded on a recording medium (not shown) with an address added, and the address is also recorded in the recording medium. Such a recording medium is supplied to the still image content/address reproducing unit 24. In this way, in the metadata generating device for reproducing the produced still image content and generating metadata for the monitored still image content, the camera 11 and the still image content recording unit 12 are not essential components.
The still image content reproduced by the still image content/address reproducing unit 24 is supplied to the monitor 25. Similarly, the reproduced address information is supplied to the metadata address adding section 19 through the output terminal 24b and the input terminal 19b. The person in charge of metadata generation confirms the content of the still image displayed on the monitor 25, and then vocally inputs the vocabulary necessary for generating the metadata through the microphone 16. In this way, information related to the still image captured by the camera 11 is input to the voice recognition unit 17 through the microphone 16. The related information of the still image includes the title, the shooting time, the photographer, the shooting location (where), the subject (who), and the subject (what), etc. The operation thereafter is the same as the description of the structure shown in FIG. 3.
(Embodiment 3) FIG. 6 is a structural block diagram of a metadata generating device according to Embodiment 3 of the present invention. This embodiment is an example in which general digital data content is the target of metadata generation. In order to identify the digital data content, the address or number of the content is used to associate the digital data content with the generated metadata.
In FIG. 6, 31 is a content database (hereinafter referred to as a content DB), and the output reproduced by the content DB 31 is supplied to the sound input unit 32, the file processing unit 35, and the recording unit 37. The output of the voice input unit 32 is supplied to the voice recognition unit 33. In addition, the data of the dictionary database (hereinafter referred to as the dictionary DB) 34 may be supplied to the voice recognition unit 33. The metadata is output by the voice recognition unit 33 and input to the file processing unit 35. The file processing unit 35 uses the time code value supplied from the content DB 31 to append predetermined data to the metadata output by the voice recognition unit 33 to file the metadata. The metadata file output by the file processing unit 35 is supplied to the recording unit 37 and recorded together with the content output by the content DB 31. The voice input unit 32 and the dictionary DB 34 are respectively provided with a voice input terminal 39 and a dictionary area selection input terminal 40. The reproduction output of the content DB 31 and the reproduction output of the recording unit 37 can be displayed on the image monitor 41.
The content DB31 has functions such as video and sound signal reproducing devices such as VTR, hard disk devices, and optical disk devices, video and sound signal reproducing devices using storage units such as semiconductor memory as recording media, or video and sound signal reproducing devices that are provided by transmission or broadcasting. ·Images that are recorded and reproduced at once by sound signals. Contents produced by sound reproducing devices, etc., generate time codes corresponding to the contents and reproduce them.
The operation of the above-mentioned metadata generating device will be described below. The image signal with the time code reproduced by the content DB 31 is supplied to the image monitor 41 and shown. After the operator uses a microphone to input a voice signal for narration based on the content shown on the image monitor 41, the voice signal is input to the voice input unit 32 through the voice input terminal 39.
At this time, the operator preferably confirms the content or time code displayed on the image monitor 41, and reads out the content management keywords extracted on the basis of the script, the commentary manuscript, or the content. By using keywords defined in advance by a script or the like as the voice signal input as described above, the recognition rate of the subsequent voice recognition unit 33 can be improved.
The sound input unit 32 converts the sound signal input from the sound input terminal 39 into data using a clock synchronized with the vertical synchronization signal output from the content DB1. The voice signal data converted into data by the voice input section 32 is input to the voice recognition section 33, and at the same time, a dictionary necessary for voice recognition is supplied from the dictionary DB34. The voice recognition dictionary used in the dictionary DB 34 can be set by the dictionary area selection input terminal 40.
For example, as shown in FIG. 7, assuming that the dictionary DB 34 is structured for each field, the field to be used is set by the dictionary field selection input terminal 40 (for example, a keyboard terminal that can perform key input). For example, in the case of a cooking program, the field of the dictionary DB 34 can be set to cooking-Japanese cuisine-cooking method-vegetable stir-frying method by terminal 40. By setting the dictionary DB 34 in this way, the words used and the words that can be voice-recognized can be restricted, and the recognition rate of the voice recognition unit 33 can be improved.
In addition, it is possible to input keywords extracted from the content of a script, script manuscript, or content through the dictionary field selection terminal 40 in FIG. 6. For example, when the content is a cooking program, the recipe shown in FIG. 8 is input through the terminal 40. In consideration of the content, the word described in the recipe is highly likely to be input as a voice signal. Therefore, the dictionary DB 34 clearly indicates the recognition priority of the recipe word input from the terminal 40, and priority is given to voice recognition. For example, when "tomato and "oyster shell are in the dictionary, and the recipe word input from terminal 40 is only "oyster shell, priority 1 is given to "shellfish oyster. When the voice recognition unit 33 recognizes a voice such as "oyster, it recognizes as "oyster shell in which the priority of the word set in the dictionary DB 34 is described as 1.
In this way, in the dictionary DB 34, the words are defined in the field input from the terminal 40, and the priority of the words is clearly indicated after the script is input from the terminal 40, so that the recognition rate of the voice recognition unit 33 can be improved.
The voice recognition unit 33 in FIG. 6 recognizes the voice signal data input from the voice input unit 32 based on the dictionary provided by the dictionary DB 34 and generates metadata. The metadata output by the voice recognition unit 33 is input to the file processing unit 35. As described above, the sound input unit 32 converts the sound signal into data in synchronization with the vertical synchronization signal reproduced by the content DB 31. In this way, the file processing unit 35 uses the synchronization information from the voice input unit 32 and the time code value supplied from the content DB 31 to generate a metadata file in the TEXT format shown in FIG. 9 in the case of the aforementioned cooking program, for example. That is, the file processing unit 35 adds TM_ENT (seconds) as the reference time per second after the start of the file (file), and TM_OFFSET representing the number of offset frames from the reference time to the metadata output by the voice recognition unit 33. And time code, file processing in this form.
The recording unit 37 records the metadata file output by the file processing unit 35 and the content output by the content DB 31. The recording unit 37 is composed of an HDD, a memory, an optical disc, etc., and the content output from the content DB 31 is also recorded in a file format.
(Embodiment 4) FIG. 10 is a block diagram showing the structure of a metadata generating device according to Embodiment 4 of the present invention. Compared with the structure of the third embodiment, the device of this embodiment adds a content information file processing unit 36. The content information file processing unit 36 generates a control file indicating the recording position relationship of the content recorded in the recording unit 37 and records it in the recording unit 37.
That is, the content information file processing unit 36 uses the content output from the content DB 31 and the recording location information of the content output from the recording unit 37 as the basis to generate the time axis information of the content and the address relationship between the content recorded in the recording unit 37 Information, and output as a control file after being converted into data.
For example, as shown in FIG. 11, with respect to the recording medium address indicating the content recording position, TM_ENT#j indicating the time axis reference of the content is directed to an equal time axis interval. For example, TM_ENT#j is pointed to the recording medium address every 1 second (30 frames in the case of the NTSC signal). With such a mapping, even if the content is distributed and recorded in units of one second, the recording address of the recording unit 37 can be uniquely obtained from TM_ENT#j.
Moreover, as shown in Figure 9, the metadata file is recorded in the form of TEXT as TM_ENT (seconds) as the reference time per second after the start of the file, TM_OFFSET representing the number of offset frames from the reference time, and time code And metadata. Therefore, as long as the metadata 1 is specified in the metadata file, the time code, reference time, and frame offset value can be known, so that the recording position in the recording unit 37 can be immediately known from the control file shown in FIG. 11.
The equal time axis interval of TM_ENT#j is not limited to the above-mentioned pointing every 1 second, and may be described in accordance with the GOP unit used in MPEG2 compression or the like.
The vertical synchronization signal in the NTSC of the TV image signal is 60/1.001Hz, therefore, two types can be used, that is, the time code according to the drop-frame mode is used to match the absolute time; and the vertical synchronization signal is used according to the vertical The non-drop-timecode of the synchronization signal (60/1.001Hz). In this case, for example, TM_ENT#j represents the non-missing time code, and TC_ENT#j represents the time code corresponding to the dropped frame.
In addition, the digitization of the control file can also be digitized using existing languages such as SMIL2. If the SMIL2 function is used, it can be digitized corresponding to the file name of the related content and metadata file, and then stored in the control file.
Figure 11 shows the structure that directly indicates the recording address of the recording unit, but it can also indicate the data capacity from the beginning of the content file to the time code instead of the recording address. The recording is calculated and detected based on the data capacity and the recording address of the file system Record address of the time code in the department.
In addition, as mentioned above, instead of storing the correspondence table of TM_ENT#j and time code in the metadata file, it is also possible to store the correspondence table of TM_ENT#j and time code in the control file, which can also be obtained. The same effect.
(Embodiment 5) Fig. 12 is a structural block diagram of a metadata retrieval device according to Embodiment 5 of the present invention. Compared with the structure of the fourth embodiment, the device of this embodiment has a search unit 38 added. The search unit 38 selects and sets the keywords of the scene to be searched from the same dictionary DB 34 used when the metadata is detected by performing voice recognition.
Then, the search unit 38 searches for metadata items in the metadata file, and displays a list of title names and content scene positions (time codes) that match the keywords. When a specific scene is set in the list display, the recording medium address in the control file is automatically detected based on the reference time TM_ENT (seconds) of the metadata file and the offset frame number TM_OFFSET, and is set in the recording unit 37. The recording unit 37 reproduces the content scene recorded in the address of the recording medium and displays it on the monitor 41. Through the above structure, the scene you want to see can be detected immediately after the metadata is detected.
If a thumbnail file linked to the content is provided, it is possible to reproduce and display a representative thumbnail of the content when the list of content names matching the aforementioned keywords is displayed.
(Embodiment 6) The foregoing embodiments 3 to 5 describe devices for adding metadata to pre-recorded content. This embodiment relates to the extension of the present invention to cameras and other systems that add metadata during shooting, and particularly relates to the application of the present invention. It is extended to an example of a device that adds a shooting location as metadata when shooting a landscape whose content is pre-defined. Fig. 13 is a structural block diagram of a metadata generating apparatus according to Embodiment 6 of the present invention.
The imaging output of the camera 51 is recorded in the content DB 54 as image content. At the same time, the GPS 52 detects the location taken by the camera, and its location information (longitude and latitude values) is converted into a voice signal by the voice synthesis unit 53, and then recorded as location information in the voice channel of the content DB 54. The camera 51, the GPS 52, the voice synthesis unit 53, and the content DB 54 may be integrated as a camera 50 with a recording unit. The content DB 54 inputs the position information of the audio signal recorded in the audio channel to the audio recognition unit 56. The dictionary DB55 supplies the dictionary data to the voice recognition unit 56. The dictionary DB55 can select and restrict the domain names and landmarks, etc. by keyboard input from the terminal 59, and then output to the voice recognition unit 56.
The voice recognition unit 56 uses the recognized latitude and longitude data and the data of the dictionary DB 55 to detect the domain names and landmarks and output them to the file processing unit 57. The file processing unit 57 converts the time code output from the content DB 54 and the domain name and landmark text (TEXT) output from the voice recognition unit 56 into metadata, and then generates a metadata file. The metadata file is supplied to the recording unit 58, and the recording unit 58 records the metadata file and the content data output from the content DB 54.
Through the above structure, metadata such as domain names and landmarks can be automatically added to each scene captured.
The above embodiment describes the structure in which the keywords recognized by the voice recognition unit are filed together with the time code into a metadata file. It is also possible to add relevant keywords to the keywords recognized by the voice recognition unit and then file them. For example, if it is recognized as "Yodo River" (note: the name of a river in Japan) by voice, additional keywords with general attributes such as terrain and rivers are added and filed. In this way, additional keywords such as terrain and rivers can also be used when searching, which can improve the search performance.
The voice recognition unit of the present invention adopts a word recognition method that performs voice recognition in units of words, and can improve the voice recognition rate by limiting the number of voice input words and the number of words in the recognition dictionary used.
In general, any misrecognition may occur in voice recognition. Each of the above-mentioned embodiments has an information processing unit such as a computer including a keyboard. In the case of misrecognition, the generated metadata or mark can be corrected by keyboard operation.
Industrial Applicability According to the metadata generation device of the present invention, in order to generate content-related metadata or add tags, it uses voice recognition to generate metadata through voice input, and establish metadata and content regulations Part of the association, compared with the existing keyboard input, can effectively generate metadata or implement additional tags.
13 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10644666B2 | Cited by | United States of America | Applicant |
| CN107808197A | Cited by | China | Search report |
| US11882333B2 | Cited by | United States of America | Applicant |
| CN102930888A | Cited by | China | Search report |
| CN105389350A | Cited by | China | Search report |
| CN102063461A | Cited by | China | Search report |
| US9769294B2 | Cited by | United States of America | Applicant |
| CN105103222A | Cited by | China | Search report |
| US7831598B2 | Cited by | United States of America | Applicant |
| US8862473B2 | Cited by | United States of America | Applicant |
| US9325381B2 | Cited by | United States of America | Applicant |
| US11057674B2 | Cited by | United States of America | Applicant |
| US10356471B2 | Cited by | United States of America | Applicant |
| US10785519B2 | Cited by | United States of America | Applicant |
11 members in 6 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 1825062002 | Japan | – | |
| 2002182506 | Japan | A | |
| 3197562002 | Japan | – | |
| 3197572002 | Japan | – | |
| 2002319756 | Japan | A | |
| 2002319757 | Japan | A | |
| 3348312002 | Japan | – | |
| 2002334831 | Japan | A |
Members11
| Document | Office | Kind | |
|---|---|---|---|
| WO2004002144A1 | World Intellectual Property Organization (WIPO) | A1 | |
| JP2004086124A | Japan | A | |
| WO2004002144B1 | World Intellectual Property Organization (WIPO) | B1 | |
| JP2004153764A | Japan | A | |
| JP2004153765A | Japan | A | |
| MXPA04012865A | Mexico | A | |
| EP1536638A1 | European Patent Office (EPO) | A1 | |
| CN1663249AThis record | China | A | |
| US2005228665A1 | United States of America | A1 | |
| EP1536638A4 | European Patent Office (EPO) | A4 | |
| JP3781715B2 | Japan | B2 |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Deemed withdrawal of patent application after publication (patent law 2001)C02 | C02 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 1663249
- Application
- 38149028
Titles2
- Chinese
- 元数据生成装置、其生成方法以及检索装置
- English
- Metadata generating device, its generating method and retrieval device
Classification
- CPC, 10
- H04N21/435
- H04N21/235
- H04N21/42203
- H04N21/4223
- H04N21/4334
- H04N21/439
- H04N21/440236
- H04N21/8106
- H04N21/84
- G06F16/7867
- IPC, 3
- G11B27 11
- G11B27 28
- H04N7 24