Voice recognition device and method of voice data recognition
Abstract
(57) A summary and subject Without inviting enlargement of a storage medium, Solution means which was made to carry out suitable to an unrestricted number's of unspecified persons use Press the voice register key on a final controlling element, and an ID number (S1) is inputted (S2), Subsequently, the data under voice corresponding to this key is inputted, pressing a desired key (S3->S4), The degree of similar is computed by checking that the keystroke and the data-under-voice input have been made simultaneously, setting equipment as data-under-voice register mode (S5), and carrying out pattern matching of the data under voice by which dictionary registration is carried out to input voice data (S6). And it carries out, when there is no similar command, dictionary registration of the input voice data is carried out (S8), and when there is a similar command, a predetermined message is displayed on a final controlling element (S9).
Term
Term ended
Projected expiry passed 11 March 2019, 7.5 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
10 claims: 4 independent, 6 dependent
- 1[Claims] 1. Identification information input means for inputting user identification information, storage means for storing and storing voice data, voice input means for inputting voice data, and identification input by the identification information input means. It is determined whether or not the associating means for associating the information with the voice data input by the voice input means and the input voice data associated with the identification information are already stored in the storage means as stored voice data. The determination means and the first registration instruction means for instructing new registration of the input voice data in the storage means when the determination means determines that the input voice data is not stored in the storage means. A voice recognition device characterized by being equipped. 【特許請求の範囲】 【請求項1】 ユーザの識別情報を入力する識別情報入力手段と、音声データを記憶して蓄積する蓄積手段と、音声データを入力する音声入力手段と、前記識別情報入力手段により入力された識別情報と前記音声入力手段により入力された音声データとを対応付ける対応付け手段と、前記識別情報に対応付けられた入力音声データが蓄積音声データとして前記蓄積手段に既に蓄積されているか否かを判断する判断手段と、該判断手段により前記入力音声データが前記蓄積手段に蓄積されていないと判断されたときは前記入力音声データの前記蓄積手段への新規登録を指示する第1の登録指示手段とを備えていることを特徴とする音声認識装置。
- 3The determination means includes a display command means for issuing a display command for similar voice data when the similarity calculated by the similarity calculation means is equal to or higher than a predetermined value. The voice recognition device described in item 2. 【請求項3】 前記判断手段は、前記類似度算出手段により算出された類似度が所定値以上のときは類似音声データの表示指令を発する表示指令手段を有していることを特徴とする請求項2記載の音声認識装置。
- 6An identification information input step for inputting user identification information, a voice input step for inputting voice data, identification information input by the identification information input step, and voice data input by the voice input step. A mapping step for associating with, a determination step for determining whether or not the input audio data associated with the identification information has already been accumulated in the storage means as stored audio data, and the determination step for the input audio data. A method for recognizing voice data, which includes a first registration instruction step for instructing new registration of the input voice data in the storage means when it is determined that the input voice data is not stored in the storage means. .. 【請求項6】 ユーザの識別情報を入力する識別情報入力ステップと、音声データを入力する音声入力ステップと、前記識別情報入力ステップにより入力された識別情報と前記音声入力ステップにより入力された音声データとを対応付ける対応付けステップと、前記識別情報に対応付けられた入力音声データが蓄積音声データとして蓄積手段に既に蓄積されているか否かを判断する判断ステップと、該判断ステップにより前記入力音声データが前記蓄積手段に蓄積されていないと判断されたときは前記入力音声データの前記蓄積手段への新規登録を指示する第1の登録指示ステップとを含んでいることを特徴とする音声データの認識方法。
- 8The determination step 7 includes a display command step that issues a display command for similar voice data when the similarity calculated in the similarity calculation step is equal to or higher than a predetermined value. How to recognize voice data. 【請求項8】 前記判断ステップは、前記類似度算出ステップで算出された類似度が所定値以上のときは類似音声データの表示指令を発する表示指令ステップを含むことを特徴とする請求項7記載の音声データの認識方法。
Independent claims4
95 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
The present invention relates to a voice recognition device and a method of recognizing voice data, and more particularly to a voice recognition device that recognizes the contents of a control command input by voice and a method of recognizing voice data.
【0002】
[Conventional technology]
In recent years, in the fields of personal computers (hereinafter referred to as "personal computers") and car navigation systems (hereinafter referred to as "car navigation systems"), it is possible to input voice data as a command and perform desired information processing by recognizing the voice data. Models that can be used are becoming widespread.
【0003】
Such voice recognition has conventionally been performed by separating voiceprints and voices and pattern matching the pitch of the voices. That is, conventionally, the command to be used and the voice data when the command is uttered are associated with each other in advance and registered as a dictionary in a storage medium such as a memory, and the input voice data and the voice of the registered command are registered. Data similarity is calculated, pattern matching is performed, the maximum value of the similarity is selected as a desired command, and desired information processing is performed based on the selected command.
【0004】
[Problems to be Solved by the Invention]
However, in the above-mentioned conventional voice data recognition method, in order to input voice data and recognize voice, all the commands to be used must be registered in the storage medium as a dictionary in advance. It was necessary for the user to say the command for all the commands that could be used, read the command, and register it in correspondence with the registration key.
【0005】
That is, in the conventional voice data recognition method, in a personal computer or car navigation system purchased by an individual, the purchased specific person mainly uses the voice recognition function, so usually, only the voice data related to the specific person is registered. be able to.
【0006】
However, in the case of equipment used by an unspecified number of people, such as commercial copiers, facsimile machines, printers, or digital multifunction devices that combine these functions, the voice data of many people who may use it is used. There is a problem that it is necessary to register it as a dictionary in the storage medium, and therefore a large-capacity storage medium is required.
【0007】
In addition, in such a device used by an unspecified number of people, the frequency of use of registered commands differs depending on the user, and therefore, even commands that are not used by a specific user at all can be used by other users. Since there is a problem, it is necessary to register the voice data of such a command in the storage medium, and there is a problem that the storage medium cannot be used efficiently.
【0008】
For this reason, the current situation is that there is no model equipped with a voice recognition function in devices used by an unspecified number of people such as the above-mentioned digital multifunction devices.
【0009】
The present invention has been made in view of such circumstances, and it has been announced that a voice recognition device and a voice data recognition method suitable for use by an unspecified number of people will be provided without inviting an increase in the size of a storage medium. Target.
【0010】
[Means for solving problems]
In order to achieve the above object, the voice recognition device according to the present invention includes identification information input means for inputting user identification information, storage means for storing and storing voice data, and voice input means for inputting voice data. , The associating means for associating the identification information input by the identification information input means with the voice data input by the voice input means, and the input voice data associated with the identification information are stored as the stored voice data. A determination means for determining whether or not the input voice data has already been accumulated in the means, and when the determination means determines that the input voice data is not accumulated in the storage means, the input voice data is newly added to the storage means. It is characterized by having a first registration instruction means for instructing registration.
【0011】
Further, the voice data recognition method according to the present invention includes an identification information input step for inputting user identification information, a voice input step for inputting voice data, and identification information and voice input by the identification information input step. A matching step for associating the voice data input by the input step, and a determination step for determining whether or not the input voice data associated with the identification information has already been stored in the storage means as the stored voice data. When it is determined by the determination step that the input voice data is not stored in the storage means, the first registration instruction step for instructing new registration of the input voice data in the storage means is included. It is characterized by that.
【0012】
It should be noted that other features of the present invention will be clarified from the description of the embodiments of the present invention below.
【0013】
BEST MODE FOR CARRYING OUT THE INVENTION
Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
【0014】
FIG. 1 is a block configuration diagram showing an embodiment of a copying machine as a voice recognition device according to the present invention. The copying machine has an operation unit 1 for inputting command information related to copy operation, and analog voice. A voice input unit 2 composed of a microphone or the like that inputs data and converts the analog voice data into digital voice data, and a voice recognition unit 3 that performs predetermined voice recognition processing on the digital voice data from the voice input unit 2. An image input unit 4 composed of a CCD or the like that reads a manuscript image and converts it into digital image data, and an ASIC (Application Specific Integrated Circuit) or the like that performs predetermined image processing on the digital image data from the image input unit 4. Image processing unit 5 equipped with the hardware circuit and software processing circuit of the above, and printers (laser beam printers, inkjet printers, etc.) and monitors (CRT, LCD, etc.) that output image data processed by the image processing unit 5 It is composed of an image output unit 6 of the above and a driver unit 7 connected to each of the above components and controlling each of these components.
【0015】
Further, as shown in FIG. 2, the voice recognition unit 3 includes a voice data input unit 8 for inputting digital voice data from the voice input unit 2 and an operation unit information input unit for inputting command information from the operation unit 1. A dictionary data storage unit 10 composed of 9 and a storage medium such as a RAM or a hard disk for storing operation unit information and digital audio data in correspondence with each other, and a digital audio data and dictionary data storage unit from the audio data input unit 8. A pattern matching unit 11 that performs pattern matching with the dictionary voice data stored in 10 and calculates the similarity between each command, and a driver information input unit that inputs various command information from the driver unit 7. 12, the command processing unit 13 that performs predetermined command processing based on the information from the pattern matching unit 11, the operation unit information input unit 9, and the driver information input unit 12, and the commands output from the command processing unit 13. It is equipped with a command output unit 14 that sends information to the driver unit 7 and the operation unit 1 as appropriate.
【0016】
FIG. 3 is a plan view of the operation unit 1, which has various key groups 15 and a mode display unit 16.
【0017】
Specifically, the various key groups 15 include a numeric keypad 17 having a numeric keypad, an ID key 17a, etc., a copy key 18 to be operated when executing a copy operation, and a stop key to be operated when the copy operation is interrupted. It has 19, a voice registration key 20 to be operated when registering voice data of a command, and a reset key 21. A liquid crystal display panel 16a is provided at an appropriate position above the mode display unit 16.
【0018】
FIG. 4 is a flowchart showing the procedure for registering voice commands.
【0019】
First, in step S1, the voice registration key 20 is pressed. As a result, as shown in FIG. 5, the message "Enter the ID number with the numeric keypad" is displayed on the liquid crystal panel display unit 16a of the operation unit 1.
【0020】
Next, in step S2, the numeric keypad 17 is operated to input a predetermined ID number (identification information), and then one key selected from the various key groups 15 is pressed (step S3), and the key is pressed. Input the "reading" of the key pressed while pressing it as voice data (step S4). For example, when the copy key 18 is pressed, "copy" is uttered, the voice data "copy" is input to the voice input unit 2, and when "1" in the ten key 17 is pressed, "Ichimai" is used. Is uttered and the voice data "Ichimai" is input to the voice input unit 2. Then, when it is determined that the key input and the voice input are performed at the same time, the copier is set to the voice data registration mode in the following step S5, then proceeds to the step S6, and the voice data by the pattern matching unit 11 Pattern matching is performed, and the similarity between the input command and the same command registered in the dictionary data storage unit 10 is calculated.
【0021】
Next, the process proceeds to step S7 to determine whether or not a similar command is registered in the dictionary data storage unit 10. Here, whether or not a similar command is registered is determined by the calculation result of the similarity, and in the present embodiment, when the input command and the registration command registered in the dictionary data storage unit 10 completely match. Is set to "100", and if the similarity is "90" or more, it is judged that the input command and the registration command are almost the same, and if the similarity is "80 or more and less than 90", the input command is It is determined that a candidate command that can be a candidate has been registered, and if the similarity is less than "80", it is determined that an unregistered command has been input, and it is determined whether or not a similar command has already been registered. For example, when "Ichimai", "Hachimai", "Sanmai", and "Kopee" are registered in the dictionary data storage unit 10 as dictionary audio data, and the audio data "Ichimai" is input. Is judged to have a similarity with the dictionary voice data "Ichimai" as "100", a similarity with the dictionary voice data "Hachimai" as "85", and the dictionary voice data "Sanmai". The similarity with "" is judged to be "30", the similarity with the dictionary voice data "copy" is judged to be "5", and the similar command is registered when the similar command is 80 or more. Judge.
【0022】
Then, when it is determined that a similar command is not registered, that is, when the answer in step S7 is affirmative (Yes), the input voice data is registered in the dictionary data storage unit 10 as command information (step S8). The voice command registration process ends.
【0023】
On the other hand, if the answer in step S7 is negative (No), for example, if the voice data "Ichimai" is input and the matching result in the pattern matching unit 11 is the similarity "90", there is a similar command. Then, the process proceeds to step S9, and as shown in FIG. 6, the message "Is the current command" one "?" Is displayed on the liquid crystal display panel 16a, and the "YES" and "NO" selection keys are displayed. Display 16b (step S9). Then, in the case of "YES", the YES key is pressed, the user confirms that "Ichimai" has already been registered, and ends the voice command registration process.
【0024】
When the "NO" key is selected on the liquid crystal display panel 16a, the registration procedure is repeated from step S1.
【0025】
FIG. 7 is a flowchart of the voice command recognition procedure.
【0026】
When the user presses the ID key 17a of the operation unit 1 in step S11, the message "Enter the ID number with the numeric keypad" is displayed on the liquid crystal panel display unit 16a of the operation unit 1 as in FIG. 5 described above. Is displayed.
【0027】
Next, in step S12, the numeric keypad 17 is operated to input a predetermined ID number, voice data is input to the voice input unit 2 (step S13), and the mode is set to the voice data execution mode (step S14).
【0028】
Next, the process proceeds to step S15, pattern matching of voice data is performed by the pattern matching unit 11, and the degree of similarity between the input command and the registration command is calculated.
【0029】
In the following step S16, it is determined whether or not a similar command is registered, and if a similar command is registered, it is determined in step S17 whether or not a plurality of similar commands are registered. Then, if the answer is negative (No), that is, if there is one similar command, the desired copy operation corresponding to the input voice command is executed (step S18), and the process is terminated.
【0030】
If the answer in step S17 is affirmative (Yes), that is, when a plurality of similar commands are searched, a predetermined message is displayed on the liquid crystal display panel 16a of the operation unit 1. For example, when the voice data "Ichimai" is input, the similarity with "Hachimai" is "85", so there are two candidate commands, "Ichimai" and "Hachimai". As shown in FIG. 8, the liquid crystal display panel 16a displays the message "Which command is the current command?" And two candidate commands 16c, that is, the candidate commands "1" and "8" are displayed on the liquid crystal display. Display on display panel 16a. Then, in step S20, a voice command, for example, pressing the candidate command "1 sheet" key to select the voice command "Ichimai", and then proceeding to step S18, the driver unit 7 performs a copy process based on such a command. End the process.
【0031】
If it is determined in step S16 that there is no similar command, that is, if unregistered voice data is input, the process proceeds to step S21, and as shown in FIG. 9, "Unregistered" is displayed on the liquid crystal display panel 16a. Please press the key to register. Is displayed.
【0032】
Next, in step S22, a desired key is selected from the various key groups 15 and operated to register the desired voice command in the dictionary data storage unit 10, and then in step S18, the driver unit 7 issues such a command. Based on this, the copy process is performed and the process is terminated.
【0033】
As described above, according to the present embodiment, since the desired voice data is registered each time by associating the ID number of each individual with the voice command, all the commands that can be used are registered in advance. In addition to saving the trouble of storing, only voice commands frequently used by each individual can be arbitrarily registered at their own discretion, and the capacity of the storage medium can be reduced.
【0034】
[Effect of the invention]
As described in detail above, according to the present invention, since desired voice data is registered as needed in association with the user's identification information, only the voice data considered to be necessary by the user according to the frequency of use of each user is registered. It is possible to deal with a model used by an unspecified number of people, such as a commercial copier, even if the storage medium has a relatively small capacity.
[Simple explanation of drawings]
[Figure 1]
It is a block block diagram which shows one Embodiment of the copying machine as a voice recognition apparatus which concerns on this invention.
[Figure 2]
It is a block block diagram which shows the detail of a voice recognition part.
[Fig. 3]
It is a top view which shows the detail of the operation part.
[Fig. 4]
It is a flowchart which shows the registration procedure of a voice command.
[Fig. 5]
It is a top view which shows an example of the operation part at the time of registration of a voice command.
[Fig. 6]
It is a main part plan view which shows other example of the operation part at the time of registration of a voice command.
[Fig. 7]
It is a flowchart which shows the recognition procedure of a voice command.
[Fig. 8]
It is a main part plan view which shows an example of the operation part at the time of recognition of a voice command.
[Fig. 9]
It is a main part plan view which shows other example of the operation part at the time of recognition of a voice command.
[Explanation of symbols]
1 Operation unit (input means) 4 Voice input unit (voice input means) 10 Dictionary data registration unit (storage means) 11 Pattern matching unit (similarity calculation means) 13 Command processing unit (association means, registration availability determination means, operation processing execution means) 14 Command output section (1st and 2nd registration instruction means, display command means) 17a ID key (identification information input means)
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2009211103A | Cited by | Japan | Examiner |
| US7835913B2 | Cited by | United States of America | Applicant |
| US7516077B2 | Cited by | United States of America | Applicant |
| JP2008003371A | Cited by | Japan | Examiner |
| JP2006514753A | Cited by | Japan | Search report |
| US6948757B2 | Cited by | United States of America | Search report |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 6465399 | Japan | A | |
| JP19990064653 | – | – | – |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Application deemed to be withdrawn because no request for examination was validly filedWithdrawnA300 | A300 | |
| Notification of appointment of power of attorneyRD03 | RD03 |
Numbers
- Publication
- 2000-259172
- Publication, DOCDB
- 2000259172
- Publication, EPODOC
- JP2000259172
- Application
- 11064653
- Application, DOCDB
- 6465399
- Application, EPODOC
- JP19990064653
Titles2
- Japanese
- 【発明の名称】音声認識装置と音声データの認識方法
- English
- Description: A voice recognition device and a method for recognizing voice data.
Classification
- IPC, 5
- G10L15 10
- G10L15 06
- G10L15 22
- G10L15 24
- G10L17 00