Device and method for understanding user intent
Claim Score by NHIP
Abstract
A voice recognizer 3 generates plural voice recognition results from one input speech 2. For each of the voice recognition results, an intent understanding processor 7 estimates an intent to thereby output one or more candidates of intent understanding results and scores of them. A weight calculator 11 calculates standby weights using setting information 9 of a control target apparatus. An intent understanding corrector 12 corrects the scores of the candidates of intent understanding result, using the standby weights, to thereby calculate their final scores, and then selects one from among the candidates of intent understanding result, as an intent understanding result 13, on the basis of the final scores.

Term
7.5 yearsto projected expiry
Projected expiry 31 March 2034, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
10 claims: 3 independent, 7 dependent
- 11An intent understanding device, comprising:a voice recognizer that recognizes one speech spoken in a natural language by a user, to thereby generate plural voice recognition results of highly ranked recognition scores;a morphological analyzer that converts the respective voice recognition results into morpheme strings;an intent understanding processor that estimates an intent about the speech by the user on the basis of each of the morpheme strings, to thereby output from each one of the morpheme strings, one or more candidates of intent understanding result and scores indicative of degrees of likelihood of the candidates and generate the candidates of intent understanding result in descending order of likelihoods of the plural voice recognition results;a weight calculator that calculates respective weights for the candidates of intent understanding result;and an intent understanding corrector that corrects the scores of the candidates of intent understanding result, using the weights, to thereby calculate their final scores, and then selects the candidate of intent understanding result with the final score that satisfies a preset condition first, as the intent understanding result.
- 19Broadest claimClaim Score 49, average(NHIP)An intent understanding method, comprising:recognizing one speech spoken in a natural language by a user, to thereby generate plural voice recognition results of highly ranked recognition scores;converting the respective voice recognition results into morpheme strings;estimating an intent about the speech by the user on the basis of each of the morpheme strings, to thereby output from each one of the morpheme strings, one or more candidates of intent understanding result and scores indicative of degrees of likelihood of the candidates and generate the candidates of intent understanding result in descending order of likelihoods of the plural voice recognition results;calculating respective weights for the candidates of intent understanding result;and correcting the scores of the candidates of intent understanding result, using the weights, to thereby calculate their final scores, and then select the candidate of intent understanding result with the final score that satisfies a preset condition first, as the intent understanding result.
Independent claims2
156 paragraphs in 8 sections, as filed
TECHNICAL FIELD
0001The present invention relates to an intent understanding device for estimating a user intent from a voice recognition result, and a method therefor.
BACKGROUND ART
0002In recent years, attention has been paid to a technology in which an operation of an apparatus is executed using a voice recognition result about a language spoken by a person. This technology is applied to in the voice interfaces in mobile phones, car-navigation devices and the like. As a conventional basic method, there is a method in which, for example, the apparatus stores beforehand a correspondence relationship between an estimated voice recognition result and an operation, and then, when a recognition result of a speech spoken by the user is the estimated one, the operation corresponding to that recognition result is executed.
0003According to this method, in comparison with the case where the user manually causes an operation, the operation can be directly executed through phonetic speech, and thus, this method serves effectively as a short-cut function. At the same time, the user is required to speak a language that the apparatus is waiting for, in order to execute the operation, so that, as the functions to be concerned by the apparatus increase, the languages that the user has to keep in mind increase. Further, generally, among the users, a few of them use the apparatus after fully understanding its operation manual. Thus, the users not understanding the manual do not know how to talk what language for an operation, so that there is a problem that, actually, the user cannot cause an operation through voice without using a command of the function kept in his/her mind.
0004In this respect, as a technology improved in the above problem, the following method is proposed: even if the user does not keep in mind a command for accomplishing the purpose, an apparatus interactively guides the user to thereby lead the user to accomplishment of the purpose. As one important technology for realizing that method, for example, Patent Document 1 discloses a technology for properly estimating the user intent from the speech of the user.
0005The voice processing device in Patent Document 1 has a linguistic dictionary database and a grammar database, for each of plural pieces of intent information indicative of respective plural intents, and further retains information of the commands executed so far, as pre-scores. For each of the plural pieces of intent information, the voice processing device calculates an acoustic score, a language score and a pre-score, each as a score indicative of a degree of conformity, to each piece of intent information, of the voice signal inputted based on the speech of the user, followed by totalizing these scores to obtain a total score, and then selects the intent information with the largest total score. Further, it is disclosed that, based on the total score, the voice processing device puts the selected intent information into execution, puts it into execution after making confirmation with the user, or delete it.
0006However, in Patent Document 1, the defined intents are uniquely identifiable intents in a form, such as “Tell me weather” or “Tell me clock time”, and there is no mention about processing of intents assuming that the intents include a variety of facility names each required for setting, for example, a destination point in a navigation device.
CITATION LIST
0007Patent Document <ul id="ul0001" list-style="none"><li id="ul0001-0001" num="0008">Patent Document 1: Japanese Patent Application Laid-open No. 2011-33680</li></ul>
SUMMARY OF THE INVENTION
Problems to be Solved by the Invention
0009In the voice processing device according to Patent Document 1, a voice recognition result is planned to be provided for each intention, so that merely the selection from among the plural different intents and the determination of execution or deletion of the finally selected intent information, are performed, and thus a next candidate of the voice recognition result is not concerned.
0010For example, in a state of listening to music, when the user speaks “I don't want to listen to music” and as the result, the first candidate of intent “I want to listen to music” and the second candidate of intent “I don't want to listen to music” are obtained, the first candidate of intent “I want to listen music” is selected.
0011Furthermore, even in a state where “‘∘∘’-center” is already set as a destination point of the navigation device, when the user speaks “Stop at ‘∘x’-center” in order to add a route point, and as the result, the first candidate of intent “Stop at ‘∘∘’-center” and the second candidate of intent “Stop at ‘∘x’-center” are provided, the first candidate of intent “Stop at ‘∘∘’-center” is selected.
0012In this manner, the conventional device does not concern the next candidate, and thus there is a problem that it is difficult to properly understand a user intent. As a result, the user has to cancel the selected first candidate and then to speak again.
0013This invention has been made to solve the problems as described above, and an object thereof is to provide an intent understanding device and an intent understanding method by which a user intent is properly understood using an input speech.
Means for Solving the Problems
0014An intent understanding device according to the invention comprises: a voice recognizer that recognizes one speech spoken in a natural language by a user, to thereby generate plural voice recognition results; a morphological analyzer that converts the respective voice recognition results into morpheme strings; an intent understanding processor that estimates an intent about the speech by the user on the basis of each of the morpheme strings, to thereby output from each one of the morpheme strings, one or more candidates of intent understanding result and scores indicative of degrees of likelihood of the candidates; a weight calculator that calculates respective weights for the candidates of intent understanding result; and an intent understanding corrector that corrects the scores of the candidates of intent understanding result, using the weights, to thereby calculate their final scores, and then selects one from among the candidates of intent understanding result, as an intent understanding result, on the basis of the final scores.
0015An intent understanding method comprises: recognizing one speech spoken in a natural language by a user, to thereby generate plural voice recognition results; converting the respective voice recognition results into morpheme strings; estimating an intent about the speech by the user on the basis of each of the morpheme strings, to thereby output from each one of the morpheme strings, one or more candidates of intent understanding result and scores indicative of degrees of likelihood of the candidates; calculating respective weights for the candidates of intent understanding result; and correcting the scores of the candidates of intent understanding result, using the weights, to thereby calculate their final scores, and then selecting one from among the candidates of intent understanding result, as an intent understanding result, on the basis of the final scores.
Effect of the Invention
0016According to the invention, the plural voice recognition results are generated from one speech; the candidates of intent understanding result are generated from each of the voice recognition results; the final scores are calculated by correcting the scores of the candidates of intent understanding result using the weights; and the intent understanding result is selected from among the plural candidates of intent understanding result on the basis of the final scores. Thus, a final intent understanding result can be selected from among the results including not only those for the first candidate of the voice recognition result for the input speech, but also those for the next candidate of the voice recognition result. Accordingly, it is possible to provide an intent understanding device which can properly understand a user intent.
0017According to the invention, the plural voice recognition results are generated from one speech; the candidates of intent understanding result are generated from each of the voice recognition results; the final scores are calculated by correcting the scores of the candidates of intent understanding result using the weights; and the intent understanding result is selected from among the plural candidates of intent understanding result on the basis of the final scores. Thus, a final intent understanding result can be selected from among the results including not only those for the first candidate of the voice recognition result for the input speech, but also those for the next candidate of the voice recognition result. Accordingly, it is possible to provide an intent understanding method by which a user intent can be properly understood.
BRIEF DESCRIPTION OF THE DRAWINGS
0018<figref idref="DRAWINGS">FIG. 1</figref> is a block diagram showing a configuration of an intent understanding device according to Embodiment 1 of the invention.
0019<figref idref="DRAWINGS">FIG. 2</figref> is a block diagram showing a configuration of a navigation device in which the intent understanding device according to Embodiment 1 is incorporated as a voice interface.
0020<figref idref="DRAWINGS">FIG. 3</figref> includes diagrams for illustrating operations of the intent understanding device according to Embodiment 1: an example of setting information is shown at <figref idref="DRAWINGS">FIG. 3(<i>a</i>)</figref>; and an example of a dialogue is shown at <figref idref="DRAWINGS">FIG. 3(<i>b</i>)</figref>.
0021<figref idref="DRAWINGS">FIG. 4</figref> includes diagrams showing output results at respective parts in the intent understanding device according to Embodiment 1: examples of voice recognition results are shown at <figref idref="DRAWINGS">FIG. 4(<i>a</i>)</figref>; and examples of respective candidates of intent understanding result and the like with respect to first-ranked to third-ranked voice recognition results are respectively shown at <figref idref="DRAWINGS">FIG. 4(<i>b</i>)</figref> to <figref idref="DRAWINGS">FIG. 4(<i>d</i>)</figref>.
0022<figref idref="DRAWINGS">FIG. 5</figref> is a table to be used by a weight calculator in the intent understanding device according to Embodiment 1, in which correspondence relationships between constraint conditions and standby weights are defined.
0023<figref idref="DRAWINGS">FIG. 6</figref> is a flowchart showing the operations of the intent understanding device according to Embodiment 1.
0024<figref idref="DRAWINGS">FIG. 7</figref> is a block diagram showing a configuration of an intent understanding device according to Embodiment 2 of the invention.
0025<figref idref="DRAWINGS">FIG. 8</figref> is a diagram for illustrating operations of the intent understanding device according to Embodiment 2, and shows an example of a dialogue.
0026<figref idref="DRAWINGS">FIG. 9</figref> includes diagrams showing output results at respective parts in the intent understanding device according to Embodiment 2: examples of voice recognition results are shown at <figref idref="DRAWINGS">FIG. 9(<i>a</i>)</figref>; and examples of respective candidates of intent understanding result and the like with respect to first-ranked to third-ranked voice recognition results are respectively shown at <figref idref="DRAWINGS">FIG. 9(<i>b</i>)</figref> to <figref idref="DRAWINGS">FIG. 9(<i>d</i>)</figref>.
0027<figref idref="DRAWINGS">FIG. 10</figref> is a diagram showing an example of a hierarchical tree of the intent understanding device according to Embodiment 2.
0028<figref idref="DRAWINGS">FIG. 11</figref> is a list of intents at respective nodes in the hierarchical tree in <figref idref="DRAWINGS">FIG. 10</figref>.
0029<figref idref="DRAWINGS">FIG. 12</figref> is a diagram showing examples of standby weights calculated by a weight calculator in the intent understanding device according to Embodiment 2.
0030<figref idref="DRAWINGS">FIG. 13</figref> is a flowchart showing the operations of the intent understanding device according to Embodiment 2.
0031<figref idref="DRAWINGS">FIG. 14</figref> is a flowchart showing specific operations in Step ST<b>20</b> in <figref idref="DRAWINGS">FIG. 13</figref>.
0032<figref idref="DRAWINGS">FIG. 15</figref> is a block diagram showing a configuration of an intent understanding device according to Embodiment 3 of the invention.
0033<figref idref="DRAWINGS">FIG. 16</figref> is a diagram showing an example of a keyword table of the intent understanding device according to Embodiment 3.
0034<figref idref="DRAWINGS">FIG. 17</figref> is a diagram showing an example of a keyword-corresponding intent of the intent understanding device according to Embodiment 3.
0035<figref idref="DRAWINGS">FIG. 18</figref> includes diagrams showing output results at respective parts in the intent understanding device according to Embodiment 3: examples of voice recognition results are shown at <figref idref="DRAWINGS">FIG. 18(<i>a</i>)</figref>; and examples of respective candidates of intent understanding result and the like with respect to first-ranked to third-ranked voice recognition results are respectively shown at <figref idref="DRAWINGS">FIG. 18(<i>b</i>)</figref> to <figref idref="DRAWINGS">FIG. 18(<i>d</i>)</figref>.
0036<figref idref="DRAWINGS">FIG. 19</figref> is a flowchart showing operations of the intent understanding device according to Embodiment 3.
0037<figref idref="DRAWINGS">FIG. 20</figref> is a flowchart showing specific operations in Step ST<b>31</b> in <figref idref="DRAWINGS">FIG. 19</figref>.
0038<figref idref="DRAWINGS">FIG. 21</figref> is a block diagram showing a modified example of an intent understanding device according to the invention.
0039<figref idref="DRAWINGS">FIG. 22</figref> is a diagram for illustrating operations of an intent understanding device according to the invention, and shows an example of a dialogue.
MODES FOR CARRYING OUT THE INVENTION
0040Hereinafter, for illustrating the invention in more detail, embodiments for carrying out the invention will be described according to the accompanying drawings.
Embodiment 1
0041As shown in <figref idref="DRAWINGS">FIG. 1</figref>, an intent understanding device <b>1</b> according to Embodiment 1 of the invention includes: a voice recognizer <b>3</b> that performs voice recognition of an input speech <b>2</b> spoken by a user, and converts it into texts; a voice recognition dictionary <b>4</b> used for the voice recognition by the voice recognizer <b>3</b>; a morphological analyzer <b>5</b> that decomposes a voice recognition result into morphemes; a morphological analysis dictionary <b>6</b> used for morphological analysis by the morphological analyzer <b>5</b>; an intent understanding processor <b>7</b> that generates candidates of intent understanding result from a morphological analysis result; an intent understanding model <b>8</b> used for estimating an intent of the user by the intent understanding processor <b>7</b>; a setting information storage <b>10</b> in which setting information <b>9</b> of a control target apparatus is stored; a weight calculator <b>11</b> that calculates weights using the setting information <b>9</b> in the setting information storage <b>10</b>; and an intent understanding corrector <b>12</b> that corrects the candidates of intent understanding result using the weights, and then selects to output one from among these candidates, as a final intent understanding result <b>13</b>.
0042The intent understanding device <b>1</b> is configured with an unshown CPU (Central Processing Unit), and when the CPU executes a program stored in an internal memory, the functions as the voice recognizer <b>3</b>, the morphological analyzer <b>5</b>, the intent understanding processor <b>7</b>, the weight calculator <b>11</b>, and the intent understanding corrector <b>12</b>, are implemented.
0043The voice recognition dictionary <b>4</b>, the morphological analysis dictionary <b>6</b>, the intent understanding model <b>8</b> and the setting information storage <b>10</b>, are configured with an HDD (Hard Disk Drive), a DVD (Digital Versatile Disc), a memory, and/or the like.
0044<figref idref="DRAWINGS">FIG. 2</figref> is a block diagram showing a configuration of a navigation device <b>100</b> in which the intent understanding device <b>1</b> is incorporated as a voice interface. The navigation device <b>100</b> is a target to be controlled by voice. A voice input unit <b>101</b> is configured with a microphone and/or the like, and converts the speech spoken by the user into signals followed by outputting them as the input speech <b>2</b> to the intent understanding device <b>1</b>. A navigation controller <b>102</b> is configured with a CPU, etc., and executes searching, guiding and like functions about a route from a current point to a destination point. The setting information <b>9</b> of the destination point and the like is outputted from the navigation controller <b>102</b> to the intent understanding device <b>1</b>. Further, the navigation controller <b>102</b> receives the intent understanding result <b>13</b> from the intent understanding device <b>1</b>, thereby to execute an operation indicated by the intent understanding result <b>13</b> or to output a voice signal about the intent understanding result <b>13</b> to a voice output unit <b>103</b>. The voice output unit <b>103</b> is configured with a speaker and/or the like, and reproduces the voice signal inputted from the navigation controller <b>102</b>.
0045Note that the intent understanding device <b>1</b> and the navigation controller <b>102</b> may be configured using their respective different CPUs, or may be configured using a single CPU.
0046An intent is represented, for example, in such a form as “<main intent>[<slot name>=<Slot value>, . . . ]”. In a specific example, it is represented as “Destination Point Setting [Facility=?]”, “Destination Point Setting [Facility=$Facility$ (=‘∘∘’ Shop)]”, or the like [‘∘∘’ means some name in Japanese]. “Destination Point Setting [Facility=?]” shows a state where the user wants to set a destination point but has not yet determined a specific facility name. “Destination Point Setting [Facility=$Facility$ (=‘∘∘’ Shop)]” shows a state where the user sets a specific facility of “‘∘∘’ Shop” as a destination point.
0047As an intent understanding method performed by the intent understanding processor <b>7</b>, a method such as, for example, a maximum entropy method or the like, may be utilized.
0048Specifically, the intent understanding model <b>8</b> retains therein many sets of words as independent words (hereinafter, referred to as features) such as “Destination point, Setting” and the like, and their correct intents such as “Destination Point Setting [Facility=?]” and the like. The intent understanding processor <b>7</b> extracts the features of “Destination Point, Setting” from the morphological analysis result of the input speech <b>2</b> “I want to set a destination point”, for example, and then estimates which one in the intention understanding model <b>8</b> has how much likelihood, using a statistical method. The intent understanding processor <b>7</b> outputs sets of intents as candidates of intent understanding result and scores indicative of likelihoods of that intents, as a list.
0049In the following, description will be made assuming that the intent understanding processor <b>7</b> executes an intent understanding method utilizing a maximum entropy method.
0050<figref idref="DRAWINGS">FIG. 3(<i>a</i>)</figref> shows an example of the setting information <b>9</b> in Embodiment 1, and <figref idref="DRAWINGS">FIG. 3(<i>b</i>)</figref> shows an example of a dialogue.
0051In the case where the target to be controlled by voice is the navigation device <b>100</b>, in the setting information <b>9</b>, the following information is included: whether a destination point or a route point is set or not; if it is set, the name of the destination point or the route point; and other than those, the type of a displayed map, and the like. The setting information storage <b>10</b> in the intent understanding device <b>1</b> stores the setting information <b>9</b> outputted by the navigation controller <b>102</b> in the navigation device <b>100</b>. In the example in <figref idref="DRAWINGS">FIG. 3(<i>a</i>)</figref>, in the setting information <b>9</b>, the information of “Destination Point: ‘ΔΔ’” and “Route Point:‘∘∘’” is included [‘ΔΔ’ means some name in Japanese].
0052<figref idref="DRAWINGS">FIG. 3(<i>b</i>)</figref> shows that a dialogue proceeds between the navigation device <b>100</b> and the user in order from the top. In the contents of the dialogue, at beginning of each line, “U:” represents an input speech <b>2</b> spoken by the user and “S:” represents a response from the navigation device <b>100</b>.
0053<figref idref="DRAWINGS">FIG. 4</figref> shows examples of output results at respective parts in the intent understanding device <b>1</b>.
0054<figref idref="DRAWINGS">FIG. 4(<i>a</i>)</figref> shows examples of voice recognition results outputted by the voice recognizer <b>3</b>. The voice recognition results are provided as a list in which each of the voice recognition results, such as “‘∘∘’ is as for destination”, is arranged as a set with a likelihood indicative of the degree of likelihood of that voice recognition result, in descending order of the likelihood.
0055<figref idref="DRAWINGS">FIG. 4(<i>b</i>)</figref> shows the candidates of intent understanding result, their scores, their standby weights and their final scores, with respect to the first ranked voice recognition result “‘∘∘’ is as for destination” in the voice recognition results in <figref idref="DRAWINGS">FIG. 4(<i>a</i>)</figref>;
0056<figref idref="DRAWINGS">FIG. 4(<i>c</i>)</figref> shows those with respect to the second ranked voice recognition result “Do not go to ‘∘∘’”; and <figref idref="DRAWINGS">FIG. 4(<i>d</i>)</figref> shows those with respect to the third ranked voice recognition result “Search ‘∘Δ’”. The intent understanding processor <b>7</b> outputs a list including each set of an intent such as “Route Point Setting [Facility=$Facility$]” and its score, as a candidate of intent understanding result. These candidates of intent understanding result are arranged in descending order of the scores. The weight calculator <b>11</b> calculates the standby weight for each of the candidates of intent understanding result outputted by the intent understanding processor <b>7</b>. The intent understanding corrector <b>12</b> calculates the final score using the standby weight, for each of the candidates of intent understanding result outputted by the intent understanding processor <b>7</b>.
0057<figref idref="DRAWINGS">FIG. 5</figref> is a table in which correspondence relationships between constraint conditions and the standby weights are defined.
0058For example, in the case where a destination point of the navigation device <b>100</b> is already set to “ΔΔ”, it is thought that the user is less likely to make once again a speech intended to “Set the destination point to ‘ΔΔ’” as the next speech. Thus, with respect to this constraint condition, the standby weight for the intent “Destination Point Setting [Facility=$Facility$ (=‘ΔΔ’)]” is set to “0.0”. Meanwhile, because there is a possibility that the user changes the destination point to “?” (a place other than ‘ΔΔ’), the standby weight for the intent “Destination Point Setting [Facility=Facility=$Facility$ (=?)]” is set to “1.0”. Further, because the user is less likely to make a speech intended to set a route point to “∘∘” that is the same as the destination point, the standby weight for the intent “Route Point Setting [Facility=$Facility$ (=‘∘∘’)]” is set to “0.0”. Furthermore, because there is a case where the user deletes an already-set route point “∘∘”, the standby weight for the intent “Destination Point Deletion [Facility=$Facility$ (=‘∘∘’)]” is set to “1.0”.
0059As described above, the weight calculator <b>11</b> retains the information of the standby weights each defined beforehand from the probability of occurrence of intent, and selects the standby weight corresponding to the intent on the basis of the setting information <b>9</b>.
0060The intent understanding corrector <b>12</b> corrects the candidates of intent understanding result from the intent understanding processor <b>7</b>, using the following formulae (1). Specifically, the intent understanding corrector <b>12</b> multiplies the likelihood of the voice recognition result acquired from the voice recognizer <b>3</b>, by an intent understanding score of the candidate of intent understanding result acquired from the intent understanding processor <b>7</b>, to thereby calculate a score (This corresponds to “Score” shown in <figref idref="DRAWINGS">FIG. 4(<i>b</i>)</figref> and the like), and then multiplies this score by the standby weight acquired from the weight calculator <b>11</b>, to thereby obtain the final score (This corresponds to “Final Score” shown in <figref idref="DRAWINGS">FIG. 4(<i>b</i>)</figref> and the like). In Embodiment 1, although intent understanding correction is performed using multiplication as in the formulae (1), the correction is not limited to this method.
0000<br />(Likelihood)×(Intent Understanding Score)=(Score) (Score)×(Standby Weight)=(Final Score)
0061Next, operations of the intent understanding device <b>1</b> will be described with reference to the flowchart in <figref idref="DRAWINGS">FIG. 6</figref>.
0062Here, it is assumed that the intent understanding device <b>1</b> is incorporated in the navigation device <b>100</b> as a control target, and a dialogue is started when the user presses down a dialog start button that is not explicitly shown. Further, assuming that the setting information <b>9</b> shown in <figref idref="DRAWINGS">FIG. 3(<i>a</i>)</figref> is stored in the setting information storage <b>10</b>, intent understanding steps with respect to the contents of the dialogue in <figref idref="DRAWINGS">FIG. 3(<i>b</i>)</figref> will be described in detail.
0063The navigation controller <b>102</b>, upon detection that the user presses down the dialogue start button of the navigation device <b>100</b>, causes the voice output unit <b>103</b> to output a speech “Please talk after beep” that is a response for prompting starting of dialogue, and successively to output a beep sound. Further, the intent understanding device <b>1</b> puts the voice recognizer <b>3</b> into a recognizable state, so that it goes into a user-speech waiting state.
0064Then, when the user makes a speech “Do not go to ‘∘∘’” as shown in <figref idref="DRAWINGS">FIG. 3(<i>b</i>)</figref>, the voice input unit <b>101</b> converts the speech into voice data and outputs it to the voice recognizer <b>3</b> in the intent understanding device <b>1</b>. The voice recognizer <b>3</b> in the intent understanding device <b>1</b> converts the input speech <b>2</b> using the voice recognition dictionary <b>4</b> into each text and calculates its likelihood, and then outputs them to the morphological analyzer <b>5</b> (Step ST<b>11</b>).
0065Then, the morphological analyzer <b>5</b> morphologically analyzes each voice recognition result using the morphological analysis dictionary <b>6</b> and outputs the resultant to the intent understanding processor <b>7</b> (Step ST<b>12</b>). For example, the voice recognition result “‘∘∘’ is as for destination” provides a morphological analysis result of “‘∘∘’/noun, ‘is’/postpositional particle in Japanese, ‘destination’/noun, and ‘as for’/postpositional particle in Japanese”.
0066Then, the intent understanding processor <b>7</b> estimates each of the intents from the morphological analysis result using the intent understanding model <b>8</b> and calculates its score, and then outputs them as a candidate of intent understanding result to the intent understanding corrector <b>12</b> (Step ST<b>13</b>). On this occasion, the intent understanding processor <b>7</b> extracts the features used for intent understanding from the morphological analysis result, and estimates the intent by collating the features with the intent understanding model <b>8</b>. From the morphological analysis result about the voice recognition result “‘∘∘’ is as for destination” in <figref idref="DRAWINGS">FIG. 4(<i>a</i>)</figref>, the features “‘∘∘’, destination” are extracted as a list, so that, as shown in <figref idref="DRAWINGS">FIG. 4(<i>b</i>)</figref>, a candidate of intent understanding result of “Route Point Setting [Facility=$Facility$ (=‘∘∘’)]” and its score “0.623”, and a candidate of intent understanding result of “Facility Search [Facility=$Facility$ (=‘∘∘’)]” and its score “0.286”, are obtained.
0067Then, the weight calculator <b>11</b> reads the setting information <b>9</b> from the setting information storage <b>10</b>, and selects the standby weight for each of the intents on the basis of the setting information <b>9</b> and the table as shown in <figref idref="DRAWINGS">FIG. 5</figref>, and then outputs it to the intent understanding corrector <b>12</b> (Step ST<b>14</b>).
0068Then, the intent understanding corrector <b>12</b> calculates by the above formulae (1), the final score of each of the candidates of intent understanding result, using the likelihood of the voice recognition result calculated by the voice recognizer <b>3</b>, the score of the candidate of intent understanding result calculated by the intent understanding processor <b>7</b> and the standby weight selected by the weight calculator <b>11</b> (Step ST<b>15</b>). On this occasion, the intent understanding corrector <b>12</b> calculates each final score in descending order of the likelihoods of the voice recognition results and in descending order of the scores of the candidates of intent understanding result for a common voice recognition result, and evaluates the final score at every calculation. For example, at the time the candidate of intent understanding result with a final score X=0.5 or more is found, the intent understanding corrector <b>12</b> determines that candidate as the final intent understanding result <b>13</b>.
0069In the example in <figref idref="DRAWINGS">FIG. 4</figref>, with respect to the first ranked voice recognition result “‘∘∘’ is as for destination” about the input speech <b>2</b> “Do not go to ‘∘∘’”, the final score becomes “0.0” for the first ranked candidate of intent understanding result “Route Point Setting [Facility=$Facility$ (=‘∘∘’)]” in <figref idref="DRAWINGS">FIG. 4(<i>b</i>)</figref>, and the final score becomes “0.286” for the second ranked candidate “Facility Search [Facility=$Facility$ (=‘∘∘’)]” therein, so that neither of the candidates of intent understanding result satisfies the condition that the final score is X or more (Step ST<b>16</b> “NO”).
0070Accordingly, for the second ranked voice recognition result “Do not go to ‘∘∘’”, the intent understanding device <b>1</b> repeats the processing of Steps ST<b>12</b> to ST<b>15</b> and as the result, obtains the final score “0.589” for the first ranked candidate of intent understanding result “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” in <figref idref="DRAWINGS">FIG. 4(<i>c</i>)</figref> and the final score “0.232” for the second ranked candidate “Facility Search [Facility=$Facility$ (=‘∘∘’)]” therein. Because the final score “0.589” of “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” that is the first ranked candidate of intent understanding result for the second ranked voice recognition result, satisfies the condition of X or more (Step ST<b>16</b> “YES”), at this time, the intent understanding corrector <b>12</b> sends a reply “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” as the final intent understanding result <b>13</b>, to the navigation controller <b>102</b>, and then terminates the processing.
0071Upon receiving the intent understanding result <b>13</b> “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” from the intent understanding device <b>1</b>, the navigation controller <b>102</b> gives an instruction to the voice output unit <b>103</b> to thereby causing it to output, as shown in <figref idref="DRAWINGS">FIG. 3(<i>b</i>)</figref>, a speech of “Will Delete Route Point ‘∘∘’. All right?”. When the user speaks “Yes” in response thereto, the intent understanding device <b>1</b> receives the input speech <b>2</b> about that speech through the voice input unit <b>101</b>, and determines that voice recognition and intent understanding have been properly performed. Further, the intent understanding device <b>1</b> performs voice recognition and intent understanding for the input speech <b>2</b> of “Yes”, and outputs the intent understanding result <b>13</b> to the navigation controller <b>102</b>. The navigation controller <b>102</b> executes operation of deleting the route point ‘∘∘’ according to the intent understanding result <b>13</b>.
0072Accordingly, in the navigation controller <b>102</b>, “Route Point Setting [Facility=$Facility$ (=‘∘∘’)]” having the largest score in the intent understanding results for the largest likelihood in the voice recognition results is not subjected to execution, but “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” is subjected to execution, and thus, an influence of false recognition can be eliminated.
0073Consequently, according to Embodiment 1, the intent understanding device <b>1</b> is configured to include: the voice recognizer <b>3</b> that recognizes one input speech <b>2</b> spoken in a natural language by a user, to thereby generate plural voice recognition results; the morphological analyzer <b>5</b> that converts the respective voice recognition results into morpheme strings; the intent understanding processor <b>7</b> that estimates an intent about the speech by the user on the basis of each of the morpheme strings, to thereby output from each one of the morpheme strings, one or more candidates of intent understanding result and their scores; the weight calculator <b>11</b> that calculates respective standby weights for the candidates of intent understanding result; and the intent understanding corrector <b>12</b> that corrects the scores of the candidates of intent understanding result, using the standby weights, to thereby calculate their final scores, and then selects one from among the candidates of intent understanding result, as the intent understanding result <b>13</b>, on the basis of the final scores. Thus, the final intent understanding result <b>13</b> can be selected from the voice recognition results including not only the first ranked voice recognition result for the input speech <b>2</b>, but also the second or lower ranked voice recognition results therefor. Accordingly, it is possible to provide the intent understanding device <b>1</b> that can properly understand the intent of the user.
0074Further, according to Embodiment 1, the intent understanding processor <b>7</b> is configured to generate the candidates of intent understanding result in descending order of likelihoods of the plural voice recognition results, and the intent understanding corrector <b>12</b> is configured to calculate the final score at every time the intent understanding processor <b>7</b> generates the candidate of intent understanding result and to select the candidate of intent understanding result with the final score that satisfies the preset condition about X, as the intent understanding result <b>13</b>. Thus, the amount of computation by the intent understanding device <b>1</b> can be reduced.
0075Further, according to Embodiment 1, the weight calculator <b>11</b> is configured to calculate the standby weights using setting information <b>9</b> of a control target apparatus (for example, the navigation device <b>100</b>) that operates based on the intent understanding result <b>13</b> selected by the intent understanding corrector <b>12</b>. Specifically, the weight calculator <b>11</b> is configured to have the table as shown in <figref idref="DRAWINGS">FIG. 5</figref> in which the constraint conditions and the standby weights in the respective cases of satisfying said constraint conditions are defined, and to determine, based on the setting information <b>9</b>, whether or not the constrained condition is satisfied, to thereby select each of the standby weights. Thus, it is possible to estimate adequately the intent matched to a situation of the control target apparatus.
Embodiment 2
0076<figref idref="DRAWINGS">FIG. 7</figref> is a block diagram showing a configuration of an intent understanding device <b>20</b> according to Embodiment 2. In <figref idref="DRAWINGS">FIG. 7</figref>, with respect to the same or equivalent parts as in <figref idref="DRAWINGS">FIG. 1</figref>, the same reference numerals are given thereto, so that their descriptions will be omitted here. The intent understanding device <b>20</b> includes a hierarchical tree <b>21</b> in which intents are expressed in a tree structure, and a weight calculator <b>22</b> that calculates a standby weight on the basis of an activated intent among the intents in the hierarchical tree <b>21</b>.
0077<figref idref="DRAWINGS">FIG. 8</figref> shows an example of a dialogue in Embodiment 2. Like in <figref idref="DRAWINGS">FIG. 3(<i>b</i>)</figref>, at beginning of each line, “U:” represents a speech of the user and “S:” represents a response from a control target apparatus (for example, the navigation device <b>100</b> shown in <figref idref="DRAWINGS">FIG. 2</figref>).
0078<figref idref="DRAWINGS">FIG. 9</figref> shows examples of output results at respective parts in the intent understanding device <b>20</b>. At <figref idref="DRAWINGS">FIG. 9(<i>a</i>)</figref>, voice recognition results and their likelihoods outputted by the voice recognizer <b>3</b> are shown. At <figref idref="DRAWINGS">FIG. 9(<i>b</i>) to (<i>d</i>)</figref>, the following are shown: the candidates of intent understanding result and their scores outputted by the intent understanding processor <b>7</b>; the standby weights outputted by the weight calculator <b>22</b>; and the final scores outputted by the intent understanding corrector <b>12</b>. The candidates of intent understanding result for the first ranked voice recognition result of “I don't want to go to ‘∘∘’” in <figref idref="DRAWINGS">FIG. 9(<i>a</i>)</figref>, are shown in <figref idref="DRAWINGS">FIG. 9(<i>b</i>)</figref>; the candidates of intent understanding result for the second ranked voice recognition result of “I want to go through ‘∘∘’” are shown in <figref idref="DRAWINGS">FIG. 9(<i>c</i>)</figref>; and the candidates of intent understanding result for the third ranked voice recognition result of “Set ‘∘∘’ as a destination” are shown in <figref idref="DRAWINGS">FIG. 9(<i>d</i>)</figref>.
0079<figref idref="DRAWINGS">FIG. 10</figref> and <figref idref="DRAWINGS">FIG. 11</figref> show an example of the hierarchical tree <b>21</b>. In the hierarchical tree <b>21</b>, as shown in <figref idref="DRAWINGS">FIG. 10</figref>, nodes each indicative of an intent are arranged in a hierarchical structure, so that the closer the node to the root (uppermost hierarchy), the more abstract the indicated intent becomes, and the closer the node to the leaf (lowermost hierarchy), the more specific the indicated intent becomes. For example, when comparing two intents of the third hierarchical node #9 of “Destination Point Setting [Facility=?]” and the fourth hierarchical node #16 of “Destination Point Setting [Facility=$Facility$ (=‘∘∘’ Shop)]”, the node #9 indicative of a more abstract intent is present at an upper hierarchy, and under that node, the node #16 indicative of an intent filled with a specific slot value (for example, ‘∘∘’ Shop) is placed.
0080The intent “Navigation” of the node #1 placed at the first hierarchy is an abstract node indicative of a unit of navigation functions of the navigation controller <b>102</b>, and at the second hierarchy under that node, the nodes #2 to #5 indicative of the respective navigation functions are placed. For example, the intent “Destination Point Setting [ ]” of the node #4 represents a state where the user want to set a destination point but has not yet determined a specific place. A change to a state where the destination point is set causes transition from the node #4 to the node #9 or the node #16. The example in <figref idref="DRAWINGS">FIG. 10</figref> shows a state where the node #4 is activated according to the speech of the user of “Set a destination” shown in <figref idref="DRAWINGS">FIG. 8</figref>.
0081In the hierarchical tree <b>21</b>, the intent node is activated according to the information outputted by the navigation device <b>100</b>.
0082<figref idref="DRAWINGS">FIG. 12</figref> shows examples of the standby weights calculated by the weight calculator <b>22</b>.
0083Because the intent “Destination Point Setting [ ]” of the node #4 in the hierarchical tree <b>21</b> is activated according to the user speech “Set a destination”, the standby weights of the intents of the nodes #9, #10 in the side of the node #4 toward the branch/leaf are each given as 1.0, and the standby weight of another intent node is given as 0.5.
0084The calculation method of the standby weight by the weight calculator <b>22</b> will be described later.
0085<figref idref="DRAWINGS">FIG. 13</figref> is a flowchart showing the operations of the intent understanding device <b>20</b>. In <figref idref="DRAWINGS">FIG. 13</figref>, the processing in Steps ST<b>11</b> to ST<b>13</b>, ST<b>15</b> and ST<b>16</b> is the same as that in Step ST<b>11</b> to ST<b>13</b>, ST<b>15</b> and ST<b>16</b> in <figref idref="DRAWINGS">FIG. 6</figref>.
0086In Step ST<b>20</b>, with reference to the hierarchical tree <b>21</b>, the weight calculator <b>22</b> calculates the standby weights of the candidates of intent understanding result from the intent understanding processor <b>7</b>, and outputs them to the intent understanding corrector <b>12</b>.
0087<figref idref="DRAWINGS">FIG. 14</figref> is a flowchart showing specific operations in Step ST<b>20</b> in <figref idref="DRAWINGS">FIG. 13</figref>. In Step ST<b>21</b>, the weight calculator <b>22</b> compares the candidate of intent understanding result from the intent understanding processor <b>7</b> with the activated intent in the hierarchical tree <b>21</b>. When the candidate of intent understanding result from the intent understanding processor <b>7</b> is placed in the side of the activated intent toward a branch/leaf of the hierarchical tree <b>21</b> (Step ST<b>22</b> “YES”), the weight calculator <b>22</b> sets a first weight “a” as the standby weight (Step ST<b>23</b>). In contrast, when the candidate of intent understanding result from the intent understanding processor <b>7</b> is placed other than in the side of the activated intent toward a branch/leaf of the hierarchical tree <b>21</b> (Step ST<b>22</b> “NO”), the weight calculator <b>22</b> sets a second weight “b” as the standby weight (Step ST<b>24</b>). In the present Embodiment 2, “a”=1.0 and “b”=0.5 are given. Further, when there is no activated intent node, the standby weight is set to 1.0.
0088Next, the operations of the intent understanding device will be described.
0089The operations of the intent understanding device <b>20</b> are basically the same as the operations of the intent understanding device <b>1</b> in Embodiment 1 described above. The difference between the present Embodiment 2 and Embodiment 1 described above resides in how to calculate the standby weight.
0090In the following, description will be made in detail about intent understanding steps with respect to the contents of the dialogue shown in <figref idref="DRAWINGS">FIG. 8</figref>. Like in Embodiment 1 described above, a case where the intent understanding device <b>20</b> is incorporated in the navigation device <b>100</b> as a control target (shown in <figref idref="DRAWINGS">FIG. 2</figref>) is assumed. Further, it is assumed that the dialogue is started when the user presses down the dialog start button that is not shown. At the time of the first user speech “Set a destination” in <figref idref="DRAWINGS">FIG. 8</figref>, because the navigation device <b>100</b> has acquired no information from the user, the hierarchical tree <b>21</b> in the intent understanding device <b>20</b> is in a state with no activated intent node.
0091Note that, in the hierarchical tree <b>21</b>, the intent node is activated based on the intent understanding result <b>13</b> outputted by the intent understanding corrector <b>12</b>.
0092After the dialogue is started, when the user makes the speech “Set a destination”, the input speech <b>2</b> about that speech is inputted to the intent understanding device <b>20</b>. The input speech <b>2</b> is recognized by the voice recognizer <b>3</b> (Step ST<b>11</b>) and decomposed by the morphological analyzer <b>5</b> into morphemes (Step ST<b>12</b>), so that the candidates of intent understanding result are determined through calculation by the intent understanding processor <b>7</b> (Step ST<b>13</b>). Here, assuming that the user speech “Set a destination” is not falsely recognized but properly recognized and its intent is properly understood, the intent understanding corrector <b>12</b> obtains “Destination Point Setting [ ]” as the intent understanding result <b>13</b>. In order to specify a facility to be set as the destination point, the navigation controller <b>102</b> gives an instruction to the voice output unit <b>103</b> to thereby cause it to output a speech of “Will set a destination point. Please talk the place”. In addition, in the hierarchical tree <b>21</b>, the node #4 corresponding to the intent understanding result <b>13</b> “Destination Point Setting [ ]” is activated.
0093Because the navigation device <b>100</b> made such a response for prompting the next speech, the dialogue with the user continues, so that it is assumed that the user makes a speech of “Set ‘∘∘’ as a destination” as in <figref idref="DRAWINGS">FIG. 8</figref>. The intent understanding device <b>20</b> performs processing in Steps ST<b>11</b>, ST<b>12</b> for the user speech “Set ‘∘∘’ as a destination”. As a result, it is assumed that the respective morphological analysis results are obtained for the voice recognition results “I don't want to go to ‘∘∘’”, “I want to go through ‘∘∘’” and “Set ‘∘∘’ as a destination” shown in <figref idref="DRAWINGS">FIG. 9(<i>a</i>)</figref>. Then, the intent understanding processor <b>7</b> estimates the intent from the morphological analysis result (Step ST<b>13</b>). At this point, it is assumed that the candidates of intent understanding result are provided as “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” and “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” in <figref idref="DRAWINGS">FIG. 9(<i>b</i>)</figref>.
0094Then, the weight calculator <b>22</b> calculates the standby weights with reference to the hierarchical tree <b>21</b> (Step ST<b>20</b>). At this time, the node #4 in the hierarchical tree <b>21</b> is in an activated state, so that the weights are calculated by the weight calculator <b>22</b> according to this state.
0095First, in Step ST<b>21</b>, information of the activated node #4 is transferred from the hierarchical tree <b>21</b> to the weight calculator <b>22</b>, and the candidates of intent understanding result “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” and “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” are transferred from the intent understanding processor <b>7</b> to the weight calculator <b>22</b>. The weight calculator <b>22</b> compares the intent of the activated node #4 with the candidate of intent understanding result, and when the candidate of intent understanding result is placed in the side of the activated node #4 toward a branch/leaf (namely, placed at the node #9 or the node #10) (Step ST<b>22</b> “YES”), the weight calculator sets a first weight “a” as the standby weight (Step ST<b>23</b>). In contrast, when the candidate of intent understanding result is placed other than in the side of the activated node #4 toward a branch/leaf (Step ST<b>22</b> “NO”), the weight calculator <b>22</b> sets a second weight “b” as the standby weight (Step ST<b>24</b>).
0096The first weight “a” is set to a value larger than the second weight “b”. For example, when “a”=1.0 and “b”=0.5 are given, the standby weights are provided as shown in <figref idref="DRAWINGS">FIG. 9(<i>b</i>)</figref>.
0097Then, the intent understanding corrector <b>12</b> calculates by the above formulae (1), the final score of each of the candidates of intent understanding result, using: the likelihood of the voice recognition result calculated by the voice recognizer <b>3</b>; the score of the candidate of intent understanding result calculated by the intent understanding processor <b>7</b>; and the standby weight calculated by the weight calculator <b>22</b> (Step ST<b>15</b>). The final scores are provided as shown in <figref idref="DRAWINGS">FIG. 9(<i>b</i>)</figref>.
0098Then, like in Embodiment 1, the intent understanding corrector <b>12</b> determines whether or not the final score satisfies the condition of X or more (Step ST<b>16</b>). If the condition is also given with X=0.5, with respect to the first ranked voice recognition result “I don't want to go to ‘∘∘’”, neither the final score “0.314” for the candidate of intent understanding result “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” nor the final score “0.127” for the candidate “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” in <figref idref="DRAWINGS">FIG. 9(<i>b</i>)</figref>, satisfies the condition.
0099Accordingly, for the second ranked voice recognition result “I want to go through ‘∘∘’”, the intent understanding device <b>20</b> repeats the processing of Steps ST<b>12</b> to ST<b>14</b>, ST<b>20</b> and ST<b>15</b>. As the result, as shown in <figref idref="DRAWINGS">FIG. 9(<i>c</i>)</figref>, the final score “0.295” for the candidate of intent understanding result “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” and the final score “0.116” for the candidate “Facility Search [Facility=$Facility$ (=‘∘∘’)]” are obtained; however, they also do not satisfy the condition of X or more.
0100Accordingly, for the third ranked voice recognition result “Set ‘∘∘’ as a destination”, the intent understanding device <b>20</b> repeats the processing of Steps ST<b>12</b>, ST<b>13</b>, ST<b>20</b> and ST<b>15</b>, and as the result, as shown in <figref idref="DRAWINGS">FIG. 9(<i>d</i>)</figref>, the final score “0.538” for the candidate of intent understanding result “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” is obtained. This final score satisfies the condition of X or more, so that the intent understanding corrector <b>12</b> outputs “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” as the final intent understanding result <b>13</b>. In the hierarchical tree <b>21</b>, the node #16 is activated based on the intent understanding result <b>13</b>.
0101Upon receiving the intent understanding result <b>13</b> “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” from the intent understanding device <b>20</b>, the navigation controller <b>102</b> gives an instruction to the voice output unit <b>103</b> to thereby causing it to output, as shown in <figref idref="DRAWINGS">FIG. 8</figref>, a speech of “Will set ‘∘∘’ as a destination point. All right?”. When the user speaks “Yes” in response thereto, the intent understanding device <b>20</b> receives the input speech <b>2</b> about that speech through the voice input unit <b>101</b>, and determines that voice recognition and intent understanding have been properly performed. Further, the intent understanding device <b>20</b> performs voice recognition and intent understanding for the input speech <b>2</b> of “Yes”, and then outputs the intent understanding result <b>13</b> to the navigation controller <b>102</b>. The navigation controller <b>102</b> sets ‘∘∘’ as the destination point according to the intent understanding result <b>13</b>, and then causes the voice output unit <b>103</b> to output a speech of “‘∘∘’ is set as a destination point” to thereby notify the user that the destination point setting is done.
0102Consequently, according to Embodiment 2, the weight calculator <b>22</b> is configured to perform weighting so that the candidate of intent understanding result corresponding to an intent expected from a flow of dialogue with the user is likely to be selected by the intent understanding corrector <b>12</b>. Thus, it is possible to estimate adequately the intent matched to a situation of a dialogue between the user and the control target apparatus.
0103Further, according to Embodiment 2, the intent understanding device <b>20</b> is configured to include the hierarchical tree <b>21</b> in which user intents are expressed in a tree structure so that the closer the intent to its root, the more abstract the intent becomes, and the closer the intent to its leaf, the more specific the intent becomes, wherein the weight calculator <b>22</b> performs weighting on the basis of the hierarchical tree <b>21</b> so that the candidate of intent understanding result that is placed in the side, toward the branch/leaf, of the intent corresponding to the intent understanding result <b>13</b> just previously selected, is likely to be selected. In this manner, the intent about the user speech is corrected using the intent hierarchy, so that it is possible to operate the control target apparatus on the basis of the adequate voice recognition result and intent understanding result.
Embodiment 3
0104<figref idref="DRAWINGS">FIG. 15</figref> is a block diagram showing a configuration of an intent understanding device <b>30</b> according to Embodiment 3. In <figref idref="DRAWINGS">FIG. 15</figref>, with respect to the same or equivalent parts as in <figref idref="DRAWINGS">FIG. 1</figref> and <figref idref="DRAWINGS">FIG. 5</figref>, the same reference numerals are given thereto, so that their descriptions will be omitted here. The intent understanding device <b>30</b> includes a keyword table <b>31</b> in which intents and their corresponding keywords are stored; a keyword search processor <b>32</b> that searches an intent corresponding to the morphological analysis result from the keyword table <b>31</b>; and a weight calculator <b>33</b> that calculates the standby weight by comparing the intent corresponding to the keyword with the activated intent in the hierarchical tree <b>21</b>.
0105<figref idref="DRAWINGS">FIG. 16</figref> shows an example of the keyword table <b>31</b>. In the keyword table <b>31</b>, respective sets are stored, each being a set of the intent and its keyword. For example, for the intent “Destination Point Setting [ ]”, the keywords each indicative of a characteristic expression of the intent, such as, “Destination”, “Visit”, “Destination Point” and the like, are given. Each keyword is given for the intent of each of the second or lower hierarchical nodes, other than the intent of the first hierarchical node #1, in the hierarchical tree <b>21</b>.
0106In the following, the intent that corresponds to the keyword is referred to as a keyword-corresponding intent. Further, the intent corresponding to the activated intent node in the hierarchical tree <b>21</b> is referred to as a hierarchical-tree-corresponding intent.
0107<figref idref="DRAWINGS">FIG. 17</figref> shows examples of the voice recognition results outputted by the voice recognizer <b>3</b>, the keywords included in the voice recognition results, and the keyword-corresponding intents searched by the keyword search processor <b>32</b>. The keyword-corresponding intent corresponding to the keyword “Not Go” for the voice recognition result “I don't want to go to ‘∘∘’” is provided as “Route Point Deletion [ ]”; the keyword-corresponding intent corresponding to the keyword “Through” for the voice recognition result “I want to go through ‘∘∘’” is provided as “Route Point Setting [ ]”; and the keyword-corresponding intent corresponding to the keyword “Destination” for the voice recognition result “Set ‘∘∘’ as a destination” is provided as “Destination Point Setting [ ]”.
0108<figref idref="DRAWINGS">FIG. 18(<i>a</i>)</figref> shows examples of the voice recognition results and their likelihoods outputted by the voice recognizer <b>3</b>. <figref idref="DRAWINGS">FIG. 18(<i>b</i>)</figref> to <figref idref="DRAWINGS">FIG. 18(<i>d</i>)</figref> show the candidates of intent understanding result and their scores outputted by the intent understanding processor <b>7</b>, the standby weights outputted by the weight calculator <b>33</b>, and the final scores outputted by the intent understanding corrector <b>12</b>. The candidates of intent understanding result for the first ranked voice recognition result “I don't want to go to ‘∘∘’” in <figref idref="DRAWINGS">FIG. 18(<i>a</i>)</figref>, are shown in <figref idref="DRAWINGS">FIG. 18(<i>b</i>)</figref>; the candidates of intent understanding result for the second ranked voice recognition result “I want to go through ‘∘∘’” are shown in <figref idref="DRAWINGS">FIG. 18(<i>c</i>)</figref>; and the candidates of intent understanding result for the third ranked voice recognition result “Set ‘∘∘’ as a destination” are shown in <figref idref="DRAWINGS">FIG. 18(<i>d</i>)</figref>.
0109<figref idref="DRAWINGS">FIG. 19</figref> is a flowchart showing operations of the intent understanding device <b>30</b>. In <figref idref="DRAWINGS">FIG. 19</figref>, the processing in Steps ST<b>11</b> to ST<b>13</b>, ST<b>15</b> and ST<b>16</b> is the same as that in Step ST<b>11</b> to ST<b>13</b>, ST<b>15</b> and ST<b>16</b> in <figref idref="DRAWINGS">FIG. 6</figref>.
0110In Step ST<b>30</b>, the keyword search processor <b>32</b> searches from the keyword table <b>31</b>, the keyword matched to the morphological analysis result, to thereby acquire the keyword-corresponding intent associated with the searched keyword. The keyword search processor <b>32</b> outputs the acquired keyword-corresponding intent to the weight calculator <b>33</b>.
0111<figref idref="DRAWINGS">FIG. 20</figref> is a flowchart showing specific operations in Step ST<b>31</b> in <figref idref="DRAWINGS">FIG. 19</figref>. In Step ST<b>32</b>, the weight calculator <b>33</b> compares the candidate of intent understanding result from the intent understanding processor <b>7</b> with the hierarchical-tree-corresponding intent activated in the hierarchical tree <b>21</b> and with the keyword-corresponding intent from the keyword search processor <b>32</b>. When the candidate of intent understanding result is matched to neither of the keyword-corresponding intent and the hierarchical-tree-corresponding intent (Step ST<b>32</b> “NO”), the weight calculator <b>33</b> sets a third weight “c” as the standby weight.
0112When the candidate of intent understanding result is matched to the hierarchical-tree-corresponding intent (Step ST<b>32</b> “YES” and Step ST<b>34</b> “YES”), the weight calculator <b>33</b> sets a fourth weight “d” as the standby weight (Step ST<b>35</b>). Note that in the condition of Step ST<b>34</b> “YES”, such a case may arise where the candidate of intent understanding result is matched to both the hierarchical-tree-corresponding intent and the keyword-corresponding intent.
0113When the candidate of intent understanding result is not matched to the hierarchical-tree-corresponding intent but matched to the keyword-corresponding intent only (Step ST<b>34</b> “NO”), the weight calculator <b>33</b> sets a fifth weight “e” as the standby weight (Step ST<b>36</b>).
0114In Embodiment 3, c=0.0, d=1.0 and e=0.5 are given. Namely, when the candidate of intent understanding result is matched to the hierarchical-tree-corresponding intent, the standby weight is 1.0; when it is not matched to the hierarchical-tree-corresponding intent but matched to the keyword-corresponding intent, the standby weight is 0.5; and when it is matched to neither of the keyword-corresponding intent and the hierarchical-tree-corresponding intent, the standby weight is 0.0.
0115Next, the operations of the intent understanding device <b>30</b> will be described.
0116The operations of the intent understanding device <b>30</b> are basically the same as the operations of the intent understanding devices <b>1</b>, <b>20</b> in Embodiments 1, 2 described above. The difference between the present Embodiment 3 and Embodiments 1, 2 described above resides in how to calculate the standby weights.
0117In the following, description will be made in detail about intent understanding steps with respect to the user speech “Set ‘∘∘’ as a destination” in the contents of the dialogue shown in <figref idref="DRAWINGS">FIG. 8</figref>. Like in Embodiments 1, 2 described above, a case where the intent understanding device <b>30</b> is incorporated in the navigation device <b>100</b> as a control target (shown in <figref idref="DRAWINGS">FIG. 2</figref>) is assumed.
0118Further, <figref idref="DRAWINGS">FIG. 10</figref> and <figref idref="DRAWINGS">FIG. 11</figref> are applied to the hierarchical tree <b>21</b> by analogy.
0119The input speech <b>2</b> about the user speech “Set ‘∘∘’ as a destination” is recognized by the voice recognizer <b>3</b> (Step ST<b>11</b>) and decomposed by the morphological analyzer <b>5</b> into morphemes (Step ST<b>12</b>), so that the candidates of intent understanding result are determined through calculation by the intent understanding processor <b>7</b> (Step ST<b>13</b>). Then, as shown in <figref idref="DRAWINGS">FIG. 18(<i>b</i>)</figref>, the candidate of intent understanding result “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” and its score “0.623”, and the candidate “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” and its score “0.127”, are obtained.
0120Then, the keyword search processor <b>32</b> searches from among the keywords in the keyword table <b>31</b>, the keyword matched to the morphological analysis result from the morphological analyzer <b>5</b>, to thereby acquire the keyword-corresponding intent corresponding to the searched keyword. In the morphological analysis result for “I don't want to go to ‘∘∘’”, there is the keyword of “Not Go” in <figref idref="DRAWINGS">FIG. 16</figref>, so that the keyword-corresponding intent is “Route Point Deletion [ ]”.
0121Then, the weight calculator <b>33</b> calculates the standby weights (Step ST<b>31</b>). At this time, the node #4 in the hierarchical tree <b>21</b> is in an activated state, so that the hierarchical-tree-corresponding intent of the node #4 is “Destination Point Setting [ ]”.
0122First, in Step ST<b>32</b>, in the hierarchical tree <b>21</b>, the hierarchical-tree-corresponding intent “Destination Point Setting [ ]” of the activated node #4 is outputted to the weight calculator <b>33</b>. Further, the intent understanding processor <b>7</b> outputs to the weight calculator <b>33</b>, the first ranked candidate of intent understanding result “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” for the user speech “I don't want to go to ‘∘∘’”. Furthermore, the keyword search processor <b>32</b> outputs to the weight calculator <b>33</b>, the keyword-corresponding intent “Route Point Deletion [ ]”.
0123Because the first ranked candidate of intent understanding result “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” is matched to the keyword-corresponding intent “Route Point Deletion [ ]” (Step ST<b>32</b> “YES” and Step ST<b>34</b> “NO”), the weight calculator <b>33</b> sets a fifth weight “e” (=0.5) as the standby weight for the first ranked candidate of intent understanding result (Step ST<b>35</b>).
0124Here, the matching is determined by the weight calculator <b>33</b> even in the case where the intents are in a parent-child relationship in the hierarchical tree <b>21</b>. Thus, “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]”, because it is a child of “Route Point Deletion [ ]”, is determined to be matched thereto.
0125Meanwhile, because the second ranked candidate of intent understanding result “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” is matched to the hierarchical-tree-corresponding intent “Destination Point Setting [ ]”(Step ST<b>32</b> “YES” and Step ST<b>34</b> “YES”), the weight calculator <b>33</b> sets a fourth weight “d” (=1.0) as the standby weight for the second ranked candidate of intent understanding result (Step ST<b>36</b>).
0126Finally, as shown in <figref idref="DRAWINGS">FIG. 18(<i>b</i>)</figref>, with respect to the first ranked voice recognition result “I don't want to go to ‘∘∘’”, the final score “0.312” for the first ranked candidate of intent understanding result “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” and the final score “0.127” for the second ranked candidate of intent understanding result “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” are obtained. Because neither the first or second ranked candidate satisfies the condition that the final score is X or more, the intent understanding device <b>30</b> performs processing in Steps ST<b>12</b>, ST<b>13</b>, ST<b>30</b>, ST<b>31</b> and ST<b>15</b> for the second ranked voice recognition result “I want to go through ‘∘∘’”.
0127As the result, as shown in <figref idref="DRAWINGS">FIG. 18(<i>c</i>)</figref>, with respect to “I want to go through ‘∘∘’”, the standby weight “0.0” (=c) is set to each of the first ranked candidate of intent understanding result “Route Point Deletion [Facility=$Facility$ (=‘∘∘’)]” and the second ranked candidate of intent understanding result “Facility Search [Facility=$Facility$ (=‘∘∘’)]”. Thus, their final scores each become “0.0” and do not satisfy the condition of X or more, even here.
0128Thus, the target to be processed is changed to the third ranked voice recognition result “Set ‘∘∘’ as a destination”, so that, as shown in <figref idref="DRAWINGS">FIG. 18(<i>d</i>)</figref>, the first ranked candidate of intent understanding result “Destination Point Setting [Facility=$Facility$ (=‘∘∘’)]” is, because its final score satisfies the condition of X or more, outputted as the intent understanding result <b>13</b>. Accordingly, like in Embodiment 2 described above, “∘∘” is set as the destination point.
0129Consequently, according to Embodiment 3, the intent understanding device <b>30</b> is configured to include the keyword search processor <b>32</b> that searches, from among the keywords in the keyword table <b>31</b> in which correspondence relationships between the intents and the keywords are defined, the keyword matched to the morpheme string, to thereby acquire the keyword-corresponding intent corresponding to the searched keyword, wherein the weight calculator <b>33</b> calculates each of the standby weights using the hierarchical-tree-corresponding intent and the keyword-corresponding intent. Thus, it is possible to correct the intent about the user speech using the intent hierarchy and the keyword indicative of a characteristic expression of the intent, to thereby operate the control target apparatus on the basis of the adequate voice recognition result and intent understanding result.
0130Note that in Embodiments 1 to 3 described above, although the description has been made about the case of Japanese language, each embodiment can be applied to the cases of a variety of languages in English, German, Chinese and the like, by changing the extraction method of the feature related to the intent estimation for each of the respective languages.
0131Further, in the case of the language whose word is partitioned by a specific symbol (for example, a space), when its linguistic structure is difficult to be analyzed, it is also allowable to extract from the natural language text of the input speech <b>2</b>, using a pattern matching or like method, a slot value such as $Facility$, $Residence$ or the like, and thereafter to directly execute the intent understanding processing.
0132Furthermore, in Embodiments 1 to 3 described above, the text of the voice recognition result is analyzed by the morphological analyzer <b>5</b> to thereby get ready for the intent understanding processing; however, though depending on the recognition method of the voice recognizer <b>3</b>, there is a case where the voice recognition result itself includes the morphological analysis result, so that, in that case, it is allowable to omit operations using the morphological analyzer <b>5</b> and the morphological analysis dictionary <b>6</b>, to thereby execute the intent understanding processing directly after the voice recognition processing.
0133Further, in Embodiments 1 to 3 described above, although an intent understanding method has been described using the case where the application of a learning model by a maximum entropy method is assumed, the intent understanding method is not limited thereto.
0134In addition, in Embodiment 3 described above, although the weight calculator <b>33</b> is configured to calculate the standby weight using the hierarchical-tree-corresponding intent and the keyword-corresponding intent, it is also allowable that the weight calculator calculates the standby weight without using the hierarchical tree <b>21</b> in such a manner that the score of the candidate of intent understanding result is changed according to the number of times of the keyword in the keyword table <b>31</b> emerging in the morphological analysis result.
0135For example, when a word that is important for specifying the intent, such as “Not Go” or “Through”, emerges in the user speech, the intent understanding processor <b>7</b> usually performs, for the user speech “I don't want to go to ‘∘∘’”, intent understanding processing using the features of “‘∘∘’, Not Go”. Instead, when a keyword included in the keyword table <b>31</b> is repeated in a manner like “‘∘∘’, Not Go, Not Go”, this allows the intent understanding processor <b>7</b> to calculate the score weighted according to the number of words of “Not Go”, at the time of intent estimation.
0136Further, in Embodiments 1 to 3 described above, the intent understanding processing is performed in descending order of likelihoods of the plural voice recognition results, and at the time the candidate of intent understanding result with the final score that satisfies the condition of X or more is found, the processing is terminated; however, when the intent understanding device has margin for computation processing, the following method is also applicable: for all of the voice recognition results, the intent understanding processing is performed and then, the intent understanding result <b>13</b> is selected.
0137Furthermore, in Embodiment 1 to 3 described above, before execution of the operation corresponding to the intent understanding result <b>13</b>, whether the execution is allowable or not is confirmed by the user (for example, “Will Delete Route Point ‘∘∘’. All right?” in <figref idref="DRAWINGS">FIG. 3(<i>b</i>)</figref>); however, whether to make such confirmation or not may be switched according to the final score of the intent understanding result <b>13</b>.
0138Further, whether to make the confirmation or not may be switched according to the ranking, for example, in such a manner that when the candidate of intent understanding result for the first ranked voice recognition result is selected as the intent understanding result <b>13</b>, no confirmation is made, and when the candidate of intent understanding result for the second or lower ranked voice recognition result is selected as the intent understanding result <b>13</b>, the confirmation is made.
0139Further, whether to make the confirmation or not may be switched according to the magnitude of the score, for example, in such a manner that when the candidate of intent understanding result with the highest score before correction by the standby weight, is selected as the intent understanding result <b>13</b>, no confirmation is made, and when the candidate of intent understanding result with the score lower than the above, is selected as the intent understanding result <b>13</b>, the confirmation is made.
0140Here, an intent understanding device <b>40</b> as a modified example is shown in <figref idref="DRAWINGS">FIG. 21</figref>. The intent understanding device <b>40</b> includes: a voice input unit <b>41</b> that converts the speech spoken by the user into signals to thereby acquire them as an input speech; an intent confirmation processor <b>42</b> that, when the intent understanding corrector <b>12</b> excludes the candidate of intent understanding result with the highest likelihood (namely, the candidate of intent understanding result with the highest score before correction by the standby weight) and selects the candidate of intent understanding result other than the excluded one, as the intent understanding result <b>13</b>, determines acceptance or non-acceptance of that intent understanding result <b>13</b> after making confirmation with the user on whether or not to accept that result; and a voice output unit <b>43</b> that outputs a voice signal generated by the intent confirmation processor <b>42</b> and used for the confirmation of the intent understanding result. These voice input unit <b>41</b>, intent confirmation processor <b>42</b> and voice output unit <b>43</b> serve the same role as the voice input unit <b>101</b>, the navigation controller <b>102</b> and the voice output unit <b>103</b> shown in <figref idref="DRAWINGS">FIG. 2</figref>, and make confirmation with the user on whether or not to accept the intent understanding result <b>13</b>, by outputting, for example, a speech of “Will Delete Route Point ‘∘∘’. All right?”, as in <figref idref="DRAWINGS">FIG. 3(<i>b</i>)</figref>.
0141Note that the method of making confirmation with the user may be other than that by outputting a voice, and may be that by displaying a screen, or the like.
0142In addition, in Embodiments 2, 3 described above, although the intent hierarchy is expressed in a tree structure as the hierarchical tree <b>21</b>, it is not necessarily expressed in a complete tree structure, and when it is expressed in a graph structure including no loop structure, this allows the processing similar to the above.
0143Furthermore, in Embodiments 2, 3 described above, only a user speech currently made is used for the intent understanding processing; however, in the case where the speech is made in the middle of hierarchical transition in the hierarchical tree <b>21</b>, the intent understanding processing may be performed using the features extracted from plural speeches including a user speech previously made. This makes it possible to estimate an intent that is difficult to be estimated from partial information obtained by plural fragmentary speeches.
0144Here, description will be made using the contents of the dialogue shown in <figref idref="DRAWINGS">FIG. 22</figref>.
0145In the case of Embodiment 2, “Destination, Set” are extracted as features from the first user speech “Set a destination”. Further, “$Facility$ (=‘∘∘’)” is extracted as a feature from the second speech “∘∘”. As the result, for the second speech, the intent understanding processing is usually performed only using “$Facility$ (=‘∘∘’)” (Step ST<b>13</b> in <figref idref="DRAWINGS">FIG. 13</figref>).
0146In contrast, when whether the speech is in the middle of hierarchical transition or not is taken into consideration, because the first speech “Set a destination” corresponds to the node #4 in the hierarchical tree <b>21</b> and the second speech is likely to be in a parent-child relationship with the node #4, a more adequate intent understanding result is obtained in such a manner that the intent understanding processing is performed for the second speech using three features of “Destination, Set, $Facility$ (=‘∘∘’)”.
0147Further, in Embodiments 1 to 3 described above, the navigation device <b>100</b> in <figref idref="DRAWINGS">FIG. 2</figref> is cited as an example of a control target apparatus by the intent understanding device; however, the control target apparatus is not limited to a navigation device. Further, in <figref idref="DRAWINGS">FIG. 2</figref>, although the intent understanding device is incorporated in the control target apparatus, it may be provided externally.
0148It should be noted that, other than the above, unlimited combination of the respective embodiments, modification of any configuration element in the embodiments and omission of any configuration element in the embodiments may be made in the present invention without departing from the scope of the invention.
INDUSTRIAL APPLICABILITY
0149As described above, the intent understanding device according to the invention is configured to estimate the intent of the user using an input speech, and thus it is suited to be used as a voice interface in a car-navigation device or the like that is difficult to be operated manually.
DESCRIPTION OF REFERENCE NUMERALS AND SIGNS
0150<b>1</b>, <b>20</b>, <b>30</b>, <b>40</b>: intent understanding device, <b>2</b>: input speech, <b>3</b>: voice recognizer, <b>4</b>: voice recognition dictionary, <b>5</b>: morphological analyzer, <b>6</b>: morphological analysis dictionary, <b>7</b>: intent understanding processor, <b>8</b>: intent understanding model, <b>9</b>: setting information, <b>10</b>: setting information storage, <b>11</b>, <b>22</b>, <b>33</b>: weight calculator, <b>12</b>: intent understanding corrector, <b>13</b>: intent understanding result, <b>21</b>: hierarchical tree, <b>31</b>: keyword table, <b>32</b>: keyword search processor, <b>41</b>, <b>101</b>: voice input unit, <b>43</b>, <b>103</b>: voice output unit, <b>42</b>: intent confirmation processor, <b>100</b>: navigation device, <b>102</b>: navigation controller.
Contents8
16 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US11516537B2 | Cited by | United States of America | Applicant |
| US11380310B2 | Cited by | United States of America | Applicant |
| US10453443B2 | Cited by | United States of America | Applicant |
| US12009007B2 | Cited by | United States of America | Applicant |
| US10741181B2 | Cited by | United States of America | Applicant |
| US11675829B2 | Cited by | United States of America | Applicant |
| US11580990B2 | Cited by | United States of America | Applicant |
| US12393782B2 | Cited by | United States of America | Search report |
| US10438595B2 | Cited by | United States of America | Applicant |
| US11809783B2 | Cited by | United States of America | Applicant |
| US2022383854A1 | Cited by | United States of America | Search report |
| US10748546B2 | Cited by | United States of America | Applicant |
| US10474753B2 | Cited by | United States of America | Applicant |
| US11475898B2 | Cited by | United States of America | Applicant |
| US11017770B2 | Cited by | United States of America | Search report |
| US10741185B2 | Cited by | United States of America | Applicant |
| US11838734B2 | Cited by | United States of America | Applicant |
| US11496600B2 | Cited by | United States of America | Applicant |
| US11360739B2 | Cited by | United States of America | Applicant |
| US11696060B2 | Cited by | United States of America | Applicant |
| US10831996B2 | Cited by | United States of America | Search report |
| US11301477B2 | Cited by | United States of America | Applicant |
| US2024134847A1 | Cited by | United States of America | Search report |
| US10984798B2 | Cited by | United States of America | Applicant |
| KR20210044188A | Cited by | Republic of Korea | Search report |
| US11842734B2 | Cited by | United States of America | Applicant |
| US12154571B2 | Cited by | United States of America | Applicant |
| US11010127B2 | Cited by | United States of America | Applicant |
| US10657966B2 | Cited by | United States of America | Applicant |
| US2020335187A1 | Cited by | United States of America | Search report |
| US10455322B2 | Cited by | United States of America | Applicant |
| US10942703B2 | Cited by | United States of America | Applicant |
| KR20200007969A | Cited by | Republic of Korea | Search report |
| US10956666B2 | Cited by | United States of America | Applicant |
| US12367879B2 | Cited by | United States of America | Applicant |
| US10417405B2 | Cited by | United States of America | Applicant |
| US12288039B1 | Cited by | United States of America | Applicant |
| US10681212B2 | Cited by | United States of America | Applicant |
| US11810562B2 | Cited by | United States of America | Applicant |
| US12073147B2 | Cited by | United States of America | Applicant |
| US11314370B2 | Cited by | United States of America | Applicant |
| US11087759B2 | Cited by | United States of America | Applicant |
| US11947548B2 | Cited by | United States of America | Search report |
| US12165635B2 | Cited by | United States of America | Applicant |
| US11127397B2 | Cited by | United States of America | Applicant |
| US11133008B2 | Cited by | United States of America | Applicant |
| US11638059B2 | Cited by | United States of America | Applicant |
| US11170166B2 | Cited by | United States of America | Applicant |
| US12614042B2 | Cited by | United States of America | Applicant |
| US11550542B2 | Cited by | United States of America | Applicant |
| US11360577B2 | Cited by | United States of America | Applicant |
| KR20200072907A | Cited by | Republic of Korea | Search report |
| US10417344B2 | Cited by | United States of America | Applicant |
| US10323953B2 | Cited by | United States of America | Search report |
| US11675491B2 | Cited by | United States of America | Applicant |
| US11538469B2 | Cited by | United States of America | Applicant |
| US10810274B2 | Cited by | United States of America | Applicant |
| US10529332B2 | Cited by | United States of America | Applicant |
| US10390213B2 | Cited by | United States of America | Applicant |
| US12136419B2 | Cited by | United States of America | Applicant |
| US10777197B2 | Cited by | United States of America | Applicant |
| US11954405B2 | Cited by | United States of America | Applicant |
| US11798547B2 | Cited by | United States of America | Applicant |
| US11120372B2 | Cited by | United States of America | Applicant |
| US10892996B2 | Cited by | United States of America | Applicant |
| US11646025B2 | Cited by | United States of America | Applicant |
| US10580409B2 | Cited by | United States of America | Applicant |
| US11145298B2 | Cited by | United States of America | Applicant |
| US11062710B2 | Cited by | United States of America | Applicant |
| US2019362217A1 | Cited by | United States of America | Search report |
| US12067985B2 | Cited by | United States of America | Applicant |
| US12254887B2 | Cited by | United States of America | Applicant |
| US12386434B2 | Cited by | United States of America | Applicant |
| US10403283B1 | Cited by | United States of America | Applicant |
| US12277954B2 | Cited by | United States of America | Applicant |
| US11750962B2 | Cited by | United States of America | Applicant |
| US11809886B2 | Cited by | United States of America | Applicant |
| US10417266B2 | Cited by | United States of America | Applicant |
| US12100389B2 | Cited by | United States of America | Search report |
| CN112513833A | Cited by | China | Search report |
| US11423908B2 | Cited by | United States of America | Applicant |
| US11886805B2 | Cited by | United States of America | Applicant |
| US11947873B2 | Cited by | United States of America | Applicant |
| US10878809B2 | Cited by | United States of America | Applicant |
| US11010561B2 | Cited by | United States of America | Applicant |
| US11636869B2 | Cited by | United States of America | Applicant |
| US12026197B2 | Cited by | United States of America | Applicant |
| US10714117B2 | Cited by | United States of America | Applicant |
| US11671920B2 | Cited by | United States of America | Applicant |
| US10699717B2 | Cited by | United States of America | Applicant |
| US10546001B1 | Cited by | United States of America | Search report |
| US11227589B2 | Cited by | United States of America | Applicant |
| US11009970B2 | Cited by | United States of America | Applicant |
| CN109634692A | Cited by | China | Search report |
| US11699448B2 | Cited by | United States of America | Applicant |
| US10311144B2 | Cited by | United States of America | Applicant |
| US10719524B1 | Cited by | United States of America | Applicant |
| US11468282B2 | Cited by | United States of America | Applicant |
| US11126400B2 | Cited by | United States of America | Applicant |
| US11237797B2 | Cited by | United States of America | Applicant |
8 members in 5 offices
Priority claims1
| Document | Office | Kind | Date |
|---|---|---|---|
| 2014059445 | Japan | W |
Members8
| Document | Office | Kind | |
|---|---|---|---|
| WO2015151157A1 | World Intellectual Property Organization (WIPO) | A1 | |
| DE112014006542T5 | Germany | T5 | |
| US2017011742A1 | United States of America | A1 | |
| JPWO2015151157A1 | Japan | A1 | |
| CN106663424A | China | A | |
| US10037758B2 | United States of America | B2 | |
| CN106663424B | China | B | |
| DE112014006542B4 | Germany | B4 |
54 transactions on the USPTO file
Allowed after 1 non-final rejection and 1 final rejection.
- Non-final rejections
- 1
- Final rejections
- 1
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Payment of Maintenance Fee, 8th Year, Large EntityM1552 | M1552 | |
| Payment of Maintenance Fee, 4th Year, Large EntityM1551 | M1551 | |
| Post Issue Communication - Certificate of CorrectionN423 | N423 | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Email NotificationEML_NTR | EML_NTR | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Reasons for AllowanceEX.R | EX.R | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| Miscellaneous Incoming LetterLET. | LET. | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Supplemental ResponseSA.. | SA.. | |
| Mail Interview Summary - Applicant Initiated - TelephonicMEXAT | MEXAT | |
| Interview Summary - Applicant Initiated - TelephonicEXAT | EXAT | |
| Response after Non-Final ActionA... | A... | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Email NotificationEML_NTR | EML_NTR | |
| Application ready for PDX access by participating foreign officesCCRDY | CCRDY | |
| PG-Pub Issue NotificationPG-ISSUE | PG-ISSUE | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Application Is Now CompleteCOMP | COMP | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Email NotificationEML_NTR | EML_NTR | |
| Email NotificationEML_NTR | EML_NTR | |
| Filing ReceiptFLRCPT.O | FLRCPT.O | |
| Notice of DO/EO Acceptance MailedM903 | M903 | |
| Sent to Classification ContractorPGPC | PGPC | |
| FITF set to YES - revise initial settingFTFS | FTFS | |
| 371 Completion Date371COMP | 371COMP | |
| Patent Term Adjustment - Ready for ExaminationPTA.RFE | PTA.RFE | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Cleared by OIPE CSRL194 | L194 | |
| Entity Status Set To Undiscounted (Initial Default Setting or Status Change)BIG. | BIG. | |
| Initial Exam Team nnIEXX | IEXX |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Maintenance fee paymentMAFP | MAFP | |
| Maintenance fee paymentMAFP | MAFP | |
| Certificate of correctionCC | CC | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| AssignmentAS | AS |
Numbers
- Publication
- 20170011742
- Application
- 15120539
Titles
- English
- DEVICE AND METHOD FOR UNDERSTANDING USER INTENT
Patent term adjustment
- Applicant delay
- −12 days
- Net adjustment
- 0 days
Classification
- CPC, 11
- G10L15/22
- G10L15/1822
- G01C21/3608
- G10L15/1815
- G06F17/2755
- G10L15/183
- G10L2015/088
- G06F40/268
- G06F40/30
- G10L15/26
- G10L2015/223
- IPC, 4
- G10L15 22
- G06F17 27
- G01C21 36
- G10L15 18