Voice interaction method, and device
Claim Score by NHIP
Abstract
A voice dialogue method performed by a voice dialog system includes: a voice signal generation unit; a voice dialog agent unit; a voice output unit; and a voice input control unit, the method including: a step of, by the voice signal generation unit, receiving a voice input and generating a voice signal based on the received voice input; a step of, by the voice dialog agent unit, performing voice recognition processing on the voice signal and performing processing based on a result of the voice recognition processing to generate a response signal; a step of, by the voice output unit, outputting a voice based on the response signal; and a step of, when the voice output unit outputs the voice, by the voice input control unit, keeping the voice signal generation unit, for predetermined period after output of the voice, a receivable state in which a voice input is receivable.

Term
7.7 yearsleft in the term
Expires 10 June 2034.
- Priority
- Filed
- Granted
- Today
- Expires
9 claims: 1 independent, 8 dependent
- 1Broadest claimClaim Score 6, narrow(NHIP)A voice dialogue method that is performed by a voice dialogue system, the voice dialogue system including:a voice signal generation unit;a voice dialogue agent unit;an additional voice dialogue agent unit;a voice output unit;and a voice input control unit a first device;a second device;a first voice dialogue agent server;and a second voice dialogue agent server , wherein the first device is a computer-embedded device that is capable of connecting to a network and performing input and output via a voice with a user, the first voice dialogue agent server and the second voice dialogue agent server are each a voice dialogue agent server that is accessed by the first device via the network, and is capable of performing, as an agent for the first device, recognition of a voice inputted by the device and synthesizing of a voice to be outputted by the first device, and the voice dialogue method comprising comprises : a step of, by the voice signal generation unit, receiving a voice input and at the first device;generating a voice signal at the first device based on the received voice input , and transmitting the voice signal from the first device to the first voice dialogue agent server ;a step of, by the voice dialogue agent unit, performing voice recognition processing on the generated voice signal and at the first voice dialogue agent server to generate first text input;determining, at the first voice dialogue agent server based on a result of the voice recognition processing the generated first text input and agent information, which one of the voice dialogue agent unit first voice dialogue agent server and the additional voice dialogue agent unit second voice dialogue agent server is appropriate for performing voice-related processing that is processing based on the voice signal, the agent information being stored in a memory included in the voice dialogue agent unit first voice dialogue agent server and associating the additional voice dialogue agent unit second voice dialogue agent server with one or more keywords;a step of, when the voice dialogue agent unit determines that the voice dialogue agent unit is appropriate for performing the voice-related processing, by the voice dialogue agent unit, performing processing based on the result of the voice recognition processing to generate a response signal, and by the voice output unit, outputting a voice based on the response signal generated by the voice dialogue agent unit;a step of, when the voice dialogue agent unit determines that the additional voice dialogue agent unit is appropriate for performing the voice-related processing, by the voice dialogue agent unit, transferring the voice signal to the additional voice dialogue agent unit, by the additional voice dialogue agent unit, performing new voice recognition processing on the transferred voice signal and performing processing based on a result of the new voice recognition processing to generate a response signal, and by the voice output unit, outputting a voice based on the response signal generated by the additional voice dialogue agent unit;and a step of, when the voice output unit outputs a voice, by the voice input control unit, keeping the voice signal generation unit in a receivable state for a predetermined period after output of the voice, the receivable state being a state in which a voice input is receivable when the determining determines that the first voice dialogue agent server is appropriate for performing the voice-related processing, (i) generating, from the generated first text input, a first instruction set for the first device or another device associated with the first voice dialogue agent server, (ii) executing the generated first instruction set using the first device or the other device associated with the first voice dialogue agent server, (iii) generating a first response signal based on the execution of the generated first instruction set using the first device or the other device associated with the first voice dialogue agent server, (iv) transmitting the generated first response signal from the first voice dialogue agent server to the first device, and (v) outputting a voice at the first device based on the received first response signal generated at the first voice dialogue agent server;when the determining determines that the second dialogue agent server is appropriate for performing the voice-related processing, (i) transferring the voice signal from the first voice dialogue agent server to the second voice dialogue agent server, (ii) performing new voice recognition processing on the transferred voice signal at the second voice dialogue agent server to generate second text input, (iii) generating, from the generated second text input, a second instruction set for the second device, (iv) executing the generated second instruction set using the second device, (v) generating a second response signal based on the execution of the generated second instruction set using the second device, (vi) transmitting the generated second response signal from the second voice dialogue agent server to the first device, and (vii) outputting a voice at the first device based on the received second response signal generated at the second voice dialogue agent server;and displaying, on a screen of the first device or a screen of the second device, a text character string obtained by recognizing voice input from the user and a text character string indicating a response signal by the first device or the second device, while indicating a distinction between the user, the first voice dialogue agent server, and the second voice dialogue agent server .
594 paragraphs in 8 sections, as filed
0001This application claims benefit to the provisional U.S. Application No. 61/836,763, filed on Jun. 19, 2013.This application is a reissue of U.S. Pat. No. 9,564,129, which issued on Feb. 7, 2017 from application Ser. No. 14/777,920, filed Sep. 17, 2015, which is the National Stage of International Application No. PCT/JP2014/003097, filed Jun. 10, 2014, which claims the benefit of U.S. Provisional Application No. 61/836,763, filed Jun. 19, 2013.
TECHNICAL FIELD
0002The present invention relates to a voice dialogue method for performing processing based on a voice that is dialogically input.
BACKGROUND ART
0003There has conventionally been known a voice dialogue system that includes voice input interface and performs processing based on a voice that is dialogically input by a user.
0004For example, Patent Literature 1 discloses a headset that includes a microphone, performs voice recognition processing on a voice input through the microphone, and performs processing based on a result of the voice recognition processing.
0005Also, Patent Literature 2 discloses a voice dialogue system that includes an agent that performs processing based on a voice that is dialogically input by a user.
CITATION LIST
Patent Literature
0006[Patent Literature 1] Japanese Patent Application Publication No. 2004-233794
0007[Patent Literature 2] Japanese Patent Application Publication No. 2008-90545
SUMMARY OF INVENTION
Technical Problem
0008According to the headset disclosed in Patent Literature 1, it is necessary to perform an operation of pressing a voice recognition control button that is provided in the headset at a start time and an end time of a voice input. Accordingly, in the case where this headset is used as input means in a voice dialogue system that performs processing based on a dialogically input voice, a user of the headset needs to start a voice input by pressing the voice recognition control button and end the voice input by pressing the voice recognition control button for each voice input.
0009This sometimes makes the user to feel troublesome to perform the operation of pressing the voice recognition control button, which needs to be performed at a start time and an end time of each voice input.
0010The present invention was made in view of the problem, and aims to provide a voice dialogue method for reducing, in a voice dialogue system, the number of times that a user needs to perform an operation in accordance with a voice that is dialogically input, compared with a conventional technique.
Solution to Problem
0011In order to solve the above aim, one aspect of the present invention provides a voice dialogue method that is performed by a voice dialogue system, the voice dialogue system including: a voice signal generation unit; a voice dialogue agent unit; a voice output unit; and a voice input control unit, the voice dialogue method comprising: a step of, by the voice signal generation unit, receiving a voice input and generating a voice signal based on the received voice input; a step of, by the voice dialogue agent unit, performing voice recognition processing on the generated voice signal and performing processing based on a result of the voice recognition processing to generate a response signal; a step of, by the voice output unit, outputting a voice based on the generated response signal; and a step of, when the voice output unit outputs the voice, by the voice input control unit, keeping the voice signal generation unit in a receivable state for a predetermined period after output of the voice, the receivable state being a state in which a voice input is receivable.
Advantageous Effects of Invention
0012According to the above voice dialogue method, in the case where a voice generated by the voice dialogue agent unit is output, a user can input a voice without performing an operation with respect to the voice dialogue system. This reduces the number of times that the user needs to perform an operation in accordance with a voice that is dialogically input, compared with conventional techniques.
BRIEF DESCRIPTION OF DRAWINGS
0013<figref idref="DRAWINGS">FIG. 1</figref> is a system configuration diagram showing configuration of a voice dialogue system <b>100</b>.
0014<figref idref="DRAWINGS">FIG. 2</figref> is a block diagram showing functional configuration of a device <b>140</b>.
0015<figref idref="DRAWINGS">FIG. 3</figref> shows switching of a state managed by a control unit <b>210</b>.
0016<figref idref="DRAWINGS">FIG. 4</figref> is a block diagram showing functional configuration of a voice dialogue agent <b>400</b>.
0017<figref idref="DRAWINGS">FIG. 5</figref> is a data structure diagram showing a dialog DB <b>500</b>.
0018<figref idref="DRAWINGS">FIG. 6</figref> is a flow chart of first device processing.
0019<figref idref="DRAWINGS">FIG. 7</figref> is a flow chart of first voice input processing.
0020<figref idref="DRAWINGS">FIG. 8</figref> is a flow chart of first agent processing.
0021<figref idref="DRAWINGS">FIG. 9</figref> is a flow chart of first instruction execution processing.
0022<figref idref="DRAWINGS">FIG. 10</figref> is a procedure diagram in a specific example.
0023<figref idref="DRAWINGS">FIG. 11A</figref> to <figref idref="DRAWINGS">FIG. 11D</figref> are each a pattern diagram showing contents displayed by the device <b>140</b>.
0024<figref idref="DRAWINGS">FIG. 12</figref> is a pattern diagram showing contents displayed by the device <b>140</b>.
0025<figref idref="DRAWINGS">FIG. 13</figref> is a block diagram showing functional configuration of a device <b>1300</b>.
0026<figref idref="DRAWINGS">FIG. 14</figref> shows switching of the state managed by a control unit <b>1310</b>.
0027<figref idref="DRAWINGS">FIG. 15</figref> is a flow chart of second device processing.
0028<figref idref="DRAWINGS">FIG. 16</figref> is a procedure diagram schematically showing a situation in which a dialogue with a voice dialogue agent is performed.
0029<figref idref="DRAWINGS">FIG. 17</figref> is a block diagram showing functional configuration of a device <b>1700</b>.
0030<figref idref="DRAWINGS">FIG. 18</figref> shows switching of the state managed by a control unit <b>1710</b>.
0031<figref idref="DRAWINGS">FIG. 19</figref> is a flow chart of third device processing.
0032<figref idref="DRAWINGS">FIG. 20</figref> is a flow chart of second voice input processing.
0033<figref idref="DRAWINGS">FIG. 21</figref> is a procedure diagram schematically showing a situation in which a dialogue with a dialogue agent is performed.
0034<figref idref="DRAWINGS">FIG. 22</figref> is a block diagram showing functional configuration of a voice dialogue agent <b>2200</b>.
0035<figref idref="DRAWINGS">FIG. 23</figref> is a data structure diagram showing a target agent DB <b>2300</b>.
0036<figref idref="DRAWINGS">FIG. 24</figref> is a flow chart of second agent processing.
0037<figref idref="DRAWINGS">FIG. 25</figref> is a flow chart of second instruction execution processing.
0038<figref idref="DRAWINGS">FIG. 26</figref> is a flow chart of first connection response processing.
0039FIG. <b>26</b>27 is a flow chart of disconnection response processing.
0040<figref idref="DRAWINGS">FIG. 28</figref> is a flow chart of third agent processing.
0041<figref idref="DRAWINGS">FIG. 29</figref> is a procedure diagram schematically showing a situation in which a dialogue with a voice dialogue agent is performed.
0042<figref idref="DRAWINGS">FIG. 30</figref> is a block diagram showing functional configuration of a voice dialogue agent <b>3000</b>.
0043<figref idref="DRAWINGS">FIG. 31</figref> is a data structure diagram showing an available service DB <b>3100</b>.
0044<figref idref="DRAWINGS">FIG. 32</figref> is a flow chart of fourth agent processing.
0045<figref idref="DRAWINGS">FIG. 33</figref> is a flow chart of third instruction execution processing.
0046<figref idref="DRAWINGS">FIG. 34</figref> is a flow chart of second connection response processing.
0047<figref idref="DRAWINGS">FIG. 35</figref> is a procedure diagram schematically showing a situation in which a dialogue with a voice dialogue agent is performed.
0048<figref idref="DRAWINGS">FIG. 36A</figref> is a diagram schematically showing an operation situation of the voice dialogue system, <figref idref="DRAWINGS">FIG. 36B</figref> and <figref idref="DRAWINGS">FIG. 36C</figref> are each a diagram schematically showing a data center administration company <b>3610</b>.
0049<figref idref="DRAWINGS">FIG. 37</figref> is a diagram schematically showing service type 1.
0050<figref idref="DRAWINGS">FIG. 38</figref> is a diagram schematically showing service type 2.
0051<figref idref="DRAWINGS">FIG. 39</figref> is a diagram schematically showing service type 3.
0052<figref idref="DRAWINGS">FIG. 40</figref> is a diagram schematically showing service type 4.
0053<figref idref="DRAWINGS">FIG. 41</figref> is a system configuration diagram showing configuration of a voice dialogue system <b>4100</b>.
0054<figref idref="DRAWINGS">FIG. 42</figref> is a block diagram showing functional configuration of a mediation server <b>4150</b>.
0055<figref idref="DRAWINGS">FIG. 43</figref> is a block diagram showing functional configuration of a mediation server <b>4350</b>.
0056<figref idref="DRAWINGS">FIG. 44A</figref> to <figref idref="DRAWINGS">FIG. 44D</figref> each show an example of an image displayed by a display unit.
0057<figref idref="DRAWINGS">FIG. 45A</figref> and <figref idref="DRAWINGS">FIG. 45B</figref> each show an example of an image displayed by the display unit.
0058<figref idref="DRAWINGS">FIG. 46</figref> shows an example of switching of the state.
0059<figref idref="DRAWINGS">FIG. 47</figref> shows an example of switching of the state.
0060<figref idref="DRAWINGS">FIG. 48</figref> shows an example of switching of the state.
0061<figref idref="DRAWINGS">FIG. 49</figref> shows an example of switching of the state.
0062<figref idref="DRAWINGS">FIG. 50</figref> shows an example of switching of the state.
DESCRIPTION OF EMBODIMENTS
Embodiment 1
0063<Outline>
0064The following explains, as one aspect of the voice dialogue method relating to the present invention and one aspect of the device relating to the present invention, a voice dialogue system including devices that are disposed in a home, a car, and so on and a voice dialogue agent server that communicates the devices.
0065In the voice dialogue system, the voice dialogue agent server embodies a voice dialogue agent by executing a program stored therein. The voice dialogue agent makes a voice dialogue via a device (input and output via a voice) with a user of the voice dialogue system. The voice dialogue agent performs processing that reflects details of the dialogue, and performs a voice output of a result of the processing via the device of the user.
0066In the case where the user hopes to make a dialogue with the voice dialogue agent (hopes to perform a voice input with respect to the voice dialogue agent), the user performs a predetermined voice input start operation with respect to the device constituting the voice dialogue system. The device is switched to a state in which voice input is receivable for a predetermined period after the voice input start operation. While the device is in the state in which voice input is receivable, the user performs voice input with respect to the voice dialogue agent.
0067The following explains the details of the voice dialogue system with reference to the drawings.
0068<Configuration>
0069<figref idref="DRAWINGS">FIG. 1</figref> is a system configuration diagram showing configuration of a voice dialogue system <b>100</b>.
0070As shown in the figure, the voice dialogue system <b>100</b> includes voice dialogue agent servers <b>110</b>a and <b>110</b>b, a network <b>120</b>, gateways <b>130</b>a and <b>130</b>b, and devices <b>140</b>a-<b>140</b>e.
0071While the gateway <b>130</b>a and the devices <b>140</b>a-<b>140</b>c are disposed in a home <b>180</b>, the gateway <b>130</b>b and the devices <b>140</b>d and <b>140</b>e are disposed in a car <b>190</b>.
0072The gateways <b>130</b>a and <b>130</b>b are hereinafter just referred to as a gateway <b>130</b> except in the case of explicit distinction. Also, the voice dialogue agent servers <b>110</b>a and <b>110</b>b are hereinafter just referred to as a voice dialogue agent server <b>110</b> except in the case of explicit distinction. The devices <b>140</b>a-<b>140</b>e each have a function of performing a wireless or wired communication with the gateway <b>130</b> and a function of performing a wireless or wired communication with the voice dialogue agent server <b>110</b>.
0073The devices <b>140</b>a-<b>140</b>c, which are disposed in the home <b>180</b>, are each for example a television, an air conditioner, a recorder, a washing machine, a portable smartphone, or the like that is disposed in the home <b>180</b>. The devices <b>140</b>d-<b>140</b>e, which are disposed in the car <b>190</b>, are each for example a car air conditioner, a car navigation system, or the like that is disposed in the car <b>190</b>.
0074Here, explanation is provided on a virtual device <b>140</b> that has functions the devices <b>140</b>a-<b>140</b>e commonly have, instead of separate explanations of the devices <b>140</b>a-<b>140</b>e.
0075<figref idref="DRAWINGS">FIG. 2</figref> is a block diagram showing functional configuration of the device <b>140</b>.
0076As shown in the figure, the device <b>140</b> includes a control unit <b>210</b>, a voice input unit <b>220</b>, an operation reception unit <b>230</b>, an address storage unit <b>240</b>, a communication unit <b>250</b>, a voice output unit <b>260</b>, a display unit <b>270</b>, and an execution unit <b>280</b>.
0077The voice input unit <b>220</b> is for example embodied by a microphone and a processor that executes programs. The voice input unit <b>220</b> is connected to the control unit <b>210</b>, and is controlled by the control unit <b>210</b>. The voice input unit <b>220</b> has a function of receiving a voice input from a user and generating a voice signal (hereinafter, referred to also as input voice data).
0078The voice input unit <b>220</b> is in either a voice input receivable state or a voice input unreceivable state under the control by the control unit <b>210</b>. In the voice input receivable state, the voice input unit <b>220</b> is able to receive a voice input. In voice input unreceivable state, the voice input unit <b>220</b> is unable to receive a voice input.
0079The operation reception unit <b>230</b> is for example embodied by a touchpanel, a touchpanel controller, and a processor that executes programs. The operation reception unit <b>230</b> is connected to the control unit <b>210</b>, and is controlled by the control unit <b>210</b>. The operation reception unit <b>230</b> has a function of receiving a predetermined contact operation performed by the user and generating an electrical signal based on the received contact operation.
0080The predetermined contact operation performed by the user, which is received by the operation reception unit <b>230</b>, includes a predetermined voice input start operation indicating that a voice input using the voice input unit <b>220</b> is to be started.
0081The voice input start operation is for example assumed to be an operation of contacting an icon for receiving the voice input start operation that is displayed on the touchpanel which is part of the operation reception unit <b>230</b>. Also, the voice input start operation is for example assumed to be an operation of pressing a button for receiving the voice input start operation that is included in the operation reception unit <b>230</b>.
0082The address storage unit <b>240</b> is for example embodied by a memory and a processor that executes programs, and is connected to the communication unit <b>250</b>. The address storage unit <b>240</b> has a function of storing therein an IP (Internet Protocol) address of one of the voice dialogue agent servers <b>110</b> in the network <b>120</b>. Hereinafter, the one voice dialogue agent server <b>110</b> is referred to as a specific voice dialogue agent server.
0083The device <b>140</b> is associated with the specific voice dialogue agent server, which is one of the voice dialogue agent servers <b>110</b>.
0084Note that the memory included in the device <b>140</b> is for example a RAM (Random Access Memory), a ROM (Read Only Memory), a flash memory, or the like.
0085The communication unit <b>250</b> is for example embodied by a processor that executes programs, a communication LSI (Large Scale Integration), and an antenna. The communication unit <b>250</b> is connected to the control unit <b>210</b> and the address storage unit <b>240</b>, and is controlled by the control unit <b>210</b>. The communication unit <b>250</b> has a gateway communication function and a voice dialogue agent server communication function described below.
0086The gateway communication function is a function of performing a wireless or wired communication with the gateway <b>130</b>.
0087The voice dialogue agent server communication function is a function of communicating with the voice dialogue agent server <b>110</b> via the gateway <b>130</b> and the network <b>120</b>.
0088Here, in communication with any one of the voice dialogue agent servers <b>110</b>, in the case where the control unit <b>210</b> does not designate a specific one of the voice dialogue agent servers <b>110</b> as a voice dialogue agent server <b>110</b> that is a communication party, the communication unit <b>250</b> communicates with a specific voice dialogue agent server with reference to an IP address stored in the address storage unit <b>240</b>.
0089The voice output unit <b>260</b> is for example embodied by a processor that executes programs and a speaker. The voice output unit <b>220</b> is connected to the control unit <b>210</b>, and is controlled by the control unit <b>210</b>. The voice output unit <b>260</b> has a function of converting an electrical signal, which is transmitted from the control unit <b>210</b>, to a voice and outputting the voice.
0090The display unit <b>270</b> is for example embodied by a touchpanel, a touchpanel controller, and a processor that executes programs. The display unit <b>270</b> is connected to the control unit <b>210</b>, and is controlled by the control unit <b>210</b>. The display unit <b>270</b> has a function of displaying images, character strings, and the like based on the electrical signal, which is transmitted from the control unit <b>210</b>.
0091The execution unit <b>280</b> is a functional block that achieves a function the device <b>140</b> as a device originally has. In the case where the device <b>140</b> is for example a television, the function is a function of receiving and decoding a television signal, displaying television images resulting from the decoding on a display, and outputting television audio resulting from the decoding via a speaker. In the case where the device <b>140</b> is for example an air conditioner, the function is a function of blowing cool air or warm air through a duct to bring a temperature in a room in which the air conditioner is disposed to a set temperature. The execution unit <b>280</b> is connected to the control unit <b>210</b>, and is controlled by the control unit <b>210</b>.
0092In the case where the device <b>140</b> is for example a television, the execution unit <b>280</b> is embodied by a television signal receiver, a television signal tuner, a television signal decoder, a display, a speaker, and so on.
0093Also, the execution unit <b>280</b> does not necessarily need to have a configuration in which all compositional elements thereof are included in a single housing. In the case where the device <b>140</b> is for example a television, the execution unit <b>280</b> is assumed to have for example a configuration in which a remote controller and the display are included in separate housings. Similarly, functional blocks of the device <b>140</b> each do not need to have a configuration in which all compositional elements thereof are included in a single housing.
0094The control unit <b>210</b> is for example embodied by a processor that executes programs. The control unit <b>210</b> is connected to the voice input unit <b>220</b>, the operation reception unit <b>230</b>, the communication unit <b>250</b>, the voice output unit <b>260</b>, the display unit <b>270</b>, and the execution unit <b>280</b>. The control unit <b>210</b> has a function of controlling the voice input unit <b>220</b>, a function of controlling the operation reception unit <b>230</b>, a function of controlling the communication unit <b>250</b>, a function of controlling the voice output unit <b>260</b>, a function of controlling the display unit <b>270</b>, and a function of controlling the execution unit <b>280</b>. The control unit <b>210</b> further has a voice input unit state management function and a first device processing execution function described below.
0095The voice input unit state management function is a function of managing the state of the voice input unit <b>220</b>, which is either the voice input receivable state or the voice input unreceivable state.
0096<figref idref="DRAWINGS">FIG. 3</figref> shows switching of the state managed by the control unit <b>210</b>.
0097As shown in the figure, in the case where the state is the voice input unreceivable state, (1) the control unit <b>210</b> keeps the state to the voice input unreceivable state until the operation reception unit <b>230</b> receives a voice input start operation. (2) After the reception of the voice input start operation by the operation reception unit <b>230</b>, the control unit <b>210</b> switches the state to the voice input receivable state. Then, in the case where the state is the voice input receivable state, (3) the control unit <b>210</b> keeps the state to the voice input receivable state until a predetermined period T<b>1</b> (for example, five seconds) has lapsed after the switching of the state to the voice input receivable state. (4) After the lapse of the predetermined period T<b>1</b>, the control unit <b>210</b> switches the state to the voice input unreceivable state.
0098Note that upon bootup of the device <b>140</b>, the control unit <b>210</b> starts managing the state as the voice input unreceivable state.
0099Returning to <figref idref="DRAWINGS">FIG. 2</figref>, the explanation on the control unit <b>210</b> is continued.
0100The first device processing execution function is a function performed by the control unit <b>210</b> controlling the voice input unit <b>220</b>, the operation reception unit <b>230</b>, the communication unit <b>250</b>, the voice output unit <b>260</b>, the display unit <b>270</b>, and the execution unit <b>280</b> to cause the device <b>140</b> to execute the first device processing as its characteristic operation to execute a sequence of processing described below. In the sequence of processing, (1) when the user performs a voice input start operation, (2) the device <b>140</b> receives a voice input from the user, and generates input voice data, (3) transmits the generated input voice data to a voice dialogue agent, (4) receives response voice data returned from the voice dialogue agent, and (5) outputs a voice based on the received response voice data.
0101Note that the first device processing is explained in detail in section <First Device Processing> later with reference to a flow chart.
0102Referring back to <figref idref="DRAWINGS">FIG. 1</figref>, the explanation on the device <b>140</b> is continued.
0103The gateway <b>130</b> is for example embodied by a personal computer or the like having a communication function, and is connected to the network <b>120</b>. The gateway <b>130</b> has the following functions achieved by executing programs stored therein: a function of performing a wireless or wired communication with the device <b>140</b>; a function of communicating with the voice dialogue agent server <b>110</b> via the network <b>120</b>; and a function of relaying communication between the device <b>140</b> and the voice dialogue agent server <b>110</b>.
0104The voice dialogue agent server <b>110</b> is for example embodied by a server, which is composed of one or more computer systems and has a communication function. The voice dialogue agent server <b>110</b> is connected to the network <b>120</b>. The voice dialogue agent server <b>110</b> has the following functions achieved by executing programs stored therein: a function of communicating with another device which is connected to the network <b>120</b>; a function of communicating with the device <b>140</b> via the gateway <b>130</b>; and a function of embodying the voice dialogue agent <b>400</b>.
0105<figref idref="DRAWINGS">FIG. 4</figref> is a block diagram showing functional configuration of the voice dialogue agent <b>400</b> embodied by the voice dialogue agent server <b>110</b>.
0106As shown in the figure, the voice dialogue agent <b>400</b> includes a control unit <b>410</b>, a communication unit <b>420</b>, a voice recognition processing unit <b>430</b>, a dialogue DB (Date Base) storage unit <b>440</b>, a voice synthesizing processing unit <b>450</b>, and an instruction generation unit <b>460</b>.
0107The communication unit <b>420</b> is for example embodied by a processor that executes programs and a communication LSI. The communication unit <b>420</b> is connected to the control unit <b>410</b>, the voice recognition processing unit <b>430</b>, and the voice synthesizing processing unit <b>450</b>, and is controlled by the control unit <b>410</b>. The communication unit <b>420</b> has a function of communicating with another device which is connected to the network <b>120</b> and a function of communicating with the device <b>140</b> via the gateway <b>130</b>.
0108The voice recognition processing unit <b>430</b> is embodied by a processor that executes programs. The voice recognition processing unit <b>430</b> is connected to the control unit <b>410</b> and the communication unit <b>420</b>, and is controlled by the control unit <b>410</b>. The voice recognition processing unit <b>430</b> has a function of performing voice recognition processing on input voice data received by the communication unit <b>420</b> to convert the voice data to a character string (hereinafter, referred to also as an input text).
0109The voice synthesizing processing unit <b>450</b> is for example embodied by a processor that executes programs. The voice synthesizing processing unit <b>450</b> is connected to the control unit <b>410</b> and the communication unit <b>420</b>, and is controlled by the control unit <b>410</b>. The voice synthesizing processing unit <b>450</b> has a function of performing voice synthesizing processing on a character sting transmitted from the control unit <b>410</b> to convert the character string to voice data.
0110The dialogue DB storage unit <b>440</b> is for example embodied by a memory and a processor that executes programs. The dialogue DB storage unit <b>440</b> is connected to the control unit <b>410</b>, and has a function of storing therein a dialog DB <b>500</b>.
0111<figref idref="DRAWINGS">FIG. 5</figref> is a data structure diagram showing the dialog DB <b>500</b> stored in the dialogue DB storage unit <b>440</b>.
0112As shown in the figure, the dialog DB <b>500</b> includes keyword <b>510</b>, target device <b>520</b>, startup application <b>530</b>, processing details <b>540</b>, and response text <b>550</b> that are associated with each other.
0113The keyword <b>510</b> indicates a character string that is assumed to be included in an input text converted by the voice recognition processing unit <b>430</b>.
0114The target device <b>520</b> indicates information for specifying a device that is to execute processing specified by the associated processing details <b>540</b>, which are described later.
0115Here, the device specified by the target device <b>520</b> may be the voice dialogue agent <b>400</b>.
0116The startup application <b>530</b> is information for specifying an application program to be started up in a device specified by the associated target device <b>520</b> in order to cause the specified device to execute processing specified by the associated processing details <b>540</b>, which are described later.
0117The processing details <b>540</b> are information for specifying, in the case where a character string indicated by the associated keyword <b>510</b> is included in an input text that is converted by the voice recognition processing unit <b>430</b>, processing that is determined to be executed by a device that is specified by the associated target device <b>520</b>.
0118The response text <b>550</b> is information for indicating, in the case where processing specified by the associated processing details <b>540</b> is executed, a character string that is determined to be generated based on a result of the processing (hereinafter, referred to also as a response text).
0119Referring back to <figref idref="DRAWINGS">FIG. 4</figref>, the explanation on the voice dialogue agent <b>400</b> is continued.
0120The instruction generation unit <b>460</b> is for example embodied by a processor that executes programs. The instruction generation unit <b>460</b> is connected to the control unit <b>410</b>, and is controlled by the control unit <b>410</b>. The instruction generation unit <b>460</b> has a function of, upon reception of a group of the target device <b>520</b>, the startup application <b>530</b>, and the processing details <b>540</b> transmitted from the control unit <b>410</b>, starting up an application program that is specified by the startup application <b>530</b> included in a device that is specified by the target device <b>520</b>, and generating an instruction set for causing the specified device to execute processing that is specified by the processing details <b>540</b>.
0121The control unit <b>410</b> is for example embodied by a processor that executes programs. The control unit <b>410</b> is connected to the communication unit <b>420</b>, the voice recognition processing unit <b>430</b>, the dialogue DB storage unit <b>440</b>, the voice synthesizing processing unit <b>450</b>, and the instruction generation unit <b>460</b>. The control unit <b>410</b> has a function of controlling the communication unit <b>420</b>, a function of controlling the voice recognition processing unit <b>430</b>, a function of controlling the voice synthesizing processing unit <b>450</b>, and a function of controlling the instruction generation unit <b>460</b>. The control unit <b>410</b> further has an input text return function, an instruction generation function, an instruction execution function, and a first agent processing execution function described below.
0122The input text return function is a function of controlling, in the case where input voice data received by the communication unit <b>420</b> is converted to an input text by the voice recognition processing unit <b>430</b>, the communication unit <b>420</b> to return the input text to the device <b>140</b> which has transmitted the input voice data.
0123The instruction generation function is a function of, upon reception of the input text transmitted from the voice recognition processing unit <b>430</b>, controlling the instruction generation unit <b>460</b> to generate an instruction set: by (1) referring to the dialog DB <b>500</b> stored in the dialogue DB storage unit <b>440</b> to read, based on the keyword <b>510</b> included in the input text, the target device <b>520</b>, the startup application <b>530</b>, the processing details <b>540</b>, and the response text <b>550</b>, which are associated with the keyword <b>510</b>; and (2) transmitting a group of the read target device <b>520</b>, startup application <b>530</b>, and processing details <b>540</b> to the instruction generation unit <b>460</b>.
0124The instruction execution function is a function of executing an instruction set generated by the instruction generation unit <b>460</b>, generating a response text specified by the response text <b>550</b> based on an execution result of the instruction set, and transmitting the generated response text to the voice synthesizing processing unit <b>450</b>.
0125In execution of the instruction execution function, the control unit <b>410</b> generates a response text by communicating with a device specified by the target device <b>520</b> with use of the communication unit <b>420</b> to cause the specified device to execute the instruction set and transmit an execution result of the instruction set.
0126The first agent processing execution function is a function performed by the control unit <b>410</b> controlling the communication unit <b>420</b>, the voice recognition processing unit <b>430</b>, the voice synthesizing processing unit <b>450</b>, and the instruction generation unit <b>460</b> to cause the voice dialogue agent <b>400</b> to execute first agent processing that is its characteristic operation to execute a sequence of processing described below. In the sequence of processing, (1) the voice dialogue agent <b>400</b> receives input voice data transmitted from a device, (2) performs voice recognition processing on the received input voice data to generate an input text, and returns the generated input text to the device, (3) generates an instruction set based on the generated input text, and executes the generated instruction set (4) generates a response text based on an execution result of the instruction set, (5) converts the generated response text to response voice data, and (6) returns the response text and the response voice data to the device.
0127Note that the first agent processing is explained in detail in section <First Agent Processing> later with reference to a flow chart.
0128Here, assume a case for example where an input text “Where is Mr. A's address?” is transmitted from the voice recognition processing unit <b>430</b>. In this case, with reference to the dialog DB <b>500</b> stored in the dialogue DB storage unit <b>440</b>, the control unit <b>410</b> causes a device “smartphone” specified by the target device <b>520</b> to start up an application program “Contact information” specified by the startup application <b>530</b> and execute processing of “Check Mr. A's address” specified by the processing details <b>540</b>, and generates a response text “Mr. A's address is XXXX.” based on an execution result of the processing.
0129The following explains the operation of the voice dialogue system <b>100</b> having the above configuration, with reference to the drawings.
0130<Operation>
0131The voice dialogue system <b>100</b> performs, as its characteristic operation, the first device processing and the first agent processing.
0132Explanation is given below on the processing in order.
0133<First Device Processing>
0134The first device processing is processing performed by the device <b>140</b>. In the first device processing, (1) when the user performs a voice input start operation, (2) the device <b>140</b> receives a voice input from the user, and generates input voice data, (3) transmits the generated input voice data to a voice dialogue agent, (4) receives response voice data returned from the voice dialogue agent, and (5) outputs a voice based on the received response voice data.
0135<figref idref="DRAWINGS">FIG. 6</figref> is a flow chart of the first device processing.
0136Upon bootup of the device <b>140</b>, the first device processing is started.
0137At a time of bootup of the device <b>140</b>, the state managed by the control unit <b>210</b> is the voice input unreceivable state.
0138When the first device processing is started, the control unit <b>210</b> stands by until the operation reception unit <b>230</b> receives a voice input start operation performed by a user of the voice dialogue system <b>100</b> (Step S<b>600</b>: Repetition of No). When the operation reception unit <b>230</b> receives the voice input start operation (Step S<b>600</b>: Yes), the control unit <b>210</b> switches the state from the voice input unreceivable state to the voice input receivable state (Step S<b>610</b>), and causes the display unit <b>270</b> to display that the state is the voice input receivable state (Step S<b>620</b>).
0139<figref idref="DRAWINGS">FIG. 11A</figref> is a pattern diagram showing an example of a situation in which in the case where the device <b>140</b> is for example a smartphone, the display unit <b>270</b> displays that the state is the voice input receivable state.
0140In the figure, a touchpanel <b>1110</b> that constitutes the smartphone is part of the display unit <b>270</b>. The touchpanel <b>1110</b> displays that the state is the voice input receivable state by blinking a region <b>1120</b> that is positioned at the lower right in the touchpanel <b>1110</b> (for example by alternately lighting black color and white color in the region <b>1120</b>).
0141Referring back to <figref idref="DRAWINGS">FIG. 6</figref>, the explanation on the first device processing is continued.
0142After the end of the processing in Step S<b>620</b>, the device <b>140</b> executes first voice input processing (Step S<b>630</b>).
0143<figref idref="DRAWINGS">FIG. 7</figref> is a flow chart of the first voice input processing.
0144When the first voice input processing is started, the voice input unit <b>220</b> receives a voice input from a user, and generates input voice data (Step S<b>700</b>). Then, when a predetermined period T<b>1</b> has lapsed after switching of the state to the voice input receivable state (Step S<b>710</b>: Yes after repetition of No), the control unit <b>210</b> switches the state from the voice input receivable state to the voice input unreceivable state (Step S<b>720</b>), and causes the display unit <b>270</b> to stop displaying that the state is the voice input receivable state (Step S<b>730</b>).
0145Then, the control unit <b>210</b> controls the communication unit <b>250</b> to transmit the input voice data, which is generated by the voice input unit <b>220</b>, to the voice dialogue agent <b>400</b> which is embodied by a specific voice dialogue agent server (Step S<b>740</b>).
0146After the end of the processing in Step S<b>740</b>, the device <b>140</b> ends the first voice input processing.
0147Referring back to <figref idref="DRAWINGS">FIG. 6</figref> again, the explanation on the first device processing is continued.
0148After the end of the first voice input processing, the control unit <b>210</b> stands by until the communication unit <b>250</b> receives an input text that is returned from the voice dialogue agent <b>400</b> in response to the input voice data transmitted in the processing in Step S<b>740</b> (Step S<b>640</b>: Repetition of No).
0149Here, the input text is a character string resulting from conversion of the input voice data transmitted in the processing in Step S<b>740</b> performed by the voice dialogue agent <b>400</b>.
0150When the communication unit <b>250</b> receives the input text (Step S<b>640</b>: Yes), the display unit <b>270</b> displays the input text (Step S<b>650</b>).
0151<figref idref="DRAWINGS">FIG. 11B</figref> is a pattern diagram showing an example of a situation in which in the case where the device <b>140</b> is for example a smartphone, the display unit <b>270</b> displays an input text.
0152In the figure, an example is shown in which the input text is a character string “What is room temperature?”. As shown in the figure, the input text, which is the character string “What is room temperature?” is displayed on the touchpanel <b>1110</b>, which is part of the display unit <b>270</b>, together with a character string “You”.
0153Referring back to <figref idref="DRAWINGS">FIG. 6</figref> again, the explanation on the first device processing is continued.
0154After the end of the processing in Step S<b>650</b>, the control unit <b>210</b> stands by until the communication unit <b>250</b> receives a response text and response voice data that are returned from the voice dialogue agent <b>400</b> in response to the input voice data transmitted in the processing in Step S<b>740</b> (Step S<b>640</b>: Repetition of No).
0155When the communication unit <b>250</b> receives the response text and the response voice data (Step S<b>660</b>: Yes), the display unit <b>270</b> displays the response text (Step S<b>670</b>), and the voice output unit <b>260</b> converts the response voice data to a voice and outputs the voice (Step S<b>680</b>).
0156<figref idref="DRAWINGS">FIG. 11C</figref> is a pattern diagram showing an example of a situation in which in the case where the device <b>140</b> is for example a smartphone, the display unit <b>270</b> displays a response text.
0157In the figure, an example is shown in which the response text is a character string “Which room?”. As shown in the figure, the response text, which is the character string “Which room?”, is displayed on the touchpanel <b>1110</b> which is part of the display unit <b>270</b>, together with a character string “Home agent”.
0158Referring back to <figref idref="DRAWINGS">FIG. 6</figref> again, the explanation on the first device processing is continued.
0159After the end of the processing in Step S<b>680</b>, the device <b>140</b> ends the first device processing.
0160<First Agent Processing>
0161The first agent processing is processing performed by the voice dialogue agent <b>400</b>. In the first agent processing, (1) the voice dialogue agent <b>400</b> receives input voice data transmitted from a device, (2) performs voice recognition processing on the received input voice data to generate an input text, and returns the generated input text to the device, (3) generates an instruction set based on the generated input text, and executes the generated instruction set, (4) generates a response text based on an execution result of the instruction set, (5) converts the generated response text to response voice data, and (6) returns the response text and the response voice data to the device.
0162<figref idref="DRAWINGS">FIG. 8</figref> is a flow chart of the first agent processing.
0163Upon bootup of the voice dialogue agent <b>400</b>, the first agent processing is started.
0164When the first agent processing is started, the voice dialogue agent <b>400</b> stands by until the communication unit <b>420</b> receives input voice data transmitted from the device <b>140</b> (Step S<b>800</b>: Repetition of No). When the communication unit <b>420</b> receives the input voice data (Step S<b>800</b>: Yes), the voice dialogue agent <b>400</b> performs first instruction execution processing (Step S<b>810</b>).
0165<figref idref="DRAWINGS">FIG. 9</figref> is a flow chart of the first instruction execution processing.
0166When the first instruction execution processing is started, the voice recognition processing unit <b>430</b> performs voice recognition processing on the input voice data, which is received by the communication unit <b>420</b>, to convert the input voice data to an input text that is a character string (Step S<b>900</b>).
0167After the conversion to the input text, the control unit <b>410</b> controls the communication unit <b>420</b> to return the converted input text to the device <b>140</b> which has transmitted the input voice data (Step S<b>910</b>).
0168The control unit <b>410</b> controls the instruction generation unit <b>460</b> to generate an instruction set by: (1) referring to the dialog DB <b>500</b> stored in the dialogue DB storage unit <b>440</b> to read, based on the keyword <b>510</b> included in the input text, the target device <b>520</b>, the startup application <b>530</b>, the processing details <b>540</b>, and the response text <b>550</b>, which are associated with the keyword <b>510</b>; and (2) transmitting a group of the read target device <b>520</b>, startup application <b>530</b>, and processing details <b>540</b> to the instruction generation unit <b>460</b>.
0169After the generation of the instruction set, the control unit <b>410</b> executes the generated instruction set (Step S<b>930</b>), and generates a response text specified by the response text <b>550</b> based on an execution result of the instruction set (Step S<b>940</b>). Here, the control unit <b>410</b> generates a response text by communicating with a device specified by the target device <b>520</b> with use of the communication unit <b>420</b> to cause the specified device to execute part of the instruction set and transmit an execution result of the part of the instruction set.
0170After the generation of the response text, the voice synthesizing processing unit <b>450</b> performs voice synthesizing processing on the generated response text to generate response voice data (Step S<b>950</b>).
0171After the generation of the response voice data, the control unit <b>410</b> controls the communication unit <b>420</b> to transmit the generated response text and response voice data to the device <b>140</b> which has transmitted the input voice data (Step S<b>960</b>).
0172After the end of the processing in Step S<b>960</b>, the voice dialogue agent <b>400</b> ends the first instruction execution processing.
0173Referring back to <figref idref="DRAWINGS">FIG. 8</figref>, the explanation on the first agent processing is continued.
0174After the end of the first instruction execution processing, the voice dialogue agent <b>400</b> returns to the processing in Step S<b>800</b> to perform the processing in Step S<b>800</b> and the subsequent steps.
0175The following explains a specific example of the operation performed by the voice dialogue system <b>100</b> having the above configuration, with reference to the drawing.
0176<Specific Example>
0177<figref idref="DRAWINGS">FIG. 10</figref> is a procedure diagram schematically showing a situation in which the user of the voice dialogue system <b>100</b> makes a voice dialogue with the voice dialogue agent <b>400</b> with use of the device <b>140</b> (here, a smartphone), and the voice dialogue agent <b>400</b> performs processing that reflects details of the dialogue.
0178When the user performs a voice input start operation (Step S<b>1000</b>, corresponding to Step S<b>600</b>: Yes in <figref idref="DRAWINGS">FIG. 6</figref>), the state is switched to the voice input receivable state (Step S<b>1005</b>, corresponding to Step S<b>610</b> in <figref idref="DRAWINGS">FIG. 6</figref>), and the device <b>140</b> performs first voice input processing (Step S<b>1010</b>, corresponding to Step S<b>630</b><figref idref="DRAWINGS">FIG. 6</figref>).
0179<figref idref="DRAWINGS">FIG. 11A</figref> is a diagram schematically showing an example of a situation in which, while the state is the voice input receivable state in the first voice input processing, the touchpanel <b>1110</b>, which is part of the display unit <b>270</b> included in the device <b>140</b> which is a smartphone, displays that the state is the voice input receivable state by blinking the region <b>1120</b>.
0180Referring back to <figref idref="DRAWINGS">FIG. 10</figref>, the explanation on the specific example is continued.
0181In the first voice input processing, in the case where the user inputs a voice “What is room temperature?”, the device <b>140</b> transmits input voice data “What is room temperature?” to the voice dialogue agent <b>400</b> (corresponding to Step S<b>740</b> in <figref idref="DRAWINGS">FIG. 7</figref>).
0182Then, the voice dialogue agent <b>400</b> receives the input voice data (corresponding to Step S<b>800</b>: Yes in <figref idref="DRAWINGS">FIG. 8</figref>), and performs first instruction execution processing (Step S<b>1060</b>, corresponding to Step S<b>810</b> in <figref idref="DRAWINGS">FIG. 8</figref>).
0183Here, in the first instruction execution processing, in the case where the voice dialogue agent <b>400</b> generates response voice data “Which room?”, the voice dialogue agent <b>400</b> transmits the response voice data “Which room?” to the device <b>140</b> (corresponding to Step S<b>960</b> in <figref idref="DRAWINGS">FIG. 9</figref>).
0184Then, the device <b>140</b> receives the response voice data (corresponding to Step S<b>660</b>: Yes in <figref idref="DRAWINGS">FIG. 6</figref>), and outputs a voice “Which room?” (Step S<b>1015</b>, corresponding to Step S<b>680</b> in <figref idref="DRAWINGS">FIG. 6</figref>).
0185In the processing in Step S<b>1010</b>, when the predetermined period T<b>1</b> has lapsed after the switching of the state to the voice input receivable state, the state is switched again to the voice input unreceivable state (corresponding to Step S<b>720</b> in <figref idref="DRAWINGS">FIG. 7</figref>). Accordingly, the user, who has heard the voice “Which room?” which is output from the device <b>140</b>, performs a new voice input start operation with respect to the device <b>140</b> to newly input a voice (Step S<b>1020</b>, corresponding to Step S<b>600</b>: Yes in <figref idref="DRAWINGS">FIG. 6</figref>). Then, the state is switched to the voice input receivable state (Step S<b>1025</b>, corresponding to Step S<b>610</b> in <figref idref="DRAWINGS">FIG. 6</figref>), and the device <b>140</b> performs first voice input processing (Step S<b>1030</b>, corresponding to Step S<b>630</b> in <figref idref="DRAWINGS">FIG. 6</figref>).
0186<figref idref="DRAWINGS">FIG. 11C</figref> is a diagram schematically showing an example of a situation in which, while the state is the voice input receivable state in the first voice input processing, the touchpanel <b>1110</b>, which is part of the display unit <b>270</b> included in the device <b>140</b> which is a smartphone, displays that the state is the voice input receivable state by blinking the region <b>1120</b>.
0187Referring back to <figref idref="DRAWINGS">FIG. 10</figref> again, the explanation on the specific example is continued.
0188In the first voice input processing, in the case where the user inputs a voice “Living room.”, the device <b>140</b> transmits input voice data “Living room.” to the voice dialogue agent <b>400</b> (corresponding to Step S<b>740</b> in <figref idref="DRAWINGS">FIG. 7</figref>).
0189Then, the voice dialogue agent <b>400</b> receives the input voice data (corresponding to Step S<b>800</b>: Yes in <figref idref="DRAWINGS">FIG. 8</figref>), and performs first instruction execution processing (Step S<b>1065</b>, corresponding to Step S<b>810</b> in <figref idref="DRAWINGS">FIG. 8</figref>).
0190Here, in the first instruction execution processing, in the case where the voice dialogue agent <b>400</b> generates response voice data “Living room temperature is 28 degrees C. Do you need any other help?”, the voice dialogue agent <b>400</b> transmits the response voice data “Living room temperature is 28 degrees C. Do you need any other help?” to the device <b>140</b> (corresponding to Step S<b>960</b>: Yes in <figref idref="DRAWINGS">FIG. 9</figref>).
0191Then, the device <b>140</b> receives the response voice data (corresponding to Step S<b>660</b>: Yes in <figref idref="DRAWINGS">FIG. 6</figref>), and outputs a voice “Living room temperature is 28 degrees C. Do you need any other help?” (Step S<b>1035</b>, corresponding to Step S<b>680</b> in <figref idref="DRAWINGS">FIG. 6</figref>).
0192In the processing in Step S<b>1010</b>, when the predetermined period T<b>1</b> has lapsed after the switching of the state to the voice input receivable state, the state is switched again to the voice input unreceivable state (corresponding to Step S<b>720</b> in <figref idref="DRAWINGS">FIG. 7</figref>). Accordingly, the user, who has heard the voice “Living room temperature is 28 degrees C. Do you need any other help?” which is output from the device <b>140</b>, performs a new voice input start operation with respect to the device <b>140</b> to newly input a voice (Step S<b>1040</b>, corresponding to Step S<b>600</b>: Yes in <figref idref="DRAWINGS">FIG. 6</figref>). Then, the state is switched to the voice input receivable state (Step S<b>1045</b>, corresponding to Step S<b>610</b> in <figref idref="DRAWINGS">FIG. 6</figref>), and the device <b>140</b> performs first voice input processing (Step S<b>1050</b>, corresponding to Step S<b>630</b> in <figref idref="DRAWINGS">FIG. 6</figref>).
0193<figref idref="DRAWINGS">FIG. 12</figref> is a diagram schematically showing an example where, in the first voice input processing, while the state is the voice input receivable state, the touchpanel <b>1110</b>, which is part of the display unit <b>270</b> included in the device <b>140</b> which is a smartphone, displays that the state is the voice input receivable state by blinking the region <b>1120</b>.
0194Referring back to <figref idref="DRAWINGS">FIG. 10</figref> again, the explanation on the specific example is continued.
0195In the first voice input processing, in the case where the user inputs a voice “No. Thank you.”, the device <b>140</b> transmits input voice data “No. Thank you.” to the voice dialogue agent <b>400</b> (corresponding to Step S<b>740</b> in <figref idref="DRAWINGS">FIG. 7</figref>).
0196Then, the voice dialogue agent <b>400</b> receives the input voice data (corresponding to Step S<b>800</b>: Yes in <figref idref="DRAWINGS">FIG. 8</figref>), and performs first instruction execution processing (Step S<b>1070</b>, corresponding to Step S<b>810</b> in <figref idref="DRAWINGS">FIG. 8</figref>).
0197Here, in the first instruction execution processing, in the case where the voice dialogue agent <b>400</b> generates response voice data “This ends dialogue.”, the voice dialogue agent <b>400</b> transmits the response voice data “This ends dialogue.” to the device <b>140</b> (corresponding to Step S<b>960</b>: Yes in <figref idref="DRAWINGS">FIG. 9</figref>).
0198Then, the device <b>140</b> receives the response voice data (corresponding to Step S<b>660</b>: Yes in <figref idref="DRAWINGS">FIG. 6</figref>), and outputs a voice “This ends dialogue.” (Step S<b>1055</b>, corresponding to Step S<b>680</b> in <figref idref="DRAWINGS">FIG. 6</figref>).
0199<Consideration>
0200According to the voice dialogue system <b>100</b> having the above configuration, the user switches the state of the device <b>140</b> by performing a voice input start operation with respect to the device <b>140</b>, and inputs a voice. Then, when the predetermined period T<b>1</b> has lapsed, the state of the device <b>140</b> is switched to the voice input unreceivable state even if the user does not perform any operation for switching the state of the device <b>140</b> to the voice input unreceivable state.
0201According to the voice dialogue system <b>100</b>, therefore, a reduced number of operations need to be performed by the user in accordance with a voice input, compared with a voice dialogue system in which each time a voice input ends, it is necessary to perform an operation for switching the state of the device <b>140</b> to the voice input unreceivable state.
Embodiment 2
0202<Outline>
0203The following explains, as one aspect of the voice dialogue method relating to the present invention and one aspect of the device relating to the present invention, a first modified voice dialogue system that is a partial modification of the voice dialogue system <b>100</b> in Embodiment 1.
0204The voice dialogue system <b>100</b> in Embodiment 1 has been explained as an example of the configuration in which when the user performs a voice input start operation, the device <b>140</b> is in the voice input receivable state for a period from performance of the voice input start operation to lapse of the predetermined period T<b>1</b>.
0205Compared with this, the first modified voice dialogue system in Embodiment 2 is an example of configuration in which in the case where a device outputs a voice based on response voice data, the device is in the voice input receivable state for a period from output of the voice to lapse of the predetermined period T<b>1</b>, in addition to the above period.
0206The following explains the details of the first modified voice dialogue system, focusing on different points from the voice dialogue system <b>100</b> in Embodiment 1, with reference to the drawings.
0207<Configuration>
0208The first modified voice dialogue system is modified from the voice dialogue system <b>100</b> in Embodiment 1 so as to include a device <b>1300</b> instead of the device <b>14</b>.
0209The device <b>1300</b> is not modified from the device <b>140</b> in Embodiment 1 in terms of hardware, but is partially modified from the device <b>140</b> in terms of software to be stored as an execution target. Accordingly, the device <b>1300</b> is modified from the device <b>140</b> in Embodiment 1 in terms of part of functions.
0210<figref idref="DRAWINGS">FIG. 13</figref> is a block diagram showing functional configuration of the device <b>1300</b>.
0211As shown in the figure, the device <b>1300</b> is modified from the device <b>140</b> in Embodiment 1 (see <figref idref="DRAWINGS">FIG. 2</figref>) so as to include a control unit <b>1310</b> instead of the control unit <b>210</b>.
0212The control unit <b>1310</b> is modified from the control unit <b>210</b> in Embodiment 1 so as to have a first modified voice input unit state management function and a second device processing execution function, which are described below, instead of the voice input unit state management function and the first device processing execution function of the control unit <b>210</b>.
0213Similarly to the voice input unit state management function in Embodiment 1, the first modified voice input unit state management function is a function of managing the state of the voice input unit <b>220</b>, which is either the voice input receivable state or the voice input unreceivable state, and conditions for switching the state are partially modified from those in the voice input unit state management function in Embodiment 1.
0214<figref idref="DRAWINGS">FIG. 14</figref> shows switching of the state managed by the control unit <b>1310</b>.
0215As shown in the figure, in the case where the state is the voice input unreceivable state, (1) the control unit <b>1310</b> keeps the state to the voice input unreceivable state until the operation reception unit <b>230</b> receives a voice input start operation or the voice output unit <b>260</b> outputs a voice included in voices based on response voice data except a predetermined voice. (2) After the reception of the voice input start operation by the operation reception unit <b>230</b> or output of the voice included in voices based on the response voice data by the voice output unit <b>260</b>, the control unit <b>1310</b> switches the state to the voice input receivable state. Then, in the case where the state is the voice input receivable state, (3) the control unit <b>1310</b> keeps the state to the voice input receivable state until a predetermined period T<b>1</b> (for example, five seconds) has lapsed after the switching of the state to the voice input receivable state. (4) After the lapse of the predetermined period T<b>1</b>, the control unit <b>1310</b> switches the state to the voice input unreceivable state.
0216Here, the predetermined voice included in the voices based on response voice data is a voice that indicates unnecessity of a new voice input, such as a voice “This ends dialogue.”. Hereinafter, this voice is referred to also as a dialogue end voice.
0217Note that upon bootup of the device <b>1300</b>, the control unit <b>1310</b> starts managing the state as the voice input unreceivable state.
0218Referring back to <figref idref="DRAWINGS">FIG. 13</figref>, the explanation on the control unit <b>1310</b> is continued.
0219The second device processing execution function is a function performed by the control unit <b>1310</b> controlling the voice input unit <b>220</b>, the operation reception unit <b>230</b>, the communication unit <b>250</b>, the voice output unit <b>260</b>, the display unit <b>270</b>, and the execution unit <b>280</b> to cause the device <b>1300</b> to execute the second device processing that is its characteristic operation to execute a sequence of processing described below. In the sequence of processing, (1) when the user performs a voice input start operation, (2) the device <b>1300</b> receives a voice input from the user, and generates input voice data, (3) transmits the generated input voice data to a voice dialogue agent, (4) receives response voice data returned from the voice dialogue agent, (5) outputs a voice based on the received response voice data, and (6) in the case where the output voice is not a dialogue end voice, the device <b>1300</b> repeats the processing (2) and the subsequent processing even if the user does not perform a voice input start operation.
0220Note that the second device processing is explained in detail in section <Second Device Processing> later with reference to a flow chart.
0221The following explains the operation of the first modified voice dialogue system having the above configuration, with reference to the drawings.
0222<Operation>
0223The first modified voice dialogue system performs second device processing as its characteristic operation, in addition to the first agent processing in Embodiment 1. The second device processing is partially modified from the first device processing in Embodiment 1.
0224Explanation is given on the second device processing below, focusing on different points from the first device processing.
0225<Second Device Processing>
0226The second device processing is processing performed by the device <b>1300</b>. In the second device processing, (1) when the user performs a voice input start operation, (2) the device <b>1300</b> receives a voice input from the user, and generates input voice data, (3) transmits the generated input voice data to a voice dialogue agent, (4) receives response voice data returned from the voice dialogue agent, (5) outputs a voice based on the received response voice data, and (6) in the case where the output voice is not a dialogue end voice, the device <b>1300</b> repeats the processing (2) and the subsequent processing even if the user does not perform a voice input start operation.
0227<figref idref="DRAWINGS">FIG. 15</figref> is a flow chart of the second device processing.
0228Upon bootup of the device <b>1300</b>, the second device processing is started.
0229At a time of bootup of the device <b>1300</b>, the state managed by the control unit <b>1310</b> is the voice input unreceivable state.
0230In the figure, processing in Steps S<b>1500</b>-S<b>1580</b> is the same as the processing in Steps S<b>600</b>-S<b>680</b> in the first device processing in Embodiment 1 (see <figref idref="DRAWINGS">FIG. 6</figref>), and is accordingly regarded as having been already explained.
0231After the end of the processing in Step S<b>1580</b>, the control unit <b>1310</b> checks whether or not the voice, which is output from the voice output unit <b>260</b> in the processing in Step S<b>1580</b>, is a dialogue end voice (Step S<b>1585</b>). This processing is executed by for example checking whether or not the response text, which is received in the processing in Step S<b>1560</b>: Yes, is a predetermined character string (for example, a character string “This ends dialogue.”).
0232In the processing in Step S<b>1585</b>, in the case where the response text is not a dialogue end voice (Step S<b>1585</b>: No), the control unit <b>1310</b> switches the state from the voice input unreceivable state to the voice input receivable state (Step S<b>1590</b>), and causes the display unit <b>270</b> to display that the state is the voice input receivable state (Step S<b>1595</b>).
0233After the end of the processing in Step S<b>1595</b>, the device <b>1300</b> returns to the processing in Step S<b>1530</b> to perform the processing in Step S<b>1530</b> and the subsequent steps.
0234In the processing in Step S<b>1585</b>, in the case where the response text is a dialogue end voice (Step S<b>1585</b>: Yes), the device <b>1300</b> ends the second device processing.
0235The following explains a specific example of the operation performed by the first modified voice dialogue system having the above configuration, with reference to the drawing.
0236<Specific Example>
0237<figref idref="DRAWINGS">FIG. 16</figref> is a procedure diagram schematically showing a situation in which the user of the first modified voice dialogue system performs a voice dialogue with the voice dialogue agent <b>400</b> with use of the device <b>1300</b> (here, assumed to be a smartphone), and the voice dialogue agent <b>400</b> performs processing that reflects details of the dialogue.
0238Here, the explanation is given based on the assumption that a dialogue end voice is a voice “This ends dialogue.”.
0239In the figure, processing in Steps S<b>1600</b>-S<b>1615</b>, processing in Steps S<b>1630</b>-S<b>1635</b>, processing in Steps S<b>1650</b>-S<b>1655</b>, and processing in Steps S<b>1660</b>-S<b>1670</b> are respectively the same as the processing in Steps S<b>1000</b>-S<b>1015</b>, the processing in Steps S<b>1030</b>-S<b>1035</b>, the processing in Steps S<b>1050</b>-S<b>1055</b>, and the processing in Steps S<b>1060</b>-S<b>1070</b> in the specific examples in Embodiment 1 (see <figref idref="DRAWINGS">FIG. 10</figref>). Accordingly, the processing in the figure is regarded as having been already explained.
0240After the end of the processing in Step S<b>1615</b>, since a voice “Which room?” is not a dialogue end voice (corresponding to Step S<b>1585</b>: No in <figref idref="DRAWINGS">FIG. 15</figref>), the state is switched to the voice input receivable state (Step S<b>1625</b>, corresponding to Step S<b>1590</b> in <figref idref="DRAWINGS">FIG. 15</figref>). The device <b>1300</b> performs first voice input processing (Step S<b>1630</b>, corresponding to Step S<b>1530</b> in <figref idref="DRAWINGS">FIG. 15</figref>).
0241After the end of the processing in Step S<b>1635</b>, since a voice “Living room temperature is 28 degrees C. Do you need any other help?” is not a dialogue end voice (corresponding to Step S<b>1585</b>: No in <figref idref="DRAWINGS">FIG. 15</figref>), the state is switched to the voice input receivable state (Step S<b>1645</b>, corresponding to Step S<b>1590</b> in <figref idref="DRAWINGS">FIG. 15</figref>). The device <b>1300</b> performs first voice input processing (Step S<b>1650</b>, corresponding to Step S<b>1530</b> in <figref idref="DRAWINGS">FIG. 15</figref>).
0242After the end of the processing in Step S<b>1635</b>, since a voice “This ends dialogue.” is a dialogue end voice (corresponding to Step S<b>1585</b>: Yes in <figref idref="DRAWINGS">FIG. 15</figref>), the state is not switched to the voice input receivable state. The device <b>1300</b> ends the second device processing.
0243<Consideration>
0244According to the first modified voice dialogue system having the above configuration, in the case where the device <b>1300</b> outputs a voice based on response voice data transmitted from the voice dialogue agent <b>400</b> and the output voice is not a dialogue end voice, the state of the device <b>1300</b> is switched to the voice input receivable state even if the user does not perform a voice input start operation.
0245Accordingly, once the user performs a voice input start operation with respect to the device <b>1300</b>, the user can newly input a voice without newly performing a voice input operation with respect to the device <b>1300</b>, for a period from output of the voice based on the response voice data to lapse of the predetermined period T<b>1</b> until a dialogue end voice is output.
0246According to the first modified voice dialogue system, as described above, a further reduced number of operations need to be performed by the user in accordance with a voice input, compared with the voice dialogue system <b>100</b> in Embodiment 1.
Embodiment 3
0247<Outline>
0248The following explains, as one aspect of the voice dialogue method relating to the present invention and one aspect of the device relating to the present invention, a second modified voice dialogue system that is partially modified from the voice dialogue system <b>100</b> in Embodiment 1.
0249The voice dialogue system <b>100</b> in Embodiment 1 has been explained as an example of the configuration in which when the user performs a voice input start operation with respect to the device <b>140</b>, the device <b>140</b> is in the voice input receivable state for a period from performance of the voice input start operation to lapse of the predetermined period T<b>1</b>.
0250Compared with this, the second modified voice dialogue system in Embodiment 3 is an example of configuration in which once a user performs a voice input start operation with respect to a device, the device is in the voice input receivable state for a period from performance of the voice input start operation to output of a dialogue end voice.
0251The following explains the details of the second modified voice dialogue system, focusing on different points from the voice dialogue system <b>100</b> in Embodiment 1, with reference to the drawings.
0252<Configuration>
0253The second modified voice dialogue system is modified from the voice dialogue system <b>100</b> in Embodiment 1 so as to include a device <b>1700</b> instead of the device <b>140</b>.
0254The device <b>1700</b> is not modified from the device <b>140</b> in Embodiment 1 in terms of hardware, but is partially modified from the device <b>140</b> in terms of software to be stored as an execution target. Accordingly, the device <b>1700</b> is modified from the device <b>140</b> in Embodiment 1 in terms of part of functions.
0255<figref idref="DRAWINGS">FIG. 17</figref> is a block diagram showing functional configuration of the device <b>1700</b>.
0256As shown in the figure, the device <b>1700</b> is modified from the device <b>140</b> in Embodiment 1 (see <figref idref="DRAWINGS">FIG. 2</figref>) so as to include the control unit <b>1710</b> instead of the control unit <b>210</b>.
0257The control unit <b>1710</b> is modified from the control unit <b>210</b> in Embodiment 1 so as to have a second modified voice input unit state management function and a third device processing execution function, which are described below, instead of the voice input unit state management function and the first device processing execution function of the functions of the control unit <b>210</b>, respectively.
0258Similarly to the voice input unit state management function in Embodiment 1 and the first modified voice input unit state management function in Embodiment 2, the second modified voice input unit state management function is a function of managing the state of the voice input unit <b>220</b>, which is either the voice input receivable state or the voice input unreceivable state, and conditions for switching the state are partially modified from those in the voice input unit state management function in Embodiment 1.
0259<figref idref="DRAWINGS">FIG. 18</figref> shows switching of the state managed by the control unit <b>1710</b>.
0260As shown in the figure, in the case where the state is the voice input unreceivable state, (1) the control unit <b>1710</b> keeps the state to the voice input unreceivable state until the operation reception unit <b>230</b> receives a voice input start operation, and (2) after the reception of the voice input start operation by the operation reception unit <b>230</b>, the control unit <b>210</b> switches the state to the voice input receivable state. Then, in the case where the state is the voice input receivable state, (3) the control unit <b>1710</b> keeps the state to the voice input receivable state until the voice output unit <b>260</b> outputs a dialogue end voice (for example, a voice “This ends dialogue.”), and (4) after the output of the dialogue end voice by the voice output unit <b>260</b>, the control unit <b>1710</b> switches the state to the voice input unreceivable state.
0261Referring back to <figref idref="DRAWINGS">FIG. 17</figref>, the explanation on the control unit <b>1710</b> is continued.
0262The third device processing execution function is a function performed by the control unit <b>1710</b> controlling the voice input unit <b>220</b>, the operation reception unit <b>230</b>, the communication unit <b>250</b>, the voice output unit <b>260</b>, the display unit <b>270</b>, and the execution unit <b>280</b> to cause the device <b>1700</b> to execute the third device processing, as its characteristic operation, to execute a sequence of processing described below. In the sequence of processing, (1) when the user performs a voice input start operation, (2) the device <b>1700</b> receives a voice input from the user, and generates input voice data, (3) transmits the generated input voice data to a voice dialogue agent, (4) receives response voice data returned from the voice dialogue agent, (5) outputs a voice based on the received response voice data, and (6) in the case where the output voice is not a dialogue end voice, repeats the processing (2) and the subsequent processing even if the user does not perform a voice input start operation.
0263Note that the third device processing is explained in detail in section <Third Device Processing> later with reference to a flow chart.
0264The following explains the operation of the second modified voice dialogue system having the above configuration, with reference to the drawings.
0265<Operation>
0266The second modified voice dialogue system performs third device processing as its characteristic operation, in addition to the first agent processing in Embodiment 1. The third device processing partially modified from the first device processing in Embodiment 1.
0267Explanation is given on the third device processing below, focusing on different points from the first device processing.
0268<Third Device Processing>
0269The third device processing is processing performed by the device <b>1700</b>. In the third device processing, (1) when the user performs a voice input start operation with respect to the device <b>1700</b>, (3) the device <b>1700</b> receives a voice input from the user, and generates input voice data, (3) transmits the generated input voice data to a voice dialogue agent, (4) receives response voice data returned from the voice dialogue agent, (5) outputs a voice based on the received response voice data, and (6) in the case where the output voice is not a dialogue end voice, repeats the processing (2) and the subsequent processing even if the user does not perform a voice input start operation.
0270<figref idref="DRAWINGS">FIG. 19</figref> is a flow chart of the third device processing.
0271Upon bootup of the device <b>1700</b>, the third device processing is started.
0272At a time of bootup of the device <b>1700</b>, the state managed by the control unit <b>1710</b> is the voice input unreceivable state.
0273In the figure, processing in Steps S<b>1900</b>-S<b>1920</b> and processing in Steps S<b>1940</b>-S<b>1980</b> is respectively the same as the processing in Steps S<b>600</b>-S<b>620</b> and the processing in Steps S<b>640</b>-S<b>680</b> in the first device processing in Embodiment 1 (see <figref idref="DRAWINGS">FIG. 6</figref>). Accordingly, the processing in the figure is regarded as having been already explained.
0274After the end of the processing in Step S<b>1920</b>, the device <b>1700</b> executes second voice input processing (Step S<b>1930</b>).
0275<figref idref="DRAWINGS">FIG. 20</figref> is a flow chart of the second voice input processing.
0276When the second voice input processing is started, the voice input unit <b>220</b> receives a voice input from a user, and generates input voice data (Step S<b>2000</b>).
0277Then, the control unit <b>1910</b> controls the communication unit <b>250</b> to transmit the input voice data, which is generated by the voice input unit <b>220</b>, to the voice dialogue agent <b>400</b> (Step S<b>2040</b>).
0278After the end of the processing in Step S<b>2040</b>, the device <b>1700</b> ends the second voice input processing.
0279Referring back to <figref idref="DRAWINGS">FIG. 19</figref>, the explanation on the third device processing is continued.
0280After the end of the second voice input processing, the device <b>1700</b> proceeds to processing in Step S<b>1940</b> to perform the processing in Step S<b>1940</b> and processing in subsequent steps.
0281After the end of the processing in Step S<b>1980</b>, the control unit <b>1710</b> checks whether or not the voice, which is output from the voice output unit <b>260</b> in the processing in Step S<b>1980</b>, is a dialogue end voice (Step S<b>1985</b>). This processing is executed by for example checking whether or not the response text, which is received in the processing in Step S<b>1960</b>: Yes, is a predetermined character string (for example, a character string “This ends dialogue.”).
0282In the processing in Step S<b>1985</b>, in the case where the output voice is not a dialogue end voice (Step S<b>1985</b>: No), the device <b>1700</b> returns to the processing in Step S<b>1930</b> to repeat the processing in Step S<b>1930</b> and the subsequent steps.
0283In the processing in Step S<b>1985</b>, in the case where the output voice is a dialogue end voice (Step S<b>1585</b>: Yes), the control unit <b>1710</b> switches the state from the voice input receivable state to the voice input unreceivable state (Step S<b>1990</b>).
0284After the end of the processing in Step S<b>1990</b>, the device <b>1700</b> ends the third device processing.
0285The following explains a specific example of the operation performed by the second modified voice dialogue system having the above configuration, with reference to the drawing.
0286<Specific Example>
0287<figref idref="DRAWINGS">FIG. 21</figref> is a procedure diagram schematically showing a situation in which the user of the second modified voice dialogue system performs a voice dialogue with the voice dialogue agent <b>400</b> with use of the device <b>1700</b> (here, assumed to be a smartphone), and the voice dialogue agent <b>400</b> performs processing that reflects the dialogue.
0288Here, the explanation is given based on the assumption that a dialogue end voice is a voice “This ends dialogue.”.
0289In the figure, processing in Step S<b>2100</b>, processing in Step S<b>2105</b>, processing in Step S<b>2115</b>, processing in Step S<b>2135</b>, processing in Step S<b>2155</b>, and processing in Steps S<b>2160</b>-S<b>2170</b> are respectively the same as the processing in Step S<b>1000</b>, the processing in Step S<b>1005</b>, the processing in Step S<b>1015</b>, the processing in Step S<b>1035</b>, the processing in Step S<b>1055</b>, and the processing in Steps S<b>1060</b>-S<b>1070</b> in the specific examples in Embodiment 1 (see <figref idref="DRAWINGS">FIG. 10</figref>). Accordingly, the processing in the figure is regarded as having been already explained.
0290After the end of the processing in Step S<b>2105</b>, the device <b>1700</b> performs second voice input processing (Step S<b>2110</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0291In the second voice input processing, in the case where the user inputs a voice “What is room temperature?”, the device <b>1700</b> transmits input voice data “What is room temperature?” to the voice dialogue agent <b>400</b> (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0292After the end of the processing in Step S<b>2115</b>, since the voice “Which room?” is not a dialogue end voice (corresponding to Step S<b>1985</b>: No in <figref idref="DRAWINGS">FIG. 19</figref>), the device <b>1700</b> performs second voice input processing (Step S<b>2130</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0293In the second voice input processing, in the case where the user inputs a voice “Living room.”, the device <b>1700</b> transmits input voice data “Living room.” to the voice dialogue agent <b>400</b> (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0294After the end of the processing in Step S<b>2135</b>, since the voice “Living room temperature is 28 degrees C. Do you need any other help?” is not a dialogue end voice (corresponding to Step S<b>1985</b>: No in <figref idref="DRAWINGS">FIG. 19</figref>), the device <b>1700</b> performs second voice input processing (Step S<b>2150</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0295In the second voice input processing, in the case where the user inputs a voice “No. Thank you.”, the device <b>1700</b> transmits input voice data “No. Thank you.” to the voice dialogue agent <b>400</b> (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0296After the end of the processing in Step S<b>2135</b>, since a voice “This ends dialogue.” is a dialogue end voice (corresponding to Step S<b>1585</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), the state is switched to the voice input receivable state (corresponding to Step S<b>1990</b> in <figref idref="DRAWINGS">FIG. 19</figref>). The device <b>1700</b> ends the third device processing.
0297<Consideration>
0298According to the second modified voice dialogue system having the above configuration, once voice input start operation is performed, the device <b>1700</b> keeps in the voice input receivable state for a period from performance of the voice input start operation to output of a dialogue end voice.
0299Accordingly, once the user performs a voice input start operation with respect to the device <b>1700</b>, the user can newly input a voice without newly performing a voice input operation with respect to the device <b>1700</b> until a dialogue end voice is output.
0300According to the second modified voice dialogue system, as described above, a further reduced number of operations need to be performed by the user in accordance with a voice input, compared with the voice dialogue system <b>100</b> in Embodiment 1.
Embodiment 4
0301<Outline>
0302The following explains, as one aspect of the voice dialogue method relating to the present invention and one aspect of the device relating to the present invention, a third modified voice dialogue system that is partially modified from the second modified voice dialogue system in Embodiment 3.
0303The second modified voice dialogue system in Embodiment 3 has been explained as an example of the configuration in which once the device <b>1700</b> starts communication with a voice dialogue agent A, a voice dialogue agent as a communication party is limited to the voice dialogue agent A until a series of processing ends.
0304Compared with this, the third modified voice dialogue system in Embodiment 4 is an example of configuration in which in the case where a device starts communication with a voice dialogue agent A and a user of the third modified voice dialogue system inputs, with use of the device, a voice indicating that the user hopes to communicate with another voice dialogue agent B, a communication party of the device is changed from the voice dialogue agent A to the voice dialogue agent B.
0305The following explains the details of the third modified voice dialogue system, focusing on different points from the second modified voice dialogue system in Embodiment 3, with reference to the drawings.
0306<Configuration>
0307The third modified voice dialogue system is modified from the second voice dialogue system in Embodiment 3 so as to include a voice dialogue agent <b>2200</b> instead of the voice dialogue agent <b>400</b>.
0308Similarly to the voice dialogue agent <b>400</b> in Embodiment 3, the voice dialogue agent <b>2200</b> is embodied by the voice dialogue agent server <b>110</b>.
0309Software for embodying the voice dialogue agent <b>2200</b>, which is executed by the voice dialogue agent server <b>110</b>, is partially modified from the software for embodying the voice dialogue agent <b>400</b> in Embodiment 3. Accordingly, the voice dialogue agent <b>2200</b> is modified from the voice dialogue agent <b>400</b> in Embodiment 3 in terms of part of functions.
0310<figref idref="DRAWINGS">FIG. 22</figref> is a block diagram showing functional configuration of the voice dialogue agent <b>2200</b>.
0311As shown in the figure, the voice dialogue agent <b>2200</b> is modified from the voice dialogue agent <b>400</b> in Embodiment 3 (see <figref idref="DRAWINGS">FIG. 4</figref>) so as to additionally include a target agent DB storage unit <b>2220</b> and include a control unit <b>2210</b> instead of the control unit <b>410</b>.
0312The target agent DB storage unit <b>2220</b> is for example embodied by a memory and a processor that executes programs. The target agent DB storage unit <b>2220</b> is connected to the control unit <b>2210</b>, and has a function of storing therein a target agent DB <b>2300</b>.
0313<figref idref="DRAWINGS">FIG. 23</figref> is a data structure diagram showing the target agent DB <b>2300</b> stored in the target agent DB storage unit <b>2220</b>.
0314As shown in the figure, the target agent DB <b>2300</b> includes keyword <b>2310</b>, target agent <b>2320</b>, and IP address <b>2330</b> that are associated with each other.
0315The keyword <b>2310</b> indicates a character string that is assumed to be included in an input text converted by the voice recognition processing unit <b>430</b>.
0316The target agent <b>2320</b> is information for specifying, as a communication party of the device <b>140</b>, one of a plurality of voice dialogue agents <b>2200</b>. Hereinafter, this one of the voice dialogue agents <b>2200</b> is referred to as an additional voice dialogue agent.
0317In this example, the additional voice dialogue agent specified by the target agent <b>2320</b> is a car agent, a retailer agent, or a home agent.
0318Here, the car agent indicates one of voice dialogue agents <b>2200</b> that provides a relatively satisfactory service relating to devices in mounted in a car. The retailer agent indicates one of voice dialogue agents <b>2200</b> that provides a relatively satisfactory service relating to devices in mounted in a retailer. The home agent indicates one of voice dialogue agents <b>2200</b> that provides a relatively satisfactory service relating to devices in mounted in a residence (home).
0319The IP address <b>2330</b> indicates an IP address in the network <b>120</b> relating to the voice dialogue agent server <b>110</b> that embodies an additional voice dialogue agent specified by the associated target agent <b>2320</b>.
0320As shown in <figref idref="DRAWINGS">FIG. 23</figref>, each of the additional voice dialogue agents specified by the target agent <b>2320</b> is associated with one or more character strings indicated by the keyword <b>2310</b>. For example, the car agent is associated with character strings indicated by the keyword <b>2310</b>, such as character strings “in-car”, “car”, “vehicle”, and “navigation system”.
0321Since each of the additional voice dialogue agents, which is specified by the target agent <b>2320</b>, is associated with one or more character strings, which are indicated by the keyword <b>2310</b>, the voice dialogue agent <b>2200</b> can respond an ambiguous input.
0322For example, in the case where the user hopes to communicate with the car agent, the user sometimes inputs a voice “Connect to voice dialogue agent of navigation system.”, and sometimes inputs a voice “Connect to voice dialogue agent of car.”.
0323Here, the character strings indicated by the keyword <b>2310</b> “navigation system” and “car” are each associated with the car agent. Accordingly, both in the case where a voice “navigation system” is input and in the case where a voice “car” is input, it is possible to specify the car agent as the additional voice dialogue agent <b>2200</b>, which is specified by the target agent <b>2320</b>, by referring to the target agent DB <b>2300</b>.
0324Referring back to <figref idref="DRAWINGS">FIG. 22</figref>, the explanation on the voice dialogue agent <b>2200</b> is continued.
0325The control unit <b>2210</b> is modified from the control unit <b>410</b> in Embodiment 3 so as to have a second agent processing execution function and a third agent processing execution function, which are described below, instead of the first agent processing execution function of the control unit <b>410</b>.
0326The second agent processing execution function is a function performed by the control unit <b>2210</b> controlling the communication unit <b>420</b>, the voice recognition processing unit <b>430</b>, the voice synthesizing processing unit <b>450</b>, and the instruction generation unit <b>460</b> to cause the voice dialogue agent <b>2200</b> to execute second agent processing as its characteristic operation to execute a sequence of processing described below. In the sequence of processing, (1) the voice dialogue agent <b>2200</b> receives input voice data transmitted from a device, (2) performs voice recognition processing on the received input voice data to generate an input text, and returns the generated input text to the device, (3) in the case where the generated input text indicates that the user hopes to communicate with another voice dialogue agent, establishes communication between the device and the other voice dialogue agent, (4) otherwise, generates an instruction set based on the generated input text, and executes the generated instruction set, (5) generates a response text based on an execution result of the instruction set, (6) converts the generated response text to response voice data, and (7) returns the response text and the response voice data to the device.
0327Note that the second agent processing is explained in detail in section <Second Agent Processing> later with reference to a flow chart.
0328The third agent processing execution function is a function performed by the control unit <b>2210</b> controlling the communication unit <b>420</b>, the voice recognition processing unit <b>430</b>, the voice synthesizing processing unit <b>450</b>, and the instruction generation unit <b>460</b> to cause the voice dialogue agent <b>2200</b> to execute third agent processing as its characteristic operation to execute a sequence of processing described below. In the sequence of processing, (1) the voice dialogue agent <b>2200</b> starts communication with a device in response to a request from another voice dialogue agent, (2) receives input voice data transmitted from the device, (3) performs voice recognition processing on the received input voice data to generate an input text, and returns the generated input text, (4) generates an instruction set based on the generated input text, and executes the generated instruction set, (5) generates a response text based on an execution result of the instruction set, (6) converts the generated response text to response voice data, and (7) returns the response text and the response voice data to the device.
0329Note that the third agent processing is explained in detail in section <Third Agent Processing> later with reference to a flow chart.
0330The following explains the operation of the third modified voice dialogue system having the above configuration, with reference to the drawings.
0331<Operation>
0332The third modified voice dialogue system performs second agent processing and third agent processing as its characteristic operation, in addition to the first agent processing in Embodiment 1. The second agent processing and the third agent processing is partially modified from the first agent processing in Embodiment 3.
0333Explanation is given on the second agent processing and the third agent processing below, focusing on different points from the first agent processing.
0334<Second Agent Processing>
0335The second agent processing is processing performed by the voice dialogue agent <b>2200</b>. In the second agent processing, (1) the voice dialogue agent <b>2200</b> receives input voice data transmitted from a device, (2) performs voice recognition processing on the received input voice data to generate an input text, and returns the generated input text to the device, (3) in the case where the generated input text indicates that the user hopes to communicate with another voice dialogue agent, establishes communication between the device and the other voice dialogue agent, (4) otherwise, generates an instruction set based on the generated input text, and executes the generated instruction set, (5) generates a response text based on an execution result of the instruction set, (6) converts the generated response text to response voice data, and (7) returns the response text and the response voice data to the device.
0336<figref idref="DRAWINGS">FIG. 24</figref> is a flow chart of the second agent processing.
0337Upon bootup of the voice dialogue agent <b>2200</b>, the second agent processing is started.
0338When the second agent processing is started, the voice dialogue agent <b>2200</b> stands by until the communication unit <b>420</b> receives input voice data transmitted from the device <b>1700</b> (Step S<b>2400</b>: Repetition of No). When the communication unit <b>420</b> receives the input voice data (Step S<b>2400</b>: Yes), the voice dialogue agent <b>2200</b> performs second instruction execution processing (Step S<b>2410</b>).
0339<figref idref="DRAWINGS">FIG. 25</figref> is a flow chart of the second instruction execution processing.
0340In the figure, processing in Steps S<b>2500</b>-S<b>2510</b> and processing in Steps S<b>2520</b>-S<b>2560</b> is respectively the same as the processing in Steps S<b>900</b>-S<b>910</b> and the processing in Steps S<b>920</b>-S<b>960</b> in the first instruction execution processing in Embodiment 3 (see <figref idref="DRAWINGS">FIG. 9</figref>). Accordingly, the processing in the figure is regarded as having been already explained.
0341After the end of the processing in Step S<b>2510</b>, the control unit <b>2210</b> checks whether or not the input text, which is converted by the voice recognition processing unit <b>430</b>, requests to communicate with another voice dialogue agent (Step S<b>2515</b>).
0342In the processing in Step S<b>2515</b>, in the case where the input text does not request communication with another voice dialogue agent (Step S<b>2515</b>: No), the voice dialogue agent <b>2200</b> proceeds to the processing in Step S<b>2520</b> to perform the processing in Steps S<b>2520</b>-S<b>2560</b>.
0343In the processing in Step S<b>2515</b>, in the case where the input text requests to communicate with another voice dialogue agent (Step S<b>2515</b>: Yes), the control unit <b>2210</b> specifies a voice dialogue agent <b>2200</b> that is requested as a communication party, with reference to the target agent DB <b>2300</b> stored in the target agent DB storage unit <b>2220</b> (Step S<b>2517</b>). In other words, the control unit <b>2210</b> specifies, as the voice dialogue agent <b>2200</b> requested as a communication party, an additional voice dialogue agent that is specified by the target agent <b>2320</b> associated with a character string that is indicated by the keyword <b>2310</b> included in the input text, which is converted by the voice recognition processing unit <b>430</b>.
0344After the specification of the additional voice dialogue agent requested as a communication party, the control unit <b>2210</b> generates a predetermined signal indicating to start communication between the specified additional voice dialogue agent and the device <b>1700</b> which has transmitted the input voice data (Step S<b>2565</b>). Hereinafter, this signal is referred to as a connection instruction.
0345After the generation of the connection instruction, the control unit <b>2210</b> controls the communication unit <b>420</b> to transmit the generated connection instruction to the additional voice dialogue agent, with use of an IP address indicated by the IP address <b>2330</b> which is associated with the character string indicate by the keyword <b>2310</b> (Step S<b>2570</b>).
0346Then, the control unit <b>2210</b> stands by until the communication unit <b>420</b> receives a connection response (described later) that is returned from the additional voice dialogue agent in response to the connection instruction that is transmitted in the processing in Step S<b>2570</b> (Step S<b>2575</b>: Repetition of No).
0347When the connection response is received by the communication unit <b>420</b> (Step S<b>2575</b>: Yes), the voice dialogue agent <b>2200</b> executes first connection response processing (Step S<b>2580</b>).
0348<figref idref="DRAWINGS">FIG. 26</figref> is a flow chart of the first connection response processing.
0349When the first connection response processing is started, the control unit <b>2210</b> generates a predetermined response text indicating that communication becomes available between the additional voice dialogue agent and the device <b>1700</b> (Step S<b>2600</b>). The predetermined response text is for example a character string “Connection to [Additional voice dialogue agent] has been established.”.
0350Here, in part [Additional voice dialogue agent] in the character string, a name of the voice dialogue agent <b>2200</b> (here, either of the car agent, the retailer agent, or the home agent), which is specified by the target agent <b>2320</b> included in the target agent DB <b>2300</b>, is inserted.
0351After the generation of the response text, the voice synthesizing processing unit <b>450</b> performs voice synthesizing processing on the generated response text to generate response voice data (Step S<b>2610</b>).
0352After the generation of the response voice data, the control unit <b>2210</b> controls the communication unit <b>420</b> to transmit the generated response text and response voice data to the device <b>1700</b> which has transmitted the input voice data (Step S<b>2620</b>).
0353After the end of the processing in Step S<b>2620</b>, the voice dialogue agent <b>2200</b> ends the first connection response processing.
0354Referring back to <figref idref="DRAWINGS">FIG. 25</figref>, the explanation on the second instruction execution processing is continued.
0355After the end of the first connection response processing, the voice dialogue agent <b>2200</b> stands by until the communication unit <b>420</b> receives a disconnection response (described later) that is transmitted from the additional voice dialogue agent (Step S<b>2585</b>: Repetition of No).
0356When the communication unit <b>420</b> receives the disconnection response (Step S<b>2585</b>: Yes), the voice dialogue agent <b>2200</b> executes disconnection response processing (Step S<b>2590</b>).
0357<figref idref="DRAWINGS">FIG. 27</figref> is a flow chart of the disconnection response processing.
0358When the disconnection response processing is started, the control unit <b>2210</b> generates a predetermined response text indicating that the communication ends between the additional voice dialogue agent and the device <b>1700</b> (Step S<b>2700</b>). The predetermined response text is for example a character string “Connection to [Additional voice dialogue agent] has been terminated. Do you need any other help?”.
0359Here, in part [Additional voice dialogue agent] in the character string, a name of the voice dialogue agent <b>2200</b> (here, either of the car agent, the retailer agent, or the home agent), which is specified by the target agent <b>2320</b> included in the target agent DB <b>2300</b>, is inserted.
0360After the generation of the response text, the voice synthesizing processing unit <b>450</b> performs voice synthesizing processing on the generated response text to generate response voice data (Step S<b>2710</b>).
0361After the generation of the response voice data, the control unit <b>2210</b> controls the communication unit <b>420</b> to transmit the generated response text and response voice data to the device <b>1700</b> which has transmitted the input voice data (Step S<b>2720</b>).
0362After the end of the processing in Step S<b>2720</b>, the voice dialogue agent <b>2200</b> ends the disconnection response processing.
0363Referring back to <figref idref="DRAWINGS">FIG. 25</figref> again, the explanation on the second instruction execution processing is continued.
0364After the end of the disconnection response processing, or after the end of the processing in Step S<b>2560</b>, the voice dialogue agent <b>2200</b> ends the second instruction execution processing.
0365Referring back to <figref idref="DRAWINGS">FIG. 24</figref>, the explanation on the second agent processing is continued.
0366After the end of the second instruction execution processing, the voice dialogue agent <b>2200</b> returns to the processing in Step S<b>2400</b> to perform the processing in Step S<b>2400</b> and the subsequent steps.
0367<Third Agent Processing>
0368The third agent processing is processing performed by the voice dialogue agent <b>2200</b>. In the third agent processing, (1) the voice dialogue agent 2200 starts communication with a device in response to a request from another voice dialogue agent, (2) receives input voice data transmitted from the device, (3) performs voice recognition processing on the received input voice data to generate an input text, and returns the generated input text, (4) generates an instruction set based on the generated input text, and executes the generated instruction set, (5) generates a response text based on an execution result of the instruction set, (6) converts the generated response text to response voice data, and (7) returns the response text and the response voice data to the device.
0369<figref idref="DRAWINGS">FIG. 28</figref> is a flow chart of the third agent processing.
0370In the figure, processing in Steps S<b>2800</b>-S<b>2810</b> and processing in Steps S<b>2820</b>-S<b>2860</b> is respectively the same as the processing in Steps S<b>900</b>-S<b>910</b> and the processing in Steps S<b>920</b>-S<b>960</b> in the first instruction execution processing in Embodiment 1 (see <figref idref="DRAWINGS">FIG. 9</figref>). Accordingly, the processing in the figure is regarded as having been already explained.
0371Upon bootup of the voice dialogue agent <b>2200</b>, the third agent processing is started.
0372When the third agent processing is started, the voice dialogue agent <b>2200</b> stands by until the communication unit <b>420</b> receives a connection instruction transmitted from another voice dialogue agent (Step S<b>2811</b>: Repetition of No). When the communication unit <b>420</b> receives the connection instruction (Step S<b>2811</b>: Yes), the control unit <b>2210</b> controls the communication unit <b>420</b> to execute connection processing of starting communication with the device <b>1700</b> that is a communication party requested by the connection instruction.
0373Here, the connection processing includes processing of changing a transmission destination of input voice data to be transmitted from the device <b>1700</b> from the voice dialogue agent <b>2200</b>, which has transmitted the connection instruction, to the voice dialogue agent <b>2200</b>, which has received the connection instruction.
0374After the execution of the connection processing, the control unit <b>2210</b> controls the communication unit <b>420</b> to generate a connection response that is a signal indicating that communication with the device <b>1700</b> has started, and transmits the generated connection response to the voice dialogue agent which has transmitted the connection instruction (Step S<b>2813</b>).
0375Then, the control unit <b>2210</b> stands by until the communication unit <b>420</b> receives the input voice data transmitted from the device <b>1700</b> (Step S<b>2814</b>: Repetition of No). When the communication unit <b>420</b> receives the input voice data (Step S<b>2814</b>: Yes), the control unit <b>2210</b> performs the processing in Steps S<b>2800</b>-S<b>2810</b>.
0376After the end of the processing in Step S<b>2810</b>, the control unit <b>2210</b> checks whether or not the input text, which is converted by the voice recognition processing unit <b>430</b>, requests to terminate communication with the voice dialogue agent <b>2200</b> (Step S<b>2815</b>).
0377In the processing in Step S<b>2815</b>, in the case where the input text does not indicate to terminate the communication with the voice dialogue agent <b>2200</b> (Step S<b>2815</b>: No), the voice dialogue agent <b>2200</b> proceeds to the processing in Step S<b>2820</b> to perform the processing in Steps S<b>2820</b>-S<b>2860</b>. After the end of the processing in Step S<b>2860</b>, the voice dialogue agent <b>2200</b> returns to the processing in Step S<b>2814</b> to perform the processing in Step S<b>2814</b> and the subsequent steps.
0378In the processing in Step S<b>2815</b>, in the case where the input text indicates to terminate the communication with the voice dialogue agent <b>2200</b> (Step S<b>2815</b>: Yes), the control unit <b>2210</b> controls the communication unit <b>420</b> to execute disconnection processing of terminating the communication with the device <b>1700</b>.
0379Here, the disconnection processing includes processing of changing the transmission destination of input voice data to be transmitted from the device <b>1700</b> from the voice dialogue agent <b>2200</b>, which has received the connection instruction, to the voice dialogue agent <b>2200</b>, which has transmitted the connection instruction.
0380After the execution of the disconnection processing, the control unit <b>2210</b> controls the communication unit <b>420</b> to generate a disconnection response that is a predetermined signal indicating that the communication with the device <b>1700</b> has been terminated, and transmits the generated disconnection response to the voice dialogue agent which has transmitted the connection instruction (Step S<b>2890</b>).
0381After the end of the processing in Step S<b>2890</b>, the voice dialogue agent <b>2200</b> returns to the processing in Step S<b>2811</b> to perform the processing in Step S<b>2811</b> and the subsequent steps.
0382The following explains a specific example of the operation performed by the third modified voice dialogue system having the above configuration, with reference to the drawing.
0383<Specific Example>
0384<figref idref="DRAWINGS">FIG. 29</figref> is a procedure diagram schematically showing a situation in which the user of the third modified voice dialogue system starts, with use of the device <b>1700</b>, a voice dialogue with a home agent, which is one of the voice dialogue agents <b>2200</b>, and then starts communication with the car agent, which is one of the voice dialogue agents <b>2200</b>, in response to a connection instruction generated by the home agent, and performs a dialogue with the car agent.
0385Here, the explanation is given based on the assumption that a specific voice dialogue agent server for the device <b>1700</b> used by the user is the voice dialogue agent server <b>110</b> that embodies the home agent, and a dialogue end voice is a voice “This ends dialogue.”.
0386In the figure, processing in Steps S<b>2900</b>-S<b>2905</b> is respectively the same as the processing in Steps S<b>2100</b>-S<b>2105</b> in the specific example in Embodiment 3 (see <figref idref="DRAWINGS">FIG. 21</figref>). Accordingly, the processing in the figure is regarded as having been already explained.
0387After the end of the processing in Step S<b>2905</b>, the device <b>1700</b> performs second voice input processing (Step S<b>2906</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0388In the second voice input processing, in the case where the user inputs a voice “Connect to car agent.”, the device <b>1700</b> transmits input voice data “Connect to car agent.” to the home agent (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0389Then, the home agent receives the input voice data (corresponding to Step S<b>2400</b>: Yes in <figref idref="DRAWINGS">FIG. 24</figref>), and performs second instruction execution processing (corresponding to Step S<b>2410</b> in <figref idref="DRAWINGS">FIG. 24</figref>).
0390In the second instruction execution processing, since the input text requests to communicate with the car agent (corresponding to Step S<b>2515</b>: Yes in <figref idref="DRAWINGS">FIG. 25</figref>), the home agent transmits a connection instruction to the car agent (corresponding to Step S<b>2570</b> in <figref idref="DRAWINGS">FIG. 25</figref>).
0391Then, the car agent receives the connection instruction (corresponding to Step S<b>2811</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>), and starts communication with the device <b>1700</b> (corresponding to Step S<b>2812</b> in <figref idref="DRAWINGS">FIG. 28</figref>), and transmits a connection response to the home agent (Step S<b>2990</b>, corresponding to Step S<b>2813</b> in <figref idref="DRAWINGS">FIG. 28</figref>).
0392The home agent receives the connection response (corresponding to Step S<b>2575</b>: Yes in <figref idref="DRAWINGS">FIG. 25</figref>), and performs first connection response processing (Step S<b>2965</b>, corresponding to Step S<b>2580</b> in <figref idref="DRAWINGS">FIG. 25</figref>).
0393Here, in the first connection response processing, in the case where the voice dialogue agent <b>2200</b> generates response voice data “Connection to car agent has been established.”, the voice dialogue agent <b>2200</b> transmits response voice data “Connection to car agent has been established.” to the device <b>1700</b> (corresponding to Step S<b>2620</b> in <figref idref="DRAWINGS">FIG. 26</figref>).
0394Then, the device <b>1700</b> receives the response voice data (corresponding to Step S<b>1960</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), and outputs a voice “Connection to car agent has been established.” (Step S<b>2907</b>, corresponding to Step S<b>1980</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0395Since the voice “Connection to car agent has been established.” is not a dialogue end voice (corresponding to Step S<b>1985</b>: No in <figref idref="DRAWINGS">FIG. 19</figref>), the device <b>1700</b> performs second voice input processing (Step S<b>2910</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0396In the second voice input processing, in the case where the user inputs a voice “What is temperature in car?”, the device <b>1700</b> transmits input voice data “What is temperature in car?” to the car agent (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0397Then, the car agent receives the input voice data (corresponding to Step S<b>2814</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>). Since the input voice data does not request to terminate the communication (corresponding to Step S<b>2815</b>: No in <figref idref="DRAWINGS">FIG. 28</figref>), the car agent generates an instruction set corresponding to the input voice data, and executes the generated instruction set (Step S<b>2994</b>, corresponding to Step S<b>2830</b> in <figref idref="DRAWINGS">FIG. 28</figref>).
0398Here, in execution of the instruction set, in the case where the car agent generates response voice data “Temperature in car is 38 degrees C. Do you need any other help?”, the car agent transmits the response voice data “Temperature in car is 38 degrees C. Do you need any other help?” to the device <b>1700</b> (corresponding to Step S<b>2860</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>).
0399Then, the device <b>1700</b> receives the response voice data (corresponding to Step S<b>1960</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), and outputs a voice “Temperature in car is 38 degrees C. Do you need any other help?” (Step S<b>2915</b>, corresponding to Step S<b>1980</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0400Since the voice “Temperature in car is 38 degrees C. Do you need any other help?” is not a dialogue end voice (corresponding to Step S<b>1985</b>: No in <figref idref="DRAWINGS">FIG. 19</figref>), the device <b>1700</b> performs second voice input processing (Step S<b>2930</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0401In the second voice input processing, in the case where the user inputs a voice “No. Thank you.”, the device <b>1700</b> transmits input voice data “No. Thank you.” to the car agent (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0402Then, the car agent receives the input voice data (corresponding to Step S<b>2814</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>). Since the input voice data requests to terminate the communication (corresponding to Step S<b>2815</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>), the car agent terminates the communication with the device <b>1700</b> (corresponding to Step S<b>2870</b> in <figref idref="DRAWINGS">FIG. 28</figref>), and transmits a disconnection response to the home agent (Step S<b>2998</b>, corresponding to Step S<b>2890</b> in <figref idref="DRAWINGS">FIG. 28</figref>).
0403Then, the home agent receives the disconnection response (corresponding to Step S<b>2585</b>: Yes in <figref idref="DRAWINGS">FIG. 25</figref>), and performs disconnection response processing (Step S<b>2970</b>, corresponding to Step S<b>2890</b> in <figref idref="DRAWINGS">FIG. 25</figref>).
0404Here, in the disconnection processing, in the case where the voice dialogue agent <b>2200</b> generates response voice data “Connection to car agent has been terminated. Do you need any other help?”, the voice dialogue agent <b>2200</b> transmits the response voice data “Connection to car agent has been terminated. Do you need any other help?” to the device <b>1700</b> (corresponding to Step S<b>2720</b> in <figref idref="DRAWINGS">FIG. 27</figref>).
0405Then, the device <b>1700</b> receives the response voice data (corresponding to Step S<b>1960</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), and outputs a voice “Connection to car agent has been terminated. Do you need any other help?” (Step S<b>2935</b>, corresponding to Step S<b>1980</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0406Since the voice “Connection to car agent has been terminated. Do you need any other help?” is not a dialogue end voice (corresponding to Step S<b>1985</b>: No in <figref idref="DRAWINGS">FIG. 19</figref>), the device <b>1700</b> performs second voice input processing (Step S<b>2950</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0407In the second voice input processing, in the case where the user inputs a voice “No. Thank you.”, the device <b>1700</b> transmits input voice data “No. Thank you.” to the home agent (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0408Then, the home agent receives the input voice data (corresponding to Step S<b>2800</b>: Yes in <figref idref="DRAWINGS">FIG. 24</figref>), and performs second instruction execution processing (Step S<b>2975</b>, corresponding to Step S<b>2410</b> in <figref idref="DRAWINGS">FIG. 24</figref>).
0409Here, in the second instruction execution processing, in the case where the home agent generates response voice data “This ends dialogue.”, the home agent transmits the response voice data “This ends dialogue.” to the device <b>1700</b> (corresponding to Step S<b>2560</b> in <figref idref="DRAWINGS">FIG. 25</figref>).
0410Then, the device <b>1700</b> receives the response voice data (corresponding to Step S<b>1960</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), and outputs a voice “This ends dialogue.” (Step S<b>2955</b>, corresponding to Step S<b>1980</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0411Since the voice “This ends dialogue.” is a dialogue end voice (corresponding to Step S<b>1585</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), the state is switched to the voice input receivable state (corresponding to Step S<b>1990</b> in <figref idref="DRAWINGS">FIG. 19</figref>). The device <b>1700</b> ends the third device processing.
0412<Consideration>
0413According to the third modified voice dialogue system having the above configuration, in the case where the user of the third modified voice dialogue system, who is communicating with the voice dialogue agent A, hopes to cause the voice dialogue agent B rather than the voice dialogue agent A to perform processing, it is possible to change the voice dialogue agent that is appropriate for performing the processing via communication from the voice dialogue agent A to the voice dialogue agent B to cause the voice dialogue agent B to perform desired processing.
0414Also, in this case, since the voice dialogue agent A transfers input voice data that is not modified to the voice dialogue agent B, the voice dialogue agent B performs voice recognition processing on the input voice data. As a result, the user can receive a more appropriate service from the voice dialogue agent B.
Embodiment 5
0415<Outline>
0416The following explains, as one aspect of the voice dialogue method relating to the present invention and one aspect of the device relating to the present invention, a fourth modified voice dialogue system that is partially modified from the third modified voice dialogue system in Embodiment 4.
0417The third modified voice dialogue system in Embodiment 4 has been explained as an example of the configuration in which in the case where a device starts communication with the voice dialogue agent A and the user of the third modified voice dialogue system inputs, with use of the device, a voice indicating that the user hopes to communicate with another voice dialogue agent B, a communication party of the device is changed from the voice dialogue agent A to the voice dialogue agent B.
0418Compared with this, the fourth modified voice dialogue system in Embodiment 5 is an example of configuration in which in the case where a device starts communication with a voice dialogue agent A and predetermined condition is satisfied for the communication, the voice dialogue agent A determines that the voice dialogue agent B rather than the voice dialogue agent A is appropriate as a communication party, and a communication party of the device is changed from the voice dialogue agent A to the voice dialogue agent B.
0419The following explains the details of the fourth modified voice dialogue system, focusing on different points from the third modified voice dialogue system in Embodiment 4, with reference to the drawings.
0420<Configuration>
0421The fourth modified voice dialogue system is modified from the third voice dialogue system in Embodiment 4 so as to include a voice dialogue agent <b>3000</b> instead of the voice dialogue agent <b>2200</b>.
0422Similarly to the voice dialogue agent <b>2200</b> in Embodiment 4, the voice dialogue agent <b>3000</b> is embodied by the voice dialogue agent server <b>110</b>.
0423Software for embodying the voice dialogue agent <b>3000</b>, which is executed by the voice dialogue agent server <b>110</b>, is partially modified from the software for embodying the voice dialogue agent <b>2200</b> in Embodiment 3. Accordingly, the voice dialogue agent <b>3000</b> is modified from the voice dialogue agent <b>2200</b> in Embodiment 4 in terms of part of functions.
0424<figref idref="DRAWINGS">FIG. 30</figref> is a block diagram showing functional configuration of the voice dialogue agent <b>3000</b>.
0425As shown in the figure, the voice dialogue agent <b>3000</b> is modified from the voice dialogue agent <b>2200</b> in Embodiment 4 (see <figref idref="DRAWINGS">FIG. 22</figref>) so as not to include the target agent DB storage unit <b>2220</b>, and so as to additionally include an available service DB storage unit <b>3020</b> and include a control unit <b>3010</b> instead of the control unit <b>2210</b>.
0426The available service DB storage unit <b>3020</b> is for example embodied by a memory and a processor that executes programs. The available service DB storage unit <b>3020</b> is connected to the control unit <b>3010</b>, and has a function of storing therein an available service DB <b>3100</b>.
0427<figref idref="DRAWINGS">FIG. 31</figref> is a data structure diagram showing the available service DB <b>3100</b> stored in the available service DB storage unit <b>3020</b>.
0428As shown in the figure, the available service DB <b>3100</b> includes keyword <b>3110</b>, target agent <b>3120</b>, processing details <b>3130</b>, IP address <b>3140</b>, and availability <b>3150</b> that are associated with each other.
0429The keyword <b>3110</b> indicates a character string that is assumed to be included in an input text converted by the voice recognition processing unit <b>430</b>.
0430The target agent <b>3120</b> is information for specifying an additional voice dialogue agent as a communication party of the device <b>1700</b>.
0431In this example, the additional voice dialogue agents specified by the target agent <b>2320</b> include the car agent, the retailer agent, and the home agent, similarly to Embodiment 4.
0432The processing details <b>3130</b> are information for specifying, in the case where a character string indicated by the associated keyword <b>3110</b> is included in an input text that is converted by the voice recognition processing unit <b>430</b>, processing that is determined to be executed by a device that is specified by the associated target device <b>3120</b>.
0433The IP address <b>3140</b> indicates an IP address in the network <b>120</b> relating to the voice dialogue agent server <b>110</b> that embodies the additional voice dialogue agent specified by the associated target agent <b>3120</b>.
0434The availability <b>3150</b> is information for specifying whether or not the voice dialogue agent can perform processing specified by the associated processing details <b>3130</b>.
0435Referring back to <figref idref="DRAWINGS">FIG. 30</figref>, the explanation on the voice dialogue agent <b>3000</b> is continued.
0436The control unit <b>3010</b> is modified from the control unit <b>2210</b> in Embodiment 4 so as to have a fourth agent processing execution function, which is described below, instead of the second agent processing execution function of the control unit <b>2210</b>.
0437The fourth agent processing execution function is a function performed by the control unit <b>3010</b> controlling the communication unit <b>420</b>, the voice recognition processing unit <b>430</b>, the voice synthesizing processing unit <b>450</b>, and the instruction generation unit <b>460</b> to control the voice dialogue agent <b>3000</b> to execute the fourth agent processing, which is its characteristic operation, to execute a sequence of processing described below. In the sequence of processing, (1) the voice dialogue agent <b>3000</b> receives input voice data transmitted from a device, (2) performs voice recognition processing on the received input voice data to generate an input text, and returns the generated input text to the device, (3) in the case where the generated input text includes a predetermined keyword, establishes communication between the device and a target agent associated with the predetermined keyword, (4) otherwise, generates an instruction set based on the generated input text, and executes the generated instruction set, (5) generates a response text based on an execution result of the instruction set, (6) converts the generated response text to response voice data, and (7) returns the response text and the response voice data to the device.
0438Note that the fourth agent processing is explained in detail in section <Fourth Agent Processing> later with reference to a flow chart.
0439The following explains the operation of the fourth modified voice dialogue system having the above configuration, with reference to the drawings.
0440<Operation>
0441The fourth modified voice dialogue system performs fourth agent processing as its characteristic operation, in addition to the second device processing and the third agent processing in Embodiment 4. The fourth agent processing is partially modified from the second agent processing in Embodiment 3.
0442Explanation is given on the fourth agent processing below, focusing on different points from the second agent processing.
0443<Fourth Agent Processing>
0444The fourth agent processing is processing performed by the voice dialogue agent <b>3000</b>. In the fourth agent processing, (1) the voice dialogue agent <b>3000</b> receives input voice data transmitted from a device, (2) performs voice recognition processing on the received input voice data to generate an input text, and returns the generated input text to the device, (3) in the case where the generated input text includes a predetermined keyword, establishes communication between the device and a target agent associated with the predetermined keyword, (4) otherwise, generates an instruction set based on the generated input text, and executes the generated instruction set, (5) generates a response text based on an execution result of the instruction set, (6) converts the generated response text to response voice data, and (7) returns the response text and the response voice data to the device.
0445<figref idref="DRAWINGS">FIG. 32</figref> is a flow chart of the fourth agent processing.
0446Upon bootup of the voice dialogue agent <b>3000</b>, the fourth agent processing is started.
0447When the fourth agent processing is started, the voice dialogue agent <b>3000</b> stands by until the communication unit <b>420</b> receives input voice data transmitted from the device <b>1700</b> (Step S<b>3200</b>: Repetition of No). When the communication unit <b>430</b> receives the input voice data (Step S<b>3200</b>: Yes), the voice dialogue agent <b>3000</b> performs second instruction execution processing (Step S<b>3210</b>).
0448<figref idref="DRAWINGS">FIG. 33</figref> is a flow chart of the third instruction execution processing.
0449In the figure, processing in Steps S<b>3300</b>-S<b>3310</b>, processing in Steps S<b>3320</b>-S<b>3360</b>, processing in Steps S<b>3365</b>-S<b>3375</b>, and processing in Steps S<b>3385</b>-S<b>3390</b> are respectively the same as the processing in Steps S<b>2500</b>-S<b>2510</b>, the processing in Steps S<b>2520</b>-S<b>2560</b>, the processing in Steps S<b>2565</b>-S<b>2575</b>, and the processing in Steps S<b>2585</b>-S<b>2590</b> in Embodiment 4. Accordingly, the processing in the figure is regarded as having been already explained.
0450After the end of the processing in Step S<b>3310</b>, the control unit <b>3010</b> refers to the available service DB <b>3100</b> stored in the available service DB storage unit <b>3020</b> (Step S<b>3312</b>) to determine whether or not another voice dialogue agent is appropriate for performing processing corresponding to the input text data (Step S<b>3315</b>). In other words, in the case where the input text data includes a character string indicated by the keyword <b>3110</b> and an additional voice dialogue agent specified by the target agent <b>3120</b> associated with the keyword <b>3110</b> is not the voice dialogue agent <b>3000</b> which is currently performing the third instruction execution processing, the control unit <b>3010</b> determines that the other voice dialogue agent (another additional voice dialogue agent specified by the target agent <b>3120</b>) is appropriate for performing the processing. Otherwise, the control unit <b>3010</b> determines that the other voice dialogue agent is not appropriate for performing the processing.
0451In the processing in Step S<b>3315</b>, in the case where the control unit <b>3010</b> determines that the other voice dialogue agent is not appropriate for performing the processing (Step S<b>3315</b>: No), the voice dialogue agent <b>3000</b> proceeds to the processing in Step S<b>3320</b> to perform the processing in Steps S<b>3320</b>-S<b>3360</b>.
0452In the processing in Step S<b>3315</b>, in the case where the control unit <b>3010</b> determines that the other voice dialogue agent is appropriate for performing the processing (Step S<b>3315</b>: Yes), the voice dialogue agent <b>3000</b> proceeds to the processing in Step S<b>3365</b> to perform the processing in Steps S<b>3365</b>-S<b>3375</b>.
0453In the processing in Step S<b>3375</b>, when the communication unit <b>420</b> receives the connection response returned from the additional voice dialogue agent (Step S<b>3375</b>: Yes), the voice dialogue agent <b>3000</b> performs second connection response processing (Step S<b>3380</b>).
0454<figref idref="DRAWINGS">FIG. 34</figref> is a flow chart of the second connection response processing.
0455When the second connection response processing is started, the control unit <b>3010</b> controls the communication unit <b>420</b> to transfer the input voice data, which is received in the processing in Step S<b>3200</b>: Yes, to the additional voice dialogue agent, which is specified by the processing in Step S<b>3315</b>: Yes (Step S<b>3400</b>).
0456After the end of the processing in Step S<b>3400</b>, the voice dialogue agent <b>3000</b> ends the second connection response processing.
0457Referring back to <figref idref="DRAWINGS">FIG. 33</figref>, the explanation on the second instruction execution processing is continued.
0458After the end of the second connection response processing, the voice dialogue agent <b>3000</b> proceeds to Step S<b>3385</b> to perform the processing in Steps S<b>3385</b>-S<b>3390</b>.
0459After the end of the processing in Step S<b>3390</b>, or after the end of the processing in Step S<b>3360</b>, the voice dialogue agent <b>3000</b> ends the third instruction execution processing.
0460Referring back to <figref idref="DRAWINGS">FIG. 32</figref>, the explanation on the fourth agent processing is continued.
0461After the end of the third instruction execution processing, the voice dialogue agent <b>3000</b> returns to the processing in Step S<b>3200</b> to perform the processing in Step S<b>3200</b> and the subsequent steps.
0462The following explains a specific example of the operation performed by the fourth modified voice dialogue system having the above configuration, with reference to the drawing.
0463<Specific Example>
0464<figref idref="DRAWINGS">FIG. 35</figref> is a procedure diagram schematically showing a situation in which the user of the fourth modified voice dialogue system starts, with use of the device <b>1700</b>, a voice dialogue with the home agent, which is one of the voice dialogue agents <b>3000</b>, and then starts communication with the car agent in response to a connection instruction generated by the home agent, and performs a dialogue with the car agent.
0465Here, the explanation is given based on the assumption that a specific voice dialogue agent server for the device <b>1700</b> used by the user is the voice dialogue agent server <b>110</b> that embodies the home agent, and a dialogue end voice is a voice “This ends dialogue.”.
0466In the figure, processing in Steps S<b>3500</b>-S<b>3505</b> is respectively the same as the processing in Steps S<b>2900</b>-S<b>2905</b> in the specific example in Embodiment 4 (see <figref idref="DRAWINGS">FIG. 29</figref>). Accordingly, the processing in the figure is regarded as having been already explained.
0467After the end of the processing in Step S<b>3505</b>, the device <b>1700</b> performs second voice input processing (Step S<b>3506</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0468In the second voice input processing, in the case where the user inputs a voice “What is temperature in car?”, the device <b>1700</b> transmits input voice data “What is temperature in car?” to the home agent (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0469Then, the home agent receives the input voice data (corresponding to Step S<b>3200</b>: Yes in <figref idref="DRAWINGS">FIG. 32</figref>), and performs third instruction execution processing (corresponding to Step S<b>3210</b> in <figref idref="DRAWINGS">FIG. 32</figref>).
0470In the third instruction execution processing, since the input text includes keywords “temperature” and “in-car” and an additional voice dialogue agent specified by the target agent <b>3120</b> is not the home agent (corresponding to Step S<b>3315</b>: No in <figref idref="DRAWINGS">FIG. 33</figref>), the home agent transmits a connection instruction to the car agent (corresponding to Step S<b>3370</b> in <figref idref="DRAWINGS">FIG. 33</figref>).
0471Then, the car agent receives the connection instruction (corresponding to Step S<b>2811</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>), and starts communication with the device <b>1700</b> (corresponding to Step S<b>2812</b> in <figref idref="DRAWINGS">FIG. 28</figref>), and transmits a connection response to the home agent (Step S<b>3590</b>, corresponding to Step S<b>2813</b> in <figref idref="DRAWINGS">FIG. 28</figref>).
0472The home agent receives the connection response (corresponding to Step S<b>3375</b>: Yes in <figref idref="DRAWINGS">FIG. 33</figref>), and performs second connection response processing (corresponding to Step S<b>3380</b> in <figref idref="DRAWINGS">FIG. 33</figref>).
0473In the second connection response processing, the home agent transmits input voice data “What is temperature in car?” to the car agent (corresponding to Step S<b>3400</b> in <figref idref="DRAWINGS">FIG. 34</figref>).
0474Then, the car agent receives the input voice data (corresponding to Step S<b>2814</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>). Since the input voice data does not request to terminate the communication (corresponding to Step S<b>2815</b>: No in <figref idref="DRAWINGS">FIG. 28</figref>), the car agent generates an instruction set corresponding to the input voice data, and executes the generated instruction set (Step S<b>3594</b>, corresponding to Step S<b>2830</b> in <figref idref="DRAWINGS">FIG. 28</figref>).
0475Here, in execution of the instruction set, in the case where the car agent generates response voice data “Temperature in car is 38 degrees C. Do you need any other help?”, the car agent transmits the response voice data “Temperature in car is 38 degrees C. Do you need any other help?” to the device <b>1700</b> (corresponding to Step S<b>2860</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>).
0476Then, the device <b>1700</b> receives the response voice data (corresponding to Step S<b>1960</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), and outputs a voice “Temperature in car is 38 degrees C. Do you need any other help?” (Step S<b>3507</b>, corresponding to Step S<b>1980</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0477Since the voice “Temperature in car is 38 degrees C. Do you need any other help?” is not a dialogue end voice (corresponding to Step S<b>1985</b>: No in <figref idref="DRAWINGS">FIG. 19</figref>), the device <b>1700</b> performs second voice input processing (Step S<b>3510</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0478In the second voice input processing, in the case where the user inputs a voice “Turn on air conditioner with 25 degrees C. of temperature setting.”, the device <b>1700</b> transmits input voice data “Turn on air conditioner with 25 degrees C. of temperature setting.” to the car agent (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0479Then, the car agent receives the input voice data (corresponding to Step S<b>2814</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>). Since the input voice data does not request to terminate the communication (corresponding to Step S<b>2815</b>: No in <figref idref="DRAWINGS">FIG. 28</figref>), the car agent generates an instruction set corresponding to the input voice data, and executes the generated instruction set (Step S<b>3594</b>, corresponding to Step S<b>2830</b> in <figref idref="DRAWINGS">FIG. 28</figref>).
0480Here, in execution of the instruction set, in the case where the car agent generates response voice data “Air conditioner is turned on with 25 degrees C. of temperature setting. Do you need any other help?”, the car agent transmits the response voice data “Air conditioner is turned on with 25 degrees C. of temperature setting. Do you need any other help?” to the device <b>1700</b> (corresponding to Step S<b>2860</b> in <figref idref="DRAWINGS">FIG. 28</figref>).
0481Then, the device <b>1700</b> receives the response voice data (corresponding to Step S<b>1960</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), and outputs a voice “Air conditioner is turned on with 25 degrees C. of temperature setting. Do you need any other help?” (Step S<b>3525</b>, corresponding to Step S<b>1980</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0482Since the voice “Air conditioner is turned on with 25 degrees C. of temperature setting. Do you need any other help?” is not a dialogue end voice (corresponding to Step S<b>1985</b>: No in <figref idref="DRAWINGS">FIG. 19</figref>), the device <b>1700</b> performs second voice input processing (Step S<b>3530</b>, corresponding to Step S<b>1930</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0483In the second voice input processing, in the case where the user inputs a voice “No. Thank you.”, the device <b>1700</b> transmits input voice data “No. Thank you.” to the car agent (corresponding to Step S<b>2040</b> in <figref idref="DRAWINGS">FIG. 20</figref>).
0484Then, the car agent receives the input voice data (corresponding to Step S<b>2814</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>). Since the input voice data requests to terminate the communication (corresponding to Step S<b>2815</b>: Yes in <figref idref="DRAWINGS">FIG. 28</figref>), the car agent terminates the communication with the device <b>1700</b> (corresponding to Step S<b>2870</b> in <figref idref="DRAWINGS">FIG. 28</figref>), and transmits a disconnection response to the home agent (Step S<b>3598</b>, corresponding to Step S<b>2890</b> in <figref idref="DRAWINGS">FIG. 28</figref>).
0485Then, the home agent receives the disconnection response (corresponding to Step S<b>2585</b>: Yes in <figref idref="DRAWINGS">FIG. 25</figref>), and performs disconnection response processing (Step S<b>2970</b>, corresponding to Step S<b>2890</b> in <figref idref="DRAWINGS">FIG. 25</figref>).
0486Here, in the disconnection processing, in the case where the voice dialogue agent <b>2200</b> generates response voice data “This ends dialogue.”, the voice dialogue agent <b>2200</b> transmits the response voice data “This ends dialogue.” to the device <b>1700</b> (corresponding to Step S<b>2720</b> in <figref idref="DRAWINGS">FIG. 27</figref>).
0487Then, the device <b>1700</b> receives the response voice data (corresponding to Step S<b>1960</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), and outputs a voice “This ends dialogue.” (Step S<b>3555</b>, corresponding to Step S<b>1980</b> in <figref idref="DRAWINGS">FIG. 19</figref>).
0488Since the voice “This ends dialogue.” is a dialogue end voice (corresponding to Step S<b>1985</b>: Yes in <figref idref="DRAWINGS">FIG. 19</figref>), the state is switched to the voice input receivable state (corresponding to Step S<b>1990</b> in <figref idref="DRAWINGS">FIG. 19</figref>). The device <b>1700</b> ends the fourth device processing.
0489<Consideration>
0490According to the fourth modified voice dialogue system having the above configuration, in the case where the voice dialogue agent A determines that the voice dialogue agent B rather than the voice dialogue agent A is appropriate as a communication party of the user while the user of the fourth modified voice dialogue system communicates with the voice dialogue agent A, it is possible to change a voice dialogue agent as the communication party of the user from the voice dialogue agent A to the voice dialogue agent B.
0491With this configuration, even if the user does not know the type of service provided by each of the voice dialogue agents, the user can receive a service provided by a more appropriate voice dialogue agent.
0492Also, in this case, since the voice dialogue agent A transfers input voice data that is not modified to the voice dialogue agent B, the voice dialogue agent B performs voice recognition processing on the input voice data. As a result, the user can receive a more appropriate service from the voice dialogue agent B.
Embodiment 6
0493The following exemplifies an operation situation of the voice dialogue system <b>100</b> in Embodiment 1. Note that the voice dialogue system <b>100</b> in Embodiment 1 may be of course operated in an operation situation other than the operation situation exemplified here.
0494<figref idref="DRAWINGS">FIG. 36A</figref> is a diagram schematically showing an operation situation in which the voice dialogue system <b>100</b> in Embodiment 1 is operated.
0495In <figref idref="DRAWINGS">FIG. 36A</figref>, a group <b>3600</b> is for example a company, an organization, or a family, and its size is not limited. A plurality of devices <b>3601</b> (devices A and B and so on) and a home gateway <b>3602</b> are disposed in the group <b>3600</b>. The devices <b>3601</b> include not only devices that are connectable to the Internet (for example, a smartphone, a PC, and a TV) but also devices that are disconnectable from the Internet by themselves (for example, an illumination lamp, a washing machine, a refrigerator). The devices <b>3601</b> may include devices that are disconnectable from the Internet by themselves but are connectable to the Internet via the home gateway <b>3602</b>. Also, the group <b>3600</b> includes a user <b>10</b> who uses the devices <b>3601</b>. For example, the devices which are disposed in the group <b>3600</b> each correspond to the device <b>140</b> in Embodiment 1.
0496A cloud server <b>3611</b> is disposed in a data center administration company <b>3610</b>. The cloud server <b>3611</b> is a virtual server that cooperates with various devices through the Internet. The cloud server <b>3611</b> mainly manages big data that is difficult to deal with by a normal data base management tool or the like. The data center administration company <b>3610</b> performs management of data and the cloud server <b>3611</b>, and administers a data center for performing such management. Services performed by the data center administration company <b>3610</b> are described in detail later. Here, the data center administration company <b>3610</b> is not limited to a company only performing data management, administration of the cloud server <b>3611</b>, and so on. For example, a device manufacturer developing and manufacturing one type of the devices <b>3601</b> may serve as the data administration center <b>3610</b> when the device manufacturer also performs data management and administration of the cloud server <b>3611</b> (see <figref idref="DRAWINGS">FIG. 36B</figref>). Also, the data center administration company <b>3610</b> does not need to be a single company. For example, when a device manufacturer and another management company perform data management and administration of the cloud server <b>3611</b> together, then either one or both of the device manufacturer and the management company may serve as the data center administration company <b>3610</b> (see <figref idref="DRAWINGS">FIG. 36C</figref>). For example, the data center administration company <b>3610</b> provides the voice dialogue agent <b>400</b> that is associated with the device <b>140</b> (hereinafter, referred to also as a first voice dialogue agent).
0497A service provider <b>3620</b> has a server <b>3621</b>. The server <b>3621</b> here for example includes a memory embedded in a PC for individual use, and its size is not limited. Also, there is a case where the service provider <b>3620</b> does not have the server <b>3621</b>. For example, the service provider <b>3620</b> provides another voice dialogue agent <b>400</b> that is connected to the first voice dialogue agent (hereinafter, referred to also as a second voice dialogue agent).
0498Next, an explanation is given on a flow of information in the above operation situation.
0499First, the device A or B, which is disposed in the group <b>3600</b>, transmits log information to the cloud server <b>3611</b>, which is disposed in the data center administration company <b>3610</b>. The cloud server <b>3611</b> accumulates the log information transmitted from the device A or B (arrow (a) in <figref idref="DRAWINGS">FIG. 36A</figref>). Here, the log information is information indicating a driving situation, an operation time and date, and so on of the devices <b>3601</b>. The log information includes for example a viewing history of a TV, timer recording information of a recorder, a driving time and date and a laundry amount of a washing machine, and a time and date and the number of opening and closing a refrigerator. Without limiting to the information described above, the log information includes all information that is acquirable from all the devices <b>3601</b>. There is a case where the log information is provided directly from the devices <b>3601</b> to the cloud server <b>3611</b> through the Internet. Alternatively, the log information may be provided from the home gateway <b>3602</b> to the cloud server <b>3611</b> after being accumulated from the devices <b>3601</b> to the home gateway <b>3602</b>.
0500Next, the cloud server <b>3611</b>, which is disposed in the data center administration company <b>3610</b>, provides the accumulated log information to the service provider <b>3620</b> in certain units. Here, the log information may be provided in units according to which the data center administration company <b>3610</b> can organize the accumulated log information and provide the organized log information to the service provider <b>3620</b>. Alternatively, the log information may be provided in units requested by the service provider <b>3620</b>. Moreover, the log information may not be provided in certain units, and alternatively an amount of the log information to be provided sometimes varies in accordance with circumstances. The log information is stored as necessary in the server <b>3621</b> of the service provider <b>3620</b> (arrow (b) in <figref idref="DRAWINGS">FIG. 36A</figref>). Then, the service provider <b>3620</b> organizes the log information so as to be adapted to a service to be provided to a user, and provides the organized information to the user. The user to which the organized information to be is provided may be the user <b>10</b> who uses the devices <b>3601</b> or an external user <b>20</b>. The service may be provided for example from the service provider <b>3620</b> directly to the user (arrow (e) in <figref idref="DRAWINGS">FIG. 36A</figref>). Alternatively, the service may be provided for example to the user again via the cloud server <b>3611</b> of the data center administration company <b>3610</b> (arrows (c) and (d) in <figref idref="DRAWINGS">FIG. 36A</figref>). Moreover, the cloud server <b>3611</b> of the data center administration company <b>3610</b> may organize the log information so as to be adapted to a service to be provided to the user, and provide the organized information to the service provider <b>3620</b>.
0501Note that the user <b>10</b> and the user <b>20</b> may be different or the same.
0502The following exemplifies several types of service that can be provided in the above operation situation.
0503<Service Type 1: Local Data Center Type>
0504<figref idref="DRAWINGS">FIG. 37</figref> is a diagram schematically showing service type 1 (local data center type service).
0505Here, the service provider <b>3620</b> acquires information from the group <b>3600</b>, and provides a service to a user. In this type of service, the service provider <b>3620</b> has functions of a data center administration company. That is, the service provider <b>3620</b> includes a cloud server <b>3611</b> performing big data management. As such, there is no data center administration company.
0506In this type of service, the service provider <b>3620</b> administers and manages the data center (the cloud server <b>3611</b>) (<b>3703</b>). Also, the service provider <b>3620</b> manages an OS (<b>3702</b>) and an application (<b>3701</b>). The service provider <b>3620</b> performs service provision (<b>3704</b>) with use of the OS (<b>3702</b>) and application (<b>3701</b>), which are managed by thereby.
0507<Service Type 2: IaaS Type>
0508<figref idref="DRAWINGS">FIG. 38</figref> is a diagram schematically showing service type 2 (IaaS (Infrastructure as a Service) type). Here, IaaS is a model in which infrastructure for constructing and operating a computer system is provided as a cloud service through the Internet.
0509In this type of service, the data center administration company <b>3610</b> administers and manages the data center (the cloud server <b>3611</b>) (<b>3703</b>). Further, the service provider <b>3620</b> manages the OS (<b>3702</b>) and the application (<b>3701</b>). The service provider <b>3620</b> performs service provision (<b>3704</b>) with use of the OS (<b>3702</b>) and the application (<b>3701</b>), which are managed thereby.
0510<Service Type 3: PaaS Type>
0511<figref idref="DRAWINGS">FIG. 39</figref> is a diagram schematically showing service type 3 (PaaS (Platform as a Service) type). Here, PaaS is a model in which a platform for constructing and operating software is provided as a service through the Internet.
0512In this type of service, the data center administration company <b>3610</b> manages the OS (<b>3702</b>), and administers and manages the data center (the cloud server <b>3611</b>) (<b>3703</b>). Further, the service provider <b>3620</b> manages the application (<b>3701</b>). The service provider <b>3620</b> performs service provision (<b>3704</b>) with use of the OS (<b>3702</b>), which is managed by the data center administration company <b>3610</b>, and the application (<b>3701</b>), which is managed by the service provider <b>3620</b>.
0513<Service Type 4: SaaS Type>
0514<figref idref="DRAWINGS">FIG. 40</figref> is a diagram schematically showing service type 4 (SaaS (Software as a Service) type). In this model, for example, an application that is provided by a platform provider having a data center (a cloud server) is provided to a business or a person (a user) without having a data center (a cloud server) as a cloud service through a network such as the Internet.
0515In this type of service, the data center administration company <b>3610</b> manages the application (<b>3701</b>), manages the OS (<b>3702</b>), and administers and manages the data center (the cloud server <b>3611</b>) (<b>3703</b>). Further, the service provider <b>3620</b> performs service provision (<b>3704</b>) with use of the application (<b>3701</b>) and the OS (<b>3702</b>), which are managed by the data center administration company <b>3610</b>.
0516The main actor in service provision is the service provider <b>3620</b> in all of the above service types. Further, for example, the service provider <b>3620</b> or the data center administration company <b>3610</b> may develop their own OS, application, or big data database, or may outsource any of these to a third party.
0000<Supplement>
0517One aspect of the voice dialogue method relating to the present invention and one aspect of the device relating to the present invention have been explained by exemplifying the five voice dialogue systems in Embodiments 1 to 5 and the operation situation of the voice dialogue system in Embodiment 6. However, the voice dialogue method and the device relating to the present invention are not of course limited to the voice dialogue method and the device as used in the voice dialogue system and the operation situation which are exemplified in Embodiments 1 to 6.
0518(1) In Embodiment 1, the voice dialogue system <b>100</b> has been explained to include the voice dialogue agent server <b>110</b>, the network <b>120</b>, the gateway <b>130</b>, and the device <b>140</b> as shown in <figref idref="DRAWINGS">FIG. 1</figref>. A voice dialogue system as another example may include a mediation server <b>4150</b> in addition to the voice dialogue agent server <b>110</b>, the network <b>120</b>, the gateway <b>130</b>, and the device <b>140</b>. The mediation server <b>4150</b> has a function of storing therein the target agent DB <b>2300</b>, associating between the voice dialogue agents, switching a connection destination, and so on.
0519<figref idref="DRAWINGS">FIG. 41</figref> is a system configuration diagram showing configuration of a voice dialogue system <b>4100</b> that includes the mediation server <b>4150</b>.
0520<figref idref="DRAWINGS">FIG. 42</figref> is a block diagram showing functional configuration of the mediation server <b>4150</b>.
0521As shown in the figure, the mediation server <b>4150</b> includes a communication unit <b>4220</b>, a control unit <b>4210</b>, and a target agent DB storage unit <b>4230</b>.
0522Here, the target agent DB storage unit <b>4230</b> has a function of storing therein the target agent DB <b>2300</b>, similarly to the target agent DB storage unit <b>2220</b> in Embodiment 4.
0523Also, a voice dialogue system as further another example may include a mediation server <b>4350</b> instead of the mediation server <b>4150</b>. The mediation server <b>4350</b> has a function of storing therein the available service DB <b>3100</b>, associating between the voice dialogue agents, switching a connection destination, and so on.
0524<figref idref="DRAWINGS">FIG. 43</figref> is a block diagram showing functional configuration of the mediation server <b>4350</b>.
0525As shown in the figure, the mediation server <b>4350</b> includes a communication unit <b>4320</b>, a control unit <b>4310</b>, and an available service DB storage unit <b>4330</b>.
0526Here, the available service DB storage unit <b>4330</b> has a function of storing therein the available service DB <b>3100</b>, similarly to the available service DB storage unit <b>3020</b> in Embodiment 5.
0527(2) In Embodiment 1, the image shown in <figref idref="DRAWINGS">FIG. 12</figref> is exemplified as an image displayed on the display unit <b>270</b> included in the device <b>140</b>.
0528Another examples of this image are shown in <figref idref="DRAWINGS">FIG. 44A</figref> to <figref idref="DRAWINGS">FIG. 44D</figref>, <figref idref="DRAWINGS">FIG. 45A</figref>, and <figref idref="DRAWINGS">FIG. 45B</figref>.
0529In the examples in <figref idref="DRAWINGS">FIG. 12</figref>, <figref idref="DRAWINGS">FIG. 44A</figref> to <figref idref="DRAWINGS">FIG. 44D</figref>, and <figref idref="DRAWINGS">FIG. 45B</figref>, displayed response texts each include, at the beginning thereof, a character string specifying a subject outputting a voice such as “You”, “Car agent”, “Home agent”, or the like. Also, in the example in <figref idref="DRAWINGS">FIG. 45A</figref>, an icon (image) specifying a subject outputting a voice is displayed.
0530In the examples in <figref idref="DRAWINGS">FIG. 44A</figref> and <figref idref="DRAWINGS">FIG. 44B</figref>, a character string specifying a voice dialogue agent with which the user currently makes a dialogue is displayed on an upper part of the screen such that the user recognizes the voice dialogue agent with which the user currently makes a dialogue. Such a character strings displayed here are “Dialogue with home agent” and “Dialogue with car agent”.
0531In the example in <figref idref="DRAWINGS">FIG. 44D</figref>, a character string specifying a voice dialogue agent with which the user currently makes a dialogue (or has made a dialogue in the past) is included in each of the displayed response texts, such that the user recognizes the voice dialogue agent with which the user currently makes a dialogue (or has made a dialogue in the past). Such a character strings displayed here are “Dialogue party is home agent” and “Dialogue party is car agent”. Also, in the example in <figref idref="DRAWINGS">FIG. 45B</figref>, an icon (image) specifying a voice dialogue agent with which the user currently makes a dialogue (or has made a dialogue in the past) is displayed.
0532These display examples are just examples. Alternatively, a voice dialogue agent with which the user currently makes a dialogue may be indicated by color, shape of the screen, shape of part of the screen, or the like. Furthermore, each subject outputting a voice may be indicated by changing a background color, a wall paper, and the like on the display. In this way, it is only necessary to display a voice dialogue agent with which the user makes a dialogue or a subject outputting a voice so as to be recognizable by the user.
0533(3) In Embodiment 1 and the modifications, the example has been explained that a voice dialogue agent with which the user makes a dialogue or a subject outputting a voice is displayed so as to be visually recognizable by the user. However, the present invention is not necessarily limited to the example where a voice dialogue agent with which the user makes a dialogue or a subject outputting a voice is displayed so as to be visually recognizable by the user, as long as the voice dialogue agent with which the user makes a dialogue or the subject outputting a voice is recognizable by the user.
0534For example, a voice “Dialogue party is home agent” may be output, such that a voice dialogue agent with which the user makes a dialogue is recognizable by the user. Alternatively, a sound effect may be output, such that the voice dialogue agent with which the user makes a dialogue is recognizable by the user. Further alternatively, the voice dialogue agent with which the user makes a dialogue may be indicated by changing voice tone, speech rate, voice volume, or the like.
0535(4) In Embodiment 1, the explanation has been provided that the state is managed by the control unit <b>210</b> in the form as shown in the switching of the state shown in <figref idref="DRAWINGS">FIG. 3</figref>. Also, in Embodiment 2, the explanation has been provided that the state is managed by the control unit <b>1310</b> in the form as shown in the switching of the state shown in <figref idref="DRAWINGS">FIG. 14</figref>. Furthermore, in Embodiment 3, the explanation has been provided that the state is managed by the control unit <b>1710</b> in the form as shown in the switching of the state shown in <figref idref="DRAWINGS">FIG. 18</figref>.
0536Management of the state performed by the control unit is not limited to be in the above forms. Alternatively, other forms for managing the state may be employed. <figref idref="DRAWINGS">FIG. 46</figref> to <figref idref="DRAWINGS">FIG. 50</figref> each show an example of switching of the state managed by the control unit in other forms.
0537For example, according to management of the state in a form shown in switching of the state in <figref idref="DRAWINGS">FIG. 48</figref>, in the case where a voice output by the voice output unit <b>260</b> based on a response text transmitted from the voice dialogue agent <b>110</b> is a dialogue end voice, the state is switched to the voice input unreceivable state even if the predetermined period has not lapsed after the switching of the state to the voice input receivable state. Accordingly, in the case where a voice signal, which is generated by the voice dialogue agent <b>110</b> and received by the communication unit <b>250</b>, indicates unnecessity of a new voice input, the voice input unit <b>220</b> is switched to the voice input unreceivable state even if the predetermined period has not lapsed after the switching to the voice input receivable state.
0538(5) In Embodiment 1, the display unit <b>270</b> has been explained, for example, as being embodied by a touchpanel, a touchpanel controller, and a processor that executes programs, and having the configuration of displaying that the display unit <b>270</b> is in the voice input receivable state by blinking the region <b>1120</b> that is positioned at the lower right in the display unit <b>270</b> (see <figref idref="DRAWINGS">FIG. 11A</figref>, <figref idref="DRAWINGS">FIG. 11C</figref>, <figref idref="DRAWINGS">FIG. 12</figref>, and so on). However, the configuration of the display unit <b>270</b> is not limited to the above configuration example as long as the user can recognize that the display unit <b>270</b> is in the voice input receivable state. Another configuration example may be employed in which the display unit <b>270</b> is embodied by an LED (Light Emitting Diode) and a processor that executes programs, and displays that the display unit <b>270</b> is in the voice input receivable state by lighting the LED. In the other configuration example, the display unit <b>270</b> does not display a response text received by the communication unit <b>250</b> because of not including means for displaying character strings.
0539(6) In Embodiment 1, the explanation has been provided that the communication unit <b>250</b> has the configuration in which in the case where a specific one of the voice dialogue agent servers <b>110</b> is not designated as a voice dialogue agent server <b>110</b> that is a communication party, the communication unit <b>250</b> communicates with a specific voice dialogue agent server with reference to an IP address stored in the address storage unit <b>240</b>. Alternatively, another configuration example may be employed in which the address storage unit <b>240</b> does not store therein the IP address of the specific voice dialogue agent server, and the communication unit <b>250</b> communicates with a voice dialogue agent server designated by the user or a voice dialogue agent server that embodies the voice dialogue agent designated by the user.
0540(7) In Embodiment 1, the devices <b>140</b> each have been explained as communicating with the voice dialogue agent <b>110</b> via the gateway <b>130</b> and the network <b>120</b>.
0541Alternatively, another configuration may be employed in which the devices <b>140</b> may each have a function of directly connecting with the network <b>120</b> without the gateway <b>130</b> and communicate with the voice dialogue agent without the gateway <b>130</b>. In the case where all the devices <b>140</b> are directly connected to the network <b>120</b> without the gateway <b>130</b>, the gateway <b>130</b> is not necessary.
0542(8) Part or all of the elements constituting the above embodiments and modifications may be configured from a single system LSI. The system LSI is a super multifunctional LSI that is manufactured by integrating a plurality of components on a single chip. Specifically, the system LSI is a computer system composed of a microprocessor, a ROM, a RAM, and so on. Functions of the system LSI are achieved by the microprocessor operating in accordance with a computer program that is stored in the ROM, the RAM, or the like.
0543(9) Part or all of the elements constituting the above embodiments and modifications may be composed of an IC (Integrated Circuit) card detachable from a device or a module. The IC card or the module is a computer system composed of a microprocessor, a ROM, a RAM, and so on. The IC card or the module may include the above super multifunctional LSI. Functions of the IC card or the module are achieved by the microprocessor operating in accordance with a computer program that is stored in the ROM, the RAM, or the like. The IC card or the module may be each tamper-resistant.
0544(10) The computer program or the digital signal which is used in the above embodiments and modifications may be recorded in a computer-readable recording medium such as a flexible disk, a hard disk, a CD-ROM, an MD, a DVD, a DVD-ROM, a DVD-RAM, a BD, a semiconductor memory, or the like.
0545Also, the computer program or the digital signal which is used in the above embodiments and modifications may be transmitted through an electric communication network, a wireless or wired communication network, a network such as the Internet, data broadcasting, or the like.
0546The computer program or the digital signal which is used in the above embodiments and modifications can be implemented in another computer system, by transmitting the computer program or the digital signal which is recorded in the recording medium to the other computer system, or by transmitting the computer program or the digital signal to the other computer system via the network.
0547(12) The above embodiments and modifications may be combined with each other.
0548(13) The following further explains configurations, modifications, and effects of the voice dialogue method and the device relating to one aspect of the present invention.
0549(a) One aspect of the present invention provides a voice dialogue method that is performed by a voice dialogue system, the voice dialogue system including: a voice signal generation unit; a voice dialogue agent unit; a voice output unit; and a voice input control unit, the voice dialogue method comprising: a step of, by the voice signal generation unit, receiving a voice input and generating a voice signal based on the received voice input; a step of, by the voice dialogue agent unit, performing voice recognition processing on the generated voice signal and performing processing based on a result of the voice recognition processing to generate a response signal; a step of, by the voice output unit, outputting a voice based on the generated response signal; and a step of, when the voice output unit outputs the voice, by the voice input control unit, keeping the voice signal generation unit in a receivable state for a predetermined period after output of the voice, the receivable state being a state in which a voice input is receivable.
0550According to the voice dialogue method relating to one aspect of the present invention, in the case where a voice generated by the voice dialogue agent unit is output, a user can input a voice without performing an operation with respect to the voice dialogue system. This reduces the number of times that the user needs to perform an operation in accordance with a voice that is dialogically input, compared with conventional techniques.
0551(b) Also, the voice dialogue system may further include a display unit, and the voice dialogue method may further comprise a step of, while the voice signal generation unit is in the receivable state, by the display unit, displaying that the voice signal generation unit is in the receivable state.
0552This configuration allows the user to visually recognize whether or not the voice signal generation unit is in the receivable state.
0553(c) Also, the voice dialogue system may further include an additional voice dialogue agent unit, and the voice dialogue method may further comprise: a step of, by the voice dialogue agent unit, determining, based on the result of the voice recognition processing, which one of the voice dialogue agent unit and the additional voice dialogue agent unit is appropriate for performing the processing based on the result of the voice recognition processing; a step of, when the voice dialogue agent unit determines that the voice dialogue agent unit is appropriate for performing the processing based on the result of the voice recognition processing, by the voice dialogue agent unit, performing the processing based on the result of the voice recognition processing; a step of, when the voice dialogue agent unit determines that the additional voice dialogue agent unit is appropriate for performing the processing based on the result of the voice recognition processing, by the additional voice dialogue agent unit, performing voice recognition processing on a voice received by the voice signal generation unit, performing processing based on a result of the voice recognition processing performed by the additional voice dialogue agent unit to generate a response signal; and a step of, by the voice output unit, outputting a voice based on the response signal generated by the additional voice dialogue agent unit.
0554According to this configuration, it is possible to cause the additional voice dialogue agent unit to perform processing that is appropriate for being performed by the additional voice dialogue agent unit rather than the voice dialogue agent unit.
0555(d) Also, the voice dialogue method may further comprise: a step of, when the voice dialogue agent unit determines that the voice dialogue agent unit is appropriate for performing the processing based on the result of the voice recognition processing, by the display unit, displaying that the voice dialogue agent unit is appropriate for performing the processing based on the result of the voice recognition processing; and a step of, when the voice dialogue agent unit determines that the additional voice dialogue agent unit is appropriate for performing the processing based on the result of the voice recognition processing, by the display unit, displaying that the additional voice dialogue agent unit is appropriate for performing the processing based on the result of the voice recognition processing.
0556This configuration allows the user to visually recognize which one of the voice dialogue agent unit and the additional voice dialogue agent unit is appropriate for performing the processing.
0557(e) Also, the voice dialogue method may further comprise a step of, when the voice dialogue agent unit determines that the additional voice dialogue agent unit is appropriate for performing the processing based on the result of the voice recognition processing, by the voice dialogue agent unit, transferring a voice signal generated by the voice signal generation unit to the additional voice dialogue agent unit, and by the additional voice dialogue agent unit, performing voice recognition processing on the transferred voice signal.
0558This configuration allows the additional voice dialogue agent unit to perform the voice recognition processing with use of the voice signal transferred from the voice dialogue agent unit.
0559(f) Also, the voice dialogue method may further comprise a step of, when the voice signal generation unit is in the receivable state and a response signal generated by the voice dialogue agent unit indicates that a new voice input does not need to be received, by the voice input control unit, switching the voice signal generation unit to an unreceivable state even during the predetermined period, the unreceivable state being a state in which a voice input is unreceivable.
0560According to this configuration, in the case where a voice input does not need to be received, it is possible to switch the voice signal generation unit to the unreceivable state even during the predetermined period.
0561(g) One aspect of the present invention provides a device comprising: a voice signal generation unit configured to receive a voice input and generate a voice signal based on the received voice input; a transmission unit configured to transmit the generated voice signal to an external server: a reception unit configured to receive a response signal that is returned from the server, the response signal being generated by the server based on the voice signal; a voice output unit configured to output a voice based on the received response signal; and a voice input control unit configured to, when the voice output unit outputs a voice, keep the voice signal generation unit in a receivable state for a predetermined period after output of the voice, the receivable state being a state in which a voice input is receivable.
0562According to the device relating to the one aspect of the present invention, in the case where a voice generated by the server is output, the user can input a voice without performing an operation with respect to the device. This reduces the number of times that the user needs to perform an operation in accordance with a voice that is dialogically input, compared with a conventional technique.
INDUSTRIAL APPLICABILITY
0563The voice dialogue method and the device relating to the present invention are widely utilizable for a voice dialogue system that performs processing based on a voice that is dialogically input by a user.
REFERENCE SIGNS LIST
0564<b>100</b> voice dialogue system
0565<b>110</b> voice dialogue agent server
0566<b>120</b> network
0567<b>130</b> gateway
0568<b>140</b> device
0569<b>210</b> control unit
0570<b>220</b> voice input unit
0571<b>230</b> operation reception unit
0572<b>240</b> address storage unit
0573<b>250</b> communication unit
0574<b>260</b> voice output unit
0575<b>270</b> display unit
0576<b>280</b> execution unit
0577<b>400</b> voice dialogue agent
0578<b>410</b> control unit
0579<b>420</b> communication unit
0580<b>430</b> voice recognition processing unit
0581<b>440</b> dialogue DB storage unit
0582<b>450</b> voice synthesizing processing unit
0583<b>460</b> instruction generation unit
Contents8
45 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| CN101558443A | Cites | China | Applicant |
| CN101689366A | Cites | China | Applicant |
| US10181099B2 | Cites | United States of America | Search report |
| US10546441B2 | Cites | United States of America | Search report |
| US10798244B2 | Cites | United States of America | Search report |
| EP1154406A1 | Cites | European Patent Office (EPO) | Search report |
| EP1591979A1 | Cites | European Patent Office (EPO) | Search report |
| JP2001056225A | Cites | Japan | Applicant |
| US2002046023A1 | Cites | United States of America | Applicant |
| JP2002116797A | Cites | Japan | Applicant |
| US2003163309A1 | Cites | United States of America | Applicant |
| JP2003241797A | Cites | Japan | Applicant |
| US2004044516A1 | Cites | United States of America | Applicant |
| JP2004233794A | Cites | Japan | Applicant |
| JP2004240150A | Cites | Japan | Applicant |
| US2005033582A1 | Cites | United States of America | Search report |
| US2005256712A1 | Cites | United States of America | Search report |
| JP2005266192A | Cites | Japan | Applicant |
| JP2006178175A | Cites | Japan | Applicant |
| US2007265831A1 | Cites | United States of America | Search report |
| JP2008090545A | Cites | Japan | Applicant |
| US2008208584A1 | Cites | United States of America | Search report |
| WO2009145796A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2009150156A1 | Cites | United States of America | Applicant |
| US2010076751A1 | Cites | United States of America | Applicant |
| US2011172994A1 | Cites | United States of America | Search report |
| US2011208525A1 | Cites | United States of America | Applicant |
| JP2011232619A | Cites | Japan | Applicant |
| US2012221502A1 | Cites | United States of America | Search report |
| JP2013114020A | Cites | Japan | Applicant |
| US2014095173A1 | Cites | United States of America | Search report |
| US2014304365A1 | Cites | United States of America | Search report |
| US2020105082A1 | Cites | United States of America | Search report |
| US2020193748A1 | Cites | United States of America | Search report |
| US6229880B1 | Cites | United States of America | Search report |
| US6249720B1 | Cites | United States of America | Applicant |
| US6480599B1 | Cites | United States of America | Search report |
| US6636831B1 | Cites | United States of America | Search report |
| US7003079B1 | Cites | United States of America | Search report |
| US7039166B1 | Cites | United States of America | Search report |
| US7117051B2 | Cites | United States of America | Search report |
| US7177402B2 | Cites | United States of America | Search report |
| US7228275B1 | Cites | United States of America | Search report |
| US7460652B2 | Cites | United States of America | Search report |
| US7903792B2 | Cites | United States of America | Search report |
| US8150020B1 | Cites | United States of America | Search report |
| US9263058B2 | Cites | United States of America | Search report |
| US9558745B2 | Cites | United States of America | Search report |
| US9564129B2 | Cites | United States of America | Search report |
| JPH1137766A | Cites | Japan | Applicant |
| US20020046023A1 | Cites | United States of America | Applicant |
| US20030163309A1 | Cites | United States of America | Applicant |
| US20040044516A1 | Cites | United States of America | Applicant |
| US20050033582A1 | Cites | United States of America | Search report |
| US20050256712A1 | Cites | United States of America | Search report |
| US20070265831A1 | Cites | United States of America | Search report |
| US20080208584A1 | Cites | United States of America | Search report |
| US20090150156A1 | Cites | United States of America | Applicant |
| US20100076751A1 | Cites | United States of America | Applicant |
| US20110172994A1 | Cites | United States of America | Search report |
| US20110208525A1 | Cites | United States of America | Applicant |
| US20120221502A1 | Cites | United States of America | Search report |
| US20140095173A1 | Cites | United States of America | Search report |
| US20140304365A1 | Cites | United States of America | Search report |
| US20200105082A1 | Cites | United States of America | Search report |
| US20200193748A1 | Cites | United States of America | Search report |
| CN101558443 | Cites | China | Applicant |
| CN101689366 | Cites | China | Applicant |
| JP1137766 | Cites | Japan | Applicant |
| JP200156225 | Cites | Japan | Applicant |
| JP2002116797 | Cites | Japan | Applicant |
| JP2003241797 | Cites | Japan | Applicant |
| JP2004233794 | Cites | Japan | Applicant |
| JP2004240150 | Cites | Japan | Applicant |
| JP2005266192 | Cites | Japan | Applicant |
| JP2006178175 | Cites | Japan | Applicant |
| JP200890545 | Cites | Japan | Applicant |
| JP2011232619 | Cites | Japan | Applicant |
| JP2013114020 | Cites | Japan | Applicant |
| WO2009145796 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| International Search Report dated Sep. 9, 2014 in corresponding International Application No. PCT/JP2014/003097 (with English translation). | Non-patent | – | Applicant |
| Extended European Search Report dated Jun. 1, 2016 in European Application No. 14814417.3. | Non-patent | – | Applicant |
| Office Action issued Jun. 1, 2018 in Chinese Application No. 201480021678.6 (with English translation of Search Report). | Non-patent | – | Applicant |
| Lin et al., (1999). “A distributed architecture for cooperative spoken dialogue agents with coherent dialogue state and history.” In: Proc. Workshop Automatic Speech Recognition and Understanding. | Non-patent | – | Search report |
| Lin et al., (1999). “A distributed architecture for cooperative spoken dialogue agents with coherent dialogue state and history.” In: Proc. Workshop Automatic Speech Recognition and Understanding. | Non-patent | – | Search report |
| International Search Report dated Sep. 9, 2014 in corresponding International Application No. PCT/JP2014/003097 (with English translation). | Non-patent | – | Applicant |
| Extended European Search Report dated Jun. 1, 2016 in European Application No. 14814417.3. | Non-patent | – | Applicant |
| Office Action issued Jun. 1, 2018 in Chinese Application No. 201480021678.6 (with English translation of Search Report). | Non-patent | – | Applicant |
16 members in 5 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 201361836763 | United States of America | P | |
| 201361836763 | United States of America | P | |
| 2014003097 | Japan | W | |
| 2014003097 | Japan | W | |
| 201414777920 | United States of America | A | |
| 201414777920 | United States of America | A | |
| 201416268938 | United States of America | A | |
| 14777920 | – | – | – |
| 61836763 | – | – | – |
| PCTJP2014003097 | – | – | – |
| US201361836763P | – | – | – |
| US201414777920 | – | – | – |
| US201416268938 | – | – | – |
| WO2014JP03097 | – | – | – |
Members16
| Document | Office | Kind | |
|---|---|---|---|
| WO2014203495A1 | World Intellectual Property Organization (WIPO) | A1 | |
| CN105144285A | China | A | |
| EP3012833A1 | European Patent Office (EPO) | A1 | |
| EP3012833A4 | European Patent Office (EPO) | A4 | |
| US2016322048A1 | United States of America | A1 | |
| US9564129B2 | United States of America | B2 | |
| JPWO2014203495A1 | Japan | A1 | |
| JP6389171B2 | Japan | B2 | |
| CN105144285B | China | B | |
| CN108806690A | China | A | |
| JP2018189984A | Japan | A | |
| JP6736617B2 | Japan | B2 | |
| JP2020173477A | Japan | A | |
| USRE49014EThis record | United States of America | E | |
| JP7072610B2 | Japan | B2 | |
| EP3012833B1 | European Patent Office (EPO) | B1 |
91 transactions on the USPTO file
Allowed after 2 non-final rejections, 2 final rejections and 2 RCEs.
- Non-final rejections
- 2
- Final rejections
- 2
- RCEs
- 2
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Payment of Maintenance Fee, 8th Year, Large EntityM1552 | M1552 | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Email NotificationEML_NTR | EML_NTR | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Email NotificationEML_NTR | EML_NTR | |
| Mailing Corrected Notice of AllowabilityMCNOA | MCNOA | |
| Examiner's Amendment CommunicationEX.A | EX.A | |
| Reasons for AllowanceEX.R | EX.R | |
| Corrected Notice of AllowabilityCNOA | CNOA | |
| Pubs Case Remand to TCPUBTC | PUBTC | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Reasons for AllowanceEX.R | EX.R | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Advisory Action (PTOL - 303)MCTAV | MCTAV | |
| Advisory Action (PTOL-303)CTAV | CTAV | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| PILOT- Request for After Final Consideration ProgramRAFC | RAFC | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Interview Summary - Applicant Initiated - TelephonicEXAT | EXAT | |
| Interview Summary RecordEXIN | EXIN | |
| Electronic request for Examiner InterviewM865E | M865E | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Advisory Action (PTOL - 303)MCTAV | MCTAV | |
| Advisory Action (PTOL-303)CTAV | CTAV | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| PILOT- Request for After Final Consideration ProgramRAFC | RAFC | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Mail Interview Summary - Applicant Initiated - TelephonicMEXAT | MEXAT | |
| Interview Summary - Applicant Initiated - TelephonicEXAT | EXAT | |
| Electronic request for Examiner InterviewM865E | M865E | |
| Mail Post CardPST_CRD | PST_CRD | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Notice of Reissue Published in Official GazetteNRE. | NRE. | |
| Paralegal Reissue Review CompletePRIR | PRIR | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Email NotificationEML_NTR | EML_NTR | |
| Application ready for PDX access by participating foreign officesCCRDY | CCRDY | |
| Application Is Now CompleteCOMP | COMP | |
| Filing ReceiptFLRCPT.O | FLRCPT.O | |
| Application Dispatched from OIPEOIPE | OIPE | |
| FITF set to YES - revise initial settingFTFS | FTFS | |
| Cleared by OIPE CSRL194 | L194 | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| PTO/SB/69-Authorize EPO Access to Search ResultsSREXR141 | SREXR141 | |
| Applicants have given acceptable permission for participating foreignAPPERMS | APPERMS | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Entity Status Set To Undiscounted (Initial Default Setting or Status Change)BIG. | BIG. | |
| Initial Exam Team nnIEXX | IEXX |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Maintenance fee paymentMAFP | MAFP | |
| Fee payment procedureENTITY STATUS SET TO UNDISCOUNTED (ORIGINAL EVENT CODE: BIG.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP |
Numbers
- Publication
- RE049014
- Publication, DOCDB
- RE49014
- Publication, EPODOC
- USRE49014E
- Application
- 16268938
- Application, DOCDB
- 201416268938
- Application, EPODOC
- US201416268938
Titles
- English
- Voice interaction method, and device
Classification
- CPC, 6
- G10L15/22
- G10L15/222
- G06F3/167
- G10L15/32
- G10L15/08
- G10L2015/088
- IPC, 6
- G10L15 00
- G10L21 00
- G10L15 22
- G06F3 16
- G10L15 08
- G10L15 32