EP3557504A1

Intent identification for agent matching by assistant systems

Abstract

In one embodiment, a method includes receiving a user request associated with one or more domains from a client system associated with a first user, parsing the user request to identify one or more semantic-intents are associated with the one or more domains and one or more slots, identifying, based on a ranker model , one or more dialog-intents associated with the user request based on the one or more semantic-intents and slots and context information associated with the user request, wherein each dialog-intent is a sub-intent of one or more of the semantic-intents, determining one or more agents for executing one or more tasks associated with the one or more dialog-intents respectively, and sending instructions for presenting a communication content information returned from the one or more agents responsive to executing the one or more tasks responsive to the user input to the client system.

EP3557504A1, drawing sheet 1
Sheet 1 of 33

Term

Projected expiry 31 October 2038.

  1. Priority and filed
  2. Published
  3. Today
  4. Projected expiry

15 claims: 4 independent, 11 dependent

  1. 1
    A method, in particular for use in an assistant system for assisting a user to obtain information or services by enabling the user to interact with the assistant system with user input in conversations to get assistance, wherein the user input includes voice, text, image or video or any combination of them, the assistant system in particular being enabled by the combination of computing devices, application programming interfaces (APIs), and the proliferation of applications on user devices, the method comprising, by one or more computing systems:receiving, from a client system associated with a first user, a user request, wherein the user request is associated with one or more domains;parsing, by a natural-language understanding module, the user request to identify one or more semantic-intents and one or more slots, wherein the one or more semantic-intents are associated with the one or more domains;identifying, based on a ranker model, one or more dialog-intents associated with the user request based on (1) the one or more semantic-intents, (2) the one or more slots, and (3) context information associated with the user request, wherein each dialog-intent is a sub-intent of one or more of the semantic-intents;determining, by a dialog engine, one or more agents from a plurality of agents for executing one or more tasks associated with the one or more dialog-intents, respectively;and sending, to the client system associated with the first user, instructions for presenting a communication content responsive to the user request, wherein the communication content comprises information returned from the one or more agents responsive to executing the one or more tasks.
  2. 2
    The method of Claim 1, wherein the ranker model is a machine-learning model trained based on a plurality of training samples comprising one or more of:(1) a plurality of user requests, (2) a plurality of positive dialog-intents, or (3) a plurality of negative dialog-intents.
  3. 3
    The method of Claim 2, wherein the plurality of training samples are generated based on one or more dry runs, each dry run comprising:accessing a dry-run request, wherein the dry-run request is associated with the one or more domains;executing, via the plurality of agents, a plurality of tasks associated with the dry-run request;determining a plurality of dialog-intents based on information returned from each of the plurality of agents responsive to executing the plurality of tasks;selecting one or more dialog-intents from the determined plurality of dialog-intents;and annotating the selected one or more dialog-intents as positive dialog-intents and remaining non-selected dialog-intents as negative dialog-intents.
  4. 4
    The method of Claim 3, further comprising generating feature representations for the plurality of training samples based on one or more of:a plurality of capability values associated with the plurality of agents, each capability value indicating a confidence of a corresponding agent being able to execute a particular task;information returned from the plurality of agents responsive to executing the plurality of tasks;one or more dialog states associated with one or more dry-run requests corresponding to the one or more dry runs;a plurality of semantic-intents associated with the one or more dry-run requests corresponding to the one or more dry runs;one or more device contexts of one or more client systems associated with one or more users associated with the one or more dry runs;or user profile data associated with the one or more users.
  5. 5
    The method of Claim 4, wherein the ranker model is customized based on the user profile data;and/or wherein each dry-run request in each dry run is associated with a particular dialog session, and wherein each dialog state indicates one or more of a domain or a context associated with the particular dialog session at a time associated with the dialog session.
  6. 6
    The method of any of Claims 1 to 5, further comprising:sending the one or more dialog-intents and the one or more slots to the one or more agents for executing the one or more tasks;and receiving, from the one or more agents, information responsive to executing the one or more tasks;optionally, further comprising: generating feature representations for the information returned from the one or more agents responsive to executing the one or more tasks;and re-training the ranker model based on the generated feature representations.
  7. 7
    The method of any of Claims 1 to 6, wherein identifying the one or more dialog-intents associated with the user request comprises:determining a plurality of candidate dialog-intents based on the one or more domains associated with the one or more semantic-intents and the one or more slots;calculating, by the ranker model, a plurality of confidence scores for the plurality of candidate dialog-intents;and identifying one or more of the dialog-intents based on their respective confidence scores;and/or wherein each of the one or more dialog-intents is associated with one of the one or more domains.
  8. 8
    The method of any of Claims 1 to 7, wherein the communication content comprises one or more of:a character string;an audio clip;an image;or a video clip;and/or wherein the method further comprises determining one or more modalities for the communication content;optionally, wherein determining the one or more modalities for the communication content comprises: identifying contextual information associated with the first user;identifying contextual information associated with the client system;and determining the one or more modalities based on the contextual information associated with the first user and the contextual information associated with the client system.
  9. 9
    The method of any of Claims 1 to 8, wherein the one or more agents execute the one or more tasks in parallel;and/or wherein each of the one or more agents is implemented with an application programming interface (API), and wherein the API receives a command from the dialog engine for executing one of the one or more tasks;and/or wherein the one or more semantic-intents are associated with one or more confidence scores, respectively.
  10. 10
    One or more computer-readable non-transitory storage media embodying software that is operable when executed to perform a method according to any of Claims 1 to 9.
  11. 11
    An assistant system for assisting a user to obtain information or services by enabling the user to interact with the assistant system with user input in conversations to get assistance, wherein the user input includes voice, text, image or video or any combination of them, the assistant system in particular being enabled by the combination of computing devices, application programming interfaces (APIs), and the proliferation of applications on user devices, the system comprising:one or more processors;and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to perform a method according to any of Claims 1 to 9.
  12. 12
    The assistant system of Claim 11 for assisting the user by executing at least one or more of the following features or steps:- create and store a user profile comprising both personal and contextual information associated with the user - analyze the user input using natural-language understanding, wherein the analysis may be based on the user profile for more personalized and context-aware understanding - resolve entities associated with the user input based on the analysis - interact with different agents to obtain information or services that are associated with the resolved entities - generate a response for the user regarding the information or services by using natural-language generation - through the interaction with the user, use dialog management techniques to manage and forward the conversation flow with the user - assist the user to effectively and efficiently digest the obtained information by summarizing the information - assist the user to be more engaging with an online social network by providing tools that help the user interact with the online social network (e.g., creating posts, comments, messages) - assist the user to manage different tasks such as keeping track of events - proactively execute pre-authorized tasks that are relevant to user interests and preferences based on the user profile, at a time relevant for the user, without a user input - check privacy settings whenever it is necessary to guarantee that accessing user profile and executing different tasks are subject to the user's privacy settings.
  13. 13
    The assistant system of Claim 11 or 12, comprising at least one or more of the following components:- a messaging platform to receive a user input based on a text modality from the client system associated with the user and/or to receive user input based on an image or video modality and to process it using optical character recognition techniques within the messaging platform to convert the user input into text, - an audio speech recognition (ASR) module to receive a user input based on an audio modality (e.g., the user may speak to or send a video including speech) from the client system associated with the user and to convert the user input based on the audio modality into text, - an assistant xbot to receive the output of the messaging platform or the ASR module.
  14. 14
    A system comprising at least one client system (130), in particular an electronic device at least one assistant system (140) according to any of Claims 11 to 13, connected to each other, in particular by a network (110), wherein the client system comprises an assistant application (136) for allowing a user at the client system (130) to interact with the assistant system (140), wherein the assistant application (136) communicates user input to the assistant system (140) and, based on the user input, the assistant system (140) generates responses and sends the generated responses to the assistant application (136) and the assistant application (136) presents the responses to the user at the client system (130), wherein in particular the user input is audio or verbal and the response may be in text or also audio or verbal.
  15. 15
    The system of Claim 14, further comprising a social-networking system (160), wherein the client system comprises in particular a social-networking application (134) for accessing the social networking system (160).