EP2017828A1

Techniques for disambiguating speech input using multimodal interfaces

Abstract

A system for disambiguating speech input. The system comprises a speech recognition component (110) for receiving recorded audio or speech input (104) and for generating one or more tokens corresponding to said input and a confidence value for each of said one or more tokens, the confidence value being indicative of a likelihood that said token correctly represents the respective input. The system also comprises a selection component (116) for identifying, according to a selection algorithm, which of any two or more tokens generated for said input are to be presented to a user (108) as alternatives (120); one or more disambiguation components (118,124) for performing an interaction with the user, in which the alternatives are presented to the user and the user's selection (122) is received; and an output interface (126) for presenting the user's selection as an input to an application (106). The system is characterised in that said interaction with the user uses a multimodal interface and said alternatives are presented to the user as a multimodal output and the user's selection is received as a multimodal input.

EP2017828A1, drawing sheet 1
Sheet 1 of 3

Term

Term ended

Projected expiry passed 10 December 2023, 2.8 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

16 claims: 10 independent, 6 dependent

  1. 1
    A system for disambiguating speech input, the system comprising:- a speech recognition component (110) for receiving recorded audio or speech input (104) and for generating one or more tokens corresponding to said input and a confidence value for each of said one or more tokens, the confidence value being indicative of a likelihood that said token correctly represents the respective input;a selection component (116) for identifying, according to a selection algorithm, which of any two or more tokens generated for said input are to be presented to a user (108) as alternatives (120);one or more disambiguation components (118,124) for performing an interaction with the user, in which the alternatives are presented to the user and the user's selection (122) is received;and an output interface (126) for presenting the user's selection as an input to an application (106);characterised in that said interaction with the user uses a multimodal interface and said alternatives are presented to the user as a multimodal output and the user's selection is received as a multimodal input.
  2. 5
    A system as claimed in any preceding claim further comprising means for receiving parameters (114) for controlling operation of the components of the system.
  3. 8
    A system as claimed in any one of claims 5 to 6 wherein the parameters include confidence thresholds for governing unambiguous recognition and inclusion of close matches.
  4. 9
    A system as claimed in any one of claims 5 to 8 wherein the selection component is configured to filter said one or more tokens according to said parameters.
  5. 10
    A system as claimed in any preceding claim, wherein the one or more disambiguation components are configured to disambiguate the alternatives in plural iterative stages, whereby the first stage narrows the alternatives to a number of alternatives that is smaller than that initially generated by the selection component, but greater than one, and whereby the one or more disambiguation components are operative iteratively to narrow the alternatives in subsequent iterative stages.
  6. 12
    A system as claimed in any preceding claim wherein the disambiguation components and the application reside on a single computing device.
  7. 13
    A system as claimed in any one of claims 1 to 11 wherein the disambiguation components and the application reside on separate computing devices.
  8. 14
    A method of processing speech input, the method comprising:receiving recorded audio or speech input (104) from a user (108);generating one or more tokens corresponding to said input;generating a confidence value for each of said one or more tokens, the confidence value being indicative of a likelihood that said token correctly represents the respective input;determining whether said input is ambiguous;if the speech input is not ambiguous, then communicating the token representative of said input to an application (106) as input to the application;and if the speech input is ambiguous: performing an interaction with the user whereby the user is presented with plural alternatives (120) and whereby the user's selection (122) from among the plural alternatives is received;and communicating the selected alternative to the application as input to the application;characterised in that the interaction with the user uses a multimodal interface and comprises presenting the alternatives to the user as a multimodal output and receiving the user's selection as a multimodal input.
  9. 15
    A system for disambiguating speech input comprising:a speech recognition component that receives recorded audio or speech input and generates: one or more tokens corresponding to the speech input;and for each of the one or more tokens, a confidence value indicative of the likelihood that the a given token correctly represents the speech input;a selection component that identifies, according to a selection algorithm, which two or more tokens are to be presented to a user as alternatives;one or more disambiguation components that perform an interaction with the user to present the alternatives and to receive a selection of alternatives from the user, the interaction taking place in at least a visual mode;and an output interface that presents the selected alternative to an application as input.
  10. 16
    A method of processing speech input comprising:receiving a speech input from a user;determining whether the speech input is ambiguous;if the speech input is not ambiguous, then communicating a token representative of the speech input to an application as input to the application;and if the speech input is ambiguous: performing an interaction with the user whereby the user is presented with plural alternatives and selects an alternative from among the plural alternatives, the interaction being performed in at least a visual mode;communicating the selected alternative to the application as input to the application.