EP1361522A2

A system for automatically annotating training data for a natural language understanding system

Abstract

The present invention uses a natural language understanding system that is currently being trained to assist in annotating training data for training that natural language understanding system. Unannotated training data is provided to the system and the system proposes annotations to the training data. The user is offered an opportunity to confirm or correct the proposed annotations, and the system is trained with the corrected or verified annotations.

EP1361522A2, drawing sheet 1
Sheet 1 of 11

Term

Term ended

Projected expiry passed 23 April 2023, 3.4 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

28 claims: 5 independent, 23 dependent

  1. 1
    A method of generating annotated training data to train a natural language understanding (NLU) system having one or more models, comprising:generating a proposed annotation with the NLU system for each unit of unannotated training data;displaying the proposed annotations for user verification or correction to obtain a user-confirmed annotation;andtraining the NLU system with the user-confirmed annotation.
  2. 24
    A user interface for training a natural language understanding (NLU) system that has one or more models, the user interface comprising:a first portion displaying a model display representative of the one or more models;a second portion displaying unannotated training inputs;anda third portion displaying a proposed annotation for a selected one of the unannotated training inputs.
  3. 26
    A method of generating annotated training data for training a natural language understanding (NLU) system having at least one model, comprising:generating a proposed annotation for a unit of unannotated training data;calculating a confidence measure for a plurality of different portions of the proposed annotation;anddisplaying the proposed annotation by visually contrasting portions that have a corresponding confidence measure that falls below a threshold level.
  4. 27
    A method of generating annotated training data for training a natural language understanding (NLU) system, comprising:generating, with the NLU system, a proposed annotation for each of a plurality of units of unannotated training data, each proposed annotation having a type;displaying the proposed annotations in an order based on the type, with user actuable inputs for user correction or verification of the proposed annotation to obtain a user-confirmed annotation.
  5. 28
    A method of generating annotated training data for training a natural language understanding (NLU) system employing a plurality of different natural language understanding techniques, comprising:generating, with each natural language training technique, a proposed annotation for a unit of unannotated training data, to obtain a plurality of proposed annotations;selecting one of the plurality of proposed annotations;anddisplaying the selected proposed annotation, with user actuable inputs for user correction or verification of the proposed annotation to obtain a user-confirmed annotation.