EP1544748A2

Method and apparatus for evaluating machine translation quality

Abstract

Quality of machine translation of natural language is determined by computing a sequence kernel that provides a measure of similarity between a first sequence of symbols representing a machine translation in a target natural language and a second sequence of symbols representing a reference translation in the target natural 1 anguage. T he measure of similarity takes into account the existence of non-contiguous subsequences shared by the first sequence of symbols and the second sequence of symbols. When the similarity measure does not meet an acceptable threshold level, the translation model of the machine translator may be adjusted to improve subsequent translations performed by the machine translator.

EP1544748A2, drawing sheet 1
Sheet 1 of 20

Term

Term ended

Projected expiry passed 8 December 2024, 1.8 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

10 claims: 4 independent, 6 dependent

  1. 1
    A method for computing machine translation performance, comprising:receiving a sequence of natural language data in a first language;translating the sequence of natural language data to a second language to define a machine translation of the sequence of natural language data;receiving a reference translation of the sequence of natural language data in the second language;computing a sequence kernel that provides a similarity measure between the machine translation and the reference translation;outputting a signal indicating the similarity measure;wherein the similarity measure accounts for non-contiguous occurrences of subsequences shared between the machine translation and the reference translation.
  2. 2
    The method according to claim I, wherein the similarity measure penalizes identified non-contiguous occurrences of subsequences shared between the machine translation and the reference translation.
  3. 4
    The method according to claim I, wherein the similarity measure scores symbols in matching subsequences in the machine translation and the reference translation with a different weight than symbols in gaps that do not match.
  4. 10
    The method according to claim I, wherein the reference translation of the sequence of natural language data in the first language is input by a user.