US9620145B2

Context-dependent state tying using a neural network

Summary by NHIP

Context-dependent state tying

The method trains a first neural network using context-dependent states derived from a second neural network's hidden layer activations. It then receives an audio signal, provides it to the first network, and generates a transcription based on the output.

Claim Score by NHIP

Read claim 1, the broadest

Abstract

The technology described herein can be embodied in a method that includes receiving an audio signal encoding a portion of an utterance, and providing, to a first neural network, data corresponding to the audio signal. The method also includes generating, by a processor, data representing a transcription for the utterance based on an output of the first neural network. The first neural network is trained using features of multiple context-dependent states, the context-dependent states being derived from a plurality of context-independent states provided by a second neural network.

US9620145B2, drawing sheet 1
Sheet 1 of 11

Term

7.9 yearsleft in the term

Expires 5 September 2034, including 108 days of term adjustment.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Expires

19 claims: 3 independent, 16 dependent

  1. 1
    Broadest claimClaim Score 39, average(NHIP)A method performed by one or more computers, the method comprising:accessing, by the one or more computers, a first neural network that has been trained using speech data examples that are each assigned a context-dependent state by clustering context independent states based on activations at a hidden layer of a second neural network that was trained to provide outputs corresponding to context-independent states, the first neural network being configured to provide outputs corresponding to one or more context-dependent states;receiving, by the one or more computers, an audio signal encoding a portion of an utterance;providing, by the one or more computers, data corresponding to the audio signal to the first neural network that has been trained using the speech data examples that are each assigned a context-dependent state based on the activations at the hidden layer of the second neural network;generating, by the one or more computers, data indicating a transcription for the utterance based on an output of the first neural network that was generated in response to the data corresponding to the audio signal;and providing, by the one or more computers, the data indicating the transcription as output of an automated speech recognition service.
  2. 11
    A system comprising a speech recognition engine comprising one or more processors and a machine-readable storage device storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:accessing, by the one or more computers, a first neural network that has been trained using speech examples that are each assigned a context-dependent state by clustering context independent states based on activations at a hidden layer of a second neural network that was trained to provide outputs corresponding to context-independent states, the first neural network being configured to provide outputs corresponding to one or more context-dependent states;receiving, by the one or more computers, an audio signal encoding a portion of an utterance;providing, by the one or more computers, data corresponding to the audio signal to the first neural network that has been trained using the speech examples that are each assigned a context-dependent state based on the activations at the hidden layer of the second neural network;generating, by the one or more computers, data indicating a transcription for the utterance based on an output of the first neural network that was generated in response to the data corresponding to the audio signal;and providing, by the one or more computers, the data indicating the transcription as output of an automated speech recognition service.
  3. 17
    A non-transitory computer-readable storage device encoding one or more computer-readable instructions, which upon execution by one or more processors cause operations comprising:accessing, by the one or more computers, a first neural network that has been trained using speech examples that are each assigned a context-dependent state by clustering context independent states based on activations at a hidden layer of a second neural network that was trained to provide outputs corresponding to context-independent states, the first neural network being configured to provide outputs corresponding to one or more context-dependent states;receiving, by the one or more computers, an audio signal encoding a portion of an utterance;providing, by the one or more computers, data corresponding to the audio signal to the first neural network that has been trained using the speech examples that are each assigned a context-dependent state based on the activations at the hidden layer of the second neural network;generating, by the one or more computers, data indicating a transcription for the utterance based on an output of the first neural network that was generated in response to the data corresponding to the audio signal;and providing, by the one or more computers, the data indicating the transcription as output of an automated speech recognition service.