EP1580730A2

Isolating speech signals utilizing neural networks

Abstract

A speech signal isolation system configured to isolate and reconstruct a speech signal transmitted in an environment in which frequency components of the speech signal are masked by background noise. The speech signal isolation system obtains a noisy speech signal from an audio source. The noisy speech signal may then be fed through a neural network that has been trained to isolate and reconstruct a clean speech signal from against background noise. Once the noisy speech signal has been fed through the neural network, the speech signal isolation system generates an estimated speech signal with substantially reduced noise.

EP1580730A2, drawing sheet 1
Sheet 1 of 16

Term

Term ended

Projected expiry passed 23 March 2025, 1.5 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

26 claims: 4 independent, 22 dependent

  1. 1
    A speech signal isolation system for extracting a speech signal from background noise in an audio signal comprising:a background noise estimation component adapted to estimate background noise intensity of an audio signal across a plurality of frequencies;a neural network component adapted to extract a speech estimate signal from the background noise;and a blending component for generating a reconstructed speech signal from the audio signal and the extracted speech based on the background noise intensity estimate.
  2. 10
    A method of isolating a speech signal from an audio signal having a speech component and background noise, and the method comprising:transforming a time-series audio signal into the frequency domain;estimating the background noise in the audio signal across multiple frequency bands;extracting a speech signal estimate from the audio signal;blending a portion of the speech signal estimate with a portion of the audio signal based on the background noise estimate to provide a reconstructed speech signal having reduced background noise.
  3. 20
    A system for enhancing a speech signal comprising:an audio signal source providing an audio time-series signal having both speech content and background noise;a signal processor providing a frequency transform function for transforming the audio signal from the time-series domain to the frequency domain;a background noise estimator;a neural network;and a signal combiner said background noise estimator forming an estimate of the background noise in said audio signal, and said neural network extracting the speech signal estimate from said audio signal, and said signal combiner combining the speech signal estimate and the audio signal based on the background noise estimate to produce a reconstituted speech signal having substantially reduced background noise.
  4. 26
    A method of isolating a speech signal from background noise comprising:receiving an audio signal;identifying portions of the audio signal where accuracy of the signal s known with a high degree of certainty;and training a neural network to estimate a reconstructed signal having significantly reduced background noise for those portions of the audio signal where the accuracy of the audio signal is in doubt.