EP1580730B1

Isolating speech signals utilizing neural networks

Abstract

This record has no abstract on file.

EP1580730B1, drawing sheet 1
Sheet 1 of 15

Term

Term ended

Expired 23 March 2025, 1.5 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

9 claims: 4 independent, 5 dependent

  1. 1
    A speech signal isolation system for extracting a speech signal from background noise in an audio signal comprising:a frequency transform component (502) for transforming said audio signal from a time-series signal to a frequency domain signal;a compression component (506) for generating a compressed audio signal having a reduced number of frequency subbands;a background noise estimation component (504) adapted to estimate background noise intensity of an audio signal across a plurality of frequencies;a neural network component (508) adapted to extract a speech estimate signal from the background noise;a blending component (510) for generating a reconstructed speech signal from the audio signal and the extracted speech based on the background noise intensity estimate;characterized by the neural network has a first set of input nodes (908) equal to the number of frequency subbands in the compressed audio signal for receiving said compressed audio signal, and a second set of input nodes (910) equal to the number of frequency subbands for receiving said background noise estimate.
  2. 2
    A speech signal isolation system for extracting a speech signal from background noise in an audio signal comprising:a frequency transform component (502) for transforming said audio signal from a time-series signal to a frequency domain signal;a compression component (506) for generating a compressed audio signal having a reduced number of frequency subbands;a background noise estimation component (504) adapted to estimate background noise intensity of an audio signal across a plurality of frequencies;a neural network component (508) adapted to extract a speech estimate signal from the background noise;a blending component (510) for generating a reconstructed speech signal from the audio signal and the extracted speech based on the background noise intensity estimate;characterized by the neural network has a first set of input node (1002) equal to the number of frequency subbands in the compressed audio signal for receiving said compressed audio signal and a second set of input nodes (1004, 1006) equal to the number of frequency subbands in the compressed audio signal for receiving the compressed audio signal from a previous time step, the output of the neural network from a previous time step or an intermediate result from a previous time step.
  3. 4
    A method of isolating a speech signal from an audio signal having a speech component and background noise, and the method comprising:transforming a time-series audio signal into the frequency domain;estimating the background noise in the audio signal across multiple frequency bands;and being characterized by : applying the background noise estimate and the audio signal to a neural network;extracting a speech signal estimate from the audio signal as an output of the neural network;and blending a portion of the speech signal estimate with a portion of the audio signal based on the background noise estimate to provide a reconstructed speech signal having reduced background noise.
  4. 9
    A method of isolating a speech signal from an audio signal having a speech component and background noise, and the method comprising:transforming a time-series audio signal into the frequency domain;estimating the background noise in the audio signal across multiple frequency bands;applying the audio signal to a neural network;and being characterized by : applying the speech signal estimate from a previous time step, an intermediate result of the speech signal estimate from a previous time step or the audio signal from a previous time step to the neural network;extracting a speech signal estimate from the audio signal as an output of the neural network;and blending a portion of the speech signal estimate with a portion of the audio signal based on the background noise estimate to provide a reconstructed speech signal having reduced background noise.