EP1515310B1

A system and method for providing high-quality stretching and compression of a digital audio signal

Abstract

This record has no abstract on file.

EP1515310B1, drawing sheet 1
Sheet 1 of 20

Term

Term ended

Expired 22 July 2024, 2.2 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

35 claims: 16 independent, 19 dependent

  1. 1
    A system for temporal modification of segments of an audio signal including speech, comprising:a frame extraction module (205) for sequentially extracting data frames from a received audio signal;a segment type detection module (215) for determining a content type of each segment of a current frame of the sequentially extracted data frames, said content types including voiced segments, unvoiced segments, and mixed segments;and means (220-240) for temporally modifying (300-380, 400-460, 500-555, 600-670) at least one segment of the current frame by automatically selecting and applying a corresponding temporal modification process for the at least one segment of the current frame from among a voiced segment temporal modification process, an unvoiced temporal modification process, and a mixed segment temporal modification process, wherein applying the mixed segment temporal modification process comprises applying both the voiced segment temporal modification process and the unvoiced temporal modification process.
  2. 5
    The system of one of claims 2 to 4, wherein the classification is at least partially based on a periodicity of the segment.
  3. 6
    The system of one of claims 1 to 5, wherein the frames are processed sequentially.
  4. 7
    The system of one of claims 1 to 6, wherein the mixed segments include both voiced and unvoiced components.
  5. 8
    A method for temporal modification of segments of an audio signal including speech, comprising:sequentially extracting data frames from a received audio signal;determining a content type of each segment of a current frame of the sequentially extracted data frames, said content types including voiced segments, unvoiced segments, and mixed segments;and temporally modifying (300-380, 400-460, 500-555, 600-670) at least one segment of the current frame by automatically selecting and applying a corresponding temporal modification process for the at least one segment of the current frame from among a voiced segment temporal modification process, an unvoiced temporal modification process, and a mixed segment temporal modification process, wherein applying the mixed segment temporal modification process comprises applying both the voiced segment temporal modification process and the unvoiced temporal modification process.
  6. 11
    The method of one of claims 8 to 10, wherein the content type of at least one segment is a voiced segment, and wherein temporally modifying the at least one segment comprises stretching the voiced segment to increase a length of the current frame.
  7. 16
    The method of one of claims 12 to 15, further comprising alternating selection points for the template such that consecutive templates are identified at different positions within the current frame.
  8. 17
    The method one of claims 8 to 16, further comprising determining whether an average compression ratio of temporally modified segments corresponds to an overall target compression ratio, and wherein a next target compression ratio for at least one next current frame is automatically adjusted as needed for ensuring that the overall target compression ratio is approximately maintained.
  9. 18
    The method of one of claims 8 to 17, wherein the content type of at least one segment is an unvoiced segment, and wherein temporally modifying the at least one segment comprises automatically generating (410-440, 520-535, 625-645, 655) and inserting (450, 540, 650) at least one synthetic segment into the current frame to increase a length of the current frame.
  10. 20
    The method of one of claims 8 to 19, wherein the content type of at least one segment is a mixed segment, and wherein the mixed segment includes both voiced and unvoiced components.
  11. 22
    The method of one of claims 8 to 21, wherein the content type of at least one segment is a voiced segment, and wherein temporally modifying the at least one segment comprises compressing the voiced segment to decrease a length of the current frame.
  12. 24
    The method of one of claims 8 to 23, wherein the content type of at least one segment is an unvoiced segment, and wherein temporally modifying the at least one segment comprises compressing the unvoiced segment to decrease a length of the current frame.
  13. 26
    A computer-implemented process for providing dynamic temporal modification of segments of a digital audio signal, comprising using a computing device to:receive one or more sequential frames of a digital audio signal;decode each frame of the digital audio signal as it is received;determine a content type of segments of the decoded frames from a group of predefined segment content types including a voiced segment content type, an unvoiced segment content type and a mixed segment content type, each segment content type having an associated type-specific temporal modification process, the type-specific temporal modification processes comprising a voiced segment temporal modification process, an unvoiced segment temporal modification process, and a mixed segment temporal modification process;and modify (300-380, 400-460, 500-555, 600-670) a temporal scale of one or more of the segments using the associated type-specific temporal modification process specific to each segment content type, wherein using the mixed segment temporal modification process comprises using both the voiced segment temporal modification process and the unvoiced segment temporal modification process.
  14. 30
    The computer-implemented process of one of claims 26 to 29, wherein determining the content type of segments comprises computing a normalized cross correlation for sub-segments of each segment, and comparing a maximum peak of each normalized cross correlation to predetermined thresholds for determining the content type of each segment.
  15. 31
    The computer-implemented process of one of claims 26 to 30, wherein at least one segment is a voiced type segment, and wherein modifying the temporal scale of voiced type segments comprises stretching at least one voiced type segment by approximately one or more pitch periods to increase a length of the at least one voiced type segment.
  16. 33
    The computer-implemented process of one of claims 26 to 32, wherein at least one segment is an unvoiced type segment, and wherein modifying the temporal scale of unvoiced type segments comprises:automatically generating (410-440, 520-535, 625-645, 655) at least one synthetic segment from one or more sub-segments of the at least one unvoiced-type segment;and inserting (450, 540, 650) the at least one synthetic segment into the at least one unvoiced type segment to increase a length of the at least one unvoiced type segment.
Independent claims16