Nova Patents
CA2483296C

Variable rate vocoder

Abstract

An apparatus and method for performing speech signal compression, by variable rate coding of frames of digitized speech samples (10). The level of speech activity for each frame of digitized speech samples is determined and an output data packet rate is selected from a set of rates based upon the determined level of frame speech activity. A lowest rate of the set of rates corresponds to a detected minimum level of speech activity, such as background noise or pauses in speech, while a highest rate corresponds to a detested maximum level of speech activity, such as alive vocalization. Each frame is then coded according to a predetermined coding format for the selected rate wherein each rate has a corresponding number of bits representative of the coded frame. A data packet is provided for each coded frame with each output data packet of a bit rate corresponding to the selected rate.

CA2483296C, drawing sheet 1
Sheet 1 of 24

Term

Term ended

Expired 3 June 2012, 14.3 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

37 claims: 22 independent, 15 dependent

  1. 1
    CA 02483296 2004-11-05 , -74769-12D CLAIMS:1. A method of speech signal compression, by variable rate coding of flames of digitized speech samples, comprising the steps of: 5 determining a level of speech activity for a frame of digitized speech samples;selecting an encoding rate from a set of rates based upon said determined level of speech activity for said frame;coding said frame according to a coding format of a set of coding 10 formats for said selected rate wherein each rate has a corresponding different coding format and wherein each coding format provides for a different plurality of parameter signals representing said digitized speech samples (s(n)) in accordance with a speech model;and generating for said frame a data packet of said parameter signals, 15 characterised by: providing a rate command indicative of a preselected encoding rate for said frame;and modifying said selected encoding rate to provide said preselected encoding rate for coding of said frame at said preselected encoding rate.
  2. 4
    The method of any one of claims 1 to 4 wherein said step of providing said data packet of said parameter signals comprises:generating a variable number of bits to represent linear predictive 10 coefficient (LPC) vector signals of said frame of digitized speech samples, wherein said variable number of bits representing said LPC vector signals is responsive to said measured speech activity level;generating a variable number of bits to represent pitch vector signals of said frame of digitized speech samples, wherein said variable number of 15 bits representing said pitch vector signals is responsive to said measured speech activity level;and generating variable number of bits to represent codebook excitation vector signals of said frame of digitized speech samples, wherein said variable number of bits representing said codebook excitation vector 20 signals is responsive to said measured speech activity level.
  3. 5
    The method of any one of claims 1 to 5 wherein said step of coding said frame comprises:generating for said frame a variable number of linear prediction 25 coefficients wherein said variable number of said linear prediction coefficients is responsive to said selected encoding rate;generating for said frame a variable number of pitch coefficients wherein said variable number of said pitch coefficients is responsive to said selected encoding rate;and CA 02483296 2004-11-05 + *74769-12D generating for said frame a variable number of codebook excitation values wherein said variable number of said codebook excitation values is responsive to said selected encoding rate. 5 ,
  4. 6
    The method of any one of claims 1 to 5 wherein said step of determining a level of speech activity comprises summing the squares of the values of said digitized speech samples.
  5. 12
    The method of any one of claims 1 to 11 further comprising the step of premultiplying said digitized speech samples (s(n)) by a predetermined windowing function.
  6. 13
    The method of any one of claims 1 to 12 further comprising the step of 10 converting said LPC coefficients to line spectral pair (LSP) values.
  7. 14
    The method of any one of claims 1 to 13 wherein said input frame of digitized samples comprises digitized values for approximately twenty milliseconds of speech.
  8. 15
    The method of any one of claims 1 to 14 wherein said input frame of digitized samples comprises approximately 160 digitized samples.
  9. 16
    The method of any one of claims 1 to 15 wherein said output data packet 20 comprises:one hundred and seventy one bits comprising forty bits for LPC data, forty bits for pitch data, eighty bits for excitation vector data and eleven bits for error protection when said output data rate is full rate;eighty bits comprising twenty bits for LPC information, twenty bits for 25 pitch information and forty bits for excitation vector data when said output data rate is half rate;forty bits comprising ten bits for LPC information, ten bits for pitch information and twenty bits for excitation vector data when said output data rate is quarter rate;and 30 sixteen bits comprising ten bits for LPC information and six bits for excitation vector information when said output data rate is eighth rate. CA 02483296 2004-11-05 74769-12D K
  10. 17
    An apparatus for compressing an acoustical signal into variable rate data comprising:means (52) for determining a level of speech activity for an input 5 frame (10) of digitized samples of said acoustical signal;means (90,294,296) for selecting an output data rate from a predetermined set of rates based upon said determined level of speech activity within said frame;means (58,104,106,108) for coding said frame according to a coding 10 format of a set of coding formats for said selected rate to provide a plurality of parameter signals wherein each rate has a corresponding different coding format with each coding format providing a different plurality of parameter signals representing said digitized speech samples (s(n)) in accordance with a speech model;and 15 means (114) for providing for said frame a corresponding data packet (p(n)) at a data rate corresponding to said selected rate, characterised by: means for providing a rate command indicative of a preselected encoding rate for said frame;and 20 means for modifying said selected encoding rate to provide said preselected encoding rate for coding of said frame at said preselected encoding rate.
  11. 20
    The apparatus of claim 20 further comprising means for adaptively adjusting said at least one of said at least one speech activity thresholds. 20
  12. 21
    The apparatus of any of claims 17 to wherein said means for determining said energy of said input frame comprises:squaring means for squaring said digitized audio samples of a frame;and summing means for summing said squares of digitized audio samples 25 of a frame.
  13. 22
    The apparatus of any of claims 17,18 or 19 wherein said means for determining a level of speech activity comprises:means (50) for calculating a set of linear predictive coefficients for 1 30 said input frame of digitized samples of said acoustical signals;and CA 02483296 2004-11-05 74769-12D means for determining said level of speech activity in accordance with at least one of said linear predictive coefficients.
  14. 23
    The apparatus of any of claims 17 to 22 further comprising means (236, 5 238) for providing error protection bits for said data packet responsive to said selected output data rate.
  15. 24
    The apparatus of claim 24 wherein said means (236,238) for providing error protection bits provides the values of said error protection bits in 10 accordance with a cyclic block code.
  16. 25
    The apparatus of any of claims 17 to 24 further comprising means (208) for converting said LPC coefficients to line spectral pair (LSP) values. 15
  17. 26
    The apparatus of any of claims 17 to 25 wherein said set of rates comprises full rate, half rate, quarter rate and eighth rate:
  18. 27
    The apparatus of any of claims 17 to 26 wherein said set of rates comprises 16 Kbps, 8 Kbps, 4 Kbps and 2 Kbps.
  19. 29
    The apparatus of any of claims 17 to 28 wherein said input frame of digitized speech samples comprises digitized speech for a duration of approximately twenty milliseconds. CA 02483296 2004-11-05 ,, -74769-12D
  20. 30
    The apparatus of any of claims 17 to 29 wherein said input frame of digitized samples comprises 160 digitized samples.
  21. 32
    The apparatus of any of claims 17 to 31 to further comprising means (52, 200) for pre-multiplying said digitized samples by a predetermined windowing function.
  22. 34
    The apparatus of any of claims 17 to 33 wherein said output data packet 15 (p(n)) comprises:a variable number of bits to represent LPC vector signals of said frame of digitized speech samples (s(n)), wherein said variable number of bits for representing said LPC vector signals is responsive to said level of speech activity;20 a variable number of bits to represent pitch vector signals of said frame of digitized speech samples (s(n)), wherein said variable number of bits for representing said pitch vector signals is responsive to said level of speech activity;and a variable number of bits to represent codebook excitation vector 25 signals of said frame of digitized speech samples (s(n)), wherein said variable number of bits for representing said codebook excitation vector signals is responsive to said level of speech activity.
  23. 36
    The apparatus of any of claims 17 to 35 wherein said output data packet 5 comprises:one hundred and seventy one bits comprising forty bits for LPC data, forty bits for pitch data, eighty bits for excitation vector data and eleven bits for error protection when said output data rate is full rate;eighty bits comprising twenty bits for LPC information, twenty bits for 10 pitch information and forty bits for excitation vector data when said output data rate is half rate;forty bits comprising ten bits for LPC information, ten bits for pitch information and twenty bits for excitation vector data when said output data rate is quarter rate;and 15 sixteen bits comprising ten bits for LPÇ information and six bits for excitation vector information when said output data rate is eighth rate.
  24. 37
    The apparatus of any of claims 17 to 36 wherein said means (90,294,296) for selecting an encoding rate is responsive to an external rate signal. Smart & Biggar Ottawa, Canada Patent Agents
Independent claims24