Concept for encoding an audio signal and decoding an audio signal using deterministic and noise like information.
Abstract
An encoder for encoding an audio signal comprises: an analyzer (120; 320) configured for deriving prediction coefficients (122; 322) and a residual signal from an unvoiced frame of the audio signal (102); a gain parameter calculator (550; 550') configured for calculating a first gain parameter (gc) information for defining a first excitation signal (c(n)) related to a deterministic codebook and for calculating a second gain parameter (gn) information for defining a second excitation signal (n(n)) related to a noise-like signal for the unvoiced frame; and a bitstream former (690) configured for forming an output signal (692) based on an information (142) related to a voiced signal frame, the first gain parameter (gc) information and the second gain parameter (gn) information.

Term
8 yearsleft in the term
Expires 10 October 2034.
- Priority
- Filed
- Granted
- Today
- Expires
18 claims: 9 independent, 9 dependent
- 1REIVINDICACIONES ι»»πτυτο MEXICANO DI LA PKOTJ» industrial 1. Codificador para codificar una señal de audio, el codificador comprende:un analizador (120;320) configurado para derivar coeficientes de predicción (122;322) y una señal residual a partir de un cuadro no vocal de la señal de audio (102);una calculadora de parámetro de ganancia (550;550’) configurada para calcular una primera información de parámetro de ganancia (g c ) para definir una primera señal de excitación (c(n)) relativa a un libro de códigos determinista y para calcular una segunda información de parámetro de ganancia (g n ) para definir una segunda señal de excitación (n(n)) relativa a una señal con características de ruido para el cuadro no vocal;y un formador de corrientes de bits (690) configurado para formar una señal de salida (692) sobre la base de una información (142) relativa a un cuadro de señal de voz, la primera Información de parámetro de ganancia (g c ) y la segunda información de parámetro de ganancia (g n );un determinador configurado para determinar si la señal residual fue determinada a partir de un cuadro de señal de audio no vocal;en donde el codificador comprende una memoria (350n) de Predicción a Largo Plazo (LTP) y un generador de señal (850) para generar una señal de excitación adaptativa para el cuadro de voz;y en donde, cuando se compara con un esquema de codificación de la CELP, el codificador se configura para no transmitir parámetros LTP para el cuadro no vocal para ahorrar bits, en donde la señal de excitación adaptativa se establece en cero para el cuadro no vocal, y en donde el libro de códigos /1 fy nwrnvro msxicanc , ¡5 determinista está configurado para codificar más pulsos para l^tWS^^ÉSsaMé^íi^ bits utilizando los bits de ahorro.
- 2El codificador de acuerdo con la reivindicación 1, en donde la calculadora de parámetro de ganancia (550;550’) está configurada para calcular un primer parámetro de ganancia (g c ) y un segundo parámetro de ganancia (g n ) y donde el formador de corrientes de bits (690) está configurado para formar la señal de salida (692) sobre la base del primer parámetro de ganancia (g c ) y el segundo parámetro de ganancia (g n );o donde la calculadora de parámetro de ganancia (550;550’) comprende un cuantificador (170-1, 170-2) configurado para cuantificar el primer parámetro de ganancia (g c ) a fin de obtener un primer parámetro de ganancia cuantificado (g„) y para cuantificar el segundo parámetro de ganancia (g n ) a fin de obtener un segundo parámetro de ganancia cuantificado (g n ) y donde el formador de corrientes de bits (690) está configurado para formar la señal de salida (692) sobre la base del primer parámetro de ganancia cuantificado (g c ) y el segundo parámetro de ganancia cuantificado (g„).
- 3El codificador de acuerdo con la reivindicación 1 o 2, que además comprende una calculadora de Información formante (160) configurada para calcular una información de conformación espectral relacionada con la voz (162) a partir de los coeficientes de predicción (122;322) y donde la calculadora de parámetro de ganancia (550;550’) está configurada para calcular la primera información de parámetro de ganancia (g c ) y la segunda Información de parámetro de ganancia (g n ) sobre la base de la Información de conformación espectral relacionada con la voz (162). I INSTITUTO MEXICANO Í /J DB LA PROPIEDAD
- 4El codificador de acuerdo con una de las reMrahfcacr precedentes, en donde la calculadora de parámptrn Hp ganancia, (fifin’) comprende:un primer amplificador (550e) configurado para amplificar la primera señal de excitac¡ón(c(n)) aplicando el primer parámetro de ganancia g c para obtener una primera señal de excitación amplificada (550f);un segundo amplificador (350e;550g) configurado para amplificar la segunda señal de excitación (n(n)) diferente de la primera señal de excitación (c(n)) aplicando el segundo parámetro de ganancia(g n ) para obtener una segunda señal de excitación amplificada (350g;550h);un combinador (550i) configurado para combinar la primera señal de excitación amplificada (550f) y la segunda señal de excitación amplificada (350g;550h) a fin de obtener una señal de excitación combinada (550k;550k’);un controlador (550n) configurado para filtrar la señal de excitación combinada (550k;550k’) con un filtro de síntesis a fin de obtener una señal sintetizada (350Γ), para comparar la señal sintetizada (350Γ) y el cuadro de la señal de audio (102) a fin de obtener un resultado comparativo, con el objeto de adaptar el primer parámetro de ganancia (g c ) o el segundo parámetro de ganancia (g n ), sobre la base del resultado comparativo;y donde el formador de corrientes de bits (690) está configurado para formar la señal de salida (692) sobre la base de una información (g c ;g n ) relativa al primer parámetro de ganancia (g c ) y el segundo parámetro de ganancia (g n ).
- 5El codificador de acuerdo con una de las reivindicaciones precedentes, en donde el controlador de parámetros de ganancia (550;550’) además comprende por lo menos un modelador (350;550b) configurado para conformar espectralmente la primera señal de exc¡tacion T '^]^^^n derivada de la misma o la segunda señal de excitación (n(n)) o una señal derivada de la misma, sobre la base de una información de conformación espectral (162).
- 6El codificador de acuerdo con una de las reivindicaciones precedentes, en donde el codificador está configurado para codificar la señal de audio (102) cuadro por cuadro en una secuencia de cuadros y donde la calculadora de parámetro de ganancia (550;550’) está configurada para ío determinar el primer parámetro de ganancia (g c ) y el segundo parámetro de ganancia (g n ) para cada uno de una pluralidad de subcuadros de un cuadro procesado y donde el controlador de parámetros de ganancia (550;550’) está configurado para determinar un valor de energía promedio asociado al cuadro procesado.
- 7El codificador de acuerdo con una de las reivindicaciones precedentes, que además comprende:una calculadora de información formante (160) configurada para calcular por lo menos una primera información de conformación espectral relacionada 20 con la voz a partir de los coeficientes de predicción (122;322);un determinador (130) configurado para establecer si la señal residual fue determinada a partir de un cuadro de audio de señal no vocal.
- 8El codificador de acuerdo con una de las reivindicaciones 25 precedentes, en donde el controlador de parámetros de ganancia (550; 550’) comprende un controlador (550n) configurado para determinar el primer parámetro de ganancia (g c ) sobre la base de:: Σ^/ο 1 xw(n) cw(ri) 9C Σ^ζ -1 ην(η)-civ(X) donde cw(n) es una señal de excitación filtrada de un libro de códigos innovativo y xw(n) es una excitación diana perceptual computada en el codificador de CELP;donde el controlador (550n) está configurado para determinar la ganancia de ruido cuantificada (.§7) sobre la base del valor cuantificado del primer parámetro de ganancia (^) y la razón de energías cuadráticas entre la primera excitación y la segunda excitación: donde Lsf es el tamaño en muestras de un subcuadro.
- 9El codificador de acuerdo con una de las reivindicaciones precedentes, que además comprende un cuantificador (170-1, 170-2) configurado para cuantificar el primer parámetro de ganancia (g c ) a fin de obtener un primer parámetro de ganancia cuantificado (^), en donde el controlador de parámetros de ganancia (550n) está configurado para determinar el primer parámetro de ganancia (g c ) sobre la base de:Une ac· Lsf . iQfír3/20 donde gc es el primer parámetro de ganancia, Lsf es el tamaño del subcuadro en muestras, cw(n) denota la primera señal de excitación IMPI «ΠΤη/ΓΌ MOICANO OC u PftO/IIDAD INDUrrtiAL conformada, xw(n) denota una señal que codifica la Predicción Lineal Excitada por Código, donde el controlador de parámetros de ganancia (550n) o el cuantificador (170-1, 170-2) está configurado asimismo para normalizar el 5 primer parámetro de ganancia (g c ) a fin de obtener un primer parámetro de ganancia normalizado sobre la base de: 9nc 9c· . I0?ír3/2O donde g nc denota el primer parámetro de ganancia normalizado y úrg es io una medición para una energía promedio de la señal residual no vocal a través de todo el cuadro;y donde el cuantificador (170-1, 170-2) está configurado para cuantificar el primer parámetro de ganancia normalizado a fin de obtener el primer parámetro de ganancia cuantificado
- 10El codificador de acuerdo con la reivindicación 9, en donde el cuantificador (170-1, 170-2) está configurado para cuantificar el segundo parámetro de ganancia (g n ) a fin de obtener un segundo parámetro de ganancia cuantificado (g„) donde el controlador de parámetros de ganancia 20 (550; 550j está configurado para determinar el segundo parámetro de ganancia (g n ) mediante la determinación de un valor de error sobre la base de:Lsf-Í Lsf—Í xw 2 (n) - Σ (g c .cw(n) + g n nw(n)') 2 IMPIO INSTITUTO MEXICANO DE LA PROPIEDAD O*®, INDUSTRIAL donde es un factor de atenuación variable en un intervalo entre 0,5 y 1, Lsf corresponde al tamaño de un subcuadro de un cuadro de audio procesado, cw(n) denota la primera señal de excitación conformada (c(n)), xw(n) denota una señal que codifica la Predicción Lineal Excitada por Código, gn denota el segundo parámetro de ganancia y g¿ denota un primer parámetro de ganancia cuantificado;donde el controlador de parámetros de ganancia (550;550’) está configurado para determinar el error para el subcuadro actual y donde el cuantificador (170-1, 170-2) está configurado para determinar la segunda ganancia cuantificada (g„) que minimiza el error y para obtener la segunda ganancia cuantificada (g„) sobre la base de: g n = Q(indexn) g c Lsf-l n=0 n(n) n(ri) donde Q(¿ndex n ) denota un valor escalar de un conjunto finito de valores posibles. 15
- 11El codificador de acuerdo con la reivindicación 10, en donde el combinador (550¡) está configurado para combinar el primer parámetro de ganancia (g c ) y el segundo parámetro de ganancia (g n ) a fin de obtener una señal de excitación combinada (e(n)) sobre la base de:e(ri) = g c c(n) + g n n(n)
- 12El codificador de acuerdo con la reivindicación 10, en donde el cuantificador (170-2) está configurado para determinar el valor de error basado en la falta de adaptación de la energía entre la primera señal de excitación conformada (c(n)) relativa y la segunda señal de excitación, en donde el IMPI ifmmTo mexicano DE LA PROPIEDAD INDUSTRIAL cuantificador (170-1) está configurado para determinar el primer parámetro de ganancia (g n ) basado en un error cuadrático medio o un error de raíz cuáarátta promedio.
- 13El codificador de acuerdo con la cualquiera de las reivindicaciones precedentes, en donde el codificador está configurado para no transmitir parámetros de predicción a largo plazo para cuadros no vocales. Un decodificador (1000) para decodificar una señal de audio recibida (1002) que comprende una información relativa a coeficientes de predicción (122), el decodificador (1000) comprende:un primer generador de señales (1010) configurado para generar una primera señal de excitación (1012) a partir de un libro de códigos determinista para una porción de una señal sintetizada (1062);un segundo generador de señales (1020) configurado para generar una segunda señal de excitación (1022) a partir de una señal con características de ruido para la porción de la señal sintetizada (1062);un combinador (1050) configurado para combinar la primera señal de excitación (1012) y la segunda señal de excitación (1022) para generar una señal de excitación combinada (1052) para la porción de la señal sintetizada (1062);y un sintetizador (1060) configurado para sintetizar la porción de la señal sintetizada (1062) a partir de la señal de excitación combinada (1052) y los coeficientes de predicción (122). en donde el decodificador comprende una memoria (350n) de Predicción a Largo Plazo (LTP) y un generador de señal (850) para generar una señal de excitación adaptativa para el cuadro de voz;y INSTITUTO MEXICANO 'LA PROPIEDAD DELA PROPIEDAD I en donde la señal de audio recibida no comprende ρβΓέηίβίτο^ΊΤί el cuadro no vocal, en donde el decodificador está r.nnfigiiradn para establecer en cero la señal de excitación adaptativa para el cuadro no vocal, y en donde el libro de códigos determinista está configurado para proporcionar más pulsos para la misma tasa de bits debido a los bits de ahorro dado por la falta de parámetros LTP para el cuadro no vocal.
- 1415. El decodificador de acuerdo con la reivindicación 14, en donde la señal de audio recibida (1002) comprende una información relativa a un primer parámetro de ganancia (g c ) y a un segundo parámetro de ganancia (g n ), en donde el decodiflcador además comprende:un primer amplificador (254;350e;550e) configurado para amplificar la primera señal de excitación (1012) o una señal derivada de la misma aplicando el primer parámetro de ganancia (g c ) para obtener una primera señal de excitación amplificada (1012’);un segundo amplificador (254;350e;550e) configurado para amplificar la segunda señal de excitación (1022) o una señal derivada, aplicando el segundo parámetro de ganancia a fin de obtener una segunda señal de excitación amplificada (1022’).
- 1516. El decodlficador de acuerdo con la reivindicación 14 o 15, que además comprende:una calculadora de información formante (160;1090) configurada para calcular una primera información de conformación espectral (1092a) y una segunda información de conformación espectral (1092b) a partir de los coeficientes de predicción (122;322);IMPIég UWnTTUTO MEXICANO W LA PROPIEDAD un primer modelador (1070) para conformar espectralméntéW espetéffede la primera señal de excitación (1012) o una señaLjdeá)4ada-xi 4auaisoa^ usando la primera información de conformación espectral (1092a);y un segundo modelador (1080) para conformar espectralmente un espectro de la segunda señal de excitación (1022) o una señal derivada de la misma, usando la segunda información de conformación (1092b);
- 1617. Una señal de audio codificada (692; 1002) que comprende una información relativa a los coeficientes de predicción (122; 322), una información relativa a un libro de códigos determinista, una información relativa a un primer parámetro de ganancia (g c ) y un segundo parámetro de ganancia (g n ) y una información (142) relativa a un cuadro de señal de voz y uno no vocal; en donde la señal de audio codificada comprende información relacionada con una señal de excitación adaptativa para el cuadro de voz; y en donde la señal de audio codificada no comprende parámetros LTP para el cuadro no vocal, en donde la señal de excitación adaptativa se establece en cero para el cuadro no vocal. (l£p Un método (1400) para codificar una señal de audio (102), el método comprende:derivar (1410) los coeficientes de predicción (122;322) y una señal residual a partir de un cuadro no vocal de la señal de audio (102);calcular (1420) una primera información de parámetro de ganancia (^) para definir una primera señal de excitación (c(n)) relativa a un libro de códigos determinista y calcular una segunda información de parámetro de ganancia (g„ ) para definir una segunda señal de excitación (n(n)) relativa a una señal con características de ruido (n(n)) para el cuadro no vocal;y IMPI® INSTITUTO MSX1CANO , Oí LA PECYíEDAD formar (1430) una señal de salida (692;1002) sobre íaN sraeM-deT información (142) relativa a un cuadro de señal de voz, La piimeca. ,¡,nfflr.mac¡óa.. de parámetro de ganancia (5/) y la segunda información de parámetro de ganancia (gj;determinar si la señal residual se determinó a partir de un cuadro de señal de audio no vocal;generar una señal de excitación adaptativa para el cuadro vocal utilizando una memoria LTP (350n) y un generador de señal (850);y cuando se compara con un esquema de codificación de la CELP, no transmitir parámetros LTP para el cuadro no vocal para ahorrar bits, establecer la señal de excitación adaptativa en cero para el cuadro no vocal, y codificar más pulsos para la misma tasa de bits utilizando el libro de códigos determinista y los bits de ahorro. SA QSLy Un método (1500) para decodificar una señal de audio recibida (692;1002) que comprende una Información relativa a los coeficientes de predicción (122;322), la señal de audio recibida no comprendiendo parámetros LTP para el cuadro no vocal, el método comprende: generar (1510) una primera señal de excitación (1012, 1012’) a partir de un libro de códigos determinista para una porción de una señal sintetizada (1062);generar (1520) una segunda señal de excitación (1022, 1022’) a partir de una señal con características de ruido (n(n)) para la porción de la señal sintetizada (1062);combinar (1530) la primera señal de excitación segunda señal de excitación (1022, 1022’) para genjprai...Íina..señal,.,rip excitación combinada (1052) para la porción de la señal sintetizada (1062);y sintetizar (1540) la porción de la señal sintetizada (1062) a partir de la señal de excitación combinada (1052) y los coeficientes de predicción (122;322);generar una señal de excitación adaptativa para el cuadro de voz usando una memoria LTP (350n) y un generador de señal (850);y establecer a cero la señal de excitación adaptativa para el cuadro no vocal, y proporcionar más pulsos para la misma tasa de bits debido a los bits de ahorro dado por la falta de parámetros LTP para el cuadro no vocal utilizando el libro de códigos determinista.
- 1720. Un medio legible por computadora para codificar una señal de audio que comprende el método de la reivindicación 18.
- 1821. Un medio legible por computadora para decodificar una señal de audio recibida que comprende el método de la reivindicación 19.
Independent claims18
364 paragraphs in 54 sections, as filed
(54) Title: CONCEPT FOR CODING AN AUDIO SIGNAL AND DECODING AN AUDIO SIGNAL USING DETERMINISTIC AND NOISE INFORMATION.
(54) Title: CONCEPT FOR ENCODING AN AUDIO SIGNAL AND DECODING AN AUDIO SIGNAL USING DETERMINISTIC AND NOISE LIKE INFORMATION.
(57) Summary
An encoder for encoding an audio signal comprises: an analyzer (120; 320) configured to derive prediction coefficients (122; 322) and a residual signal from a non-voice frame of the audio signal (102); a gain parameter calculator (550; 550 ') configured to calculate a first gain parameter information (ge) in order to define a first drive signal (c (n)) relative to a deterministic codebook and to calculate a second gain parameter information (gn) in order to define a second drive signal (n (n)) relative to a signal with noise characteristics for the non-voice picture; and a bitstream former (690) configured to form an output signal (692) based on information (142) relating to a voice signal box, the first gain parameter information (ge) and the second gain parameter information (gn).
(57) Abstract
An encoder for encoding an audio signal comprises: an analyzer (120; 320) configured for deriving prediction coefficients (122; 322) and a residual signal from an unvoiced trame of the audio signal (102); a gain parameter calculator (550; 550 ') configured for calculating a first gain parameter (ge) Information for defining a first excitation signal (c (n)) related to a deterministic codebook and for calculating a second gain parameter (gn) Information for defining a second excitation signal (n (n)) related to a noise-like signal for the unvoiced trame; and a bitstream former (690) configured for forming an output signal (692) based on an Information (142) related to a voiced signal trame, the first gain parameter (ge) Information and the second gain parameter (gn) Information.
Yes
Μ ΡI í 1. τ, P * <i qc ^ MeSB # I »*
PATENT TITLE No. 355258
Headlines): FRAUNHOFER-GESELLSCHAFT ZUR FÓRDERUNG DER ANGEWANDTEN
FORSCHUNG EV
Address: Hansastrasse 27c, 80686, Munich, GERMANY
Denomination.
Classification:
Inventor (s):
CONCEPT TO CODE AN AUDIO SIGNAL AND DECODE AN AUDIO SIGNAL USANDQ DETERMINISTIC AND NOISE INFORMATION. G10L19 / 2BvG1 $ Lt9 / e $ I. **> (.G10L1W08; G10L19 / 06; G10L19 / T¿, 4l0Lt9 / 2O; G10L19 / 083; G10L25 / 15; ftl0L> 025/932, f
EMf & IWQEL RAVELLI; MARKUS
Number:
MX / a / 20107094922
Póté:
EP #
Validity: Vejnléf years Date of
Expedition date:
The referei patent
<img file="MX355258B_D0001.tif" />
' and. V, p- 'and f 1 1!
International:
014 / 'M *.
Number:
13189392.7
14178785.3 • 'ft fte October 2034
Who subscribes to this title is (Official Gazette of the Federation '
25/01/2006, 06/05/2009,06/01/2010,
Regulations of the Mexican Institute articles 1, 3, 4, 5 fraction V Clause a) .i
12/27/1999, amended on 10/10/2002, 07/29/20I Deputy Generals, Coordinator, Divi Departmental Directors and other subordinates of the Mexii Institute 08/04/2004 and 09/13/2007).
yg · 'Ά ty¡ ^ tte | ajjeyide la Prepfe & pi
In accordance with artw (p as of the filing date «1¼%.
<img file="MX355258B_D0002.tif" />
er valid JIMerei
Industrial.
non-extendable, counted to dates.
isí (te the Industrial Property Law 5/1999. 26/01/2004, 06/16/2005, 6 a), 4th and 12th sections I and III of 7/2004, 07/28/2004 and 7 / 09/2007); or of Industrial Property (DOF that delegates powers to the Divisional Deputy Directors, Coordinators 5/12/1999, amended on 02/04/2000, 07/29/2004,
This letter is signed with an advanced electronic signature (FIEL), based on articles 7 BIS 2 of the Industrial Property Law; 3 of its Regulations, and 1 fraction III, 2 fraction V, 26 BIS and 26 TER of the Agreement establishing the guidelines for the use of the Electronic Payment and Services Portal (PASE) of the Mexican Institute of Industrial Property, in the procedures indicated.
THE DIVISIONAL DIRECTOR OF PATENTS
<img file="MX355258B_D0003.tif" />
NAHANNY CANAL REYES
Original string:
NAHANNY MARISOL CANAL REYES | 00001000000403252793 | Tax Administration Service | 1695 || MX / 2018/30338 | MX / a / 2016/004922 | Patent title PCT | 1220 | RRGO | Page (s) 1 | GI1 ++ 07cXEM3LKgd1kjwjAdCMSo =
Digital stamp:
fd3RWM2qSggFlmXatC2 + eznyVRVad9SWS1gWwetsuSa8RvwwMwwbULvf + UvWRyvl / gf7WLOFGqt58eyWO + wfeL7 + xN
2276FCDM9N34NwrSTIk4YahvAR5scgNW7uQDAZ8 + AH + OMua6Yz + dOxeqcybE9 / 6B4n3cg1jf7ZCNoK3zAd / jiqKVHE pXOGw9606D2f + r3FqYn6ATcBfsORTI1hjMiBk6YsHLXGcO / 3MBw0XLiO6Upfkh8Nabw + BmlCbsYkXLvbDqbAG4 / P26
SzD9ONrl2OXnfqd8sZSGx0aaE + / AJE39s1ALqUwYbsEWi1GJzfKUARdvHrDc93dgJ75EadwA ==
Arenal No. 550 Floor 1, Pueblo Sania María Tepepan, Xochimilco. 16020. Mexico City.
(55) 53340700 www.gob mx / impi
<img file="MX355258B_D0004.tif" />
MX / 2018/30338
-OR
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX355258B_D0005.tif" />
CONCEPT FOR CODING AN AUDIO SIGNAL AND DECODING AN AUDIO SIGNAL USING DETERMINIST INFORMATION
AND OF NOISE TYPE
Description
The present invention relates to encoders for encoding an audio signal, in particular a voice related audio signal. The present invention further relates to decoders and methods for decoding an encoded audio signal. The present invention also relates to encoded audio signals and advanced voice non-voice encoding at low bit rates.
With a low bit rate, speech encoding can take advantage of special handling for non-voice frames to maintain voice quality while reducing bit rate. Non-vocal frames can be modeled perceptually in the form of a random excitation shaped in both the frequency and time domain. Since the waveform and excitation look and sound almost the same as Gaussian white noise, its encoding of the waveform can relax and be replaced by synthetically generated white noise. Coding will then consist of coding the shapes of the time and frequency domains of the signal.
Fig. 16 shows a schematic block diagram of a parametric non-voice coding scheme. A synthesis filter 1202 is configured to model the vocal apparatus and parameterized by the LPC (Linear Predictive Coding) parameters. From the derived LPC filter, which comprises a filter function A (z), a weighted filter can be derived
ΙΜΡΪ @>
MEXICAN INSTITUTE Ot INDUSTRIAL PROPERTY
<img file="MX355258B_D0006.tif" />
perceptual by weighting the coefficients d<sup>alpr</sup> n fiitm pgrnfintual fw (n) generally has a transfer function of the form:
Ffw (z ') =
A (z) A (z / w) where w is less than 1. The gain parameter g<sub>n</sub> is computed to obtain a synthesized energy corresponding to the original energy in the perceptual domain according to:
<sub>=</sub> ugly ^ n) <sup>9n</sup> Jz £<sub>0</sub>nw2 (n) where sw (n) and nw (n) are the input signal and generated noise, respectively, filtered by the perceptual filter fw (n). Gain g<sub>n</sub> it is computed for each sub-frame of size Ls. For example, an audio signal can be divided into frames with a length of 20 ms. Each frame can be subdivided into subframes, for example, into four subframes, each comprising a length of 5 ms.
The code excited linear prediction (CELP) encoding scheme is in widespread use in voice communications and is a very efficient way to encode speech. It allows for more natural voice quality than parametric encoding, but it also requires higher rates. The
CELP synthesizes an audio signal by transmission to a Linear Predictive filter, called the LPC synthesis filter, which can comprise a 1 / A (z) form, the sum of two excitations. An excitement comes from the decoded past, which is called the adaptive codebook. The other contribution comes from an innovative codebook populated by fixed codes. However, at low bit rates the innovative codebook is not populated enough to
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX355258B_D0007.tif" />
<img file="MX355258B_D0008.tif" />
effectively model the fine structure of the t'í '^ itf ^ i<sup>Ar> rnn</sup> noise characteristics of the non-vocal. Therefore, the perceptual quality is degraded, especially the non-vocal pictures, which then sound garish and not at all natural.
To mitigate encoding artifacts at low bit rates, different solutions have already been proposed. In G.718 [1] and [2], the innovative codebook codes are adaptively and spectrally shaped by improving the spectral regions corresponding to the formants of the current table. Formant shapes and positions can be derived directly from LPC coefficients, coefficients already available on both the encoder and decoder side. The formative improvement of the c (n) codes is done through a simple filtration according to:
c (n) * fe (ri) where * denotes the convolution operator and where fe (n) is the pulse response of the transfer function filter:
Ffe (z) =
A (z / wí) A (z / w2) where w1 and w2 are the two weighting constants that more or less emphasize the formant structure of the transfer function Ffe (z). The resulting shaped codes inherit a characteristic from the voice signal and the synthesized signal sounds cleaner.
At CELP, it is also common to add a Spectral Tilt to the innovative codebook decoder. This is done by filtering the codes with the following filter:
..r.rtr; IMPI
MEXICAN INSTITUTE! FROM THE PMOPIBDAI? INDUSTRIAL
<img file="MX355258B_D0009.tif" />
Ft (z) - 1 - βζ <sup>1</sup>
The β factor is generally related to the voicing of the previous table and is dependent, that is, it varies. Live sound can be calculated from the energy contribution of the adaptive codebook. If the preceding box is voice, it is considered that the current box is also voice and that the codes would have more energy at low frequencies, that is, they would show a negative inclination. In contrast, the added spectral bias will be positive for the voice boxes and more energy will be distributed towards the high frequencies.
io The use of spectral shaping for voice enhancement and reduction of decoder output noise is standard practice. What is called formant enhancement as a postfiltration consists of an adaptive postfiltration for which the coefficients are derived from the LPC parameters of the decoder. The postfilter is similar to (fe (n)) that is used to shape the innovative excitation in certain CELP encoders, as previously discussed here. However, in such a case, post-filtering is only applied at the end of the decoding process and not on the encoder side.
In conventional CELP (CELP = Linear Prediction)
Excited by (Code) Book), the frequency setting is modeled by the LP synthesis filter (Linear Prediction), while the time domain conformation can be approximated by the excitation gain sent to Each subframe, although the Long Term Prediction (LTP) and the innovative codebook are generally not suitable for modeling excitation with noise characteristics of the frames.
<img file="MX355258B_D0010.tif" />
vowels. CELP requires a relatively high bit rate to achieve good quality of the unvoiced voice.
A voice or non-vocal characterization can be related to the segmentation of the voice into portions and associate each one with a different original voice model. The original models, as used in the CELP voice coding scheme, are based on an adaptive harmonic excitation that simulates the air flow coming out of the glottis and a resonance filter models the vocal apparatus excited by the current. of air produced. These models can provide good results for phonemes such as the vowels, but they are an incorrect model for the portions of the voice that are not generated by the glottis, especially when the vocal cords do not vibrate, as is the case with the deaf phonemes "s" or "F".
On the other hand, parametric voice encoders are also called vocoders and adopt a unique original model for non-voice pictures. It can achieve very low bit rates, while still achieving a certain synthetic quality that is not as natural as the quality achieved with CELP encoding schemes at much higher rates.
Therefore, there is a tangible need to improve audio signals.
An object of the present invention is to increase the sound quality at low bit rates and / or reduce the bit rates, for good sound quality.
This is achieved by an encoder, a decoder, an encoded audio signal and the methods according to the independent claims.
The inventors discovered that in a first aspect, a quality of a decoded audio signal relative to a non-vocal frame of the audio signal can be increased, i.e. improved, by determining a
<img file="MX355258B_D0011.tif" />
IMPI
MRXíCan INSTITUTE.- »
GIVE IT INDUSTRIAL PROPERTY voice-related conformation information, so that a gain parameter information for signal amplification can be derived from the voice-related conformation information. Also, voice related shaping information can be used to spectrally shape a decoded signal. In this way, the frequency regions that are most important to the voice, for example low frequencies below 4 kHz, can be processed in such a way that they include fewer errors.
The inventors further discovered that, in a second aspect, generating a first excitation signal from a deterministic codebook for a frame or subframe (portion) of a synthesized signal and generating a second excitation signal from a signal with noise characteristics for the frame or subframe of the synthesized signal and combining the first drive signal and the second drive signal to generate a combined drive signal, the sound quality of the synthesized signal can be increased, that is, improved. Especially for portions of an audio signal, which comprises a voice signal with background noise, the sound quality can be improved by adding signals with noise characteristics. A gain parameter can be determined to optionally amplify the first drive signal in the encoder and information relating thereto can be transmitted with the encoded audio signal.
Alternatively or additionally, the enhancement of the synthesized audio signal can be at least partially exploited in order to reduce the bit rates to encode the audio signal.
An encoder according to the first aspect comprises an analyzer configured to derive prediction coefficients and a residual signal from a frame of the audio signal. The encoder further comprises a
<img file="MX355258B_D0012.tif" />
MEXICAN INSTITUTE Df. THE PROPERTY
INDUSTRIAL
<img file="MX355258B_D0013.tif" />
formant information calculator configured to calculate voice related spectral shaping information from prediction coefficients. The encoder further comprises a gain parameter calculator configured to calculate a gain parameter from a non-vocal residual signal and spectral shaping information and a bit stream former configured to form an output signal based on a information relating to a speech signal box, the gain parameter or a quantized gain parameter and the prediction coefficients.
Other embodiments of the first aspect provide an encoded audio signal comprising prediction coefficient information for a voice box and a non-voice box of the audio signal, additional information regarding the voice signal box and a parameter. gain or a quantized gain parameter for the non-vocal frame. This enables voice-related information to be transmitted efficiently to result in decoding of the encoded audio signal to produce a synthesized (restored) signal with high audio quality.
Other embodiments of the first aspect provide a decoder for decoding a received signal comprising prediction coefficients. The decoder comprises a formant information calculator, a noise generator, a modeler, and a synthesizer. The formant information calculator is configured to calculate speech related spectral shaping information from the prediction coefficients. The noise generator is configured to generate a signal with decoding noise characteristics. The modeler is configured to shape a spectrum of the signal with decoding noise characteristics or an amplified representation of it, using the spectral shaping information to obtain a signal with noise characteristics
<img file="MX355258B_D0014.tif" />
shaped decoder. The synthesizer is configured to synthesize a synthesized signal from the encoding signal with shaped and amplified noise characteristics and prediction coefficients.
Other embodiments of the first aspect relate to a method of encoding an audio signal, a method of decoding a received audio signal, and a computer program.
The embodiments of the second aspect provide an encoder for encoding an audio signal. The encoder comprises an analyzer configured to derive prediction coefficients and a residual signal from a non-voice frame of the audio signal. The encoder further comprises a gain parameter calculator configured to compute a first gain parameter information to define a first drive signal relative to a deterministic codebook and to compute a second gain parameter information to define a second trigger signal. excitation relative to a signal with noise characteristics for the non-voice picture. The encoder further comprises a bitstream former configured to form an output signal based on information relating to a voice signal box, the first gain parameter information and the second gain parameter information.
Other embodiments of the second aspect provide a decoder for decoding a received audio signal comprising information regarding prediction coefficients. The decoder comprises a first signal generator configured to generate a first drive signal from a deterministic codebook for a portion of a synthesized signal. The decoder further comprises a second signal generator configured to generate a second drive signal from a signal with noise characteristics for the portion of the synthesized signal. The
IMP
I
MEXICAN INSTITUTE OF THE FORMID
INDUSTRIAL
<img file="MX355258B_D0015.tif" />
decoder further comprises a combiner and a synthesizer, where the combiner is configured to combine the first drive signal and the second drive signal to generate a combined drive signal for the portion of the synthesized signal. The synthesizer is configured to synthesize the portion of the synthesized signal from the combined drive signal and prediction coefficients.
Other embodiments of the second aspect provide an encoded audio signal comprising information relating to prediction coefficients, information relating to a deterministic codebook, information relating to a first gain parameter and a second gain parameter, and a information regarding a voice and non-voice signal box.
Other embodiments of the second aspect provide methods for encoding and decoding an audio signal, a received audio signal respectively, and a computer program.
The preferred embodiments of the present invention will now be described with respect to the accompanying drawings, in which:
Fig. 1 shows a schematic block diagram of an encoder for encoding an audio signal in accordance with an embodiment of the first aspect;
Fig. 2 shows a schematic block diagram of a decoder for decoding a received input signal in accordance with an embodiment of the first aspect;
Fig. 3 shows a schematic block diagram of an additional encoder for encoding the audio signal in accordance with an embodiment of the first aspect;
Fig. 4 shows a schematic block diagram of an encoder comprising a varied gain parameter calculator when
ΙΜΡΪ
MSXICAN INSTITUTE OF INDUSTRIAL PROPERTY compares with Fig. 3 according to an embodiment of the first aspect;
Fig. 5 shows a schematic block diagram of a gain parameter calculator configured to calculate a first gain parameter information and to form a code excited signal according to an embodiment of the second aspect;
Fig. 6 shows a schematic block diagram of an encoder for encoding the audio signal and comprising the gain parameter calculator described in Fig. 5 according to an embodiment of the second aspect;
Fig. 7 shows a schematic block diagram of a gain parameter calculator comprising an additional modeler configured to form a signal with noise characteristics when compared to Fig. 5 according to an embodiment of the second aspect;
Fig. 8 shows a schematic block diagram of a non-voice coding scheme for CELP according to an embodiment of the second aspect;
Fig. 9 shows a schematic block diagram of a parametric non-voice encoding according to an embodiment of the first aspect;
Fig. 10 shows a schematic block diagram of a decoder for decoding an encoded audio signal, according to an embodiment of the second aspect;
Fig. 11a shows a schematic block diagram of a modeler implementing an alternative structure when compared to a modeler illustrated in Fig. 2, in accordance with an embodiment of the first aspect;
Fig. 11b shows a schematic block diagram of an additional modeler that implements an additional alternative when compared to the
<img file="MX355258B_D0016.tif" />
<img file="MX355258B_D0017.tif" />
ϊ
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY modeler illustrated in Fig. 2, according to an embodiment of the first aspect;
Fig. 12 shows a schematic flow diagram of a method for encoding an audio signal in accordance with an embodiment of the first aspect;
Fig. 13 shows a schematic flow diagram of a method for decoding a received audio signal comprising prediction coefficients and a gain parameter, in accordance with an embodiment of the first aspect;
Fig. 14 shows a schematic flow diagram of a method for encoding an audio signal in accordance with an embodiment of the second aspect; and Fig. 15 shows a schematic flow diagram of a method for decoding a received audio signal, in accordance with an embodiment of the second aspect.
The same or equivalent elements or elements with equal or equivalent functionality are indicated in the following description with equal or equivalent reference numbers, even if they appear in different figures.
In the following description, a plurality of details are provided in order to provide a more complete explanation of the embodiments of the present invention. However, those skilled in the art will appreciate that the embodiments of the present invention can be practiced without those specific details. In other cases, structures and devices that are well known are illustrated in block diagram form rather than in detail, so as not to hinder the description of the embodiments of the present invention. Furthermore, the characteristics of the
<img file="MX355258B_D0018.tif" />
<img file="MX355258B_D0019.tif" />
IA PRORIEDAU MEXICAN INSTITUTE
INDUSTRIAL
<img file="MX355258B_D0020.tif" />
Different embodiments to be described may be combined with each other, unless specifically indicated otherwise.
Next, reference will be made to modifying an audio signal. An audio signal can be modified by amplifying and / or attenuating portions of the audio signal. A portion of the audio signal may be, for example, a sequence of the audio signal in the time domain and / or a spectrum thereof in the frequency domain. With respect to the frequency domain, the spectrum can be modified by amplifying or attenuating the spectral values ordered in frequencies or frequency ranges. The modification of the spectrum of the audio signal may comprise a sequence of operations such as an amplification and / or attenuation of a first frequency or frequency range and subsequently an amplification and / or attenuation of a second frequency or frequency range. Modifications in the frequency domain can be represented as a calculation, for example a multiplication, division, addition, or the like, of spectral values and gain values and / or attenuation values. Modifications can be made sequentially, for example by first multiplying the spectral values with a first multiplication value and then with a second multiplication value. Doing the multiplication with the second multiplication value and then with the first multiplication value allows us to arrive at an identical or almost identical result. On the other hand, the first multiplication value and the second multiplication value can be combined first and then applied in terms of a combined multiplication value to the spectral values, while still reaching the same result, or comparable result, of the operation. Accordingly, the steps of the modifications, configured to form or modify a spectrum of the audio signal to be described below, are not restricted to the order described,
<img file="MX355258B_D0021.tif" />
I
UTο ι ITUTO MF.XICA NOT FROM THE? RÜ? IADAD industrial but can also be carried out in a different order leaving du Lluyui. to the same result and / or effect.
FIG. 1 shows a schematic block diagram of an encoder 100 for encoding an audio signal 102. Encoder 100 comprises a frame builder 110 configured to generate a frame sequence 112 based on audio signal 102. The Sequence 112 comprises a plurality of frames, where each frame of audio signal 102 comprises a length (time duration) in the time domain. For example, each frame can be 10 ms, 20 ms or 30 ms long.
Encoder 100 comprises an analyzer 120 configured to derive prediction coefficients (LPC = linear prediction coefficients) 122 and a residual signal 124 from a frame of the audio signal. Frame builder 110 or analyzer 120 are configured to determine a representation of audio signal 102 in the frequency domain.
Alternatively, the audio signal 102 may already be a representation in the frequency domain.
Prediction coefficients 122 can be, for example, linear prediction coefficients. Alternatively, nonlinear prediction can also be applied, such that predictor 120 is configured to determine nonlinear prediction coefficients. An advantage of linear prediction is a reduction in computational effort to determine prediction coefficients.
Encoder 100 comprises a speech / non-speech determiner 130 configured to determine whether residual signal 124 was determined from a non-speech audio frame. Determiner 130 is configured to supply the residual signal to a voice frame encoder 140, if the residual signal 124 was determined from a voice signal frame, and to supply the residual signal to a gain parameter calculator 150 if the signal
<img file="MX355258B_D0022.tif" />
MEXiCAN INSTITUTE. » Z DELA PI! O7JFr> A¡-, '^ A INDUSTRIAL
<img file="MX355258B_D0023.tif" />
Residual 124 was determined from a non-vocal audio chart. To determine whether the residual signal 122 was determined from a voice or non-voice signal box, the determiner 130 can use different approaches, such as autocorrelation of samples of the residual signal. A method of deciding whether a signal box was vocal or non-vocal is provided, for example in ITU (International Telecommunication Union) - T (Telecommunication Standardization Sector) standard G.718. A high amount of energy drawn at low frequencies can indicate a voice portion of the signal. Alternatively, a non-vocal signal can generate large amounts of energy at high frequencies.
Encoder 100 comprises a formant information calculator 160 configured to calculate speech related spectral shaping information from prediction coefficients 122.
The voice-related spectral shaping information may consider the formant Information, for example, by determining the frequencies or frequency ranges of the processed audio frame that comprise a higher amount of energy than in the vicinity. The spectral shaping information can segment the spectrum of magnitudes of the voice into formant frequency registers, that is, peaks, and non-formants, that is, valley. The spectrum-forming rulers can be derived, for example, using the representation of Immittance Spectral Frequencies (ISF) or Line Spectral Frequencies (LSF) of the prediction coefficients 122. In fact, the ISFs or LSFs represent the frequencies for which the synthesis filter that employs the prediction coefficients 122 resonates.
Spectrum conformation information related to speech 162 and non-vocal residuals are transmitted to the parameter calculator of
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX355258B_D0024.tif" />
gain 150 which is configured to calculate a gain parameter g<sub>n</sub> from the residual non-vocal signal and the spectral shaping information 162. The gain parameter g<sub>n</sub> it can be a scalar value or a plurality thereof, that is, the gain parameter can comprise a plurality of values relative to an amplification or attenuation of the spectral values in a plurality of frequency intervals of a spectrum of the signal that is it has to amplify or attenuate. A decoder can be configured to apply the gain parameter g<sub>n</sub> to the information of a received encoded audio signal, such that the portions of the received io encoded audio signals are amplified or attenuated on the basis of the gain parameter during decoding. The gain parameter calculator 150 can be configured to determine the gain parameter g<sub>n</sub> using one or more mathematical expressions or determination rules that produce a continuous value. Operations performed digitally, for example, through a processor, that express the result in a variable with a limited number of bits, can result in a quantized gain g<sub>n</sub>. Alternatively, the result can be further quantized according to a quantization scheme, such that quantized gain information is obtained. Encoder 100, therefore, can comprise a quantizer
170. Quantizer 170 can be configured to quantize the determined gain g<sub>n</sub>, up to a closest digital value, justified by the digital operations of encoder 100. Alternatively, quantizer 170 can be configured to apply a quantization function (linear or non-linear) to a gain factor g<sub>n</sub> already digitized and therefore quantified. A nonlinear quantization function may consider, for example, logarithmic dependencies of the human ear highly sensitive to low voice pressure levels and less sensitive to high pressure levels.
IMPI
INSTITUTE M «ICANI!
OF INDUSTRIAL PROPERTY
Encoder 100 further comprises an information deriving unit 180 configured to derive information related to prediction coefficients 182 from prediction coefficients 122. Prediction coefficients, such as linear prediction coefficients used to drive codebooks Innovative, they include a low robustness against distortions or errors. Therefore, for example, the conversion of linear prediction coefficients to interspectral frequencies (ISF) and / or the derivation of line spectral pairs (LSP) and the transmission of information related to them with the audio signal are known. encoded. The LSP and / or ISF io information comprises greater robustness against distortions in the transmission medium, for example error, or computer errors. Information deriving unit 180 may further comprise a quantizer configured to provide quantized information with respect to LSF and / or ISP.
Alternatively, the information deriving unit can be configured to transmit prediction coefficients 122. Alternatively, encoder 100 can be performed without information deriving unit 180. Alternatively, the quantizer can be a functional block of the gain parameter calculator 150 o of bit stream former 190, such that bit stream former 190 is configured to receive the gain parameter g<sub>n</sub> and derive the quantized gain g<sub>n</sub> based on it.
Alternatively, when the gain parameter g<sub>n</sub> already quantized, encoder 100 can be performed without quantizer 170.
Encoder 100 comprises a bitstream former 190 configured to receive a voice signal, voice information 142 relating to a voice box of an encoded audio signal respectively provided by voice box encoder 140, to receive the quantized gain g<sub>n</sub> and the
<img file="MX355258B_D0025.tif" />
IMPI
MEXICAN INSTITUTE ÜE THE INDUSTRIAL PROPERTY
<img file="MX355258B_D0026.tif" />
information relating to the prediction coefficients 182 and forming an output signal 192 based thereon. <sup>l</sup> ''
Encoder 100 may be part of a voice encoding apparatus such as a landline or mobile phone or an apparatus comprising a microphone for transmitting audio signals, such as a computer, a tablet PC, etc. The output signal 192 or a signal derived therefrom can be transmitted, for example, by mobile (wireless) communications or by wired communications, such as a network signal.
An advantage of encoder 100 is that output signal 192 comprises the information derived from spectral shaping information converted to the quantized gain g<sub>n</sub>. Therefore, the decoding of the output signal 192 will allow to produce or obtain more information related to the voice and, therefore, to decode the signal in such a way that the decoded signal obtained comprises a higher quality with respect to a perceived level of voice quality.
FIG. 2 shows a schematic block diagram of a decoder 200 for decoding a received input signal 202. The received input signal 202 can correspond, for example, to the output signal 192 provided by the encoder 100, where the output signal 192 can be encoded through high-level layer encoders, transmitted through a medium, received by a receiving apparatus, decoded in high layers, thereby producing input signal 202 for decoder 200.
Decoder 200 comprises a bitstream deformer (demultiplexer; DE-MUX) to receive input signal 202. Bitstream deformer 210 is configured to provide prediction coefficients 122, the quantized gain g<sub>n</sub> and the voice information 142. To obtain the prediction coefficients 122, the bitstream deformer
<img file="MX355258B_D0027.tif" />
it may comprise an inverse information bypass unit that performs an inverse operation when compared to information bypass unit 180. Alternatively, decoder 200 may comprise an inverse information bypass unit, not illustrated, configured to perform the inverse operation with respect to to the information deriving unit 180. In other words, the prediction coefficients are decoded, that is, restored.
Decoder 200 comprises a formant information calculator 220 configured to calculate voice related spectral shaping information from prediction coefficients 122, as described for formant information calculator 160. Formant information calculator 220 is configured to provide voice related spectral shaping information 222. Alternatively, input signal 202 may also comprise speech-related spectral shaping information 222, where the transmission of prediction coefficients or information relating thereto, such as, for example, quantized LSF and / or ISF, instead of speech related spectral shaping information 222, allows a lower bit rate of input signal 202.
Decoder 200 comprises a random noise generator 240 configured to generate a signal with noise characteristics, which can simply be denoted as a signal with noise characteristics. Random noise generator 240 can be configured to reproduce a signal with noise characteristics obtained, for example, by measuring and storing a signal with noise characteristics. A signal with noise characteristics can be measured and recorded, for example, by generating thermal noise in a resistor or other electrical component and storing recorded data
<img file="MX355258B_D0028.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX355258B_D0029.tif" />
in a memory. The random noise generator 240 is configured to provide the n (n) noise (type) signal.
Decoder 200 comprises a modeler 250 comprising a shaping processor 252 and a variable amplifier 254. Modeler 250 is configured to spectrally shape a signal spectrum with noise characteristics n (n). The shaping processor 252 is configured to receive the voice-related spectral shaping information and to shape the signal spectrum with noise characteristics n (n), for example by multiplying the spectral values of the signal spectrum with the noise characteristics n (n) and the values of the spectral conformation information. The operation can also be performed in the time domain by means of the convolution of the signal with noise characteristics n (n) with a filter given by the spectral conformation information. The shaping processor 252 is configured to supply a signal with shaped noise characteristics 256, a spectrum thereof respectively to the variable amplifier 254. The variable amplifier 254 is configured to receive the gain parameter g<sub>n</sub> and to amplify the spectrum of the signal with shaped noise characteristics 256 in order to obtain a signal with shaped noise characteristics amplified 258. The amplifier can be configured to multiply the spectral values of the signal with shaped noise characteristics 256 with the values of the gain parameter g<sub>n</sub>. As already indicated, the modeler 250 can be implemented in a way that the variable amplifier 254 is configured to receive the signal with noise characteristics n (n) and to provide a signal with amplified noise characteristics to the configured forming processor 252 to conform the signal with amplified noise characteristics. Alternatively, the forming processor 252 can be configured to receive the information from
ΪΜ
Tsir
IL IL
MEXICAN INSTITUTE .Zl
OF THE PROPERTY
INDUSTRIAL voice related spectral conformation 222 and the gain parameter g<sub>n </sub>and to apply consecutively, one after the other, the two Information to the signal with noise characteristics n (n) or to combine the two Information, for example, by multiplication or other calculations, and to apply a combined parameter to the signal with noise characteristics n (n).
The signal with noise characteristics n (n) or the amplified version thereof, conformed with the spectral conformation information related to the voice gives rise to the decoded audio signal 282 that comprises a sound quality more related to the voice ( natural). This enables high quality audio signals to be obtained and / or encoder-side bit rates to be reduced, while maintaining or improving output signal 282 at the decoder with reduced range.
Decoder 200 comprises a synthesizer 260 configured to receive the prediction coefficients 122 and the signal with amplified shaped noise characteristics 258 and to synthesize a synthesized signal 262 from the signal with amplified shaped noise characteristics 258 and the prediction coefficients 122 The synthesizer 260 can comprise a filter and can be configured to match the filter with the prediction coefficients. The synthesizer can be configured to filter the signal with amplified shaped noise characteristics 258 with the filter. The filter can be implemented as software or as a hardware structure and can comprise an infinite Impulse Response (IIR) or Finite Impulse Response (FIR) structure.
The synthesized signal corresponds to a non-vocal decoded frame of an output signal 282 from decoder 200. Output signal 282 comprises a frame sequence that can be converted to a continuous audio signal.
<img file="MX355258B_D0030.tif" />
MEXICAN INSTITUTE E> 2 THE PROPERTY
INDUSTRIAL
<img file="MX355258B_D0031.tif" />
The bitstream deformer 210 is configured to separate and supply the voice information signal 142 from the input signal 202. The decoder 200 comprises a voice frame decoder 270 configured to provide a voice frame on the basis of the Voice Information 142. The voice frame decoder (voice frame processor) is configured to determine a voice signal 272 based on the voice information 142. Voice signal 272 may correspond to the voice audio box and / or the voice residual of decoder 100.
Decoder 200 comprises a combiner 280 configured to combine the non-voice decoded frame 262 and the voice frame 272 to obtain the decoded audio signal 282.
Alternatively, the modeler 250 can be performed without an amplifier, such that the modeler 250 is configured to conform the signal spectrum with n (n) noise characteristics without further amplifying the obtained signal. This may result in a reduction in the amount of information transmitted by input signal 222 and, therefore, a reduction in bit rate or a shorter duration of a sequence of input signal 202. Alternatively or additionally, the decoder 200 can be configured to decode only non-voice frames or to process voice and non-voice frames, both by spectral shaping of the signal with noise characteristics n (n) and by synthesizing the synthesized signal 262 for voice and non-vocal boxes. This may allow the implementation of the decoder 200 without the voice frame decoder 270 and / or without a combiner 280 and thus results in a reduction in the complexity of the decoder 200.
Output signal 192 and / or input signal 202 comprise information regarding prediction coefficients 122, information for a voice box and a non-voice box such as a flag indicating whether the box
IVx A j:
MEXICAN INSTITUTE
DK LA wtomr.An
INDUSTRIAL
<img file="MX355258B_D0032.tif" />
processed is vocal or non-vocal, and more information regarding the r.nadrn gives voice signal such as an encoded voice signal. The output signal 192 and / or the input signal 202 further comprise a gain parameter or a quantized gain parameter for the non-vocal frame, such that the non-vocal frame can be decoded based on the prediction coefficients 122 and the gain parameter g<sub>n</sub>, g „, respectively.
FIG. 3 shows a schematic block diagram of an encoder 300 for encoding audio signal 102. Encoder 300 comprises frame builder 110, a predictor 320 configured to determine linear prediction coefficients 322, and a residual signal 324, applying a filter A (z) to frame sequence 112 provided by frame constructor 110. Encoder 300 comprises determiner 130 and voice frame encoder 140 to obtain information from voice signal 142. Encoder 300 further comprises formant information calculator 160 and gain parameter calculator 350.
The gain parameter calculator 350 is configured to provide a gain parameter g<sub>n</sub> as previously described. The gain parameter calculator 350 comprises a random noise generator 350a to generate a coding signal with noise characteristics
350b. The gain calculator 350 further comprises a modeler 350c having a shaping processor 350d and a variable amplifier 350e. The shaper processor 350d is configured to receive the voice related shaping information 162 and the noise characteristic signal 350b, and to shape a spectrum of the noise characteristic signal 350b with the voice related spectral shaping information 162 , as described for modeler 250. Variable amplifier 350e is configured to amplify a signal with shaped noise characteristics
ΙΜΡΪβΥγ
INSTITUTE Μ EX ICA NO
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
350f with a gain parameter g<sub>n</sub>(temp) which is a temporary gain parameter received from a 350k controller. The variable amplifier 350e is also configured to provide a signal with amplified shaped noise characteristics 350g, as described for the signal with amplified noise characteristics 258. As described for the modeler 250, the order of shaping and amplifying the signal with noise characteristics can be combined or modified, when compared to Fig. 3.
The gain parameter calculator 350 comprises a comparator 350h, which is configured to compare the non-vocal residual, provided by the determiner 130 and the signal with amplified shaped noise characteristics 350g. The comparator is configured to obtain a measurement for a similarity of the non-vocal residual and the signal with 350g amplified shaped noise characteristics. For example, comparator 350h can be configured to determine a cross correlation of both signals. Alternatively or additionally, the comparator 350h can be configured to compare the spectral values of both signals in some or all of the frequency bins. The comparator 350h is further configured to obtain a comparative result 350i.
The gain parameter calculator 350 comprises the controller
350k configured to determine gain parameter g<sub>n</sub>(temp) based on comparative result 350i. For example, when the comparative result 350i indicates that the signal with amplified shaped noise characteristics comprises an amplitude or magnitude less than a corresponding amplitude or magnitude of the non-vocal residual, the controller may be configured to increase one or more values of the parameter of gain g<sub>n</sub>(temp) for some or all of the signal frequencies with 350g amplified noise characteristics. Alternatively or additionally, the
-. <4 tr .AL Jfi. -your MEXICAN INSTITUTE OF THE F «O? íE> AD INDUSTRIAL controller can be configured to reduce one or more values of the gain parameter g<sub>n</sub>(temp) when the comparative result 350¡ indicates that the signal with amplified shaped noise characteristics comprises a magnitude or amplitude that is too high, that is, that the signal with amplified shaped noise characteristics is too high (in volume). Random noise generator 350a, modeler 350c, comparator 350h, and controller 350k can be configured to implement closed-loop optimization to determine the gain parameter g<sub>n</sub>(temp). When the measurement for the similarity of the non-vocal residual with the 350g amplified shaped noise characteristic signal, for example, expressed as a difference between the two signals, indicates that the similarity is above a threshold value, the 350k controller is configured to provide the gain parameter g<sub>n</sub> determined. A quantizer 370 is configured to quantize the gain parameter g<sub>n</sub> in order to obtain the quantized gain parameter g<sub>n</sub>.
The random noise generator 350a can be configured to produce a Gaussian noise. The random noise generator 350a can be configured to operate (summon) a random generator with a number of n uniform distributions between a lower limit (minimum value) such as -1 and an upper limit (maximum value) such as +1. For example, the random noise generator 350 is configured to summon the random generator three times. Since digitally implemented random noise generators can produce pseudorandom values, the addition or superposition of a plurality or a multitude of pseudorandom functions can allow obtaining a sufficiently randomly distributed function. This procedure is based on the Central Limit Theorem. The noise generator
<img file="MX355258B_D0033.tif" />
Random 350a can be configured to summon the random generator for 15 minus two, three, or more times, as indicated by the following pseudocode:
for (i = 0; i <Ls; ¡++) {n [i] = uniform_random ();
n [i] + = uniform_random ();
n [i] + = uniform_random ();
} io Alternatively, the random noise generator 350a may generate the signal with noise characteristics from a memory, as described for the random noise generator 240. Alternatively, the random noise generator 350a may comprise, for example , an electrical resistance, or some other means to generate a signal with noise characteristics by running a code or by measuring physical effects such as thermal noise.
The shaper processor 350b can be configured to add a formant structure and an inclination to signals with noise characteristics 350b by filtering the signal with noise characteristics 350b with fe (n), as already indicated. The inclination can be added by filtering the signal with a filter t (n) that comprises a transfer function based on:
Ft (z) = 1 - βζ<sup>-1</sup> where the factor β can be deduced from the voicing in the previous sub-box:
energy (4C contribution) - energy (IC contribution) voicing = -—-: --—: - renergy ^ sum of contributions)
I py
MEXICAN INSTITUTE OF PROPERTY
<img file="MX355258B_D0034.tif" />
where AC is the acronym for Adaptive Codebook and IC is the acronym for Innovative Codebook.
β = 0.25 (1 + voicing)
The gain parameter g<sub>n</sub>, the quantized gain parameter g<sub>n</sub> respectively, they provide additional information that may reduce an error or a mismatch between the encoded signal and the corresponding io decoded signal, decoded in a decoder such as decoder 200.
Regarding the determination rule
Ffe (z) =
A (z / wl) A (z / w2) The parameter w1 can comprise a positive non-zero value of 1.0 at most, preferably at least 0.7 and at most 0.8 and more preferably a value of 0 , 75. Parameter w2 may comprise a positive non-zero scalar value of 1.0 at most, preferably at least 0.8 and at most 0.93 and more preferably a value of 0.9. The parameter w2 is preferably greater than w1.
Fig. 4 shows a schematic block diagram of an encoder
400. Encoder 400 is configured to provide voice signal information 142 as described for encoders 100 and 300. Compared to encoder 300, encoder 400 comprises a varied gain parameter calculator 350 '. A comparator 350h 'is configured to compare audio box 112 and a synthesized signal 350I', in order to
<img file="MX355258B_D0035.tif" />
obtain a comparative result 350i '. The rain iigrinr<sub>fl</sub> Ho pprámotrn? gananciaο gain 350 'comprises a synthesizer 350m' configured to synthesize the synthesized signal 3501 'on the basis of the signal with amplified shaped noise characteristics 350g and the prediction coefficients 122.
Basically, the gain parameter calculator 350 'implements at least partially a decoder by synthesizing the synthesized signal 350Γ. When compared to encoder 300 comprising comparator 350h configured to compare the non-vocal residual and the signal with amplified shaped noise characteristics, encoder 400 comprises comparator 350h ', which is configured to compare the audio frame (probably complete ) and the synthesized signal. This results in much higher precision as the signal frames are compared to each other and not just its parameters. This greater precision may require an increased computational effort, since the audio box 122 and the synthesized signal 350I 'can understand a greater complexity when compared with the residual signal and with the information with shaped and amplified noise characteristics, in such a way that the comparison of both signals is also more complex. In addition, the synthesis must be calculated with the consequent requirement of computational efforts by the 350m 'synthesizer.
The gain parameter calculator 350 'comprises a memory
350n 'configured to record encoding information comprising the encoding gain parameter g<sub>n</sub> or a quantized version of it. This allows the 350k controller to obtain the stored gain value when a subsequent audio frame is processed. For example, the controller can be configured to determine a first value or set of values, that is, a first instance of the gain factor g<sub>n</sub>(temp) based on or equal to the value of g<sub>n</sub> for the preceding audio frame.
SL VAT
INSTITUTO MEXICANO Pt LA PROPIEDAD
INDUSTRIAL
<img file="MX355258B_D0036.tif" />
Fig. 5 shows a schematic block diagram of a gain parameter calculator 550 configured to calculate a first gain parameter information g<sub>n</sub> according to the second aspect. The gain parameter calculator 550 comprises a signal generator
550a configured to generate an excitation signal c (n). Signal generator 550a comprises a deterministic codebook and an index within the codebook for generating signal c (n). That is, input information such as prediction coefficients 122 produces a deterministic drive signal c (n). Signal generator 550a can be configured to or generate the excitation signal c (n) in accordance with an innovative codebook of a CELP coding scheme. The codebook can be determined or trained according to voice data measured in previous calibration steps. The gain parameter calculator comprises a modeler 550b configured to shape a spectrum of the code signal c (n) based on voice related shaping information 550c for the code signal c (n). The voice related conformation information 550c can be obtained from the formant information controller 160. The modeler 550b comprises a shaper processor 550d configured to receive the shaping information 550c to shape the code signal. The modeler 550b further comprises a variable amplifier 550e configured to amplify the shaped code signal c (n), in order to obtain an amplified shaped code signal 550f. Thus, the code gain parameter is configured to define the code signal c (n) that is relative to a deterministic codebook.
The gain parameter calculator 550 comprises the noise generator 350a configured to supply the signal (with characteristics of) noise n (n) and an amplifier 550g configured to amplify the signal with characteristics of
ΙΜΡΪ
MEXICAN INSTITUTE OF INDUSTRIAL EROE1EDAD noise n (n) based on the parameter of noise gain g<sub>n</sub>, in order to obtain a signal with 550h amplified noise characteristics. The gain parameter calculator comprises a combiner 550, configured to combine the amplified shaped code signal 550f and the signal with amplified noise characteristics 550h, to obtain a combined drive signal 550k. The combiner 550¡ can be configured, for example, to spectrally add or multiply the spectral values of the amplified shaped code signal and the signal with amplified noise characteristics 550f and 550h. Alternatively, the 550i combiner can be configured to convolve both signals 550f and 550h.
As described above with respect to the modeler 350c, the modeler 550b can be implemented in such a way that first the code signal c (n) is amplified by the variable amplifier 550e and then is formed by the conformator processor 550d. Alternatively, the shaping information 550c for the code signal c (n) can be combined with the code gain parameter information g<sub>c</sub>, so that combined information is applied to the code signal c (n).
The gain parameter calculator 550 comprises a comparator 5501 that is configured to compare the combined drive signal 550k and the residual non-vocal signal obtained for the vocal / non-vocal determiner 130. The comparator 550I may be the comparator 550h and is configured to provide a comparative result, ie a 550m measurement for a similarity of the combined 550k excitation signal and the residual non-vocal signal. The code gain calculator comprises a 550n controller configured to control the code gain parameter information g<sub>c</sub> and the noise gain parameter information g<sub>n</sub>. The g code gain parameter<sub>c</sub> and the noise gain parameter information g<sub>n</sub> can understand a
<img file="MX355258B_D0037.tif" />
<img file="MX355258B_D0038.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY plurality or a multitude of scalar or imaginary values. guide can -eer »relative to a frequency range of the signal with noise characteristics n (n) or a signal derived from it or to a spectrum of the code signal c (n) or a signal derived from it.
Alternatively, the gain parameter calculator 550 can be implemented without the forming processor 550d. Alternatively, the shaper processor 550d can be configured to shape the signal with noise characteristics n (n) and provide a signal with shaped noise characteristics to the variable amplifier 550g.
Thus, controlling both information of gain parameter g<sub>c</sub> yg<sub>n</sub>, the similarity of the combined excitation signal 550k can be increased when compared to the non-vocal residual, such that a decoder receiving information to the code gain parameter information g<sub>c</sub> and the noise gain parameter information g<sub>n</sub> It can reproduce an audio signal that comprises good sound quality. Controller 550n is configured to provide an output signal 550o comprising information relating to the g-code gain parameter information<sub>c</sub> and the noise gain parameter information g<sub>n</sub>. For example, signal 550o may comprise both gain parameter information g<sub>n</sub> yg<sub>c</sub> as scalar or qualified values or as values derived therefrom, for example, encoded values.
Fig. 6 shows a schematic block diagram of an encoder 600 for encoding audio signal 102 and comprising the gain parameter calculator 550 described in Fig. 5. Encoder 600 can be obtained, for example, by modifying the encoder. 100 or 300. Encoder 600 comprises a first quantizer 170-1 and a second quantizer 170-2. The first quantizer 170-1 is configured to quantize the information of
MBXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX355258B_D0039.tif" />
gain parameter g<sub>c</sub>, in order to obtain information with quantified profit g<sub>c</sub>. The second quantizer 170-2 is configured to quantize the noise gain parameter information g<sub>n</sub>, in order to obtain a quantized noise gain parameter information g<sub>n</sub>. A bitstream former 690 is configured to generate an output signal 692 that comprises the voice signal information 142, the information related to LPC 122, and both quantized gain parameter information g<sub>c</sub> yg<sub>n</sub>.
When compared to output signal 192, output signal 692 is extended or revalued by the quantized gain parameter information io g<sub>c</sub>. Alternatively, the quantizer 170-1 and / or 170-2 may be a part of the gain parameter calculator 550. Also, one of the quantizers
170-1 and / or 170-2 can be configured to obtain both quantized gain parameters g .and g „.
Alternatively, encoder 600 may be configured to comprise a quantizer configured to quantize the code gain parameter information g<sub>c</sub> and the noise gain parameter g<sub>n</sub>, in order to obtain the quantified parameter information g<sub>c</sub> yg<sub>n</sub>. Both Gain parameter information can be quantized, for example, consecutively.
Formant information calculator 160 is configured to calculate speech related spectral conformation information 550c from prediction coefficients 122.
Fig. 7 shows a schematic block diagram of a gain parameter calculator 550 'that is modified, compared to the gain parameter calculator 550. The gain parameter calculator 550' comprises the modeler 350 described in Fig. 3, instead of the 550g amplifier. Modeler 350 is configured to provide the signal
<img file="MX355258B_D0040.tif" />
MEXICAN INSTITUTE Say THE INDUSTRIAL PROPERTY of noise shaped amplified 350g. The combiner 5501 is configured to combine the amplified conformal code signal 550f and the amplified conformal noise signal 350g, to provide a combined drive signal 550k '. Formant information calculator 160 is configured to provide both voice related formant information 162 and 550c. The formative information related to the voice 550c and 162 can be the same. Alternatively, both information 550c and 162 may differ from each other. This allows a separate modeling, that is, the conformation of the signal generated by codes c (n) and n (n).
The 550n controller can be configured to determine the gain parameter information g<sub>c</sub> yg<sub>n</sub> for each subframe of a processed audio frame. The controller can be configured to determine, i.e. calculate, the gain parameter information g<sub>c</sub> yg<sub>n</sub> based on the details below.
First, the average energy of the subframe can be computed in the original short-term prediction residual signal, available during LPC analysis, that is, in the non-vocal residual signal. The energy is averaged over the four subframes of the current frame in the logarithmic domain according to:
<img file="MX355258B_D0041.tif" />
beef<sup>2</sup>(l Lsf + n) l = on = 0
Lsf
·) Where Lsf is the size of a subframe in samples. In this case, the table is divided into 4 sub-tables. The averaged energy can then be encoded in a number of bits, for example three, four, or five, using a previously trained stochastic codebook. The stochastic codebook may comprise a number of entries (size) according to a number of different values that can be represented by the
IMPI
MEXICAN INSTITUTE OF CURRENCY »
INDUSTRIAL number of bits, for example a size of 8 for a 3-bit number, a size of 16 for a 4-bit number or a size of 32 for a 5-bit number. A quantized gain irrg can be determined from the selected codeword from the codebook. For each sub-frame, the two information of gain g are computed<sub>c</sub> yg<sub>n</sub>. G code gain<sub>c</sub> It can be computed, for example, according to:
Ση ^ ό<sup>1</sup> xw (jT) cw (n ~) <sup>9C</sup> Σηΐζ<sup>1</sup> cw (n) · cw (n) where cw (n) is, for example, the fixed innovation selected from the fixed codebook 10, comprised of the signal generator 550a filtered by the perceptual weighted filter. The expression xw (n) corresponds to the conventional perceptual target excitation, computed in CELP encoders. G code gain information<sub>c</sub> can then be normalized to obtain a normalized gain g<sub>nc</sub>, based on:
Σ, ηΙρ<sup>1</sup> c (n) <sup>3nc 9c</sup>'Lsf * ÍO ^ /<sup>20</sup>
The normalized gain g<sub>nc</sub> it can be quantized, for example, by quantizer 170-1. Quantification can be done according to a linear or logarithmic scale. A logarithmic scale can comprise a scale of size 4, 5 or more bits. For example, the logarithmic scale comprises a size of 5 bits. Quantification can be done on the basis of:
<img file="MX355258B_D0042.tif" />
Index<sub>nc</sub> = [20 * log<sub>1Q</sub>((g<sub>nc</sub> + 20) / 1,25) + 0,5J
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY where the lndex index<sub>nc</sub> can be limited between 0 and 31, if the scale
<img file="MX355258B_D0043.tif" />
logarithmic comprises 5 bits. The lndex index<sub>nc</sub> it may be the quantized gain parameter information. The quantized gain of the g code<sub>c</sub> it can then be expressed on the basis of:
_ i <sub>n</sub>10 (index<sub>nc</sub>.l, 25-20) / 20) * 10<sup>nrg / 2</sup>°
Σ „= ο Vc (n) -c (n)
The code gain can be computed in order to minimize the mean square root error or mean square error (MSE)
Lsf-l - (xw (n) - g<sub>c</sub> cw (n))<sup>2 </sup>'n = 0 where Lsf corresponds to spectrum frequencies of lines determined from the prediction coefficients 122.
The noise gain parameter information can be determined in terms of power mismatch by minimizing an error based on i
Lsf
Lsf — l k-xw<sup>2</sup>(n) n = 0
Lsf-l
Σ • cw (n) + g „nw (ri)) 'n = 0
The variable k is an attenuation factor that can be varied depending on or based on the prediction coefficients, where the prediction coefficients allow to determine if the voice comprises a portion of low background noise or even no background noise (clear voice ). Alternatively, the signal may also be determined as a loud voice, for example when the audio signal or a frame thereof comprises
<img file="MX355258B_D0044.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY modifications between non-vocal and non-non-vocal cadres. The variable k_ can be set to a value of at least 0.85, of at least 0.95 or even up to a value of 1 for clear voice, where the high energy dynamics are perceptually important. Variable k can be set to a value of at least 0.6 and at most 0.9, preferably a value of at least 0.7 and at most 0.85 and more preferably a value of 0.8 for voice noisy, where the noise excitation is made more conservative to avoid fluctuation in output power between non-vocal and non-non-vocal frames. The error (lack of energy adaptation) can be computed for each of these io candidates of quantized gain g<sub>c</sub>. A table divided into four sub-tables can result in four candidates of quantized gain g<sub>c</sub>. That candidate that minimizes the error can be produced by the controller. The quantized noise gain (noise gain parameter information) can be computed on the basis of:
_ Σ ^ -η<sup>1</sup> >/<sup>c</sup>(n) · c (n)
Tn = (index<sub>n</sub> 0,25 + 0,25) & '
Ση = ο V ^ (n) -nCn) where the index lndex<sub>n</sub> it is limited between 0 and 3, according to the four candidates. A resulting combined drive signal, such as drive signal 550k or 550k ', can be obtained based on:
e (n) = c (n) + n (n) where e (n) is the combined excitation signal 550k or 550k '.
An encoder 600 or a modified encoder 600 comprising the gain parameter calculator 550 or 550 'may allow non-voice encoding, based on a CELP encoding scheme.
<img file="MX355258B_D0045.tif" />
MEXICAN INSTITUTE OF FROFIEDAD
INDUSTRIAL
<img file="MX355258B_D0046.tif" />
The CELP coding scheme can be modified based on the following representative details for the manipulation of non-vocal frames:
• LTP parameters are not transmitted as there is almost no periodicity in non-voice frames and the resulting encoding gain is very low. Adaptive arousal is set to zero.
• Saving bits are reported to the fixed codebook. More pulses can be encoded for the same bit rate and the quality can then be improved.
• At low rates, ie for rates between 6 and 12 kbps, the pulse coding is not sufficient to properly model excitation with target noise characteristics of the non-voice picture. A Gaussian codebook is added to the fixed codebook to build the final excitement.
Fig. 8 shows a schematic block diagram of a non-voice coding scheme for CELP, according to the second aspect. A modified 810 controller comprises both functions of the 550I comparator and the 550n controller. Controller 810 is configured to determine g code gain parameter information<sub>c</sub> and the noise gain parameter information g<sub>n</sub> on the basis of analysis by synthesis, that is to say comparing a synthesized signal with the input signal indicated as s (n) which is, for example, the non-vocal residual. Controller 810 comprises a synthesis analysis filter 820 configured to generate an excitation for the signal generator (innovative excitation) 550a and to provide the gain parameter information g<sub>c</sub> yg<sub>n</sub>. The synthesis analysis block 810 is configured to compare the combined excitation signal 550k '
ΙΜΡΪ
MEXICAN INSTITUTE OF THE PtOHSOAD
INDUSTRIAL
<img file="MX355258B_D0047.tif" />
using an internally synthesized signal by adapting uit filliu according to the parameters and information provided.
Controller 810 comprises an analysis block configured to obtain prediction coefficients, as described for analyzer 320, to obtain prediction coefficients 122. The controller further comprises a synthesis filter 840 to filter the combined drive signal 550k with the 840 synthesis filter, where the 840 synthesis filter is adapted by filter coefficients 122. An additional comparator can be configured to compare the input signal s (n) and the synthesized signal s (n), for example, the decoded (restored) audio signal. Also, memory 350n is provided, where controller 810 is configured to store the predicted signal and / or predicted coefficients in memory. A signal generator 850 is configured to provide an adaptive drive signal based on the predictions stored in memory 350n, allowing improvement of adaptive drive based on a previous combined drive signal.
Fig. 9 shows a schematic block diagram of a non-voice parametric encoding, according to the first aspect. The amplified shaped noise signal may be an input signal from a synthesis filter
910 which is adapted by the determined filter coefficients (prediction coefficients) 122. A synthesized signal 912 produced by the synthesis filter can be compared to the input signal s (n) which can be, for example, the audio signal. The synthesized signal 912 comprises an error when compared to the input signal s (n). Modifying the noise gain parameter g<sub>n</sub> By means of the analysis block 920 that can correspond to the gain parameter calculator 150 or 350, the error can be reduced or minimized. Storing the amplified shaped noise signal 350f in memory 350n,
IMPI
MEXICAN INSTITUTE PE THE PROPERTY
INDUSTRIAL
<img file="MX355258B_D0048.tif" />
An adaptive codebook update can be performed, such that the speech audio frame processing can also be improved based on the improved encoding of the non-voice audio frame.
FIG. 10 shows a schematic block diagram of a decoder 1000 for decoding an encoded audio signal, eg, encoded audio signal 692. Decoder 1000 comprises a signal generator 1010 and a noise generator 1020 configured to generate a signal with 1022 noise characteristics. The received signal 1002 comprises information related to the LPC, where a bitstream deformer 1040 is configured to provide the prediction coefficients 122 based on the information related to the prediction coefficients. For example, decoder 1040 is configured to extract prediction coefficients 122. Signal generator 1010 is configured to generate a code excited excitation signal 1012, as described for signal generator 558. A combiner 1050 of decoder 1000 is configured to combine the code excited signal 1012 and the signal with characteristics of Noise 1022, as described for combiner 550, to obtain a combined drive signal 1052. Decoder 1000 comprises a synthesizer 1060 having a filter to be matched with the prediction coefficients 122, where the synthesizer is configured to filter the combined drive signal 1052 with the matched filter to obtain a non-voice decoded frame 1062. The decoder 1000 further comprises combiner 284 which combines the non-vocal decoded frame and voice frame 272 to obtain the sequence of audio signals 282. When compared to the decoder
200, the decoder 1000 comprises a second signal generator configured to provide the code 1012 excited drive signal.
<img file="MX355258B_D0049.tif" />
IMPI
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL excitation signal with noise characteristics 1022 may be, for example, the signal with noise characteristics n (n) illustrated in Fig. 2.
Audio signal sequence 282 can comprise good quality and high similarity when compared to an encoded input signal.
Other embodiments provide decoders that enhance decoder 1000 by shaping and / or amplifying the code generated (code excited) signal 1012 and / or the signal with noise characteristics 1022. Thus, the decoder 1000 may comprise a shaping processor and / or a variable amplifier arranged between signal generator 1010 and combiner 1050, between noise generator 1020 and combiner 1050, respectively. The input signal 1002 may comprise information relative to the code gain parameter information g<sub>c</sub> and / or the noise gain parameter information, where the decoder can be configured in order to adapt an amplifier to amplify the excitation signal generated by code 1012 or a conformed version thereof, using the gain parameter information code g<sub>c</sub>. Alternatively or additionally, the decoder 1000 may be configured to adapt, i.e. to control an amplifier in order to amplify the signal with noise characteristics 1022 or a conformed version thereof, with an amplifier, using the gain parameter information of noise.
Alternatively, decoder 1000 may comprise a modeler 1070 configured to shape the code excited excitation signal 1012 and / or a modeler 1080 configured to shape the signal with noise characteristics 1022, as indicated by the dotted lines. 1070 and / or 1080 modelers can receive the gain parameters g<sub>c</sub> I g<sub>n</sub> and / or information
<img file="MX355258B_D0050.tif" />
voice-related conformation. The 1070 and / or 1080 modelers can be formed as described for the 250, 350c and / or 550b modelers described above.
Decoder 1000 may comprise a formant information calculator 1090 to provide voice related conformation information 1092 for modelers 1070 and / or 1080, as described for formant information calculator 160. Formant information calculator 1090 can be configured to provide different voice related conformation information (1092a; 1092b) to 1070 and / or modelers
1080.
Fig. 11a shows a schematic block diagram of a modeler 250 'that implements an alternative structure when compared to modeler 250. Modeler 250' comprises a combiner 257 to combine the shaping information 222 and the gain parameter related to noise g<sub>n</sub>, in order to obtain a combined information 259. A modified shaping processor 252 'is configured to shape the signal with noise characteristics n (n), using the combined information 259 to obtain the signal with amplified shaped noise characteristics 258. Since both the conformation information 222 and the gain parameter g<sub>n</sub> They can be interpreted as multiplication factors, both multiplication factors can be multiplied using combiner 257 and then applied in combination to the signal with noise characteristics n (n).
Fig. 11b shows a schematic block diagram of a modeler 250 that implements an additional alternative when compared to modeler 250. When compared to modeler 250, variable amplifier 254 is first arranged and configured to generate a signal with amplified noise characteristics, by amplifying the signal with
IMPI nmmjTo mexican 0 £ LA PROPIEDAD INDUSTRIAL noise characteristics n (n) using the gain parameter g<sub>n</sub>. Shaper processor 252 is configured to shape the amplified signal using the shaping information 222 to obtain the amplified shaped signal 258.
YES well the Flgs. 11a and 11b refer to modeler 250 and illustrate
Alternative implementations, the preceding descriptions also apply to 350c, 550b, 1070 and / or 1080 modelers.
Fig. 12 shows a schematic flow diagram of a 1200 method for encoding an audio signal according to the first aspect. The method
1210 it comprises deriving prediction coefficients and a residual signal from an audlo signal box. The method 1200 comprises a step 1230 in which a gain parameter is calculated from a residual non-vocal signal and the spectral shaping information and a step 1240 in which an output signal is formed on the basis of relative information to a voice signal box, the gain parameter or a quantized gain parameter and the prediction coefficients.
Fig. 13 shows a schematic flow diagram of a method 1300 for decoding a received audio signal comprising prediction coefficients and a gain parameter, according to the first aspect. Method 1300 comprises a step 1310 in which speech related spectral shaping information is calculated from the prediction coefficients. In a step 1320, a signal with decoder noise characteristics is generated. In a step 1330, a spectrum of the signal with decoding noise characteristics or an amplified representation thereof is formed, using the spectral shaping information to obtain a signal with conforming decoding noise characteristics. In a step 1340 of method 1300, a synthesized signal is synthesized from the signal
<img file="MX355258B_D0051.tif" />
<img file="MX355258B_D0052.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY coding with conformed noise characteristics and amplified prediction coefficients.
Fig. 14 shows a schematic flow diagram of a method 1400 for encoding an audio signal according to the second aspect. The method 1400 comprises a step 1410 in which the prediction coefficients and a residual signal are derived from a non-vocal frame of the audio signal. In a step 1420 of method 1400, a first gain parameter information is calculated to define a first drive signal relative to a deterministic codebook and a second gain parameter information to define a second drive signal relative to a signal with noise characteristics for the non-vocal frame.
In a step 1430 of method 1400 an output signal is formed on the basis of information relating to a voice signal box, the first gain parameter information and the second gain parameter information.
Fig. 15 shows a schematic flow diagram of a method 1500 for decoding a received audio signal, according to the second aspect. The received audio signal comprises information regarding the prediction coefficients. Method 1500 comprises a step 1510 in which a first drive signal is generated from a deterministic codebook for a portion of a synthesized signal. In a step 1520 of method 1500, a second drive signal is generated from a signal with noise characteristics for the portion of the synthesized signal. In a step 1530 of method 1000, the first drive signal and the second drive signal are combined to generate a combined drive signal for the portion of the synthesized signal. In a step 1540 of method 1500, the portion of
<img file="MX355258B_D0053.tif" />
IMPI
THE MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY the signal synthesized from the combined excitation signal and the prediction coefficients. '<sup>1,1 1</sup>
In other words, the aspects of the present invention propose a new way of encoding non-vocal frames by means of the conformation of a randomly generated Gaussian noise and spectrally conforming it by adding to it a formant structure and a spectral inclination. Spectral shaping is done in the excitation domain before the synthesis filter is excited. As a consequence, the shaped excitation will be updated in the memory of the long-term forecast, to generate subsequent adaptive codebooks.
Subsequent frames, which are not non-vocal, will also benefit from spectral shaping. Unlike the formative improvement in postfiltration, the proposed noise shaping is carried out on both sides, the encoder and the decoder.
Such kind of excitation can be used directly in a parametric coding scheme in pursuit of very low bit rates. However, we also propose the association of this excitation in combination with a conventional innovative codebook within a coding scheme of
CELP.
For both methods, we propose a new gain encoding, especially effective for both clear voice and background noise. We propose some mechanisms to get as close as possible to the original energy, but at the same time avoiding too fast transitions with non-vocal frames and also avoiding unwanted instabilities due to gain quantizing gain.
The first aspect points to non-voice encoding at a rate of 2.8 and kiloblts per second (kbps). Non-vocal frames are detected first. This
<img file="MX355258B_D0054.tif" />
It can be done using a usual voice classification, as it is done in the Variable Rate Multimode Bandwidth (VMR-WB), as it is known under [3].
Doing the spectral shaping at this stage exhibits two main advantages. First, the spectral conformation is taken into account for the calculation of the excitation gain. Since the computation of gain is the only non-blind module during generation of the excitation, it is a huge advantage to have it at the end of the chain after shaping. Second, it allows you to save the enhanced excitation in the memory of the LTP. The enhancement will then also serve frames that are not subsequent non-vowels.
Although the quantizers 170, 170-1 and 170-2 were described as configured to obtain the quantized parameters g<sub>c</sub> and g „, the quantized parameters may be provided as related information, for example, an index or an identifier of a database entry, where the input comprises the quantized gain parameters g<sub>c</sub>yg<sub>n</sub>Although some aspects have been described in the context of an apparatus, it is evident that such aspects also represent a description of the corresponding method, where a block or device corresponds to a method stage or a characteristic of a method stage. Similarly, the aspects described in the context of a method step also represent a description of a corresponding block or article or feature of a corresponding apparatus.
The encoded audio signal of the invention can be stored on a digital storage medium or it can be transmitted on a
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL \ S ™ ss
<img file="MX355258B_D0055.tif" />
transmission such as a wireless transmission medium or a wired transmission medium such as the Internet. ....... '
Depending on certain requirements of the implementation, the embodiments of the invention can be implemented in hardware or in software. The implementation can be done using a digital storage medium, for example a floppy disk, a DVD, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, which has stored electronic read control signals, which cooperate (or may cooperate) with a programmable computer system, such that the respective method is implemented.
Some embodiments in accordance with the invention comprise an information carrier having electronic read control signals, which have the ability to cooperate with a programmable computer system, such that one of the methods described herein is implemented.
In general, the embodiments of the present invention can be implemented in the manner of a computer program product with a program code, the program code is operated to execute one of the methods when the computer program product is executed in a computer. The program code can be stored, for example, in a machine-readable carrier.
Other embodiments comprise the computer program for executing one of the methods described herein, which is stored in a possible machine-readable carrier.
In other words, an embodiment of the method of the invention is, therefore, a computer program that has a program code to execute one of the methods described herein, when the computer program is executed on a computer. .
<img file="MX355258B_D0056.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX355258B_D0057.tif" />
A further embodiment of the methods of the invention is, therefore, an Information carrier (or a digital storage medium or a computer-readable medium) comprising, self-recorded, the computer program for executing one of the methods described herein.
A further embodiment of the method of the invention is, therefore, a data stream or a sequence of signals representing the Computer program for executing one of the methods described herein. The data stream or signal sequence can be configured, for example, to be transferred over a data communication connection, for example over the Internet.
A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured or adapted to execute one of the methods described herein.
A further embodiment comprises a computer that has the computer program installed to execute one of the methods described herein.
In some embodiments, a programmable logic device (for example a programmable field gate network) may be used to execute some or all of the functionalities of the methods described herein. In some embodiments, a programmable field gate network can cooperate with a microprocessor in order to execute one of the methods described herein. In general, the methods are preferably practiced with any hardware equipment.
The embodiments described above are merely illustrative of the principles of the present invention. It is understood that the modifications and
<img file="MX355258B_D0058.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX355258B_D0059.tif" />
Variations of the arrangements and details described herein will be obvious to those skilled in the art. Therefore, they are intended to be limited only by the scope of the claims of the present patent and not by the specific details presented by way of description and explanation of the embodiments of the present.
<img file="MX355258B_D0060.tif" />
Bibliography [1] ITU-T Recommendation G.718: “Frame error robust narrow-band and wideband embedded variable bit-rate coding of speech and audio from 8-32 kbit / s” [2] US Patent Number US 5,444. 816, “Dynamic codebook for efficient speech coding based on algébrale codes” [3] Jelinek, M.¡ Salami, R., Wideband Speech Coding Advances in VMRWB Standard, Audio, Speech, and Language Processing, IEEE Transactions on, vol.15 , No. 4, pp. 1167,1179, May 2007
Contents54
76 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76
81 members in 19 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 13189392 | European Patent Office (EPO) | A | |
| 13189392 | European Patent Office (EPO) | A | |
| 131893927 | European Patent Office (EPO) | – | |
| 14178785 | European Patent Office (EPO) | A | |
| 14178785 | European Patent Office (EPO) | A | |
| 141787853 | European Patent Office (EPO) | – | |
| 2014071769 | European Patent Office (EPO) | W | |
| 2014071769 | European Patent Office (EPO) | W | |
| 131893927 | – | – | – |
| 141787853 | – | – | – |
| EP20130189392 | – | – | – |
| EP20140178785 | – | – | – |
| PCTEP2014071769 | – | – | – |
| WO2014EP71769 | – | – | – |
Members81
| Document | Office | Kind | |
|---|---|---|---|
| CA2927716A1 | Canada | A1 | |
| CA2927722A1 | Canada | A1 | |
| WO2015055531A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015055532A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201523588A | Taiwan Province of China | A | |
| TW201528255A | Taiwan Province of China | A | |
| AR098072A1 | Argentina | A1 | |
| AR098073A1 | Argentina | A1 | |
| AU2014336356A1 | Australia | A1 | |
| AU2014336357A1 | Australia | A1 | |
| SG11201603000SA | Singapore | A | |
| SG11201603041YA | Singapore | A | |
| KR20160070147A | Republic of Korea | A | |
| KR20160073398A | Republic of Korea | A | |
| CN105723456A | China | A | |
| CN105745705A | China | A | |
| MX2016004922A | Mexico | A | |
| MX2016004923A | Mexico | A | |
| US2016232908A1 | United States of America | A1 | |
| US2016232909A1 | United States of America | A1 | |
| EP3058568A1 | European Patent Office (EPO) | A1 | |
| EP3058569A1 | European Patent Office (EPO) | A1 | |
| JP2016533528A | Japan | A | |
| JP2016537667A | Japan | A | |
| TWI575512B | Taiwan Province of China | B | |
| TWI576828B | Taiwan Province of China | B | |
| AU2014336356B2 | Australia | B2 | |
| AU2014336357B2 | Australia | B2 | |
| BR112016008544A2 | Brazil | A2 | |
| BR112016008662A2 | Brazil | A2 | |
| RU2016118979A | Russian Federation | A | |
| RU2016119010A | Russian Federation | A | |
| ZA201603158B | South Africa | B | |
| RU2644123C2 | Russian Federation | C2 | |
| RU2646357C2 | Russian Federation | C2 | |
| KR20180021906A | Republic of Korea | A | |
| MX355091B | Mexico | B | |
| MX355258BThis record | Mexico | B | |
| KR101849613B1 | Republic of Korea | B1 | |
| JP6366705B2 | Japan | B2 | |
| JP6366706B2 | Japan | B2 | |
| CA2927722C | Canada | C | |
| KR101931273B1 | Republic of Korea | B1 | |
| US10304470B2 | United States of America | B2 | |
| US2019228787A1 | United States of America | A1 | |
| US10373625B2 | United States of America | B2 | |
| US2019333529A1 | United States of America | A1 | |
| CN105723456B | China | B | |
| CN105745705B | China | B | |
| US10607619B2 | United States of America | B2 | |
| CN111370009A | China | A | |
| US2020219521A1 | United States of America | A1 | |
| CA2927716C | Canada | C | |
| MY180722A | Malaysia | A | |
| EP3058569B1 | European Patent Office (EPO) | B1 | |
| PT3058569T | Portugal | T | |
| EP3058568B1 | European Patent Office (EPO) | B1 | |
| US10909997B2 | United States of America | B2 | |
| EP3779982A1 | European Patent Office (EPO) | A1 | |
| PT3058568T | Portugal | T | |
| US2021098010A1 | United States of America | A1 | |
| EP3806094A1 | European Patent Office (EPO) | A1 | |
| PL3058569T3 | Poland | T3 | |
| ES2839086T3 | Spain | T3 | |
| PL3058568T3 | Poland | T3 | |
| ES2856199T3 | Spain | T3 | |
| MY187944A | Malaysia | A | |
| BR112016008544B1 | Brazil | B1 | |
| BR112016008662B1 | Brazil | B1 | |
| US11798570B2 | United States of America | B2 | |
| CN111370009B | China | B | |
| US11881228B2 | United States of America | B2 | |
| EP3779982B1 | European Patent Office (EPO) | B1 | |
| EP3779982C0 | European Patent Office (EPO) | C0 | |
| EP3806094B1 | European Patent Office (EPO) | B1 | |
| EP3806094C0 | European Patent Office (EPO) | C0 | |
| EP4632735A2 | European Patent Office (EPO) | A2 | |
| ES3042587T3 | Spain | T3 | |
| PL3779982T3 | Poland | T3 | |
| ES3044088T3 | Spain | T3 | |
| EP4632735A3 | European Patent Office (EPO) | A3 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 355258
- Publication, DOCDB
- 355258
- Publication, EPODOC
- MX355258
- Application
- 2016004922
- Application, DOCDB
- 2016004922
- Application, EPODOC
- MX20160004922
Titles2
- Spanish
- CONCEPTO PARA CODIFICAR UNA SEÑAL DE AUDIO Y DECODIFICAR UNA SEÑAL DE AUDIO USANDO INFORMACION DETERMINISTA Y DE TIPO RUIDO.
- English
- CONCEPT FOR ENCODING AN AUDIO SIGNAL AND DECODING AN AUDIO SIGNAL USING DETERMINISTIC AND NOISE LIKE INFORMATION.
Classification
- CPC, 11
- G10L19/08
- G10L19/083
- G10L19/20
- G10L19/0017
- G10L19/008
- G10L19/12
- G10L19/06
- G10L2025/932
- G10L25/15
- G10L19/07
- G10L2019/0016
- IPC, 2
- G10L19 20
- G10L19 08