Improving initial coding using duplicating band
Abstract
FIELD: radio engineering; initial coding systems. ^ SUBSTANCE: band width is reduced in front of coder or within coder followed by duplicating spectrum band in decoder. This is made by using new methods of transposition jointly with spectrum envelope. Proposed invention can be implemented in hardware or software codec, or can be used as self-contained processor in combination with codec. Improvement is ensured irrespective of type of codec and state-of-the-art. ^ EFFECT: reduced bit transfer speed at desired perception quality or enhanced perception quality at desired bit transfer speed. ^ 20 cl, 34 dwg
Term
Term ended
Expired 9 June 2018, 8.3 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
21 claims: 5 independent, 16 dependent
- 1A method for decoding an encoded signal (201, 301), wherein the encoded signal is received from the original signal and represents only a portion of frequency bands included in the original signal, the frequency content of the encoded signal (201, 301) is represented by the subband samples for a plurality of subbands or a plurality of spectral represented coefficients comprising that emit a first bandpass signal (203, 303), said first bandpass signal has a sampling sub-bands of a predetermined number of the analyzed subbands or a predetermined number of the analyzed spectral coefficients, wherein the first band signal has a bandwidth less the bandwidth encoded signal (201, 301) carried transposition (205, 305) samples of the subband of subband analysis or spectral coefficients of the analysis are included in the bandpass signal (203, 303), the second band signal (205, 305) having a frequency content that included in the original signal and that is not included in the encoded signal, and a second bandpass signal having subband synthesis or spectral coefficients of the synthesis, and for implementing transposition attached subband analysis to subbands synthesizing or combining the spectral coefficients of the analysis to the spectral coefficients of the synthesis, with the selected sampling sub-bands or spectral coefficients included in the second band signal (203, 303) before or after the transposition will adapt according to the spectral envelope using the envelope of the spectrum obtained from the original signal or the encoded signal to obtain rigged on the spectral envelope of the transposed samples subband or rigged on the envelope spectrum transposed spectral coefficients, and the combined sample subbands and rigged transposed sampling subbands or spectral coefficients and rigged transposed spectral coefficients to produce an output signal (209, 309), wherein the output signal has a frequency content including the frequency content of the encoded signal and the frequency content of the second band signal. 1. Способ декодирования кодированного сигнала (201, 301), причем кодированный сигнал получен из исходного сигнала и представляет только часть полос частот, включенных в исходный сигнал, частотное содержимое кодированного сигнала (201, 301) представлено выборками субполос для множества субполос или представлено множеством спектральных коэффициентов, заключающийся в том, что выделяют первый полосовой сигнал (203, 303), причем первый полосовой сигнал имеет выборки субполос из предварительно определенного количества анализируемых субполос или имеет предварительно определенное количество анализируемых спектральных коэффициентов, при этом первый полосовой сигнал имеет полосу частот меньшую полосы частот кодированного сигнала (201, 301), осуществляют транспозицию (205, 305) выборок субполос из субполос анализа или спектральных коэффициентов анализа, включенных в полосовой сигнал (203, 303), во второй полосовой сигнал (205, 305), имеющий частотное содержимое, которое включено в исходный сигнал и которое не включено в кодированный сигнал, причем второй полосовой сигнал имеет субполосы синтеза или спектральные коэффициенты синтеза, а при осуществлении транспозиции присоединяют субполосы анализа к субполосам синтеза или присоединяют спектральные коэффициенты анализа к спектральным коэффициентам синтеза, при этом выбранные выборки субполос или спектральные коэффициенты, включенные во второй полосовой сигнал (203, 303), перед или после осуществления транспозиции подстраивают по огибающей спектра с использованием информации огибающей спектра, полученной из исходного сигнала или кодированного сигнала, для получения подстроенных по огибающей спектра транспонированных выборок субполос или подстроенных по огибающей спектра транспонированных спектральных коэффициентов, и объединяют выборки субполос и подстроенные транспонированные выборки субполос или спектральные коэффициенты и подстроенные транспонированные спектральные коэффициенты для получения выходного сигнала (209, 309), причем выходной сигнал имеет частотное содержимое, включающее в себя частотное содержимое кодированного сигнала и частотное содержимое второго полосового сигнала.
- 17A method of providing a transposed signal which is transposed a factor M, in the input signal from which rejected at least one frequency band, comprising the steps that an input signal are filtered by using parallel sets of L filter with the impulse response 17. Способ обеспечения транспонированного сигнала, который транспонирован с коэффициентом М, из входного сигнала, из которого отброшена по меньшей мере одна полоса частот, заключающийся в том, что осуществляют фильтрацию входного сигнала с использованием параллельного набора из L фильтров с импульсными характеристиками вида where k = 0,1, ..., L-1, K - constant p0 (n) - type low pass filter of length N, forming a set of L complex signals is carried out down-sampling said set of coefficient signals L L / M to generate a set of L complex signals subband, multiplies the phase angles of said set of L complex signals subbands to M to form a new set of signal subbands implement selection of the real parts of said new set signal subband to generate a set of L valid signal subband, carried oversampling said set of L real subband signals by a factor of L 'to generate the set of valid signals are filtered said set of valid signals via parallel sets of L' filters with impulse responses form где k=0,1,...,L-1, К - константа, р0(n) - модель фильтра нижних частот длины N, формирующего набор из L комплексных сигналов, осуществляют дискретизацию с пониженной частотой упомянутого набора из L сигналов с коэффициентом L/M для формирования набора из L комплексных сигналов субполос, осуществляют умножение фазовых углов упомянутого набора из L комплексных сигналов субполос на М для формирования нового набора сигналов субполос, осуществляют выделение действительных частей упомянутого нового набора сигналов субполос для формирования набора из L действительных сигналов субполос, осуществляют дискретизацию с повышенной частотой упомянутого набора из L действительных сигналов субполос с коэффициентом L’ для формирования набора действительных сигналов, осуществляют фильтрацию упомянутого набора действительных сигналов посредством параллельного набора из L’ фильтров с импульсными откликами вида where k = 0,1, ..., L'-1, K - constant p'0 (n) - type low pass filter of length N ', forming a set of L' filtered signals, and summing said set carried out L 'filtered signals and the input signal to generate a transposed signal. где k=0,1,...,L’-1, К’ - константа, р’0(n) - модель фильтра нижних частот длины N’, формирующего набор из L’ отфильтрованных сигналов, и осуществляют суммирование упомянутого набора из L’ отфильтрованных сигналов и входного сигнала для формирования транспонированного сигнала.
- 19A decoder for decoding an encoded signal (201, 301), wherein the encoded signal is received from the original signal or the encoded signal represents only a portion of frequency bands included in the original signal, the frequency content of the encoded signal (201, 301) is represented by the subband samples for a plurality of subbands or represented by a set of spectral coefficients containing the device isolation to isolate a first band signal (203, 303), said first bandpass signal has a sampling sub-bands of a predetermined number of subbands analysis or has a predetermined number of spectral coefficients of the analysis, the first band signal has a frequency band less than the frequency content of the encoded signal (201, 301), the device transposition for the transposition (205, 305) of selected samples subband of the subband analysis or spectral coefficients of the analysis are included in the first band signal (203, 303), the second band signal (205 , 305) having a frequency content that is included in the original signal and that is not included in the encoded signal, and a second bandpass signal having subband synthesis or spectral coefficients of the synthesis and transposition includes joining subband analysis to subbands synthesis or attachment of spectral coefficients Analysis for spectral coefficients of the synthesis, with selected samples of subbands or spectral coefficients included in the second band signal (203, 303) before or after the transposition will adapt (207, 307) on the spectrum envelope using the envelope of the spectrum obtained from the original signal or the encoded signal for rigged by the spectral envelope of the transposed samples of sub-bands, or rigged by the spectral envelope of the transposed spectral coefficients and device combination for combining samples of sub-bands and rigged transposed samples subbands or spectral coefficients and rigged transposed spectral coefficients to produce an output signal (209, 309), and output signal has a frequency content including the frequency content of the encoded signal and the frequency content of the second band signal. 19. Декодер для декодирования кодированного сигнала (201, 301), причем кодированный сигнал получен из исходного сигнала или кодированного сигнала и представляет только часть полос частот, включенных в исходный сигнал, частотное содержимое кодированного сигнала (201, 301) представлено выборками субполос для множества субполос или представлено множеством спектральных коэффициентов, содержащий устройство выделения для выделения первого полосового сигнала (203, 303), причем первый полосовой сигнал имеет выборки субполос из предварительно определенного количества субполос анализа или имеет предварительно определенное количество спектральных коэффициентов анализа, при этом первый полосовой сигнал имеет полосу частот меньшую, чем частотное содержимое кодированного сигнала (201, 301), устройство транспозиции для транспозиции (205, 305) выбранных выборок субполос из субполос анализа или спектральных коэффициентов анализа, включенных в первый полосовой сигнал (203, 303), во второй полосовой сигнал (205, 305), имеющий частотное содержимое, которое включено в исходный сигнал и которое не включено в кодированный сигнал, причем второй полосовой сигнал имеет субполосы синтеза или спектральные коэффициенты синтеза, а транспозиция включает в себя присоединение субполос анализа к субполосам синтеза или присоединение спектральных коэффициентов анализа к спектральным коэффициентам синтеза, при этом выбранные выборки субполос или спектральные коэффициенты, включенные во второй полосовой сигнал (203, 303), перед или после выполнения транспозиции подстраивают (207, 307) по огибающей спектра с использованием информации огибающей спектра, полученной из исходного сигнала или кодированного сигнала, для получения подстроенных по огибающей спектра транспонированных выборок субполос или подстроенных по огибающей спектра транспонированных спектральных коэффициентов и устройство объединения для объединения выборок субполос и подстроенных транспонированных выборок субполос или спектральных коэффициентов и подстроенных транспонированных спектральных коэффициентов для получения выходного сигнала (209, 309), причем выходной сигнал имеет частотное содержимое, включающее в себя частотное содержимое кодированного сигнала и частотное содержимое второго полосового сигнала.
- 20An apparatus for the transposed signal which is transposed a factor M, in the input signal from which rejected at least one frequency band, comprising:a filter for filtering the input signal using a parallel set of L filters with impulse responses form 20. Устройство для обеспечения транспонированного сигнала, который транспонирован с коэффициентом М, из входного сигнала, из которого отброшена по меньшей мере одна полоса частот, содержащее фильтр для фильтрации входного сигнала с использованием параллельного набора из L фильтров с импульсными откликами вида where k = 0,1, ..., L-1, K - constant;p0 (n) - a model low-pass filter length N;M - factor, forming a set of L complex signals, the apparatus downsampling for downsampling said set of L signal by a factor L / M to generate a set of L complex signals subband multiplier for multiplying a phase angle of said set of complex signals subband to M to form a new set of signals subband assignment apparatus for allocation of the real parts of said new set signal subband to generate a set of L actual signals subband device oversampling frequency for oversampling said set of L valid signal subband by a factor L ' to generate the set of valid signals, a filter for filtering said set of valid signals via parallel sets of L 'filters with impulse responses form где k=0,1,...,L-1, К – константа;р0(n) - модель фильтра нижних частот длины N;М - коэффициент, формирующий набор из L комплексных сигналов, устройство дискретизации с пониженной частотой для дискретизации с пониженной частотой упомянутого набора из L сигналов с коэффициентом L/M для формирования набора из L комплексных сигналов субполос, умножитель для умножения фазовых углов упомянутого набора комплексных сигналов субполос на М для формирования нового набора сигналов субполос, устройство выделения для выделения действительных частей упомянутого нового набора сигналов субполос для формирования набора из L действительных сигналов субполос, устройство дискретизации с повышенной частотой для дискретизации с повышенной частотой упомянутого набора из L действительных сигналов субполос с коэффициентом L’ для формирования набора действительных сигналов, фильтр для фильтрации упомянутого набора действительных сигналов посредством параллельного набора из L’ фильтров с импульсными откликами вида where k = 0,1, ..., L'-1, K - constant, p'0 (n) - type low pass filter of length N ', forming a set of L' filtered signals, and an adder for summing said set from L 'signal and the filtered input signal to generate a transposed signal. где k=0,1,...,L’-1, К’ - константа, p’0(n) - модель фильтра нижних частот длины N’, формирующего набор из L’ отфильтрованных сигналов, и сумматор для суммирования упомянутого набора из L’ отфильтрованных сигналов и входного сигнала для формирования транспонированного сигнала. Priorities: Приоритеты:
Independent claims5
206 paragraphs in 5 sections, as filed
TECHNICAL FIELD
In systems coding digital source data is compressed before transmission or recording, to reduce the required bit rate or memory size. The present invention relates to a novel method and apparatus for improvement of coding systems by duplicating the original spectral band (DSP). A significant reduction of data transmission rate without degrading the perceptual quality or conversely, improvement in perceptual quality is achieved at a predetermined rate. This is achieved by reducing the bandwidth of the spectrum on the encoding side and subsequent spectral band overlap in the decoder, i.e. The invention exploits new concepts of signal redundancy in the spectral domain.
BACKGROUND
Methods sound source coding can be divided into two classes: natural audio coding and speech coding. Natural audio coding is widely used for music or arbitrary signals at medium speed of data transmission and, in principle, provides a wide band audio frequencies. Speech coders are basically limited to speech reproduction, but, on the other hand, can be used at very low bitrates, albeit with a narrow band of audio frequencies. A wideband speech signal provides a very significant increase in quality compared to narrow band speech signal. Expanding the bandwidth not only improves intelligibility and naturalness of speech, but also facilitates speaker recognition. Wideband speech coding is an important problem facing the next generation telephone systems. Furthermore, due to the rise of multimedia applications, transmission of music and other non-speech signals in telephone systems is a desirable quality.
Linear signal with pulse code modulation (PCM), characterized by high reliability, is not effective for the transmission rate depending on the perceptual entropy. Standard compact discs (CD) requires the sampling rate of 44.1 kHz, 16 bit resolution per sample and stereo. This corresponds to a transmission rate of 1411 kbit / s. To substantially reduce the transmission rate of the original encoding may be performed using a perceptual audio codec with splitting the spectrum. These natural audio codecs use perceptual irrelevancy and statistical redundancy in the signal. By using the best codec technology can be achieved by reduction of the data volume by about 90% for the standard CD-signal format without any degradation of the intelligibility. It is thus possible very high quality stereo sound at a rate of about 96 kbit / s, i.e. the compression ratio is about 15: 1. Some perceptual codecs provide even higher compression ratios. To achieve this, it is generally necessary to reduce the sampling rate and thus bandwidth of audio frequencies. It is common to reducing the number of quantization levels that allows random sound distortion due to quantization and the use of the stereo field degradation due to intensive coding. The widespread use of such methods result in poor perception. Existing technology codec itself is almost exhausted, and further progress in the preparation of coding gain is not expected. To further improve the performance of the encoding requires a new approach.
Human speech and the majority of musical instruments form a quasi-stationary signals received at the output of generation systems. According to Fourier theory, any periodic signal may be expressed as a sum of sinusoidal signals with frequencies f, 2f, 3f, 4f, 5f etc. where f - the fundamental frequency. These frequencies form a harmonic sequence. Limiting the bandwidth of the signal equivalent to the truncated sequences f harmonics. Such a truncation alters the perceived timbre, tone color of a musical instrument or voice and gives the audio signal that will sound "muffled" or "monotonous", and intelligibility may be reduced. High frequencies, so important to the subjective sensation of sound quality.
Methods known from the prior art, mainly intended to improve the codec performance and in particular intended for High Frequency Regeneration (RVCH), which is a problem in encoding the speech signal. Such methods employ broadband linear frequency shifts, nonlinearity or aliasing (US Patent No. 5.127.054), resulting in the generation of intermodulation products or other non-harmonic frequency components that create a strong dissonance when applied to music signals. Such dissonance is described in the speech coding literature as "sharp" and "rough" sounding. Other methods for synthesizing a speech signal generated sinusoidal harmonics that are based on fundamental pitch estimation and thus limited tonal stationary sounds (U.S. Patent 4.771.465). Such methods are known from the prior art, while useful for low-quality speech applications, do not applicable for high quality speech or music signals. Several methods are aimed at improving the characteristics of high quality codec sound sources. One uses synthetic noise signals generated at the decoder to substitute noise-like signals in speech or music previously excluded by the encoder (see. "Improving Audio Codecs by Noise Substitution" D.Schultz, JAES, Vol.44, number 7/8, 1996). This is performed within the high frequency band, otherwise normally transmitted on an intermittent basis when noise. Another method recreates some lost high-frequency harmonics that were lost in the coding process (see. "Audio Spectral Coder" AJS Ferreira, AES Preprint 4201, 100th Convention, May 11-14 1996, Copenhagen), and also depends on the tone and the detection pitch. Both methods are based on the low duty cycle, providing comparatively limited coding gain or efficiency.
SUMMARY OF THE INVENTION
The present invention provides a novel method and apparatus for substantial improvements of digital source coding systems and more particularly to improvements in audio codecs. The invention reduces the data rate or to improve the perceptual quality, or to realize a combination of these properties. The invention is based on new methods of using the harmonic redundancy, offering the opportunity to drop the signal bandwidth before transmission or recording. It does not feel the deterioration of perception, if the decoder performs high quality repetition (redundancy) of the spectrum according to the invention. The discarded bits represent the coding gain at a fixed perceptual quality. Alternatively, more bits can be allocated for encoding the lowband information at a fixed rate, thus achieving higher quality perception.
The present invention postulates that a truncated harmonic sequence can be extended based on the direct relation between the spectral components of lowband and highband. This extended sequence similar to the original in terms of perception, if certain rules. First, the extrapolated spectral components must be harmonically related sequence and the truncated harmonic dissonance to avoid distortion. The present invention uses transposition as a means for the procedure of spectral overlap, which guarantees the satisfaction of this criterion. However, successful operation is not necessary that the spectral components lowband harmonics sequence formed as new redundant components, harmonically related components lowband, will not change the noise-like or non-stationary nature of the signal. Transposition is defined as the transfer of partial tones from one position on a musical scale to another while maintaining the frequency ratios for these partial tones. Second, the spectral envelope, i.e. rough spectrum allocation duplicated high frequency band should be good enough to repeat such a distribution of the original signal. The present invention provides two modes of operation, the DSP-1 and DSP-2, which differ in the way the spectral envelope adjustment.
The first mode overlapping spectral bands (DSP-1) adapted to improve the average quality codec applications, is a single-channel process, which uses only information contained in the received signal in the low band decoder. The spectral envelope of this signal is determined and extrapolated, for instance using polynomials together with a set of rules or a code book. This information is used to continuously adjust and equalize the duplicated highband. DSP-1 method offers the advantage of post-processing, ie, No modifications on the encoding side. The owner of a radio transmitting station will receive a prize in the use of channels, or will be able to improve the quality of perception, or provide a combination of these qualities. Existing standard syntax and data stream can be used without modification.
Mode DSP-2, intended for the improvement of high quality codec applications, is a two-channel process in which, in addition to the transmitted lowband signal according to the mode of the DSP-1 is encoded and transmitted spectral envelope of highband. Since changes in the envelope of the spectrum are much lower speed than the change of the high band signal, it is required to transmit only a limited amount of information in order to successfully represent the spectral envelope. DSP-2 mode can be used to improve the efficiency of existing technologies codec with minimal or without changing the existing syntax or protocols, and as a very valuable tool for the development of future codecs.
Modes DSP-1 and DSP-2 can be used to replicate smaller passbands lowband when such bands are eliminated by the encoder as stipulated by the psycho-acoustic model in a bit failure. This leads to improved perceptual quality by spectral overlap in the lower frequency band in addition to the spectral overlap is lowband. Furthermore, modes DSP-1 and DSP-2 can also be used in codecs employing scaling rate, where the perceptual quality of the signal at the receiver varies depending on transmission channel conditions. This usually involves changing the bandwidth of the audio signal receiver. Under these conditions, chipboard modes could be successfully used to maintain a constant high band, which further improves the perceptual quality.
The present invention operates on a continuous basis, performing duplicate content of any type of signals, i.e. tonal or non-tonal (noise-like and transient signals). Furthermore, this method creates an exact duplicate of the spectrum on the perception of a copy of the discarded bands from available frequency bands at the decoder.
Consequently, the method provides substantially chipboard higher coding gain or perceptual quality improvement compared to methods known in the art. This invention can be used in conjunction with the methods of improving the codec of the prior art; However, from such combinations should not expect any increase in efficiency.
DSP method comprises the following steps:
- Encoding of a signal derived from an original signal, where frequency bands of the signal are removed, the removal is performed before or during the encoding, wherein a first signal is generated,
- Transposition of frequency bands of the first signal during or after decoding, to form a second signal,
- Performance tuning spectral envelope and
- Combining the decoded signal and the second signal to generate an output signal.
The bandwidth of the second signal may be set so as not to overlap or partly overlap with a frequency band of the first signal, and can be set depending on the temporal characteristics of the original signal and / or the first signal, or transmission channel conditions. Spectral envelope adjustment is performed based on estimation of the original spectral envelope of said first signal or the transmitted envelope information of the original signal.
The present invention comprises two main types of devices transposition: multiband devices and devices transposition transposition with forecasting with time-varying circuit search, having different properties. The basic multiband transposition may be performed according to the present invention as follows:
- Filtering the signal to be transposed through a set of N≥ 2 bandpass filters with passbands comprising the frequency (f1, ..., fn), respectively, for generating signals N passbands
- Shift signal bandwidths in the frequency domain comprising frequency M (f1, ..., fn), where M ≠ 1 is the transposition factor, and
- The union shifted signal bandwidths with the formation of the transposed signal.
Alternatively, this base mnogopolosovaya transposition may be performed in accordance with the invention as follows:
- Bandpass filtering the signal to be transposed, using a set of analysis filters or transmitter for generating low frequency signals of real or complex subband
- An arbitrary number of channels k of said set with analysis filters or the inverter connected to channels Mk, where M ≠ 1, in a set of synthesis filter or converter and
- Transposed signal is formed using a set of synthesis filter or converter.
An improved multiband transposition according to the present invention includes adjusting phase characteristics are improved multiband transposition base.
Transposition with the prediction with a time-varying search pattern according to the present invention can be performed as follows:
- Detecting transition in the first signal,
- Determining which segment of the first signal to be used when duplicating portions of the first signal depending on the outcome of the transient detection,
- Tuning the properties of the state vector and a set of codes, depending on the result of detection of the transient and
- Search for synchronization points in chosen segment of the first signal based on the synchronization point found in the previous search synchronization point.
Methods chipboard and apparatus of the present invention provide the following qualities:
1. These methods and devices use the new concepts of signal redundancy in the spectral domain.
2. The methods and signals applied to an arbitrary signal.
3. Each harmonic set is individually created and regulated.
4. All duplicate harmonics are generated in such a way as to form a continuation of the existing harmonic sequence.
5. The process of duplication of the spectrum is based on the transposition and creates no interference or minor interference.
6. Duplicating spectrum may overlap to provide a plurality of smaller bands and / or a wide frequency range.
7. The method DSP-1 processing is performed on the decoder side only, i.e. all standards and protocols can be used without modification.
8. A method DSP-2 may be used in accordance with most standards and protocols with no or minor modifications.
9. The method provides a DSP-2 codec designer a new powerful compression tool.
10. Coding provides significant gains. The most effective application relates to the improvement of various types of low-rate codecs, such as MPEG 1/2 Layer I / II / III (US Patent No. 5.040.217), MPEG 2/4 AAC, Dolby AC-2/3, NTT Twin VQ (U.S. 5.684.920), AT & T / Lucent PAC etc. This invention is also useful for high-quality speech codecs such as wide band CELP and SB-ADPCM G.722 etc. to enhance the perceptual quality. The above codecs are widely used in multimedia, in the telephone industry, on the Internet as well as in professional systems. System T-DAB (Terrestrial Digital Sound Broadcasting System) use the low-speed protocols, which provide gain in using the channels using the present method, or improve the quality of FM and AM digital broadcasting. Satellite S-DAB system can obtain a significant gain due to the high costs of using the system of the present invention to increase the number of channels multiplexed in a digital audio broadcasting. In addition, the first audio stream in real-time full range available through the Internet using a low-speed telephone modem.
BRIEF DESCRIPTION OF DRAWINGS
The present invention is explained below by examples of its implementation, not limiting the scope or spirit of the invention, with reference to the drawings, in which:
Figure 1 - schematic representation of the particle board in a coding system according to the present invention;
2 - presentation of overlapping spectrum of upper harmonics according to the present invention;
3 - range view of duplication of medium harmonics according to the present invention;
4 - block diagram of an embodiment of a time-domain transposition device according to the present invention;
5 - block diagram of a device in the working cycle transposition feedforward search pattern;
6 - a flowchart of operations for searching a synchronization point according to the present invention;
7a-7b - codesets positioning during transients according to the present invention;
8 - is a block diagram illustrating the use of multiple devices of transposition in the time domain in association with a suitable filterbank, for operation according to the present invention particleboard;
9a-9c - a block diagram showing an apparatus for the analysis and synthesis using Fourier transform for a short time interval PFKV configured to generate harmonics of order 2 of the present invention;
10a-10b - is a block diagram for one sub-band with a linear frequency shift in the device according to the present invention PFKV;
11 - the scheme for one sub-band using fazoumnozhitelya the present invention;
12 - illustration of generation of harmonics of order 3 of the present invention;
13 - illustration of generating harmonics 2nd and 3rd order of the present invention;
14 - illustration of generating non-overlapping combination of several harmonic series of the present invention;
15 - illustration of generating a combination of several alternating harmonic series of the present invention;
16 - illustration of generation of broadband linear frequency shifts;
17 - illustration of generating sub-harmonics of the present invention;
18a-18b - flowcharts perceptual codec;
19 - the basic structure of a set of filters with maximum decimation;
Figure 20 - Illustration generate harmonics of order 2 in the set of filters with maximum decimation of the present invention;
21 - a block diagram of the improved multiband transposition in a set with a maximum decimation filters for subband signals according to the present invention;
Figure 22 - a block diagram showing the improved multiband transposition in a set with a maximum decimation filters for subband signals according to the present invention;
23 - presentation of sub-bands and the scale factors for typical codec;
Figure 24 - representation of sub-bands and the envelope mode DSP-2 according to the present invention;
Figure 25 - Illustration of secure communication envelope mode DSP-2 according to the present invention;
Figure 26 - Illustration of redundant coding mode DSP-2 according to the present invention;
27 - embodiment of the method of using the codec DSP-1 according to the present invention;
28 - an embodiment of the codec method using DSP-2 according to the present invention;
29 - block-diagram "Pseudo" generator according to the present invention.
DESCRIPTION OF PREFERRED EMBODIMENTS
In describing the embodiment of a particular emphasis on the problems of natural audio source coding. However, it should be understood that the present invention is applicable to a wide range of source coding tasks differing from the tasks of encoding and decoding audio signals.
Fundamentals of transposition
Transposition as defined according to the present invention, it is the ideal method of spectral overlap and has several important advantages over the prior art, including no need of detecting a pitch is achieved equally high quality attribute for monochrome and polyphonic program material, and transposition is realized equally well for tonal and tones. In contrast to other methods of transposition according to the invention can be used in systems source coding arbitrary audio signals of arbitrary type.
Exact transposition factor M of a discrete time signal x (n) in the form of a sum of cosines with time varying amplitudes determined by the relation
<img he="11" wi="92" file="00000002.tif" img-content="undefined" img-format="tif" />
<img he="11" wi="87" file="00000003.tif" img-content="undefined" img-format="tif" />
where N - number of sinusoids is further defined as the partial tones, fi, ei (n), α i - individual input frequencies, time envelopes and phase constants respectively, β i - arbitrary output phase constants and fs - sampling frequency, and O≤ MFI ≤ fs / 2.
2 illustrates the generation of harmonics of order M, where M is an integer ≥ 2. The term "harmonic M-th order" is used for simplicity, albeit the process generates Mth harmonic order for all signals in a certain frequency range, in which most cases are themselves harmonics of unknown order. The input signal represented in the frequency domain X (f) is limited to 201 to strip the range 0 to fmax. The contents of signals in the range fmax / M to Qfmax / M, where Q is the desired bandwidth expansion factor 1 <Q≤ M is allocated through the band pass filter to form a bandpass signal with spectrum 203 CWS (f). This bandpass signal is transposed a factor M, forming a second band signal with a spectrum of 205 Am (f), covering the range from fmax to Qfmax. The envelope of the spectrum of this signal is adjusted by means of software with an escape-equalizer, forming a signal with spectrum XE 207 (f). This signal is then combined with a delayed version of the input signal to compensate for the delay caused by the bandpass filter and the transposition device, whereby an output signal 209 with a spectrum Y (f), covering the range from 0 to Qfmax. Alternatively, the bandwidth allocation may be performed after the transposition M, using cut-off frequencies and fmah Ofmax. When using multiple devices transposition is possible, of course, the simultaneous generation of different harmonic series. The above scheme can also be used to "fill in" stop band of the input signal, as shown in Figure 3, where the input signal has a stopband 301 from f0 to Qf0. Bandwidth [f0 / M, Qf0 / M], then released (303), transposed by a factor of M to [f0, Qf0] (305), adjusted for envelope (307) and combined with the delayed input signal forming the output signal 309 from spectrum Y (f).
Can be used approximation is the exact transposition. According to the present invention, the quality of such approximations is determined using dissonance theory. A criterion for dissonance is presented in "Tonal Consonance and Critical Bandwidth" R.Plomp, WJM Levelt JASA, Vol.38, 1965 and consists in that two partial tones are considered as dissonant if the frequency difference is within approximately 5 to 50% of the critical bandwidth of the frequency bands in which these partial tones. Critical bandwidth for a given frequency can be estimated by the relation
<img he="14" wi="85" file="00000004.tif" img-content="undefined" img-format="tif" />
with f, and cb in Hz. Furthermore, in the aforementioned paper it states that human hearing can not separate the two partial tone if they differ in frequency by an amount less than about 5 percent of the critical bandwidth in which they are located. The exact transposition in equation (2) can be approximated by a
<img he="11" wi="112" file="00000005.tif" img-content="undefined" img-format="tif" />
where f - the deviation from the exact transposition. If the input partial tones form a harmonic series, the hypothesis of the invention states that the deviations from the harmonic series transposable partial tones shall not exceed five percent of the critical bandwidth in which they are located. This could explain why the methods known from the prior art, give unsatisfactory "rough" results, since broadband linear frequency shifts create a much larger deviation than acceptable. When the methods known from the prior art, is formed over a partial tone to only one input partial tones, these partial tone must nevertheless be within the established limits of deviations to be perceived as one partial tone. This again explains the poor results obtained in the methods of the prior art using nonlinearities etc, since they form a partial intermodulation tones outside the limits of deviations.
When using the above method, the duplication of the spectrum on the basis of the transposition according to the present invention achieves the following important properties.
- As there was no overlap in the frequency domain between the duplicated harmonic and existing partial tones.
- Duplicate partial tones are harmonic partial tones of the input signal and does not lead to increased dissonance or distortion.
- Spectral envelope duplicate harmonics forms a smooth continuation of the input signal spectral envelope, perceptually matching the original envelope.
Transposition based prediction changing over time scheme search
There are different ways to create the required devices transposition. Typical implementation of the time domain signal is expanded in time by duplicating signal segments based on pitch period. This signal is sequentially read at different speeds. Unfortunately, such methods strongly depend on the pitch detection and require accurate time signal coupling segments. Furthermore, the need to work with signal segments based on pitch period makes them sensitive to transients. Since the detected pitch period can be much longer than the actual process of transition is obvious risk of duplication complete the transition process rather than simply expand it in time. Another type of algorithms in the time domain implements a temporary expansion / contraction of the speech signal using the prediction output signal search pattern (see. "Pattern Search Prediction of Speech" R.Bogner, T.Li, Proc.ICASSP'89, Vol.1, May 1989 , "Time-Scale Modification of Speech based on a nonlinear Oscillator Model" G.Kubin, WBKleijn, IEEE, 1994). This is a form of granular synthesis, where the input signal is divided into small parts, granules, used to synthesise the output signal. This synthesis is usually done by performing correlation of signal segments in order to identify the best point of docking. This means that the segments used to form the output signal are not dependent on the pitch period, and thus not required to solve a non-trivial task of pitch detection ground. However, in these methods there are problems with rapidly changing signal amplitudes, and the need for high transposition growing requirements for computing. The invention provides an improved apparatus and shifting the pitch of transposition in the time domain, where the use of transient detection and dynamic system parameters produces a more accurate transposition for high transposition factors for both stationary (tonal and non-tonal) and transient sounds at low computational cost.
4 illustrates the following modules: transient detector 401, headlight window 403, a set generator 405, synchronization signal selector 407, the synchronization position memory 409, a minimum difference estimator 411, an output segment memory 413, the mixing unit 415 and the sampling device 417. The low frequency signal is supplied as an input to the generator set 405, and the transient detector 401. If transient is detected, information about its position is sent to the window position module 403. This module sets the size and position of the window that is multiplied by the input signal to create a set of codes. Code set generator 495 crush position data synchronization from the synchronization data allocation module 407, provided it is connected to another device transposition. If position data synchronization codes available in a set, they are used and an output segment is produced. Otherwise, the set of codes is sent to the minimum difference estimator 411 which produces a new synchronization position. The new output segment is attached to the window, along with the previous output segment in the mix module 415 and then sampled in the module 417.
To illustrate the idea entered the field conditions. Here, the state vectors, or granules are the input and output signals. The input signal is represented by the state vector x (n):
<img he="5" wi="137" file="00000006.tif" img-content="undefined" img-format="tif" />
which is obtained from N delayed samples of the input signal, where N - the dimension of the state vector, a D - the delay between the input samples used to construct the vector. Granular reflection gives sample x (n) corresponding to each state vector x (n-1). As a result, we obtain the equation (6), where a (*) - Display:
<img he="5" wi="51" file="00000007.tif" img-content="undefined" img-format="tif" />
In the present process a granular mapping is used to determine the next output based on the result of the previous output result, using a set of state transition of codes. Set code length L is constantly being rebuilt, including the state vectors and the next sample following each state vector. Each state vector is separated from the next-to-sample; This allows the system to adjust the time resolution depending on the characteristics of the currently processed signal, where K equal to one represents the best resolution. The segment of the input signal is used to build a set of codes is selected based on the provisions of a possible transient and the synchronization position in the previous set of codes.
This means that the mapping a (*), theoretically, is evaluated for all transitions included in the set of codes
<img he="38" wi="94" file="00000008.tif" img-content="undefined" img-format="tif" />
C this set of code transitions in the output of the new (n) is computed state vector search in a set of codes, most similar to the current state vector y (n-1). This search for the closest neighbor is carried out by calculating the difference between the minimum and gives the new output sample
<img he="5" wi="51" file="00000009.tif" img-content="undefined" img-format="tif" />
However, the system is not limited to work on the basis of the samples, it is preferably operated on the basis of the segments. The new output segment is introduced into the box and added, mixed with the previous output segment and then sampled. Pitch transposition factor is determined by the ratio of the length of the input segment represented by a set of codes, and the output segment length read from the set of codes.
5 and 6 are block diagrams showing the cycle of operation of transposition. Stage 501 is input; at step 503, detection is made on the transition process of the input signal segment; search for transients is performed on a segment length equal to the length of the output segment. If in step 505 found the transition process, then in step 507 the position of the transient is recorded and the parameters L (representing the length of a set of codes), K (representing the distance between the state vectors in quanta) and D (representing the delay between the quanta in each state vector) mounted on step 509. The position of the transition process is compared with the position of the previous output segment in step 511 to determine whether to process this transition process. If a positive result of checking at step 513, the set position code (window L), and the parameters K, L and D are set in step 515. After setting the required parameters, based on a detection result of the transition process, there is a search of a new synchronization point or interface (step 517). This procedure is shown in Figure 6. First, in step 601, a new synchronization point is calculated according to the preceding relation
<img he="5" wi="110" file="00000010.tif" img-content="undefined" img-format="tif" />
where - there are old and new position synchronization, respectively, S - the length of the input segment being processed, and M - the coefficient of transposition. Synchronization point is used to compare the accuracy of the new junction point with the accuracy of the old junction point in step 603. If step 605 determined that the matching is equal to or better than the previous one, this new synchronization point is issued at step 607, provided that it is within the a set of codes. If not, then a new search is performed in a cycle synchronization point 609. This is performed in a similar manner, in this case a minimum difference function (611), however, also possible to use correlation in the time or frequency domain. If at step 613 it is determined that this position gives a better match than the previous position found, the synchronization position is stored at step 615. When all positions are checked (step 617), the system returns (619) according to the procedure of the flowchart of Figure 5 . The new synchronization point obtained is stored at step 519 and a new segment is read from the code set in step 521, starting with the given synchronization point. This segment is added to the window and added to the previous step 523, quantized coefficient transposition in step 525 and stored in the output buffer at step 527.
<IMG>
<IMG>
7 illustrates the operation of the system during the transition process, flattening into account the position of the set of codes. Before the transition process code set 1 representing the input segment 1 is set to "left" of the segment 1. Correlation segment 1 represents a part of the previous result and the output turns to find synchronization point 1 in a set of codes 1. When the transient is detected and the point of the transition process is treated, code set of Figure 7 is moved and remains stationary until the input segment currently being processed again becomes "right" in the code set. This makes it impossible to duplicate the transient since the system is not allowed to search for synchronization point to the transition process.
<img he="3" wi="146" file="00000015.tif" img-content="undefined" img-format="tif" />
Device transposition in the time domain, as explained above, the systems used to implement the DSP-1 and DSP-2 according to the following examples, illustrative but not limiting. 8 uses three expansion unit time to generate a second harmonic, third and fourth order. As in this example, each time domain expansion / transposition device operates using a wideband signal, it is advantageous to adjust the spectral envelope of the source frequency range prior to transposition, considering that there is no means to accomplish this after the transpositions, without adding a separate equalizer system. Spectral envelope adjusters 801, 803 and 805 each operates on several filterbank channels. Gain of each channel in the envelope adjusters must be set so that the sum, 813, 815, 817 at the output, after transposition, would give the desired spectral envelope. Transposition device 807, 809 and 811 are interconnected to share information on the status of the synchronization data. This is based on the fact that under certain conditions there will be a high correlation between the synchronization positions found in the code set during correlation in the separate blocks transposition. Propose, as an example, without any limitations on the scope of the invention that the device of the fourth order harmonic transposition works based on a time interval equal to half the interval of the transposition device the second order harmonics, but with a duty cycle twice as large. Assume further that the code sets are used for the two expansion units are the same and that the synchronization positions of the two extension units in the time domain as denoted respectively. This gives the following relationship:
<img he="4" wi="128" file="00000016.tif" img-content="undefined" img-format="tif" />
<IMG>
<IMG>
Where
<IMG>
a S - is the length of the input segment represented by a set of codes. This is valid as long as neither of the synchronization position pointers reaches the end of dialing codes. During normal operation n is increased by one for each time-frame processed by the second order harmonic transposition, and when the end inevitably is reached by any code set of the pointers, the counter n is set to n = 0, and u are calculated individually. Similar results are obtained for the device of the third order harmonic transposition when connecting to a device of the fourth order harmonic transposition.
<IMG>
<IMG>
The above use of several interconnected devices transposition in the time domain to generate higher order harmonics results in a substantial reduction of the calculation amount. Furthermore, the proposed use of devices of transposition in the time domain in conjunction with an appropriate set of filters provides the ability to adjust the envelope of the created spectrum while maintaining the simplicity and low computational cost devices transposition in the time domain, since these devices are more or less, may be implemented using arithmetic Fixed-point and only the operations of addition / subtraction.
Other, illustrative but not limiting, examples of the present invention are:
- Use of the device in the time domain transposition in each subband in the set of subband filter bank, thus reducing the complexity of the signal for each device transposition;
<img he="12" wi="92" file="00000019.tif" img-content="undefined" img-format="tif" />
- Use of the device in the time domain transposition in wideband speech codec, operating, for example, the residual signal obtained after linear extrapolation.
<img he="12" wi="107" file="00000020.tif" img-content="undefined" img-format="tif" />
Transposition based on a set of filters
<img he="12" wi="72" file="00000021.tif" img-content="undefined" img-format="tif" />
KVPF for N discrete points in time signal x (n) is defined by
<IMG>
where k = 0, 1, ..., N-1 and ω k = 2π k / N and h (n) is a window. If the window satisfies the following conditions:
<img he="21" wi="119" file="00000022.tif" img-content="undefined" img-format="tif" />
There is an inverse transform, and it is given by the equation
<IMG>
Direct conversion can be interpreted as an analyzer, see. 9a, consisting of a set of N bandpass filters with pulse signal h (n) exp (jω kn) 901 followed by a set of N multipliers with carriers exp (-jω kn) with 903 shear band signals in the region around 0 Hz, forming the N signals Xk (n) analysis. This window acts like a low pass filter. Xk (n) have small bandwidth and down-sampled (block 905). Equation (12) thus only evaluated at n = rR, where R - is the decimation factor and r - new time variable. Xk (n) can be recovered from Xk (rR) by oversampling, see figure 9b, i.e. entering zeros (block 907) after a low pass filter filtering 909. The inverse transform may be interpreted as a synthesizer consisting of a set of N multipliers with carriers 911 (1 / N exp (jω Kn), which shifts the signals Xk (n) up to the original frequency followed by spending 913 (9c) which is added components yk (n) from all channels. KVPF converse KVPF (OKVPF) can be rearranged to use Discrete Fourier Transform (DFT) and inverse DFT (IDFT), which allows the use of fast Fourier transform (FFT) (see. "Implementation of the Phase Vocoder using the Fast Fourier Transform" MRPortnoff, IEEE ASSP, Vol.24, No.3, 1976).
9c shows the connection 915 for generating second harmonics, M = 2, with N = 32. For simplicity shown only channels 0 through 16. The central frequency band 16 is equal to the Nyquist frequency, channels 17 through 31 correspond to negative frequencies. Blocks marked P 917, and the gain blocks 919 will be described later, and is now to be regarded as abbreviations. The input signal is in this example bandlimited so that only channels 0 through 7 contain signals. Channels analyzer 8 to 16, thus empty and need not be displayed in the synthesizer. Analyzer channels 0 through 7 are connected to synthesizer channels 0 through 7, corresponding to an input signal delay path. Analysis channels k, where 4≤ k≤ 7 are connected to synthesis channels Mk, M = 2, which shift the signals to frequency domain with the center frequency of bandpass filters relative to two-fold k. Thus, the signals are shifted up to their original ranges as well as transposed one octave up. To explore the harmonic generation in terms of the actual output response of the filter and modulator, should also be considered negative frequencies, see the lower branch 10a. Therefore, the combined output of the inverse transform corresponds to the mapping k → Mk 1001 and Nk → N-Mk 1003 where 4≤ k≤ 7.
This gives
<IMG>
wherein M = 2. Equation (15) can be interpreted as a band-pass filtering the input signal, followed by a linear frequency shift or Upper sideband modulation, i.e. single sideband modulation using the upper side band (see. 10b), where 1005 and 1007 form a Hilbert transformer, 1009 and 1011 are multipliers with cosine m sinusoidal carrier, and 1013 - a cascade of differentiation, which highlights the upper sideband. Clearly, such a multiband method bandpass filtering one sideband can be applied explicitly, i.e. without binding the filter set in the time or frequency domain, allowing arbitrary selection of individual implement passbands and oscillator frequencies.
According to equation (15) is a sinusoid with the frequency ω i in the frequency analysis channel k yields a harmonic at the frequency Mω k + (ω i-ω k). Hence this method is called basic multiband transposition, only generates exact harmonics for input signals with frequencies ω i = ω k, where 4≤ k≤ 7. However, if the number of filters is sufficiently large, the deviation from an exact transposition slightly (see. Eq (4 )). Furthermore, the transposition is performed exactly for quasi-stationary tonal signals of arbitrary frequencies by inserting the blocks denoted P 917 (9c), provided every analysis channel contains maximum one partial tone. In this case, Xk (rR) are complex exponentials with frequencies equal to the differences between the partial frequencies ω i and tones center frequencies ω k analysis filterbank. To obtain the exact transposition frequencies must be increased by a factor M, modifying the above frequency ratio to the mean ω i → Mω k + M (ω i-ω k) = Mω i. Frequencies Xk (rR) are equal to the time derivatives of their respective deployed phase angles and may be estimated using first order differences of successive phase angles. Frequency estimates are multiplied by M and synthesis phase angles are calculated using those new frequencies. However, the same result, except for the phase constant, is obtained in a simplified manner, by multiplying the analysis arguments M directly, eliminating the need for frequency estimation. This is described in Figure 11, representing the blocks 917. Thus Xk (rR), where 4≤ k≤ 7 in this example, is converted from rectangular to polar coordinates as shown in blocks R → P (1101). The arguments are multiplied by M = 2 (block 1103), and the amplitude does not change. The signals are then converted back to rectangular coordinates (P → R) in the block 1105 forming the signals YMk (rR), and fed to synthesizer channels according to 9c. This improved multiband transposition method thus has two stages: a coarse transposition binding provides as a basic method, and fazoumnozhiteli provide accurate frequency adjustment. The above multiband transposition methods differ from traditional pitch shifting techniques using KVPF where generators are used for synthesis on the basis of the conversion tables, or when used OKVPF synthesis signal which is extended in time and thinned, i.e. binding is not used.
<img he="6" wi="76" file="00000025.tif" img-content="undefined" img-format="tif" />
It is also possible to combine amplitude and phase information from different analyzer channels. Amplitude signals [Xk (rR)] can be linked according to 16, whereas the phase signals arg {Xk (rR)} are connected according to rule 16 on. Thus, lower frequencies will still be transported, whereby the envelope of the generated periodic repetition of the source area, instead of the extended envelope that results from a transposition according to equation (2). Gating or other means may be used to avoid amplification of "empty" source channels. 17 illustrates another application, the generation of sub-harmonics relative highpass filtered or bass limited signal, using the compounds from the upper to the lower subband. When using the above transpositions it may be advantageous to use controlled switching of connections based on the signal characteristics.
In the above description, it is assumed that the highest frequency contained in the input signal is considerably lower than the Nyquist frequency. Thus, it may perform a bandwidth expansion without an increase in sample rate. This, however, does not always occur, so that it may be necessary to increase the sampling frequency beforehand. By using methods based on a set of filters for transposition may be included in the processing procedure of oversampling.
Most perceptual codecs used filter sets with maximum decimation when displaying the time frequency ["Introduction to Perceptual Coding" K.Brandenburg, AES, Collected Papers on Digital Audio Bitrate Reduction, 1996]. 18a shows the basic structure of a perceptual coding system. Analysis filterbank 1801 splits the input signal into several subband signals. These samples are individually quantized sub-bands (1803), using a reduced number of bits, where the number of quantization levels is determined by the perceptual model (1807), which estimates the minimum masking threshold. These sub-bands are normalized, coded with optional encoding methods combined with redundancy and additional information consisting of the normalization factors, bit allocation information and other codec specific data (1805) to generate a serial bit stream. This bit stream is then stored or transmitted. The decoder (18b) the coded bitstream is demultiplexed (1809), is decoded and the subband samples repeatedly kvantiruyutsya equal to the number of bits (1811). Synthesis filterbank combines subband samples to restore the original signal (1813). Embodiments using filter sets with maximum decimation significantly reduce the computational cost. In the following description focuses on cosine modulated filter sets. However, it should be understood that the present invention may be implemented using other types of filters, or sets of transducers, including interpreting a set of filters with low intensity wavelength conversion, filterbanks or other transducers with unequal bandwidths and multidimensional filter sets or converters.
In the illustrative, but not limiting description below assumes that the L-channel cosine modulated filterbank breaks the input signal x (n) into L subband. The overall structure of a filter set with a maximum decimation is shown in Figure 19. The analysis filters are denoted Hk (z) 1901, where k = 0, 1, ..., L-1. Ν subband signals to (n) maximally decimated (1903), each of sampling frequency fs / L, where fs - sampling frequency x (n). Synthesis unit reconnects back subband signals after interpolation (1905) and filter (1907) for generating x (n). Synthesis filters denoted Fk (z). Furthermore, the present invention performs a spectral overlap at x (n), to form a resulting signal y (n).
Synthesis subband signals using the QL-channel filter set that uses only the L channel lower frequencies and bandwidth expansion factor Q is chosen so that QL - integer, resulting in the output bit stream with sampling frequency Qfs. Consequently, an extended set of filters will act as if it were a set of L-channel filter device followed by oversampling. Since in this case the L (Q-1) high-pass filters are unused (fed to them zeros), the audio bandwidth will not change - the filter set simply reproduce version x (n) with higher sampling frequency. If, however, the subband signals associated with high-pass filters, the bandwidth increases by a factor Q, to form y (n), this version is set with a maximum decimation filter device multiband transposition according to the invention. Using this scheme, the sampling process upsampled integrated into the process of synthesis filtering as explained earlier. It should be noted that it may be used by the synthesis filterbank to any size, resulting in different sampling rates of the output signal, and hence different bandwidth expansion factors. Performing overlapping spectrum to the present invention corresponds to the basic multi-band transposition method with an integer transposition factor M, is executed as a binding subband signals
<IMG>
<IMG>
<IMG>
wherein k∈ [0, L -1] and chosen so that Mk∈ [L, QL -1], emk (n) - the envelope correction and (-1) (M-1) kn - correction factor for spectral inverted sub-bands. Spectral inversion results from decimation subband signals and inverted signals can be repeatedly reversed by changing the sign of every second sample in those channels. Referring to Figure 20, consider a 16-channel synthesis filterbank connected (2009) to the transposition factor M = 2, with Q = 2. Blocks 2001 and 2003 denote the analysis filters HK (z) and decimation by units 19, respectively. Similarly, 2005 and 2007 are the interpolators and synthesis filters Fk (z). Equation (16) then simplifies binding signals respectively four upper frequency subband data obtained in each second group of the eight uppermost channels in the synthesis filterbank. Due to spectral inversion, every second associated subband signal must be frequency inverted before the synthesis. Furthermore, the amplitudes of signals connected must be adjusted (2011) according to the rules of the DSP-DSP-1 or 2.
When using the basic multiband transposition method according to the invention generated harmonics in general are not exactly multiples of the fundamental frequency. All frequencies but the lowest in every subband differs in some extent from an exact transposition. Furthermore, the duplicated spectrum contains zeros since the resulting band interval covers a wider frequency range than the source interval range. Further, spurious signal suppression properties cosine modulated filterbank disappears as the subband signals are separated in frequency in the output range. That is, neighboring subband signals do not overlap in the high frequency region. However, methods for reducing the spurious signals known to those skilled in the art can be used to reduce this type of interference. Advantages of this transposition method consist in the simplicity of implementation, and the very low computational cost.
To ensure the accuracy of transposition of sinusoids, it presents a solution based on a set of filters effective maximum thinning improved method multiband transposition. The system uses an additional modified analysis filter set, while the synthesis filterbank is cosine modulated as described in "Multi-rate Systems and Filter Banks", PPVaidyanathan, Prentice Hall, Englewood Cliffs, New Jersey, 1993, ISBN 0- 13-605718-7. Stages multiband transposition method according to the present invention based filterbanks with maximum decimation are shown schematically in Figure 21 and the flowchart of Figure 22 and is as follows:
1. L received subband signals are synthesized using a QL-channel filterbank 2101, 2201, 2203, where L (Q-1) upper channels are fed zeros, to form signal x (n), which is thus redundantly sampled expansion coefficient width Q. band
<img he="13" wi="74" file="00000026.tif" img-content="undefined" img-format="tif" />
3. Select the integer value of K as the size of a synthesis filterbank, limited so that T = KM / Q - an integer, where T - the size of the modified analysis filterbank, and M - the transposition factor 2207, 2209, 2211. K should preferably be Select more stationary (tonal) signals, and smaller for dynamic (transient) signals.
<img he="7" wi="94" file="00000027.tif" img-content="undefined" img-format="tif" />
5. Signals ν k (M) (n '') are converted to a polar representation (magnitude and phase angle). The phase angles are multiplied by the factor M, and these signals are converted back into a rectangular coordinates according to Scheme 11. Taken real component of the complex signal, thereby producing the signals sk (M) (n '') 2109, 2215. After this operation, the signals sk (M) (n '') critically sampled.
6. The gains of the signals sk (M) (n '') is governed by the rules of DSP-1 or DSP-2 (2111, 2217).
<img he="12" wi="117" file="00000028.tif" img-content="undefined" img-format="tif" />
8. x3 (M) (n) finally added to x1 (n), to give y (n) 2223, which is the desired signal of the duplicated spectrum.
<img he="10" wi="133" file="00000029.tif" img-content="undefined" img-format="tif" />
<IMG>
<img he="10" wi="47" file="00000030.tif" img-content="undefined" img-format="tif" />
<IMG>
<img he="12" wi="116" file="00000031.tif" img-content="undefined" img-format="tif" />
The modified set of filters for the analysis stage (4) is obtained according to the theory of the cosine-modulated filter sets, where the modulated lapped transform (see. "Lapped Transforms for Efficient Transform / Subband Coding" HSMalvar, IEEE Trans ASSP, vol.38, no.6, 1990) is a special case. Impulse responses hk (n) filters in a T-channel cosine modulated filter set can be written as
<img he="12" wi="136" file="00000032.tif" img-content="undefined" img-format="tif" />
where k = 0, 1, ..., T-1, N - length of the lowpass prototype filter ro (n), C - constant and Fk - phase angle, which ensures exclusion of interference between adjacent channels. Restrictions on FC following:
<img he="10" wi="144" file="00000033.tif" img-content="undefined" img-format="tif" />
which can be simplified to the closed form expression in
<IMG>
With this choice fc using a synthesis filterbank may be prepared accurate reconstruction systems or approximate reconstruction systems (systems psevdoQMF) with impulse responses as
<img he="10" wi="49" file="00000034.tif" img-content="undefined" img-format="tif" />
Consider filters
<IMG>
<img he="15" wi="146" file="00000035.tif" img-content="undefined" img-format="tif" />
<IMG>
yields filters that have the same shape as the amplitude response as a Hk (z) for positive frequencies but are zero for negative frequencies, use a set of filters with impulse responses as in Eq (24) gives a set signal subband, which may be interpreted as a signal analysis (complex) corresponding subband signals obtained from a set of filters with impulse responses as in Eq (19). Analyzing signals suitable for manipulation, since the sample values of a full machining can be written in polar form, i.e.
z (n) = r (n) + ji (n) = | z (n) | exp {j arg (z (n))}.
<img he="34" wi="146" file="00000036.tif" img-content="undefined" img-format="tif" />
<IMG>
<img he="6" wi="145" file="00000037.tif" img-content="undefined" img-format="tif" />
Combining equality and equity 25 24 gives
<img he="17" wi="86" file="00000038.tif" img-content="undefined" img-format="tif" />
<img he="17" wi="146" file="00000039.tif" img-content="undefined" img-format="tif" />
Here are some explanations regarding the stage (5). Downsampling the complex sub-band signals leads to oversampling on M, which is an essential criterion when the phase angles subsequently are multiplied by the transposition factor M. oversampling causes that the number of subband samples per bandwidth, after transposition to a range of destination becomes equal to number of samples of the original sub-band range. Individual bandwidth of the transposed subband signals M times larger than the bandwidth of the source range, due to the action fazoumnozhitelya. This leads to the fact that the subband signals are critically sampled after step (5), and in addition, the spectrum will not be zero when the tone transposition.
<img he="17" wi="145" file="00000040.tif" img-content="undefined" img-format="tif" />
<IMG>
where | ν K (M) (n '') | absolute value ν K (M) (n ''), the following trigonometric relationship is used:
<IMG>
<img he="6" wi="145" file="00000041.tif" img-content="undefined" img-format="tif" />
<IMG>
<IMG>
and
<IMG>
calculating for step (5) can be accomplished without trigonometric calculations, reducing computational complexity.
When using transpositions even-M may be problems for the phase-multiplier, depending on the characteristics of the lowpass filter po (n). All applicable filters have zeros on the unit circle in the plane Z. A zero on the unit circle creates a 180-degree shift in the phase response of the filter. For even M fazoumnozhitel translates these shifts in the shifts of 360 degrees, ie, phase shifts vanish. Partial tones so arranged in frequency that such phase shifts vanish will lead to interference in the synthesized signal. The worst case in this situation occurs when a partial tone corresponds to a point in frequency corresponding to the vertex of the first side lobe response of a filter assay. Depending on the easing of the tab in the amplitude response of the interference will be more or less audible. As an example, the first side lobe of the filter used for the layer 1 and 2 standard ISO / MPEG, attenuated by 96 dB, while the attenuation of the first side lobe for the sine window is used in the circuit MDCT layer 3 standard ISO / MPEG only 23 dB. It is clear that this type of disturbance, using sinusoidal window will be heard. The following is a solution to this problem, defined as the relative phase synchronization.
Filters ha to (n) have linear phase responses. Phase angles FC administered relative phase differences between adjacent channels, and the zeros on the unit circle introduce a phase shift of 180 degrees at positions in frequency that may differ for the different channels. By controlling the phase difference between neighboring subband signals, before starting fazoumnozhitelya easily identify the channels that contain phase-inverted information. For tone phase difference is approximately π / 2 M according to equation (25) for non-inverted signals and, accordingly, is approximately π (1-1 / 2 M) for signals, if any of the signals is inverted. Isolation of inverted signals may be accomplished by calculating the scalar product signals in neighboring subbands as
<IMG>
If the product in equation (32) is negative, the phase difference is greater than 90 degrees, and the phase inversion condition is present. Phase angles of the complex subband signals are multiplied by M circuit according to step (5) and, finally, the signals are labeled as inverses subtracted. Relative phase locking method thus forces the 180 ° shifted to subband signals retain this shift after the phase multiplication and thereby maintain the property of suppression.
Adjusting spectral envelope.
Most of sounds such as speech and music, are characterized by the works of slowly varying envelopes and rapidly varying carriers with constant amplitude, as described in "The Application of Generalized Linearity to Automatic Gain Control" TGStockham, Jr, IEEE Tans on Audio and Electroacoustics, Vol .AU-16, No.2, June 1968 and in equation (1).
In perceptual audio encoder with split band audio signal is segmented into blocks and split into multiple frequency bands using subband filters or a transform from the time domain frequency domain. In most types of codecs tone is divided into two major signal components for transmission or storage, spectral envelope representation and the normalized subband samples or coefficients. In the following description, the term "subband samples" or "coefficients" refers to sample values obtained from subband filters as well as coefficients obtained for a transformation from the time domain to the frequency domain. The term "spectral envelope" or "scale factors" represent values for the subbands based on a time frame, such as the average or maximum amplitude in each subband, used for normalization of the subband samples. However, the spectral envelope may also be obtained using linear prediction (U.S. Patent 5,684,920). In a typical codec normalized sampling subband require coding at a high bit rate (using approximately 90% of the available transmission rate) compared to the envelopes slowly time-varying, and thereby, spectral envelopes, that may be coded at a much slower rate (using approximately 10 % of the available bit rate).
Accurate spectral envelope redundant bandwidth is important if the quality should be kept tone of the original signal. The perceived timbre of a musical instrument or voice is determined primarily by the spectral distribution below the frequency flim (boundary) located in the highest octaves of the audible range. Portion of the spectrum above flim, thus, have a minimal value and, accordingly, highband fine structures obtained by the above transposition methods require no adjustment, while the coarse structures generally require. To ensure that this adjustment is useful to filter the spectral representation of the signal to separate the envelope coarse structure from the fine structure.
In the embodiment using DSP-1 according to the present invention, coarse spectral envelope is estimated by highband lowband information available at the decoder. This assessment is carried out continuous monitoring of the lower band of the envelope and the spectral envelope adjustment of the upper band, in accordance with special rules. A new method of calculating the envelope uses asymptotes in a logarithmic frequency-amplitude domain, which is equivalent to curve fitting with polynomials using alternating order in the linear region. Calculate the slope of the upper level and the lower-band spectrum and estimates are used to determine the level and slope of one or several segments representing the new highband envelope. Assimptot points of intersection are fixed in frequency and act as pivot points. However not always necessary, although it is advantageous to set limits to keep the highband envelope deviation in actual borders. An alternative approach to estimation of the spectral envelope is to use vector quantization, VQ, of more typical spectrum envelopes and storing them in a lookup table or a set of codes. Vector quantization is performed by training the desired number of vectors on a large amount of training data, in this case audio spectral envelopes. Education is usually performed using the generalized Lloyd algorithm (see Ref. "Vector Quantization and Signal Compression" A.Gersho, RMGray, Kluwer Academic Publishers, USA 1992, ISBN 0-7923-9181-0) and gives vectors which optimally grasp the contents of the data training. Considering the set of codes VQ, consisting of A spectral envelope, trained by B envelopes (B >> A), then A envelopes represent the A most likely transitions from the lowband envelope to envelope the upper band on the basis of observations in a wide variety of audio signals. This, in theory, is A rule for predicting the envelope based on the B observations. When assessing a new envelope of the spectrum envelope of the original upper band lower band is used to find a set of codes, and part of the upper band most closely matching records set of codes used to create a new range of top bands.
Figure 23 subband samples normalization reference numeral 2301 and the spectral envelope represented by the scaling factors 2305. For purposes of illustration, the transmission to decoder 2303 is shown in parallel form. The method DSP-2 (Figure 24), the spectral envelope information is generated and transmitted according to 23, wherein only transmitted lowband subband samples. Transmitted ratios thus cover the full frequency range while the subband samples only span a limited frequency range, excluding the upper band. In the decoder, sampling lowband subband transposed 2401 (2403) and combined with the received information of the spectral envelope 2405 highband. Thus, the synthesized highband spectral envelope is identical to the original envelope, maintaining at the same time a significant decrease in the data rate.
In some codecs may transmit the scale factors for the complete spectral envelope dropping while highband subband samples, as shown in Figure 24. Other codec standards stipulate that scale factors and subband samples must cover the same frequency band, ie, scale factors can not be transmitted if the subband samples are omitted. In such cases, there are several solutions: information about the spectral envelope highband can be transmitted in separate frames, where the frames have their own headers and optional error protection, followed by data.
Contents5
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| RU2638748C2 | Cited by | Russian Federation | Search report |
| US8886346B2 | Cited by | United States of America | Applicant |
| US10186280B2 | Cited by | United States of America | Applicant |
| US8818541B2 | Cited by | United States of America | Applicant |
| US11562755B2 | Cited by | United States of America | Applicant |
| US9799346B2 | Cited by | United States of America | Applicant |
| US10192565B2 | Cited by | United States of America | Applicant |
| US11837246B2 | Cited by | United States of America | Applicant |
| US11100937B2 | Cited by | United States of America | Applicant |
| USRE49801E | Cited by | United States of America | Applicant |
| US10586550B2 | Cited by | United States of America | Applicant |
| US11031025B2 | Cited by | United States of America | Applicant |
| US9236061B2 | Cited by | United States of America | Applicant |
| US11993817B2 | Cited by | United States of America | Applicant |
| RU2493618C2 | Cited by | Russian Federation | Search report |
| US10584386B2 | Cited by | United States of America | Applicant |
| US11591657B2 | Cited by | United States of America | Applicant |
| US11935555B2 | Cited by | United States of America | Applicant |
| USRE47180E | Cited by | United States of America | Applicant |
| US10200974B2 | Cited by | United States of America | Applicant |
| RU2646314C1 | Cited by | Russian Federation | Search report |
| RU2495505C2 | Cited by | Russian Federation | Search report |
| RU2487428C2 | Cited by | Russian Federation | Search report |
| RU2494478C1 | Cited by | Russian Federation | Search report |
| US11682410B2 | Cited by | United States of America | Applicant |
| US10947594B2 | Cited by | United States of America | Applicant |
| RU2667629C1 | Cited by | Russian Federation | Search report |
| RU2598035C1 | Cited by | Russian Federation | Search report |
| US9830928B2 | Cited by | United States of America | Applicant |
| US10043526B2 | Cited by | United States of America | Applicant |
| US8612214B2 | Cited by | United States of America | Applicant |
| US11935551B2 | Cited by | United States of America | Applicant |
| RU2507572C2 | Cited by | Russian Federation | Search report |
| RU2596594C2 | Cited by | Russian Federation | Search report |
| US9384750B2 | Cited by | United States of America | Applicant |
| US10600427B2 | Cited by | United States of America | Applicant |
47 members in 15 offices
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 9702213 | Sweden | A | |
| 9800268 | Sweden | A | |
| 97022131 | – | – | – |
| 98002686 | – | – | – |
| SE19970002213 | – | – | – |
| SE19980000268 | – | – | – |
Members47
| Document | Office | Kind | |
|---|---|---|---|
| SE9702213D0 | Sweden | D0 | |
| SE9704634D0 | Sweden | D0 | |
| SE9800268D0 | Sweden | D0 | |
| SE9800268L | Sweden | L | |
| WO9857436A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU7446598A | Australia | A | |
| BR9805989A | Brazil | A | |
| EP0940015A1 | European Patent Office (EPO) | A1 | |
| WO9857436A3 | World Intellectual Property Organization (WIPO) | A3 | |
| SE512719C2 | Sweden | C2 | |
| CN1272259A | China | A | |
| HK1030843A1 | Hong Kong, China | A1 | |
| JP2001521648A | Japan | A | |
| EP1367566A2 | European Patent Office (EPO) | A2 | |
| EP0940015B1 | European Patent Office (EPO) | B1 | |
| AT257987T | Austria | T | |
| ATE257987T1 | Austria | T1 | |
| US6680972B1 | United States of America | B1 | |
| DE69821089D1 | Germany | D1 | |
| HK1057815A1 | Hong Kong, China | A1 | |
| US2004078194A1 | United States of America | A1 | |
| US2004078205A1 | United States of America | A1 | |
| DK0940015T3 | Denmark | T3 | |
| PT940015E | Portugal | E | |
| US2004125878A1 | United States of America | A1 | |
| ES2213901T3 | Spain | T3 | |
| EP1367566A3 | European Patent Office (EPO) | A3 | |
| DE69821089T2 | Germany | T2 | |
| CN1206816C | China | C | |
| CN1629937A | China | A | |
| JP2005173607A | Japan | A | |
| RU2256293C2This record | Russian Federation | C2 | |
| US6925116B2 | United States of America | B2 | |
| EP1367566B1 | European Patent Office (EPO) | B1 | |
| AT303679T | Austria | T | |
| ATE303679T1 | Austria | T1 | |
| DE69831435D1 | Germany | D1 | |
| DK1367566T3 | Denmark | T3 | |
| PT1367566E | Portugal | E | |
| ES2247466T3 | Spain | T3 | |
| DE69831435T2 | Germany | T2 | |
| JP3871347B2 | Japan | B2 | |
| CN1308916C | China | C | |
| US7283955B2 | United States of America | B2 | |
| US7328162B2 | United States of America | B2 | |
| JP4220461B2 | Japan | B2 | |
| BR9805989B1 | Brazil | B1 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Correction of name of patent ownerPD4A | PD4A |
Numbers
- Publication, DOCDB
- 2256293
- Publication, EPODOC
- RU2256293
- Application
- 9910481409
- Application, DOCDB
- 99104814
- Application, EPODOC
- RU19990104814
Titles2
- English
- IMPROVING INITIAL CODING USING DUPLICATING BAND
- Russian
- УСОВЕРШЕНСТВОВАНИЕ ИСХОДНОГО КОДИРОВАНИЯ С ИСПОЛЬЗОВАНИЕМ ДУБЛИРОВАНИЯ СПЕКТРАЛЬНОЙ ПОЛОСЫ
Classification
- IPC, 1
- H04J3 18