Method and apparatus for conversion between multichannel audio formats
Abstract
FIELD: information technology. SUBSTANCE: present invention relates to a method for conversion between different multichannel audio formats with the highest possible quality, without limitation to specific multichannel representations, i.e., the present invention relates to a method which enables to perform conversion between arbitrary multichannel formats. An input multichannel representation is converted into a different output multichannel representation of a spatial audio signal, and an intermediate representation of the spatial audio signal is derived therein; the intermediate representation having direction parameters indicating the direction of origin of a portion of the spatial audio signal; and the output multichannel representation of the spatial audio signal is generated therein using the intermediate representation of the spatial audio signal. EFFECT: high quality of reproducing a spatial audio signal. 20 cl, 6 dwg
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
20 claims: 3 independent, 17 dependent
- 1A device for converting the input multi-channel representation (102) into the output multi-channel representation (110) spatial audio signal other than the input, including an input interface for receiving an input multi-channel representation (102), the analyzer (104) to obtain an intermediate representation (106) of spatial the sound signal having direction parameters (40) indicating the direction of the origin of spatial audio signal;analyzer (104) adapted to obtain the audio downmix channel based on a combination of audio channels corresponding to the speakers (26-34) associated with the input multi-channel representation (102) and the signal synthesizer (108) for generating an output multi-channel representation (110) spatial audio signal using the downmix channel in accordance with the parameters of the direction the intermediate representation (106) spatial audio signal. 1. Устройство для преобразования входного многоканального представления (102) в выходное многоканальное представление (110) пространственного звукового сигнала, отличное от входного, включающее входной интерфейс для получения входного многоканального представления (102), анализатор (104) для получения промежуточного представления (106) пространственного звукового сигнала, имеющего параметры направления (40), указывающие направления происхождения области пространственного звукового сигнала;анализатор (104) выполнен с возможностью получения звукового канала понижающего микширования, основанного на объединении звуковых каналов, соответствующих громкоговорителям (26-34), связанным с входным многоканальным представлением (102), и синтезатор сигналов (108) для генерирования выходного многоканального представления (110) пространственного звукового сигнала с использованием канала понижающего микширования в соответствии с параметрами направления промежуточного представления (106) пространственного звукового сигнала. 1. Устройство для преобразования входного многоканального представления (102) в выходное многоканальное представление (110) пространственного звукового сигнала, отличное от входного, включающее входной интерфейс для получения входного многоканального представления (102), анализатор (104) для получения промежуточного представления (106) пространственного звукового сигнала, имеющего параметры направления (40), указывающие направления происхождения области пространственного звукового сигнала;анализатор (104) выполнен с возможностью получения звукового канала понижающего микширования, основанного на объединении звуковых каналов, соответствующих громкоговорителям (26-34), связанным с входным многоканальным представлением (102), и синтезатор сигналов (108) для генерирования выходного многоканального представления (110) пространственного звукового сигнала с использованием канала понижающего микширования в соответствии с параметрами направления промежуточного представления (106) пространственного звукового сигнала.
- 19A method for converting an input into an output multi-channel representation of the spatial representation of multi-channel audio signal different from the input;characterized in that it further comprises receiving input multi-channel representation, obtaining an intermediate representation (74;76) a spatial audio signal;wherein the intermediate representation has direction parameters indicating a direction of origin of spatial audio signal;wherein the downmix audio channel is obtained based on the combining of audio channels corresponding to the speakers (26-34) associated with the input multi-channel representation, and generating (82) an output multi-channel representation of spatial audio signal using the downmix channel in accordance with the parameters of the direction the intermediate representation spatial audio signal. 19. Способ преобразования входного многоканального представления в выходное многоканальное представление пространственного звукового сигнала, отличное от входного;характеризующийся тем, что дополнительно включает получение входного многоканального представления, получение промежуточного представления (74;76) пространственного звукового сигнала;при этом промежуточное представление имеет параметры направления, указывающие направление происхождения области пространственного звукового сигнала;где звуковой канал понижающего микширования получен, базируясь на объединении звуковых каналов, соответствующих громкоговорителям (26-34), связанным с входным многоканальным представлением, и генерирование (82) выходного многоканального представления пространственного звукового сигнала с использованием канала понижающего микширования в соответствии с параметрами направления промежуточного представления пространственного звукового сигнала. 19. Способ преобразования входного многоканального представления в выходное многоканальное представление пространственного звукового сигнала, отличное от входного;характеризующийся тем, что дополнительно включает получение входного многоканального представления, получение промежуточного представления (74;76) пространственного звукового сигнала;при этом промежуточное представление имеет параметры направления, указывающие направление происхождения области пространственного звукового сигнала;где звуковой канал понижающего микширования получен, базируясь на объединении звуковых каналов, соответствующих громкоговорителям (26-34), связанным с входным многоканальным представлением, и генерирование (82) выходного многоканального представления пространственного звукового сигнала с использованием канала понижающего микширования в соответствии с параметрами направления промежуточного представления пространственного звукового сигнала.
- 20The computer readable medium with stored thereon a computer program which, when running on a computer, implements a method for converting multi-channel representation in the spatial representation of the output multi-channel audio signal different from the input;the method comprising receiving an input multi-channel representation;forming an intermediate representation of spatial audio signal;intermediate representation has direction parameters indicating a direction of origin of spatial audio signal;wherein the downmix audio channel is obtained based on the combining of audio channels corresponding to the speakers (26-34) associated with the input multi-channel representation;and generating a spatial representation of the output multichannel audio signal using the downmix channel in accordance with the parameters of the intermediate representation of the spatial directions of the audio signal. 20. Машиночитаемый носитель с сохраненной на нем компьютерной программой, которая будучи запущенной на компьютере, реализует способ преобразования многоканального представления в выходное многоканальное представление пространственного звукового сигнала, отличное от входного;при этом способ включает получение входного многоканального представления;получение промежуточного представления пространственного звукового сигнала;промежуточное представление имеет параметры направления, указывающие направление происхождения области пространственного звукового сигнала;в котором звуковой канал понижающего микширования получен на основании объединения звуковых каналов, соответствующих громкоговорителям (26-34), связанным с входным многоканальным представлением;и генерирование выходного многоканального представления пространственного звукового сигнала с использованием канала понижающего микширования в соответствии с параметрами направления промежуточного представления пространственного звукового сигнала. 20. Машиночитаемый носитель с сохраненной на нем компьютерной программой, которая будучи запущенной на компьютере, реализует способ преобразования многоканального представления в выходное многоканальное представление пространственного звукового сигнала, отличное от входного;при этом способ включает получение входного многоканального представления;получение промежуточного представления пространственного звукового сигнала;промежуточное представление имеет параметры направления, указывающие направление происхождения области пространственного звукового сигнала;в котором звуковой канал понижающего микширования получен на основании объединения звуковых каналов, соответствующих громкоговорителям (26-34), связанным с входным многоканальным представлением;и генерирование выходного многоканального представления пространственного звукового сигнала с использованием канала понижающего микширования в соответствии с параметрами направления промежуточного представления пространственного звукового сигнала.
Independent claims3
82 paragraphs, as filed
The present invention relates to a method for conversion between different formats of multi-channel sound with the best possible quality, but are not limited to specific multi-channel representation. That is, the present invention relates to a method that allows for the conversion between arbitrary multichannel formats.
Normally, when a multi-channel playback and listening to the listener is surrounded by numerous speakers. There are various methods of capturing audio signals for certain installations. The overall objective in the reproduction is to reproduce the spatial composition of the originally recorded sound, then there is the origin of individual sound sources, such as the location of the tube in an orchestra. Using multiple acoustic settings is quite common and can create different spatial experiences. No special techniques using Layout commonly known two-channel stereo systems can only recreate auditory events on the line between the two loudspeakers. This is mainly achieved by the so-called "amplitude panning", where the amplitude of a signal associated with a sound source is distributed between the two loudspeakers, depending on the position of the sound source relative to the loudspeakers. This is usually done during recording or subsequent mixing. That is the source of the sound coming from the far left position relative to the listener, will be mainly played left speaker and the sound source to the listener position will be reproduced with identical amplitude (level) two speakers. However, the sound coming from other directions can not be reproduced.
Consequently, by using more loudspeakers which are distributed around the listener, a larger number of directions can be covered and can be created a more natural spatial impression. Probably the most well known multichannel loudspeaker arrangement - this is the standard 5.1 (ITU-R775-1), which consists of five speakers, azimuthal angles are defined to be 0 °, ± 30 ° and ± 110 ° relative to the position of the listener. This means that during recording or mixing signal adapts to this particular configuration of the speaker and the deviation from the standard playback settings will reduce the playback quality.
Numerous other systems with varying number of loudspeakers located at different directions have also been proposed. Professional and special systems, especially in theaters and sound installations, also include loudspeakers arranged at different heights.
It has recently been proposed a universal sound playback system called DirAC, which can record and play back sound for any acoustic systems. DirAC aim is to reproduce the spatial impression of an existing acoustical environment as precisely as possible using a multichannel loudspeaker system having an arbitrary geometric structure. Within the environment of feedback sound recording environment (which can be continuously recorded sound or impulse response) is measured using the omnidirectional microphone (W) and a set of microphones that measure the direction of arrival of sound and diffuseness of the sound. In the following sections and through use, the term "diffuseness" should be understood as a measure for undirected sound. That is the sound arriving at the listening or recording with equal force in all directions, the most absent-minded. The usual method of measuring the diffusion is to use diffusivity values in the interval [0, ..., 1], where the maximum value of 1 describes the diffuse sound, and the value 0 describes perfectly directional sound, i.e. sound coming from only one direction is clearly discernible. One well-known method of measuring the direction of arrival of sound microphones involves the use of 3 "eights» (XYZ), oriented along the axes of a Cartesian coordinate system. It has been specially designed microphones, so-called "sound field microphones" which directly result in all of the desired response. However, as mentioned above, the signals W, X, Y and Z may also be calculated from a set of discrete omnidirectional microphones.
Another method of storing audio formats for any number of channels to one or two downmix channel recording with the accompanying directional characteristics was recently proposed by Goodwin and Jot. This format can be applied to arbitrary playback system. The directional characteristics, i.e., specifications, containing information about the direction of sound sources are calculated by using "vectors Gerzon," which consist of a velocity vector and a vector of energy. The velocity vector - a weighted sum of vectors pointing to the speakers with the listening position, where each weight - the value of the frequency spectrum at a given time / frequency for a given speaker. Vector energy - similar to a weighted vector sum. However, the weight - a short-term power estimation signal loudspeaker, that is, they describe several smoothed signal or integral signal power contained in the signal within the time interval of finite length. These vectors have the same disadvantage as the case of absence, depending on the physical and perceptual quantities in sound manner. For example, the relative phase of the speakers on each other are not adequately taken into account. This means for example that if a broadband signal is fed to loudspeakers stereophonic installation situated before the listening position with opposite phase, the listener will perceive the sound from the surrounding areas, and the sound field at the listening position will have a sonic energy vibrations from side to side (e.g., the left side to the right side). In this scenario, Gerson vectors would indicate the front direction, which is obviously not a physical or perceptual difference.
Of course, having the numerous multi-channel format or presentation on the market, there is a need to be able to convert between the different views, so that the individual representations could be reproduced plants, originally intended for the reconstruction of an alternative multi-channel representation. That is, for example, the transformation between the channels 5.1 and 7.1 or 7.2 channels may require the use of an existing 7.1 or 7.2 channel playback settings for playback of multi-channel representations 5.1, commonly used on DVD. A wide variety of audio formats audio content makes production difficult, as all formats require a specific format mixing and storage / transmission. It is therefore necessary conversion between different formats of recording for playback on a variety of playback settings.
Many methods for converting audio material in a specific audio format to another audio format. However, these methods are always adapted to the specific multi-channel format or representation. That is, they are applicable only to the transformation from one predetermined multi-channel representation to another specific multi-channel representation.
Generally, reducing the number of playback channels (so - called "downmixed") is easier than increasing the number of playback channels ("upmixing"). For some standard acoustic reproduction systems include recommendations ITU e.g., to implement the down mix plants reproduction with a smaller number of playback channels. In these so-called «ITU» downmix equations outputs are extracted as simple static linear combinations of the input signals. Generally, reducing the number of playback channels leads to deterioration of the perceived spatial image, that is, deterioration of playback quality spatial audio signal.
For the possible benefits of using a large number of playback channels or loudspeakers reproducing techniques have been developed downmix for certain types of transformations. Most researched problem is that the conversion of dual-channel stereo record for playback on a circular five-channel speaker system. One approach or the performance of the up-mix to 2 channels to 5 should use the so-called "matrix" decoder. Such a proliferation of decoders for multi-channel downmix 5.1 sound through stereo transmission infrastructure, particularly in the early stages of development of surround sound for cinema and home theater. The main idea is to reproduce sound components that are in phase with the signal in the stereo sound image in the front and in the room out of phase components in the rear speakers. An alternative method of up-mix channels 2 to 5 proposes to extract the surrounding components of a stereo signal and reproduce these components through the rear speakers settings 5.1. The approach of pursuing the same basic ideas on perceptually more informed basis and used mathematically more elegant version, was recently proposed K.Follerom in "Parametric multi-channel audio coding: Synthesis of coherence cues», IEEE On the processing of speech and audio signals., Edition of 14 Number 1, January 2006
Recently published MPEG standard performs upmixing one or two transmitted downmix channels at the end used in reproduction or playback, which is usually by mixing 5.1. This is done either by using space more information (additional information is similar to the BCC technique) or without additional information by using the phase relations between the two channels of the stereo downmix ("unmanaged method" or "advanced matrix method").
All techniques for transforming the format described in the preceding sections, intended for application to certain configurations of both the source and target format recording playback and thus they are not universal. That is the conversion between the arbitrary "input multi-channel representation and arbitrary output multi-channel representation can not be done. That is the prototype of transformation methods specifically adapted to the number of speakers and their exact location for the input multi-channel audio performance, as well as the output multi-channel representation.
International Patent Application 2004/077884 suggests using DirAC-encoding recording pulse characteristics of audio signals within the listening environment. Using such impulse responses recorded, the audio signals can be reproduced with the spatial perception of the listening environment.
AES-6658 agreement is intended for audio coding DirAC and offers a method of creating an effective representation of the encoded signals recorded microphones b-format.
International patent application 01/82651 relates to a method of multi-channel surround recording method of original and reproduction. Special spatial encoding technique proposed for the transmission of a compact encoded representation. The encoded representation may then be decoded by a decoder specifically designed at the receiving end.
Naturally, it is desirable to have the concept of a multi-channel transform is applicable to arbitrary combinations of input and output multi-channel representation.
According to one embodiment of the invention, a device for converting the input multi-channel representation into an output multi-channel representation is different from the input, the spatial audio signal comprising: a parser for receiving the intermediate representation of spatial audio signal; intermediate representation having direction parameters indicating a direction of origin of spatial audio signal; and the synthesizer output signal for the production of multi-channel representation of the spatial audio signal using the intermediate representation of the spatial audio signal.
This uses an intermediate representation, which has a direction parameters indicating a direction of origin of spatial audio signal; conversion can be achieved between arbitrary multi-channel view, if we know the configuration of the acoustic output multi-channel representation. It is important to note that the acoustic configuration of the output multi-channel representation need not be known in advance, that is, during the design of the device to convert. Since the conversion apparatus and method for universal multi-channel representation provided as an input multi-channel representation and designed for a certain acoustic setup can be changed at the receiving end to match the available installation playback so that the playback quality spatial audio signal increases.
According to a further embodiment of the present invention, the direction of origin of spatial audio signal is analyzed within different frequency bands. Thus, the direction of the various parameters obtained for the finite width of the space-frequency sound. To obtain a finite width frequency domain, can be used, for example, the filter unit or the Fourier transform. According to another embodiment, the frequency domain or frequency ranges for which the analysis is carried out individually, are selected to correspond to the frequency resolution of human hearing threshold. These embodiments may have the advantage that the direction of origin of pieces of spatial audio signal performed as well as the human auditory system, and can determine the direction of origin of sound signals. Therefore, the analysis is performed without potential loss of precision in determining the origin of the sound of an object or part of the signal when a signal is analyzed restored or reproduced through any acoustic setting.
According to a further embodiment of the present invention, one or more downmix channels derived further intermediate representation and belong. That is, the downmix channels derived from the audio channels corresponding to the speakers connected to the input multi-channel representation, which can then be used to form an output or multi-channel representation for audio channels corresponding to the speakers connected to the output multi-channel representation.
For example, a monophonic downmix channel can be produced from the input channels of conventional 5.1 5.1 channel audio. This could, for example, be performed by calculating the sum of all the individual audio channels. Based on such received monophonic downmix channel synthesizer signal may distribute those parts monophonic downmix channel corresponding to the analyzed part of the input multi-channel representation of the output channels of multi-channel representation, as indicated by the direction parameters. That is, the analyzed frequency domain / time or part of the signal, which must come from extreme left surround audio signal to be redistributed to the speakers the output multi-channel representation, which are located on the left side relative to the seating position.
Typically, some of the present invention allow the parts to distribute the spatial audio signal with a higher intensity on the channel corresponding to the speaker that is closer to the direction indicated by the direction parameters, and not on the channel located farther from this direction. That is, no matter how the location of loudspeakers used for reproduction is defined in the output multi-channel representation spatial redistribution is reached as possible quality, applicable to existing installations playback.
In some embodiments of the invention, the spatial resolution at which can be determined the direction of origin of spatial audio signal, much higher than the angle of the three-dimensional space associated with a single input multi-speaker presentation. That is, the direction of origin of spatial audio signal may be obtained with greater accuracy than the spatial resolution that can be obtained by a simple redistribution of audio channels from one installation to another individual unit, such as the trunking installation to installation 5.1 7.1 or 7.2.
Summarizing, we can say that some of the invention allow the use of an advanced method for converting a format that is universally applicable and does not depend on the specific location of the desired title / speaker configuration. Some embodiments transform the input multi-channel audio format (representation) to the output channels N1 multichannel format (representation) having the N2 channels, by retrieving direction parameters (DirAC similar), which are then used for synthesizing an output signal having channels N2. Furthermore, in some embodiments, many N0 downmix channels are calculated from the input signal N1 (audio channels corresponding to the speakers according to the input multi-channel representation), which are then used as the basis for the decoding process, using the extracted parameters directions.
Several embodiments of the present invention will hereinafter be described with reference to the accompanying drawings.
1 illustrates descent direction parameters indicating a direction of origin domain audio signal; and
2 shows a further implementation of the origin direction parameters based on the representation of the channel 5.1;
3 shows an example of generating an output multi-channel representation;
4 shows an example of the sound conversion unit with a channel 5.1 setup channel 8.1; and
5 shows an example of the inventive device for carrying out conversion between multi-channel audio formats.
Some embodiments of the present invention produce an intermediate representation of spatial audio signal having direction parameters indicating a direction of origin of spatial sound. One possibility is to obtain the velocity vector, indicating a direction of origin of spatial sound. An example of this will be described in the following paragraphs with reference to Figure 1.
Before detailing the concept, it should be noted that the following analysis may be applied to multiple individual frequency region or a time base spatial audio signal simultaneously. For simplicity, however, the analysis will be described only for one specific frequency or time or a time domain / frequency. The assay is based on the analysis of the sound field energy recorded at the recording position 2 located in the center of the coordinate system, as shown in Figure 1.
The coordinate system - Cartesian coordinate system having the X-axis and Y-axis 4 6 perpendicular to each other. Using a right-handed system, Z axis, not shown in Figure 1, it indicates the direction of the drawing area.
For the analysis of the direction taken that recorded 4 signal (known as B-format signals). Logged one omnidirectional signal w, i.e. the signal, receives signals from all directions with (ideally) equal sensitivity. Furthermore, recorded signals are three-dimensional X, Y and Z, with the sensitivity distribution indicating the direction of the axes of a Cartesian coordinate system. Examples of possible samples of the sensitivity of the microphones used are given in Figure 1 shows two sample "figure-eight" 8a and 8b, indicates the direction of the axes. Two possible audio sources 10 and 12, moreover, the design illustrated in the two-dimensional coordinate system shown in Figure 1.
For the analysis of instantaneous velocity vector direction (at time index n) is made for the different frequency regions (described index i) using:
<img file="00000001.tif" he="6" wi="139" img-format="tif" img-content="undefined" />
That is, the generated vector having microphones individually recorded signals associated with the axis of the coordinate system as components. In the previous and following equations values indexed in time (n), as well as the frequency (i) two indices (n, i). I.e
EX, eu and ez are the unit vectors of the Cartesian.
Using both recorded omnidirectional signal w, the instantaneous intensity I is calculated as
<img file="00000002.tif" he="5" wi="128" img-format="tif" img-content="undefined" />
instantaneous energy is obtained according to the following formula:
<img file="00000003.tif" he="7" wi="124" img-format="tif" img-content="undefined" />
where it denotes a vector norm.
<IMG>
That is, the intensity value obtained corrected for the possibility of interference between the two signals (since there may be both positive and negative amplitude). Further, the obtained energy value, which, naturally, does not consider the interference between the two signals, as the energy value does not include negative values, taking into account the canceling signal.
These properties of the signal intensity and energy can be advantageously used to obtain the direction of origin of the signal portions with high accuracy, while maintaining the actual audio channel correlation (relative phase between the channels), as will be described in more detail below.
<img file="00000005.tif" he="11" wi="129" img-format="tif" img-content="undefined" />
<IMG>
where W2 - Hanning window for short-term averaging D.
That is, optionally, may be obtained by short-time average direction vector having parameters indicating a direction of origin of the spatial audio signal.
<img file="00000006.tif" he="16" wi="135" img-format="tif" img-content="undefined" />
<IMG>
wherein W1 (m) - window function defined between -M / 2 and M / 2 for a short averaging.
It should again be noted that the differentiation is designed in such a way as to keep the actual correlation of audio channels. That is, the phase information is properly taken into account, which is not the case in the areas of assessment based solely on estimates of energy (such as vectors Gerson).
The following simple example will explain this in more detail. Consider a perfectly diffused light, which is reproduced by two loudspeakers stereo system. If dispelled signal (derived from all directions) it should be played by both speakers with equal intensity. However, since the perception it will be dissipated require a phase shift of 180 degrees. In this scenario, estimate the direction, based solely on energy will lead to a direction vector indicating exactly midway between the two loudspeakers, which, of course, the result is undesirable, do not reflect reality.
According to the idea of the invention described in detail above, the actual correlation of audio channels stored in the evaluation direction parameters (vector direction). In this particular example, the direction vector will be zero, indicating that the sound does not come from one certain direction, that is not actually the case. Accordingly, the diffuseness parameter of equation (5) - 1, which perfectly corresponds to the real situation.
Hanning window in the above equations can also have different lengths for different frequency bands.
As a result of this analysis for each time interval a frequency domain, a vector direction or direction parameters indicating a direction of origin of spatial audio signal, for which analysis was performed. Optionally, it may be obtained by the diffuseness parameter indicating the direction of the field diffuseness spatial audio signal. As described earlier, the value of the diffusion parameter is obtained according to Equation (4) describes the maximum signal diffuseness, i.e. originating from all directions with equal intensity.
Conversely, small values assigned to the areas diffuseness signal originating predominantly from one direction.
2 shows an example of receiving direction parameters from the input multi-channel representation having five channels according to ITU-775-1. The multi-channel audio input signal, i.e. the input multi-channel representation is first converted into B-format recording nereverberiruyuschey by simulating the corresponding multi-channel sound system. The center of the Cartesian coordinate system 20 having, x-axis and y-axis 22, 24, back-right speaker 26 is located at an angle of 110 °. Forward-right speaker 28 is angled by 30 °, the center speaker at an angle of 0 °, the front-left speaker 32 at an angle of -31 ° and rear-left speaker 34 at an angle of -110 °. In practice, nereverberiruyuschaya record can be modeled by the use of simple operations matrixing; the geometric structure of the input multi-channel representation is known.
<img file="00000007.tif" he="13" wi="60" img-format="tif" img-content="undefined" />
<IMG>
When the loudspeaker signal Cn denotes n-th channel, and N - the number of channels. The term angle must be interpreted as an operator for calculating a spatial angle between these two vectors. That is, for example, the angle 40 (Θ) between Y axis 24 and the front-left speaker 32 in the two-dimensional case illustrated in Figure 2.
Further receiving direction parameters could for example be performed as illustrated in Figure 1 and detailed in the following descriptions, i.e. the audio signals X, Y, and Z may be divided into frequency bands according to the frequency resolution of the human auditory system. The direction of sound, i.e. direction of origin of spatial audio signal, and, optionally, diffuseness analyzed versus time in each frequency channel. Optionally, change the sound diffusivity with another non-diffuseness, dissimilarity index signal can also be used, for example, the coherence between (stereo) channels associated with a spatial audio signal.
If, as a simplified example, there is a sound source 44, as indicated in Figure 2, where this source contributes only signals within a certain frequency range is received direction vector 46 pointing to the sound source direction vector 44 represented parameters Direction ( vector components) indicating the direction of spatial audio signal originating from the sound source 44. In Figure 2, the playback of such a signal is reproduced mainly anterior-left speaker 32, as illustrated by symbolic waveforms associated with this loudspeaker. However, small signal region will also be reproduced with the rear-left speaker 32. Consequently, the directional microphone signal associated with the X coordinate of 22, will receive the signal components from the front 32, left channel (audio channel associated with the front-left speaker 32) and posterior left channel 34.
Since, according to the above implementation, directional signal Y, associated with the axis Y, will also have field signal reproduced by the front-left speaker 32, directional analysis based on directional signals X and Y, will be able to restore the sound coming from the direction vector 46 with high accuracy.
For the final conversion to the desired multi-channel representation (multi-format) using direction parameters indicating a direction of origin of the areas of audio signals. Optionally, may be used one or more (N0) additional audio downmix channels. This downmix channel may for example be omnidirectional W channel or any other channel monaural. However, the spatial distribution, using only a single channel connected with the intermediate representation has little adverse impact. That is, multiple downmix channels such as stereo channels mixed W, X and Y, or all channels in the format can be used as long as the direction parameters sent or received data and can be used to reconstruct or generate an output multi-channel representation. Alternatively, it is also possible to use five channels 2 directly, or any combination of channels associated with the input multi-channel representation as a possible replacement for the downmix channels. When only one stored channel quality deterioration may occur during playback of the scattered sound.
3 shows an example of a playback signal source 44 by setting the speaker is significantly different from the speaker installation 2, which has an input multi-channel representation from which the parameters were obtained. 3 shows, as an example, six speakers 50a-50f, are equally distributed along the front of the listening position 60, defining the center of the coordinate system having the X-axis 22 and Y axis 24, as shown in Figure 2. As the preceding analysis has provided direction parameters describing the direction of the direction vector 46 pointing to the sound source 44, the output multi-channel representation, adapted to install the speaker 3, it can easily be obtained through redeployment of spatial sound to be played on the loudspeakers, It is close to the direction of the sound source 44, i.e. those loudspeakers that are located close to the direction indicated by the direction parameters. That is, the audio channels corresponding to loudspeakers in the direction indicated by the direction parameters, gives special importance with respect to the audio channels corresponding to loudspeakers located far away from this area. That is, the speakers 50a and 50b may be adjusted (e.g., using a pan amplitude) for reproducing signal area, despite the fact that the speakers 50c-50f do not reproduce the specific area of the signal, while they may be used for reproducing the scattered sound or other signal region of different frequency bands.
Using a synthesizer for generating an output signal representation of spatial audio multi-channel signal using parameters directions it can also be interpreted as being a decoding intermediate signal in the desired multi-channel output format having N2 output channels. Sound channel downmix signals are typically generated or processed in the same frequency range in which they were analyzed. Decoding may be performed in a manner similar DirAC. In a further diffuse sound playback audio for presentation use unscattered flux is typically one or more of the signals N0 downmix channels or a linear combination.
To further create a scattered flux, there are several variants of synthesis to create a diffuse part of the output signal or output channels corresponding to the speakers according to the output multi-channel representation. If there is only one transmitted downmix channel, the channel must be used to create unscattered signals to each speaker. If you have a larger number of transmitted channels, there are more options for creating diffuse sound. If, for example, using a stereo downmix in the conversion process, the most appropriate method - use the left downmix channel to the speakers on the left and the right downmix channel to the loudspeakers on the right side. If more downmix channels are used to convert (ie N0> 1), the scattered flux for each speaker can be calculated as a weighted sum of these differentially downmix channels. One possibility, for example, the signal transmission format (channel X, Y, Z and w, as previously described) and calculating the actual signal cardioid microphone signal for each speaker.
The following text describes a possible procedure to convert the input multi-channel representation in the output multi-channel representation in the form of a list. In this example, the sound is recorded with the help of simulated in-formatted microphone and then subjected to further processing sound synthesizer to listen to or play with the help of a multi-channel or monaural speaker setup. The individual steps are explained with reference to Figure 4, showing the transformation of the input multi-channel representation in a 5.1 channel representation of a multi-channel output with 8 channels. Reason - N1-format audio channels (N1 = 5 in the specific example). To convert the input multi-channel representation to another output multi-channel representation, the following steps.
1. not simulated reverberant recording an arbitrary multi-channel audio representation having audio channels N1 (5 channels), as illustrated in the segment recording 70 (using a simulated B-aspect microphone in the center of the circuit 72).
2. In step 74 the analysis of the modeled microphone signals are divided into frequency bands, and in step 76 analyzes the directional direction of origin areas obtained modeled microphone signals. Also, optionally, diffuseness (or coherence) can be determined in step 78 diffuseness termination.
As previously mentioned, the directional analysis can be performed without the use of an intermediate stage in the format. That is, typically, the intermediate representation of spatial audio signal to be received, based on the input multi-channel representation, where the intermediate representation has direction parameters indicating a direction of origin of spatial sound.
3. In step 80 the downmix, N0 downmix audio signals are obtained, to be used as a base to convert / generate the output multi-channel representation. In step 82 the compound, N0 downmix audio signals are decoded or are upmixed to arbitrary acoustic installations requiring N2 audio channels, using an appropriate synthesis method (e.g., amplitude panning or using analogous methods).
The result can be reproduced in a multi-channel speaker system having, for example, the speaker 8 as shown in playback scripts 84 4. However, due to flexibility of the concept, the conversion may also be performed for monaural speaker setup, providing an effect as if the spatial audio signal was recorded with a directional microphone.
5 shows a schematic diagram of an apparatus for converting between 100 multichannel audio formats.
The device 100 is designed for the input multi-channel representation 102.
The apparatus 100 includes a parser 104 to obtain an intermediate representation of spatial audio signal 106, 106 has an intermediate representation direction parameters indicating a direction of origin of spatial sound.
Device 100 further includes a signal synthesizer 108 to generate an output multi-channel representation of spatial audio signal 110 using the intermediate representation (106) spatial audio signal.
Summarizing, we can say that the implementation of the previously described apparatus and method provide significant advantages conversion. First of all, virtually any format of input audio may be processed in this way. In addition, the conversion process can generate output for any scheme the speakers, including non-standard location / speaker configuration without the need to specifically establish new connections for new combinations of input location / speaker configuration and output location / speaker configuration. Furthermore, the spatial resolution of audio reproduction increases as the number of loudspeakers, contrary to prior art processes.
Depending on the specific requirements of the invention process can be accomplished in vehicles or in the instrument software. Execution may be performed using a digital storage medium, in particular disks, DVD- or CD-ROM drive, preserving electronically readable control signals which cooperate with a programmable computer system such that the inventive methods allows. In general, the present invention - a computer program product with a program code stored on a computer readable medium; a control program necessary for performing the inventive methods when the computer program product runs on a computer. In other words, the inventive methods - a computer program having a program code for performing at least one of the inventive methods when the computer program is run on a computer.
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| RU2630754C2 | Cited by | Russian Federation | Search report |
| US9852735B2 | Cited by | United States of America | Applicant |
| US9756448B2 | Cited by | United States of America | Applicant |
| US10499176B2 | Cited by | United States of America | Applicant |
| US10770087B2 | Cited by | United States of America | Applicant |
| RU2685997C2 | Cited by | Russian Federation | Search report |
| US9892737B2 | Cited by | United States of America | Applicant |
| WO2004077884A1 | Cites | World Intellectual Property Organization (WIPO) | – |
| RU2129336C1 | Cites | Russian Federation | – |
| EP1275272A1 | Cites | European Patent Office (EPO) | – |
| US20060004583A1 | Cites | United States of America | – |
| RU2234819C2 | Cites | Russian Federation | – |
| US5812674A | Cites | United States of America | – |
37 members in 12 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 60896184 | United States of America | – | |
| 89618407 | United States of America | P | |
| 11742502 | United States of America | – | |
| 74250207 | United States of America | A | |
| 11742502 | – | – | – |
| 60896184 | – | – | – |
| US20070742502 | – | – | – |
| US20070896184P | – | – | – |
Members37
| Document | Office | Kind | |
|---|---|---|---|
| US2008232601A1 | United States of America | A1 | |
| US2008232616A1 | United States of America | A1 | |
| WO2008113427A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2008113428A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200841326A | Taiwan Province of China | A | |
| TW200845801A | Taiwan Province of China | A | |
| KR20090117897A | Republic of Korea | A | |
| KR20090121348A | Republic of Korea | A | |
| EP2130204A1 | European Patent Office (EPO) | A1 | |
| EP2130403A1 | European Patent Office (EPO) | A1 | |
| CN101658052A | China | A | |
| CN101669167A | China | A | |
| JP2010521909A | Japan | A | |
| JP2010521910A | Japan | A | |
| US2010166191A1 | United States of America | A1 | |
| US2010169103A1 | United States of America | A1 | |
| EP2130403B1 | European Patent Office (EPO) | B1 | |
| AT476835T | Austria | T | |
| HK1138977A1 | Hong Kong, China | A1 | |
| DE602008002066D1 | Germany | D1 | |
| RU2416172C1 | Russian Federation | C1 | |
| RU2009134474A | Russian Federation | A | |
| KR101096072B1 | Republic of Korea | B1 | |
| RU2449385C2This record | Russian Federation | C2 | |
| TWI369909B | Taiwan Province of China | B | |
| JP4993227B2 | Japan | B2 | |
| US8290167B2 | United States of America | B2 | |
| KR101195980B1 | Republic of Korea | B1 | |
| CN101658052B | China | B | |
| JP5455657B2 | Japan | B2 | |
| BRPI0808217A2 | Brazil | A2 | |
| BRPI0808225A2 | Brazil | A2 | |
| TWI456569B | Taiwan Province of China | B | |
| US8908873B2 | United States of America | B2 | |
| US9015051B2 | United States of America | B2 | |
| BRPI0808225B1 | Brazil | B1 | |
| BRPI0808217B1 | Brazil | B1 |
Numbers
- Publication
- 2449385
- Publication, DOCDB
- 2449385
- Publication, EPODOC
- RU2449385
- Application
- 200913447408
- Application, DOCDB
- 2009134474
- Application, EPODOC
- RU20090134474
Titles2
- English
- METHOD AND APPARATUS FOR CONVERSION BETWEEN MULTICHANNEL AUDIO FORMATS
- Russian
- ?????? ? ?????????? ??? ????????????? ?????????????? ????? ??????????????? ????????? ?????????
Classification
- CPC, 4
- H04S3/02
- G10L19/008
- G10L19/173
- H04S2420/11