Method and apparatus for conversion between multi-channel audio formats
Abstract
method and equipment for converting between multichannel audio formats ". a multichannel input representation is converted to a multichannel output representation other than a spatial audio signal, in which an intermediate representation of the spatial audio signal is derived, the intermediate representation having direction parameters indicating an origin direction of a portion of the spatial audio signal; and in which the multichannel output representation of the spatial audio signal is generated using the intermediate representation m of the spatial audio signal.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
22 claims: 7 independent, 15 dependent
- 1Equipment for converting an input multichannel representation (102) into an output multichannel representation (110) different from a spatial audio signal, characterized in that it comprises:input representation decoder for deriving some audio channels corresponding to the highs -speakers associated with the input multichannel representation;analyzer (104) for derivation, using the number of audio channels corresponding to the speakers associated with the input multichannel representation, an intermediate representation (106) of the spatial audio signal, the intermediate representation (106) having direction parameters (40) indicating an origin direction of a portion of the spatial audio signal;and signal composer to generate the output multi-channel representation (110) of the spatial audio signal using the intermediate representation (106) of the spatial audio signal. 1. Equipamento para conversão de uma representação multicanal de entrada (102)em uma representação multicanal de saída (110) diferente de um sinal de áudio espacial, caracterizado pelo fato de que compreende: decodificador de representação de entrada para derivar alguns canais de áudio correspondentes aos alto-falantes associados à representação multicanal de entrada;analisador (104) para derivação, usando o número de canais de áudio correspondentes aos alto-falantes associados à representação multicanal de entrada, uma representação intermediária (106) do sinal de áudio espacial, sendo que a representação intermediária (106) possui parâmetros de direção (40) que indicam uma direção de origem de uma porção do sinal de áudio espacial;e compositor de sinal para gerar a representação multicanal de saída (110) do sinal de áudio espacial usando a representação intermediária (106) do sinal de áudio espacial.
- 4Equipment, according to the claim 4. Equipamento, de acordo com a reivindicação 1, caracterizado pelo fato de que o analisador (104) funciona derivando diferentes parâmetros de direção (40) para porções de freqüência de largura finita do sinal de áudio espacial. 1, characterized in that the analyzer (104) works by deriving different direction parameters (40) for finite-width frequency portions of the spatial audio signal.
- 5Equipment, according to the claim 5. Equipamento, de acordo com a reivindicação 1, caracterizado pelo fato de que o analisador (104) funciona derivando diferentes parâmetros de direção (40) para porções de tempo de extensão finita do sinal de áudio espacial. 1, characterized in that the analyzer (104) works by deriving different direction parameters (40) for finite length time portions of the spatial audio signal.
- 6Equipment, according to the claim 6. Equipamento, de acordo com a reivindicação 4, caracterizado pelo fato de que o analisador (104) funciona derivando (76) os diferentes parâmetros de direção (40) para porções de tempo de extensão finita do sinal de áudio espacial associado às porções de freqüência, onde a extensão de uma primeira porção de tempo associada a uma primeira porção de freqüência difere da extensão de uma associação de uma segunda porção de tempo a uma segunda porção de freqüência diferente do sinal de áudio espacial. 4, characterized by the fact that the analyzer (104) works by deriving (76) the different direction parameters (40) for time portions of finite extension of the spatial audio signal associated with the frequency portions, where the extension of a first portion of time associated with a first frequency portion differs from the extent of an association of a second time portion with a different second frequency portion of the spatial audio signal.
- 7Equipment, according to the claim 7. Equipamento, de acordo com a reivindicação 1, caracterizado pelo fato de que o analisador (104) funciona para derivar parâmetros (76) de direção que descrevem um vetor que aponta para a direção de origem da porção do sinal de áudio espacial. 1, characterized in that the analyzer (104) functions to derive direction parameters (76) that describe a vector that points to the source direction of the portion of the spatial audio signal.
- 8Equipment, according to the claim 8. Equipamento, de acordo com a reivindicação 1, caracterizado pelo fato de que o analisador (104) funciona também derivando (80) um ou mais canais de áudio associados à 1, characterized in that the analyzer (104) also works by deriving (80) one or more audio channels associated with the Petition 870210000624, of 01/04/2021, p. 6/20 Petição 870210000624, de 04/01/2021, pág. 6/20 3/6 intermediate representation (106). 3/6 representação intermediária (106).
- 21Method for converting a 21. Método para conversão de uma Petition 870210000624, of 01/04/2021, p. 9/20 Petição 870210000624, de 04/01/2021, pág. 9/20 6/6 multi-channel representation of input (102) in a multi-channel representation of output (110) different from a spatial audio signal, the method characterized by the fact that it comprises:deriving a number of audio channels corresponding to the speakers associated with the input multichannel representation;derive, using the number of audio channels corresponding to the speakers associated with the input multichannel representation, an intermediate representation (106) of the spatial audio signal, the intermediate representation having direction parameters (40) that indicate a direction of origin of a portion of the spatial audio signal;and generating the output multichannel representation of the spatial audio signal using the intermediate representation (106) of the spatial audio signal. 6/6 representação multicanal de entrada (102)em uma representação multicanal de saída (110)diferente de um sinal de áudio espacial, sendo que o método caracterizado pelo fato de que compreende: derivar um número de canais de áudio correspondente aos alto-falantes associados à representação multicanal de entrada;derivar, usando o número de canais de áudio correspondente aos alto-falantes associados à representação multicanal de entrada, uma representação intermediária (106) do sinal de áudio espacial, sendo que a representação intermediária possui parâmetros de direção (40) que indicam uma direção de origem de uma porção do sinal de áudio espacial;e gerar a representação multicanal de saída do sinal de áudio espacial usando a representação intermediária (106) do sinal de áudio espacial.
Independent claims7
94 paragraphs in 3 sections, as filed
(45) Concession Date: 4/6/2021
National Institute of Industrial Property (54) Title: METHOD AND EQUIPMENT FOR CONVERSION BETWEEN MULTI-CHANNEL AUDIO FORMATS (51) Int.CI .: H04S 3/02; G10L 19/16; G10L 19/008.
(52) CPC: H04S 3/02; G10L 19/173; G10L 19/008.
(30) Unionist Priority: 3/21/2007 US 60 / 896,184; 04/30/2007 US 11 / 742,502.
(73) Holder (s): FRAUNHOFER-GESELLSCHAFT ZUR FÕRDERUNG DER ANGEWANDTEN FORSCHUNG EV.
(72) Inventor (s): HERRE, JURGEN; VILLE PULKKI.
(86) PCT Application: PCT EP2008000830 of 02/01/2008 (87) PCT Publication: WO 2008/113428 of 09/25/2008 (85) Date of the Beginning of the National Phase: 09/18/2009 (57) Summary: METHOD AND EQUIPMENT FOR CONVERSION BETWEEN MULTI-CHANNEL AUDIO FORMATS. A multichannel input representation is converted to a multichannel output representation other than a spatial audio signal, in which an intermediate representation of the spatial audio signal is derived, the intermediate representation having direction parameters indicating an origin direction of a portion of the spatial audio signal; and in which the multichannel output representation of the spatial audio signal is generated using the intermediate representation m of the spatial audio signal.
METHOD AND EQUIPMENT FOR CONVERSION BETWEEN
MULTI-CHANNEL AUDIO FORMATS
Field of the Invention
The present invention relates to a technique of how to convert between different multichannel audio formats with the highest possible quality without limiting itself to specific multichannel representations, that is, the present invention refers to a technique that allows the conversion between multichannel / arbitrary formats.
M 10 History of the Invention and previous method _ In general, in multi-channel reproduction and listening, the listener is surrounded by multiple speakers. There are several methods for capturing audio signals for specific configurations. A general objective in reproduction is to reproduce the spatial composition of the originally recorded sound event, that is, the origins of the individual audio sources, such as the location of a trumpet within an orchestra. Various speaker configurations are quite common, and can create different spatial impressions. Without using special post-production techniques, the commonly known two-channel stereo configurations can only recreate auditory events on a line between the two speakers. This is achieved • mainly by the so-called panoramic amplitude, where ”the amplitude of the signal associated with an audio source is distributed between the two speakers, depending on the position of the audio source in relation to the speakers. This is usually done during subsequent recording or mixing. That is, an audio source that comes from the left end in relation to the listening position, will be played mainly by the left speaker, while an audio source in front of the listening position.
<img file="BRPI0808217B1_D0001.tif" />
listening will be played with identical amplitude (level) by both speakers. However, the sound that emanates from other directions cannot be played.
Consequently, when using more speakers that are distributed around the listener, more directions can be covered, and a more natural spatial impression can be created. The probably best known multichannel speaker layout is the 5.1 standard (ITU-R775-1), which is composed of 5 speakers whose azimuth angles in relation to the listening position are predetermined at 0<sup>O</sup>, ± 30 ° and ± 110 °. This means that during recording or mixing, the signal is customized for that specific speaker configuration, and deviations from the standard of a playback setting will result in a reduction in playback quality.
Several other systems have also been proposed, with varying numbers of speakers located in different directions. Professional and special systems, especially in theaters and sound installations, also include speakers at different heights.
A universal audio reproduction system called DirAC was recently proposed, which is capable of recording and reproducing sound for arbitrary speaker configurations. The purpose of DirAC is to reproduce the spatial impression of an existing acoustic environment as precisely as possible, using a multichannel speaker system with arbitrary geometric configuration. Within the recording environment, the responses of the<sup>3</sup> ambient (which can be continuous recorded sound or impulse responses) are measured with an omnidirectional microphone (W), and with a set of microphones that measure the direction of arrival of the sound and the diffusibility of the sound. In the following paragraphs and within the application, the term diffusibility must be understood as a measure for the non-directivity of the sound, that is, the sound that reaches the listening or recording position with equal power from all * directions, is maximally diffuse. A common way of quantifying the <sub>;</sub> diffusion is to use diffusibility values of the interval [0, ..., 1], where a value of 1 describes sound that is maximally diffuse and a value of 0 describes sound that is perfectly directional, that is, sound that emanates from only one clearly distinguishable direction. A commonly known method of measuring the direction of arrival of the sound is to apply 3 figure eight microphones (XYZ) aligned with Cartesian coordinate axes. Special microphones, the so-called SoundField microphones, have been designed, which directly produce all the desired responses. However, as mentioned above, the W, X, Y and Z signals can also be computed from a set of discrete omnidirectional microphones.
Another method for storing audio formats for an arbitrary number of channels in one or two downmix audio channels with tracked directional data was recently proposed by Goodwin and Jot. This format can be applied to arbitrary, reproductive systems. Directional data, that is, 25 data containing information about the direction of audio sources, are computed using Gerzon vectors, which are composed of a velocity vector and an energy vector. The velocity vector is a weighted sum of vectors facing the listening position's speakers, where each weight is the magnitude of a frequency spectrum at a given moment / frequency tile of a speaker. The energy vector is a similarly weighted vector sum. However, weights are short-term energy estimates of the speaker signals, that is, they describe a signal in some way smoothed or the entire signal energy contained in the signal within time intervals. <sup>r</sup> finite extent. These vectors share the disadvantage of not being related to a physical or perceptual quantity in a well-founded way. For example, the relative phase of the speakers in relation to each other is not properly taken into account. This means, for example, that if a broadband signal is supplied to the speakers of a stereo set in front of a listening position with an opposite phase, a listener would perceive the sound from the ambient direction, and the sound field in the listening position would have oscillations of sound energy from one side to the other (for example, from the left side to the right side). Under these conditions, the Gerzon vectors would be pointing to the frontal direction, which. it is obviously not representing the physical or perceptual situation.
Naturally, with multiple multichannel formats or representations on the market, there is a requirement for the ability to convert between different representations, so that individual representations can be reproduced with sets, originally developed for the reconstruction of an alternative multichannel representation. That is, for example, a transformation between channels 5.1 and channels 7.1 or 7.2 may be necessary to use an existing playback configuration of channel 7.1 or 7.2 to reproduce the multichannel representation 5.1 r 5 commonly used on DVD. The wide variety of audio formats makes the production of audio content difficult, as all formats require specific mixes and storage / transmission formats. Thus, it is necessary to convert between different 5 recording formats for playback in different playback configurations.
There are some proposed methods for converting audio from a specific audio format to another audio format. '<sub>φ</sub> However, these methods are always customized for multichannel formats or specific representations. That is, they are only applicable to the conversion of a specific predetermined multichannel representation into another specific multichannel representation.
. In general, a reduction in the number of reproduction channels (called downmix) is simpler to implement than an increase in the number of reproduction channels (upmix). For * some standard speaker playback settings, requirements are made, for example, the ITU on how to downmix playback settings with a smaller number of playback channels. In these so-called ITU downmix equations, the output signals are derived as simple static linear combinations of input signals. Normally, a reduction in the number of reproduction channels leads to a degradation of the perceived spatial image *, that is, a degraded reproduction quality of a spatial audio signal.
For a possible benefit of a high number of reproduction channels or reproduction speakers, upmixing techniques have been developed for specific types of conversions. A frequently investigated problem is how to convert 2-channel stereo audio for playback with 5-channel surround speaker systems. One approach or implementation for this type of 2-to-5 upmix is to use a so-called matrix decoder. These decoders have become commonplace for providing or upmixing multichannel 5.1 sound in stereo transmission infrastructures, especially in the beginning of surround sound for cinema and home theaters. The basic idea is<sub>?</sub> reproduce sound components that are in phase on the stereo signal in front of the sound image, and put the components out of phase on the rear speakers. An alternative 2-to-5 upmixing method
V proposes to extract the ambient components of the stereo signal and reproduce these components through the rear speakers of the 5.1 configuration. An approach that follows the same basic ideas 15 in a perceptually more justified way and using a mathematically more elegant implementation was recently proposed by C. Faller in Parametric Multi-channel Audio Coding: Synthesis of Coherence Cues, IEEE Trans. On Speech and Audio Proc., Vol. 14, no. 1, January 2006.
The recently published MPEG surround standard performs an upmix from one or two channels with downmix and transmitted, to the final channels used in playback or playback, which is normally 5.1. This is implemented using<sub>;</sub> spatial side information (side information similar to the BBC technique) or without side information, using the phase relationships between the two channels of a stereo downmix (unguided mode or extended matrix mode).
All of the '7 format conversion methods described in the previous paragraphs are specialized to be applied to specific reproduction format settings for both source and destination, and are therefore not universal. That is, a conversion between arbitrary multichannel input representations into arbitrary multichannel output representations cannot be performed. This means that the transformation techniques of the previous method are specifically designed 'for the number of speakers and their exact position for the' $ multichannel audio representation as well as for the multichannel output representation.
The international patent application 2004/077884 f proposes to use DirAC encoding to record impulse responses of audio signals within listening environments. Using these recorded impulse responses, audio signals can be reproduced with the spatial impression of the listening environment.
The work of the AES 6658 convention is aimed at DirAC audio coding and proposes a method for creating an efficient coded representation of signals recorded by b-format microphones.
International patent application 01/82651 refers to surround mastering and multichannel reproduction techniques. A particular spatial coding technique is - proposed, in order to enable the transmission of a compact coded representation *. The encoded representation can then be decoded by a specially designed decoder at the receiving end.
Naturally, it is desirable to have a concept for multichannel transformation that is applicable to arbitrary combinations of multichannel input and output representations. Summary of the Invention In accordance with a configuration of the present invention, equipment for converting a multichannel input representation to a multichannel output representation other than a spatial audio signal, composed of: analyzer to derive an intermediate representation of the audio signal ' spatial, and the intermediate representation contains <sub>?</sub> direction parameters indicating an origin direction of a portion of the spatial audio signal; and a signal composer to generate the multichannel representation of the spatial audio signal using the intermediate representation of the spatial audio signal.
Since an intermediate representation is used, which has direction parameters that indicate a direction of origin for a portion of the spatial audio signal, conversion can be achieved between arbitrary multichannel representations, as long as the speaker configuration of the multichannel representation output is known. It is important to note that the speaker configuration of the multichannel output representation does not need to be known in advance, that is, during the design of the conversion equipment. As the conversion equipment and method are universal, a multichannel representation provided as «multichannel input representation and designed for a specific speaker configuration can be changed on the receiving side, to suit the available reproduction configuration, so that the quality of a reproduction of a spatial audio signal is improved.
According to another configuration of the present invention, the origin direction of a portion of the spatial audio signal is analyzed within different frequency bands. Thus, different direction parameters are derived for finite 5 with frequency portions of the spatial audio signal. To derive the finite width frequency portions, for example, a filter bank or a Fourier transformer can be used. According to another configuration, the frequency portions or frequency bands, for which the analysis is performed .10 individually, are chosen in order to correspond to the frequency resolution of the human auditory process. These configurations can have the advantage that the direction of origin of the portions of the spatial audio signal is performed as well as the human auditory system itself is able to determine the direction of origin of the audio signals. Therefore, the analysis is performed without a potential loss of precision in determining the origin of an audio object or a portion of the signal, when that analyzed signal is reconstructed and reproduced through a configuration. of arbitrary speaker.
According to another embodiment of the present invention, one or more downmix channels are also derived, belonging to the intermediate representation. That is, the downmix channels are derived from audio channels corresponding to the • speakers associated with the multichannel input representation, which can then be used to generate the multichannel output representation, or to generate audio channels corresponding to the speakers associated with the multichannel output representation.
For example, a single channel monophonic downmix can be generated by the 5.1 input channels of a common 5.1 channel audio signal. This could, for example, be accomplished by computing the sum of all individual audio channels. Based on this derived monophonic downmix channel, a signal composer 5 can distribute those portions of the monophonic downmix channel corresponding to the analyzed portions of the multichannel input representation in the channels of the multichannel output representation, as indicated by the direction parameters. That is, a frequency / time or signal portion analyzed as coming from ^ 0 from the left end of a spatial audio signal will be redistributed to the speakers of the multichannel output representation, which are located on the left side in relation to the listening position. .
In general, some configurations of the present invention allow to distribute portions of the spatial audio signal with greater intensity in a channel corresponding to a speaker closer to the direction indicated by the direction parameters than to a channel further away from that direction. That is, regardless of how the location of the speakers used for reproduction is defined in the multichannel output representation, spatial redistribution will be obtained by adapting the reproduction configuration available in the best possible way.
According to some configurations of the present invention, a spatial resolution, with which a direction of origin of a portion of the spatial audio signal can be determined, is much higher than the angle of three-dimensional space associated with a single speaker. of the multichannel input representation. That is, the direction of origin of a portion of the spatial audio signal can be derived with better precision than a spatial resolution that can be obtained simply by redistributing the audio channels from one distinct configuration to another specific configuration, such as for example , redistributing the 5 channels of a 5.1 configuration in a 7.1 or 7.2 configuration.
In summary, some configurations of the invention allow the application of an improved method for format conversion, which is universally applicable and does not depend on a desired target speaker layout / configuration. ^ 0 Some configurations convert an input multichannel audio format (representation) with NI channels to an output multichannel w format (representation) with N2 channels by extracting direction parameters (similar to DirAC), which are then used to synthesize the signal output with N2 channels. In addition, according to some configurations, some NO downmix channels are computed from NI input signals (audio channels corresponding to speakers according to the multichannel input representation), which are then k used as basis for a decoding process using 20 extracted direction parameters.
Brief description of the drawings
Various configurations of the present invention will be described below, with reference to the accompanying drawings.
Fig. 1 shows an illustration of the derivation of direction parameters that indicate an origin direction of a portion of an audio signal; and
Fig. 2 shows another configuration of derivation of direction parameters based on a 5.1 channel representation;
Fig. 3 shows an example of generating a multichannel output representation;
Fig. 4 shows an example of converting audio from a 5.1 channel configuration to an 8.1 channel configuration; and
Fig. 5 shows an example of an inventive device for converting between multichannel audio formats.
Some configurations of the present invention 4, 0 derive an intermediate representation of a spatial audio signal with direction parameters indicating an origin direction of a portion of the spatial audio signal. One possibility is to derive a velocity vector that indicates the direction of origin of a portion of a spatial audio signal. An example for doing this will be described in the following paragraphs, with reference to Fig.
1.
Before detailing the concept, it can be seen that the following analysis can be applied to multiple portions of. individual frequency or time of the underlying spatial audio signal simultaneously. For simplicity, however, the analysis will be described for only a specific frequency or time or portion of time / frequency. The analysis is based on an energetic analysis of the sound field recorded in a - recording position 2, located in the center of a 25-coordinate system, as shown in Fig. 1.
The coordinate system is a Cartesian Coordinate System, with an x 4 axis and a y 6 axis, perpendicular to each other. Using a right hand system,:<sup>; 13</sup> the ζ axis, not shown in Fig. 1, points to the direction outside the drawing plane.
For direction analysis, it is assumed that signals 4 (known as format B signals) are recorded. An omnidirectional signal w is recorded, that is, a signal that receives signals from all directions with (ideally) equal sensitivity. In addition, three directional signals X, Y and Z are recorded, with a 'sensitivity distribution pointing in the direction of the axes of the' * Cartesian Coordinate System. Examples of possible ^ LO sensitivity patterns of the microphones used are given in Fig. 1, showing two figure patterns of eight 8a and 8b, pointing in the axes directions. Two possible audio sources 10 and 12 are further illustrated in the two-dimensional projection of the coordinate system shown in Fig. 1.
For direction analysis, an instantaneous velocity vector (in time index n) is composed for different frequency portions (described by index i) by v (n, i) = X (n, i) and<sub>x</sub>+ Y (n, i) and<sub>y</sub>+ Z (n, i) and<sub>z</sub>. (1)
That is, a vector is created with the 20 microphone signals recorded individually from the microphones associated with the axis of the coordinate system as components. In the previous and next equations, the Quantities are indexed in Time (n) * 'and also in frequency (i) by two indices (n, l). That is, 'and<sub>x</sub>, and<sub>y</sub> and is<sub>z</sub> represent Cartesian unity vectors.
Using the simultaneously recorded omnidirectional signal W, an instantaneous intensity I is computed as
I (n, i) = w (n, i) v (n, i), (2) instantaneous energy is derived according to the following formula:
E (n, i) = w<sup>2</sup> (n, i) + || v || (h, z), (3) where || || denotes vector norm.
That is, an amount of intensity is derived, allowing a possible interference between two signals (as positive and negative amplitudes can occur). In addition, an amount of energy is derived, which naturally does not allow interference between two signals, as the amount of energy does not contain negative values that allow a signal to be canceled.
These intensity properties and energy signals can be used advantageously to derive an origin direction of signal portions with high precision, preserving a virtual correlation of audio channels (a relative phase between the channels), as will be detailed below.
On the other hand, the instantaneous intensity vector can be used as a vector that indicates the direction of origin of a portion of the spatial audio signal. However, this vector can undergo rapid changes, thus causing artifacts within the signal reproduction. Therefore, alternatively, an instantaneous direction can be computed using a short-term average, using a Hanning W window.<sub>2</sub> according to the following formula:
M / 2
5 D (n, i) = ~ + (4) m = -M / 2 where W<sub>2</sub> is Hanning's window for making the short-term average D.
That is, optionally, a direction vector can be derived with short term average .with parameters that indicate a direction of origin of the spatial audio signal.
Optionally, a diffusivity measure ψ can be computed as follows:
I (n + m, i<sup>2</sup> W. (m) ψ (η, z) = 1 - ---- T773 ------------------------- (5) where Wi ( m) is a window function defined between. -M / 2 and M / 2 for short term average.
It should again be noted that the derivation is /
AO performed in order to preserve the virtual correlation of the audio channels. That is, the phase information is properly considered, which is not the case for direction estimates based only on energy estimates (such as Gerzon vectors).
0 The simple example below should serve to explain this in more detail. Consider a perfectly diffused signal that is reproduced by two speakers in a stereo system. As the signal is diffuse (it originates from all directions), it must be reproduced by both speakers with equal. 20 intensity. However, as the perception will be diffuse, it is<sub>=</sub> 180 degree phase shift is required. In this scenario, a direction estimate based purely on energy would produce a direction vector that would point exactly in the middle, between the two speakers, which is certainly an undesirable result 25 that does not reflect reality.
According to the inventive concept detailed above, the virtual correlation of the audio channels is preserved, while estimating the direction parameters (direction vectors). In this particular example, the direction vector would be zero, indicating that the sound does not originate from a different direction, which is clearly not the case in reality. Correspondingly, the diffusivity parameter of equation (5) is 1, corresponding perfectly to the real situation.
The Hanning windows in the above equations can also have different extensions for different frequency bands.
slO As a result of this analysis, for each time slice of a frequency portion, a direction vector or direction parameters are derived, indicating an origin direction of the spatial audio signal portion, for which the analysis was performed. Optionally, a diffusibility parameter can be derived, indicating the diffusibility of the direction of a portion of the spatial audio signal. As previously described, a diffusion value of a derivative according to equation (4) describes a maximum diffusibility signal, that is, originating from k in all directions with equal intensity.
Conversely, small diffusibility values are attributed to portions of signal originating predominantly from one direction.
Fig. 2 shows an example for the derivation of direction parameters of a multichannel input representation with five channels, according to ITU-775-1. The multichannel input audio signal, that is, the multichannel input representation, is first transformed into B format, simulating an anechoic recording of the corresponding multichannel audio configuration. With respect to a center 20 of the Cartesian Coordinate System with an x 22 and 24 x axis, a right rear speaker 26 is located at an angle of 110 °. A front right speaker 28 is located at + 30 °, a center speaker at 0 °, a front left speaker 32 at 31 °, and a left rear speaker 34 at -110 °. In practice, an anechoic recording can be simulated by applying simple matrix operations, the geometric configuration of the multichannel representation. input is known.
Ά0 An omnidirectional signal w can be obtained by making a direct sum of all the speaker signals, that is, all the audio channels corresponding to the speakers associated with the multichannel input representation. The dipole or figure signals of eight X, Y and Z can be formed 15 by adding the speaker signals weighted by the cosine of the angle between the speaker and the corresponding Cartesian axes, that is, the direction of maximum sensitivity of the dipole microphone to be simulated. Suppose Ln is the 2-D or 3-D Cartesian vector k that points in the direction of the umpteenth speaker20 and V is the unit vector that points in the direction of the Cartesian axis corresponding to the dipole microphone. Thus, the weighting factor is cos (angle (Ln, V)). The directional sign X would, for example, be written as ► N
Χ = Σ £<sub>η</sub>· Cos (angle (L<sub>n</sub>, V)), n-1 when C<sub>n</sub> denotes the speaker signal of the nth channel and N is the number of channels. The term angle must be interpreted as an operator, computing the spatial angle between the two given vectors. That is, for example, the angle 40 (Θ) between the Y axis 24 and the left front speaker 32 in the two-dimensional case illustrated in Fig. 2.
The additional derivation of direction parameters 5 could, for example, be done according to the illustration of Fig.
1, and detailed in the corresponding description, that is, the audio signals X, Y and Z can be divided into frequency bands according to the frequency resolution of the human auditory system. The 'direction of sound, that is, the direction of origin of the portions of the spatial audio signal * 10 and, optionally, the diffusibility, are analyzed, depending on the time in each frequency channel. Optionally, a substitution for sound diffusibility using another measure of signal dissimilarity other than diffusibility can also be used, for example, the coherence between channels (stereo) associated with the spatial audio signal.
If, in a simplified example, an audio source 44 is present, as shown in Fig. 2, where that source | only contribute to the signal within a specific frequency band, a direction vector 46 that points to the audio source 44 would be derived. The direction vector is represented by direction parameters (vector components) that indicate the direction of the portion of the spatial audio signal originating from the audio source '44. In the reproduction configuration of Fig. 2, this signal would be reproduced mainly by the left front speaker 32, as illustrated by the symbolic wave associated with this speaker. However, small portions of the signal will also be reproduced by the left rear speaker 32. Thus, the directional signal from the microphone associated with the X 22 coordinate would receive the signal components from the left front channel 32 (the audio channel associated with the speaker. left front 32) and left rear channel 34.
As, according to the implementation above, the directional signal Y associated with the y axis will also receive portions of the signal reproduced by the left front speaker 32, a directional analysis based on directional signals X and Y can> reconstruct the sound that comes from the vector direction 46 with high Ί0 accuracy.
For the final conversion to the desired multichannel representation (multichannel format), the direction parameters that indicate the direction of origin of portions of the audio signals are used. Optionally, one or more (NO) additional audio downmix channels of 15 can be used. This downmix channel can, for example, be the omnidirectional channel W or any other monophonic channel. However, for spatial distribution, the use of only a single channel associated with the intermediate representation is k with a small negative impact. That is, several downmix channels, such as a stereo mix, channels W, X and Y or all channels of a B format can be used, as long as the direction parameters or directional data have been derived and can be used for the reconstruction or generation of the * multichannel output representation. It is also alternatively possible to use the 5 channels of Fig. 2 directly, or any combination of channels associated with the multichannel input representation as a replacement for possible downmix channels.
When only one channel is stored, there may be a degradation in the quality of the diffused sound reproduction.
Fig. 3 shows an example of reproducing the signal from the audio source 44 with a speaker configuration that differs significantly from the speaker configuration of Fig.
2, which was the multichannel input representation from which the parameters had been derived. Fig. 3 shows, as an example, six speakers 50a to 50f, equally distributed along a line in front of a listening position 60, defining the center of a coordinate system with an x 22 axis and a y axis
24, as introduced in Fig. 2. As a previous analysis provided direction parameters that describe the direction of the ♦ direction vector 46 that points to the audio signal source 44, a multichannel output representation adapted to the speaker configuration of Fig . 3 it can be easily derived by redistributing the portion of the spatial audio signal to be reproduced to the speakers near the direction of the audio source 44, that is, through the speakers near the direction indicated by the direction parameters. That is, the audio channels corresponding to the loud | speakers in the direction indicated by the direction parameters are emphasized in relation to the audio channels corresponding to the speakers that are distant from this direction. That is, speakers 50a and 50b can be oriented (for example, using amplitude panning) to reproduce the signal portion, 'while speakers 50c and 50f do not reproduce that specific portion of the signal, but can be used for reproducing diffuse sound or other signal portions of different frequency bands.
The use of a signal composer to generate the multichannel output representation of the spatial audio signal using the direction parameters can also be interpreted as decoding the intermediate signal in the desired multichannel output format, with N2 output channels. . The channels of 5 audio downmix or generated signals are typically processed in the same frequency band in which they were analyzed. Decoding can be performed in a manner similar to DirAC. In the optional reproduction of diffuse sound, the use of audio stops. representing a non-diffuse current is typically one of the two optional 4.0 downmix NO channel signals or linear combinations of them. ··
For the optional creation of a diffuse current, there are several synthesis options to create the diffuse part of the output signals or the output channels corresponding to the speakers15 according to the multichannel output representation. If there is only one downmix channel transmitted, that channel must be used to create non-diffused signals for each speaker. If more channels are transmitted, there are more k options for the way in which diffused sound can be created. If, for example, a stereo downmix is used in the conversion process, an obviously suitable method is to apply the left downmix channel to the speakers on the left, and the right downmix channel to the speakers on the right. If several downmix channels are used for the conversion (ie, NO> 1), the diffuse current of each speaker can be computed as a differently weighted sum of these downmix channels. One possibility would be, for example, to transmit a signal of format B (channels X, Y, Z and W as described previously) and to compute the signal of a virtual cardioid microphone for each speaker.
The following text describes a possible procedure for converting an input multichannel representation to an output multichannel representation as a list.
In this example, the sound is recorded with a simulated B-format microphone and then continues to be processed by a signal composer for listening or reproduction with a multichannel or monophonic speaker configuration. The unique steps are explained with reference to Fig. 4, showing the conversion of a
5.0 multichannel representation of channel input 5.1 in a multichannel representation of channel output 8. The base is an NI channel audio format (NI being 5 in the specific example). To convert the multichannel input representation to a different multichannel output representation, the following steps 15 must be performed.
1. Simulate an anechoic recording of an arbitrary multichannel audio representation with NI audio channels (5 channels), as illustrated in recording section 70 (with a k format B microphone simulated in a center 72 of the layout).
two. In an analysis step 74, the simulated microphone signals are divided into frequency bands, and in a directional analysis step 76, the direction of origin of portions of the simulated microphone signals is derived. In addition, 'optionally, diffusibility (or coherence) can be determined at a stage of terminating diffusibility 78.
As previously mentioned, a direction analysis can be performed without the use of an intermediate step of format B. That is, in general, an intermediate representation of the spatial audio signal has to be derived based on a multichannel input representation, where the intermediate representation has direction parameters that indicate an origin direction of a portion of the spatial audio signal 5.
3. In a downmix step 80, audio signals are derived from NO downmix, to be used as a basis for the conversion / creation of the multichannel output representation. In a composition step 82, the downmix audio signals are 40 decoded or upmixed to an arbitrary speaker configuration that requires N2 audio channels by an appropriate synthesis method (for example, using amplitude panning or equally suitable techniques ).
The result can be reproduced by a multichannel speaker system 15, for example having 8 speakers, as shown in reproduction example 84 in Fig. 4. However, thanks to the universality of the concept, a conversion can also be made for a monophonic speaker configuration, | providing an effect as if the spatial audio signal was recorded with a single directional microphone.
Fig. 5 shows an outline of an example of an equipment for converting between multichannel audio formats 100.
* Equipment 100 receives an input 25 multichannel representation 102.
Equipment 100 is composed of an analyzer 104 to derive an intermediate representation 106 of the spatial audio signal, the intermediate representation 106 having direction parameters indicating an origin direction of a portion of the spatial audio signal.
Equipment 100 is further composed of a signal composer 108 to generate a multichannel representation of 5 output 110 of the spatial audio signal using the intermediate representation (106) of the spatial audio signal.
In summary, the conversion equipment configurations and conversion methods described earlier provide some great advantages. First, virtually any input audio format can be processed in this way. In addition, the conversion process can output to any speaker layout, including non-standard speaker layout / settings, without the need to specifically customize new ratios for new input speaker layout / configuration combinations and output speaker layout / settings. Furthermore, the spatial resolution of audio reproduction increases when the number of speakers is increased, unlike the implementations of the previous method.
k Depending on certain requirements for implementing the inventive methods, the inventive methods can be implemented in hardware or in software. The implementation can be done using a digital storage medium, in particular a disk, DVD or CD with readable control signals * electronically stored on them, which work in conjunction with a programmable computer system so that the inventive methods are executed . In general, the present invention is therefore a computer program product with a program code stored in a machine-readable carrier, the program code working to execute inventive methods when the computer program runs on a computer . In other words, the inventive methods are, therefore, a computer program with a program code to execute at least one of the inventive methods when the computer program runs on a computer.
Although the above disclosure has been particularly demonstrated and described with reference to particular configurations, it will be understood by those skilled in the art that various other changes in form and details can be made without departing from the spirit and scope of the invention. It must be understood that several changes can be made in adapting to different configurations without leaving the broader concepts revealed in this document and covered by the following claims.
Contents3
6 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6
37 members in 12 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 60896184 | United States of America | – | |
| 89618407 | United States of America | P | |
| 11742502 | United States of America | – | |
| 74250207 | United States of America | A | |
| 2008000830 | European Patent Office (EPO) | W | |
| 11742502 | – | – | – |
| 60896184 | – | – | – |
| PCTEP2008000830 | – | – | – |
| US20070742502 | – | – | – |
| US20070896184P | – | – | – |
| WO2008EP00830 | – | – | – |
Members37
| Document | Office | Kind | |
|---|---|---|---|
| US2008232601A1 | United States of America | A1 | |
| US2008232616A1 | United States of America | A1 | |
| WO2008113427A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2008113428A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200841326A | Taiwan Province of China | A | |
| TW200845801A | Taiwan Province of China | A | |
| KR20090117897A | Republic of Korea | A | |
| KR20090121348A | Republic of Korea | A | |
| EP2130204A1 | European Patent Office (EPO) | A1 | |
| EP2130403A1 | European Patent Office (EPO) | A1 | |
| CN101658052A | China | A | |
| CN101669167A | China | A | |
| JP2010521909A | Japan | A | |
| JP2010521910A | Japan | A | |
| US2010166191A1 | United States of America | A1 | |
| US2010169103A1 | United States of America | A1 | |
| EP2130403B1 | European Patent Office (EPO) | B1 | |
| AT476835T | Austria | T | |
| HK1138977A1 | Hong Kong, China | A1 | |
| DE602008002066D1 | Germany | D1 | |
| RU2416172C1 | Russian Federation | C1 | |
| RU2009134474A | Russian Federation | A | |
| KR101096072B1 | Republic of Korea | B1 | |
| RU2449385C2 | Russian Federation | C2 | |
| TWI369909B | Taiwan Province of China | B | |
| JP4993227B2 | Japan | B2 | |
| US8290167B2 | United States of America | B2 | |
| KR101195980B1 | Republic of Korea | B1 | |
| CN101658052B | China | B | |
| JP5455657B2 | Japan | B2 | |
| BRPI0808217A2 | Brazil | A2 | |
| BRPI0808225A2 | Brazil | A2 | |
| TWI456569B | Taiwan Province of China | B | |
| US8908873B2 | United States of America | B2 | |
| US9015051B2 | United States of America | B2 | |
| BRPI0808225B1 | Brazil | B1 | |
| BRPI0808217B1This record | Brazil | B1 |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent or certificate of addition of invention granted [chapter 16.1 patent gazette]GrantedPRAZO DE VALIDADE: 10 (DEZ) ANOS CONTADOS A PARTIR DE 06/04/2021, OBSERVADAS AS CONDICOES LEGAIS.B16A | B16A | |
| Decision: intention to grant [chapter 9.1 patent gazette]B09A | B09A | |
| Patent application procedure suspended [chapter 6.1 patent gazette]B06A | B06A | |
| Application suspended after technical examination (opinion) [chapter 7.1 patent gazette]B07A | B07A | |
| Preliminary requirement: requests with searches performed by other patent offices: procedure suspended [chapter 6.21 patent gazette]B06U | B06U | |
| Others concerning applications: alteration of classificationAS CLASSIFICACOES ANTERIORES ERAM: G10L 19/00 , H04S 3/02B15K | B15K | |
| Objections, documents and/or translations needed after an examination request according [chapter 6.6 patent gazette]B06F | B06F |
Numbers
- Publication
- PI0808217
- Publication, DOCDB
- PI0808217
- Publication, EPODOC
- BRPI0808217
- Application
- 8217
- Application, DOCDB
- PI0808217
- Application, EPODOC
- BR2008PI08217
Titles2
- Portuguese
- método e equipamento para conversão entre formatos de áudio multicanal
- English
- METHOD AND EQUIPMENT FOR CONVERSION BETWEEN MULTI-CHANNEL AUDIO FORMATS
Classification
- CPC, 4
- H04S3/02
- G10L19/008
- G10L19/173
- H04S2420/11
- IPC, 3
- H04S3 02
- G10L19 16
- G10L19 008