Method and device for decoding an audio soundfield representation for audio playback
Abstract
This record has no abstract on file.
Term
4.5 yearsto projected expiry
Projected expiry 25 March 2031, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
15 claims: 7 independent, 8 dependent
- 1Zastrzeżenia patentowe 1. Sposób dekodowania odwzorowania pola dźwiękowego audio do odtwarzania audio, obejmujący etapy:- obliczania (110) funkcji (W) panoramowania, dla ka ż dego z wielu gło ś ników, przy zastosowaniu metody geometrycznej, w oparciu o położenia głośników i wiele kierunków źródeł;- obliczania (120) macierzy modowej ( Ξ N) z kierunków ź ródła;- obliczania (130) pseudo-odwrotnej macierzy modowej (Ξ + ) macierzy modowej (Ξ);oraz - dekodowania (140) odwzorowania pola dźwiękowego audio, w którym dekodowanie oparte jest na macierzy (D) dekodowania, którą otrzymuje się z co najmniej funkcji (W) panoramowania i uogólnionej pseudo-odwrotnej macierzy modowej (Ξ + ).
- 2Sposób według zastrzeżenia 1, w którym metoda geometryczna, zastosowana na etapie obliczania funkcji panoramowania jest panoramowaniem amplitudy bazy wektora (VBAP).
- 3Sposób według zastrzeżenia 1 albo 2, w którym odwzorowanie pola dźwiękowego jest formatu Ambisonics co najmniej 2-go rzędu.
- 4Sposób według dowolnego z zastrzeżeń 1-3, w którym pseudo-odwrotną macierz modową (Ξ+) otrzymuje się według Ξ Η [ ΞΞ Η ] -1 , w którym Ξ jest macierzą modową wielu kierunków źródła.
- 5Sposób według zastrzeżenia 4, w którym macierz (DN) dekodowania otrzymuje się (135) według D =W Ξ Η [ΞΞ Η ] -1 = W Ξ+, w którym W jest zestawem funkcji panoramowania dla każdego głośnika.
- 6Urządzenie do dekodowania odwzorowania pola dźwiękowego audio do odtwarzania audio, zawieraj ące - pierwszy element obliczeniowy (210) do obliczania funkcji (W) panoramowania, dla każdego z wielu głośników, przy zastosowaniu metody geometrycznej, w oparciu o położenia głośników i wiele kierunków źródeł;- drugi element obliczeniowy (220) do obliczania macierzy modowej (Ξ) z kierunków źródła;- trzeci element obliczeniowy (230) do obliczania pseudo-odwrotnej macierzy modowej (Ξ + ) macierzy modowej (Ξ);oraz - element dekoduj ą cy (240) do dekodowania odwzorowania pola d ź wi ę kowego, w którym dekodowanie oparte jest na macierzy (D) dekodowania a element dekodujący wykorzystuje do otrzymania macierzy (D) dekodowania co najmniej funkcję (W) panoramowania i pseudo-odwrotną macierz modową (Ξ + ).
- 7Urządzenie według zastrzeżenia 6, w którym urządzenie do dekodowania zawiera ponadto element (235) do obliczania macierzy (D) dekodowania z funkcji (W) panoramowania i pseudo-odwrotnej macierzy modowej (Ξ + ).
- 8Urządzenie według zastrzeżenia 6 albo 7, w którym metoda geometryczna, użyta na etapie obliczania funkcji panoramowania jest panoramowaniem amplitudy bazy wektora (VBAP).
- 9Urządzenie według dowolnego z zastrzeżeń 6-8, w którym odwzorowanie pola dźwiękowego jest formatu Ambisonics co najmniej 2-go rzędu.
- 10Urządzenie według dowolnego z zastrzeżeń 6-9, w którym pseudo-odwrotną macierz modową Ξ+ otrzymuje się według Ξ+ = Ξ Η [ΞΞ Η ] -1 , w którym Ξ jest macierzą modową wielu kierunków źródła.
- 11Urządzenie według zastrzeżenia 10, w którym macierz (DN) dekodowania jest otrzymywana w elemencie (245) do obliczania macierzy dekodowania, według D =W Ξ η [ΞΞ η ] -1 = W Ξ+, w którym W jest zestawem funkcji panoramowania dla każdego głośnika.
- 12Nośnik czytelny dla komputera, z zapisanymi na nim instrukcjami wykonawczymi do realizowania przez komputer sposobu dekodowania odwzorowania pola dźwiękowego audio do odtwarzania audio, przy czym sposób obejmuje etapy - obliczania (110) funkcji (W) panoramowania, dla każdego z wielu głośników, przy zastosowaniu metody geometrycznej, w oparciu o położenia głośników i wiele kierunków źródeł;- obliczania (120) macierzy modowej (Ξ) z kierunków źródła;- obliczania (130) pseudo-odwrotnej macierzy modowej (Ξ + ) macierzy modowej (Ξ);oraz - dekodowania (140) odwzorowania pola dźwiękowego audio, w którym dekodowanie oparte jest na macierzy (D) dekodowania, którą otrzymuje się z co najmniej funkcji (W) panoramowania i pseudo-odwrotnej macierzy modowej (Ξ + ).
- 13Nośnik czytelny dla komputera według zastrzeżenia 12, w którym metoda geometryczna, użyta na etapie obliczania funkcji panoramowania jest panoramowaniem amplitudy bazy wektora (VBAP).
- 14Nośnik czytelny dla komputera według zastrzeżenia 12 albo 13, w którym odwzorowanie pola dźwiękowego jest formatu Ambisonics co najmniej 2-go rzędu.
- 15Nośnik czytelny dla komputera według dowolnego z zastrzeżeń 12-14, w którym pseudo-odwrotną macierz modową Ξ+ otrzymuje się według Ξ+ = E H [EE H ] -1 , w którym Ξ jest macierzą modową wielu kierunków źródła. Uprawniony:Thomson Licensing Pełnomocnik: mgr inż. Marta Skrobot Rzecznik patentowy Fig.l 102 103 Fig.2 Dopasowanie modów Amb. — Panoramowanie pełne Amb. » Panoramowanie VBAP Amb. 100 ra c Φ en O 60 O Fig.6 avp avp avp avp 102 .103 AC ciec Fig.7
Independent claims15
56 paragraphs, as filed
[0001] The present invention relates to a method and apparatus for decoding audio sound field mapping, and in particular Ambisonics audio mapping, for audio reproduction.
Background of the invention [0002] The purpose of this chapter is to familiarize the reader with various aspects of the technique that relate to the various aspects of the present invention that are described and / or claimed below. It is believed that this discussion will be helpful in providing the reader with basic information to facilitate a better understanding of various aspects of the present invention. Accordingly, it is assumed that if the source is not precisely indicated, these statements will be read in this light, and not as a state of the art.
[0003] For any spatial audio reproduction system, the precise location is a key goal. Such playback circuits are widely used in conference systems, games, or other virtual environments that take advantage of 3D sound. 3D sound scenes can be synthesized or captured as a natural sound field. Sound field signals such as Ambisonics carry a reproduction of the desired sound field. The Ambisonics format is based on the distribution of spherical sound field components. Although the zero-order and first-order spherical components are used in the basic Ambisonics or B-format, in the higher order format, the so-called Higher Order Ambisonics (HOA), successive at least second-order spherical components are also used. The decoding process is required to obtain individual speaker signals. For the synthesis of audio scenes, panning functions are required, which refer to the spatial arrangement of the speakers, to obtain the spatial location of the sound source. When a natural sound field needs to be recorded, a microphone system is required to capture spatial information. To achieve this, the known Ambisonics method is the right tool. Signals in the Ambisonics format carry a mapping of the desired sound field. To obtain individual speaker signals from these signals, the Ambisonics decoding process is required. Since also in this case panning functions can be obtained from the decoding function, panning functions are the main issue when describing the task of spatial location. The spatial speaker arrangement is referred to herein as the speaker configuration.
[0004] Commonly used speaker configurations are: a stereo configuration that uses two speakers, a standard surround configuration with five speakers, and an extension of the surround configuration with more than five speakers. These configurations are well known. However, they are limited to two dimensions (2D), i.e. the height information is not reproduced.
[0005] Speaker configurations for three-dimensional (3D) playback are described, for example, in "Wide listening area with exceptional spatial sound quality of a 22.2 multichannel sound system", K. Hamasaki, T. Nishiguchi, R. Okumaura, and Y. Nakayama at Audio Engineering Society Preprints, Vienna, Austria, May 2007, which are a proposal for NHK television with very high resolution with 22.2 format, or Dabringhaus 2 + 2 + 2 (mdg-musikproduktion dabringhaus und grimm, www.mdg.de) and configuration 10.2 in "Sound for Film and Television", T. Holman in 2nd ed. Boston: Focal Press, 2002. One of the few known systems for spatial reproduction and panning strategies is the vector base amplitude panning (VBAP) method in Virtual Sound Source Positioning Using Vector Base Amplitude Panning, Journal of Audio Engineering Society, Volume 45, No. 6, pp. 456-466, June 1997, in this document Pulkki. The VBAP (panning vector base amplitude) system was used by Pulkki to play virtual acoustic sources using any speaker configuration. A pair of speakers is required to place a virtual source in a 2D plane, while for 3D, three speakers are required. For each virtual source, selected loudspeakers with full configuration are fed with a monophonic signal with different amplifications (depending on the position of the virtual source). Then the speaker signals for all virtual sources are added together. For panning between the speakers in VBAP, a geometric method was used to calculate the gain of the speaker signals.
[0006] An example 3D speaker configuration, taken into account and proposed again in this document, has 16 speakers, which are positioned as shown in Fig. 2. The spacing has been made for practical reasons, with four columns of three speakers each and additional speakers between these speakers. To be more precise, eight speakers are evenly spaced around the head of the listener, covering 45 degrees. Additional four speakers are located at the top and bottom, determining azimuth angles equal to 90 degrees. Compared to Ambisonics, this configuration is irregular and leads to problems related to decoder construction, as mentioned in "An format Ambisonics for flexible playback layouts" by H. Pomberger and F. Zotter in Proceedings of the 1st Ambisonics Symposium, Graz, Austria, July 2009.
[0007] In traditional Ambisonics decoding, as described in EP 2094032 and in "Three-dimensional surround sound systems based on spherical harmonics" by M. Poletti in J. Audio Eng. Soc., Vol. 53, no. 11, pp. 1004-1025, November 2005, the commonly known mod matching process is used. Mods are described by fashion vectors that contain values of spherical components for a clear direction of incidence. The combination of all directions provided by individual speakers leads to a fashion matrix of speaker configurations so that the fashion matrix represents the speaker positions. To reproduce a clear source signal mod, the speaker modes are weighted in such a way that the superimposed single speaker modes add up the desired mod. To obtain the necessary weights, calculate the mapping of the inverse matrix to the speaker's matrix matrix. From the point of view of signal decoding, the weights form the speaker control signal, and the inverse matrix of the speaker modes is called the "decoding matrix", which is used to decode the signal mapping formatted in the Ambisonics system. In particular, for many speaker configurations, e.g., the configuration shown in Fig. 2, it is difficult to obtain the inverse of the fashion matrix.
[0008] As mentioned above, commonly used speaker configurations are limited to 2D, i.e. height information is not reproduced. In commonly known technical methods, decoding sound field mapping to speaker configurations, with mathematically irregular spatial distribution, leads to problems of location and timbre. Decoding matrix (i.e. decoding matrix matrix) is used to decode Ambisonics signal. In traditional Ambisonics signal decoding, especially HOA signals, at least two problems arise. The first refers to the correct decoding in which to obtain the decoding matrix it is necessary to know the directions of the signal source. The second refers to the mapping to the existing speaker configuration, which is subject to a systematic error due to the following mathematical problem: mathematically correct decoding will result not only in positive, but also in some negative loudspeaker amplitudes. However, they are incorrectly reproduced as positive signals and thus lead to the abovementioned problems.
Summary of the Invention [0009] The present invention describes a method of decoding sound field mapping for irregular spatial distributions with greatly improved sound localization and timbre properties. It presents another way of obtaining a decoding matrix for sound field data, e.g. in the Ambisonics format, which uses the process of assessing the system. Considering the set of possible incidence directions, panning functions related to the desired speakers are calculated. Panning functions are assumed as the output of the decoding process in the Ambisonics system. The required input signal is a fashion matrix of all considered directions. Thus, as shown below, the decoding matrix is obtained by right-multiplying the weight matrix by the inverted form of the input matrix's fashion matrix.
[0010] Regarding the second problem mentioned above, it has also turned out that it is also possible to obtain a decoding matrix from the inverse of a so-called fashion matrix which represents the positions of the speakers, and the position dependent weight functions W ("panning functions"). One aspect of the invention is that these W-panning functions can be obtained using a method other than commonly used. Preferably, a simple geometric method is used. This method does not require any knowledge of any direction of the signal source, so this solves the first problem mentioned above. One such method is known as panning of the Vector-Based Amplitude Panning (VBAP) vector base amplitude. According to the invention, VBAP is used to calculate the required panning functions, which are then used to calculate the decoding matrix in the Ambisonics system. Another problem is that you need to know the inverse of the fashion matrix (which represents the speaker configuration). However, it is difficult to get the exact inverse, which also leads to incorrect audio playback. Thus, an additional aspect is that a pseudo-inverse fashion matrix is calculated to obtain the decoding matrix, which is obtained much easier.
[0011] The invention uses a two-step method. The first step is to get panning functions that depend on the configuration of the speakers used for playback. In the second stage, the decoding matrix in the Ambisonics system is calculated from all the panning functions for all speakers.
[0012] An advantage of the invention is that a description of the parameters of the sound sources is not required; instead, a sound field description such as Ambisonics can be used.
[0013] According to the invention, the method of decoding audio field mapping for audio reproduction comprises the steps of calculating the panning function for each of a plurality of loudspeakers, using a geometric method, based on loudspeaker positions and multiple source directions, calculating a fashion matrix from source directions, calculating pseudo-inverse fashion matrix from a fashion matrix and decoding audio sound field mapping, wherein the decoding is based on a decoding matrix that is obtained from at least the pan function and the pseudo-reverse fashion matrix.
[0014] According to another aspect, the audio sound field decoding device for audio reproduction comprises a first computing element for calculating the panning function for each of the multiple speakers using a geometric method based on the position of the speakers and multiple source directions, the second computing element for calculating the fashion matrix from source directions, a third calculation element for calculating the pseudo-inverse fashion matrix and a decoding element for decoding the sound field mapping, in which the decoding is based on the decoding matrix and the decoding element uses at least the pan function and the pseudo-inverse fashion matrix to obtain the decoding matrix. The first, second and third computing elements can be a single processor or two or more separate processors.
[0015] According to yet another aspect, a computer readable medium with performance instructions written thereon for computer performing a method of decoding an audio sound field mapping for audio reproduction, comprises the steps of calculating the panning function for each of the multiple speakers using a geometric method in based on speaker positions and multiple source directions, calculating the fashion matrix from source directions, calculating a pseudo-inverse fashion matrix and decoding an audio field mapping in which the decoding is based on a decoding matrix that is obtained from at least the pan function and the pseudo-inverse fashion matrix.
[0016] Preferred embodiments of the invention are disclosed in the dependent claims, the following description and the figures.
Brief Description of the Drawing Figures [0017] Exemplary embodiments of the invention are described with reference to the accompanying drawing figures, which show:
Fig. 1 - block diagram of the method;
Fig. 2 - example 3D configuration with 16 speakers;
Fig. 3 - directional characteristics resulting from decoding using unordered mode matching;
Fig. 4 - directional characteristics resulting from decoding using an ordered fashion matrix;
Fig. 5 directional characteristics resulting from decoding using a decoding matrix obtained from VBAP;
Fig. 6 - listening test results; and Fig. 7 - block diagram of the device.
Detailed description of the invention [0018] As shown in Fig. 1, the method of decoding SFc mapping of an audio sound field for audio reproduction comprises the steps of calculating 110 W panning functions for each of the multiple speakers using a geometric method based on 102 speaker positions (L is the number of speakers) and multiple source directions 103 (S is the number of source directions), calculating 120 fashion matrix Ξ from source and given order N sound field mapping, calculating 130 pseudo-inverse fashion matrix Ξ<sup>+</sup> fashion matrix Ξ, and decoding 135, 140 mapping the SFc audio field in which the decoded audio AUdec data is obtained. The decoding is based on the decoding matrix D, which is obtained by 135 from at least the W panning function and the pseudo-inverse fashion matrix Ξ<sup>+</sup>. In one embodiment, the pseudo-inverse fashion matrix is obtained according to Ξ + = Ξ<sup>Η </sup>[ΞΞ<sup>Κ</sup>]'<sup>1</sup>. The row N of the sound field mapping can be pre-determined or can be extracted 105 from the SFc input signal.
[0019] As shown in Fig. 7, the audio sound field decoding device for audio reproduction comprises a first computing element 210 for calculating the W panning function for each of a plurality of loudspeakers, using a geometric method based on 102 loudspeaker positions and multiple source 103 directions, second calculation element 220 for calculating the fashion matrix Ξ from source directions, third calculation element 230 for calculating the pseudo-inverse fashion matrix Ξ<sup>+</sup> from the fashion matrix Ξ, and a decoding element 240 for decoding sound field mapping. Decoding is based on the decoding matrix D, which is obtained from at least the W panning function and the pseudo-inverse fashion matrix Ξ<sup>+</sup> by computing element 235 of a decoding matrix (e.g., multiplier). The decoder element 240 uses a decoding matrix D to obtain the AUdec decoded audio signal. The first, second and third calculation element 220,
230, 240 can be a single processor, or two or more separate processors. The row N of the sound field mapping can be predefined or obtained by means of the element 205 for extracting the row from the SFc input signal. [0020] A particularly useful 3D speaker configuration has 16 speakers. As shown in Fig. 2, there are four columns with three speakers in each and additional speakers between these columns. Eight speakers are evenly spaced in a circle around the listener's head, covering 45 degrees. Additional four speakers are located at the top and bottom, determining azimuth angles equal to 90 degrees. Compared to Ambisonics, this configuration is irregular and usually leads to problems with decoder construction.
[0021] Below, the vector base panning (VBAP) is described in detail. In one embodiment, VBAP is used herein to place virtual audio sources in any speaker configuration where the same speaker distances from the listening position are assumed. Three speakers are used to place the virtual source in 3D space in VBAP. For each virtual source, a monophonic signal with different amplifications is supplied to the speakers used. Gains for different speakers depend on the location of the virtual source. VBAP is a geometric method of calculating speaker gain for panning between speakers. In the case of 3D, three speakers placed in the shape of a triangle form the base of the vector. Each vector base is identified by k, m, n loudspeaker numbers and Ik, Im, In vectors of the position of the loudspeakers given in Cartesian coordinates, related to the unit length. The vector base for loudspeakers k, m, n is determined by the formula
Lkmn "{Iki Im, In} (1) [0022] The desired direction Ω = (θ, φ) of the virtual source is given by the azimuth angle φ and the elevation angle θ. The vector μ (Ω) of the position with the unit length of the virtual source in Cartesian coordinates is therefore determined by the formula ρ (Ω) = {cos <t> sinO, είηφ sin0, cos0}<sup>T</sup> (2) [0023] The location of the virtual source can be represented by the vector base and gain coefficients g ^) = (~ gk, ~ gm, ~ gn)<sup>T</sup> using the formula ρ (Ω) = L<sub>km</sub>"Ρ (Ω) = ~ g<sub>k</sub> l<sub>k</sub>+ ~ g<sub>m</sub> l<sub>m</sub>+ ~ g<sub>n</sub> l<sub>n</sub> (3) [0024] By inverting the vector base matrix, the required gain factors can be calculated using the formula <sub>9</sub>(Ω) = ΙΛπ, πΡ (Ω) (4) [0025] The vector base used is determined according to the Pulkki development: At the beginning, the gain is calculated according to the Pulkki for all vector bases. Then, the minimum value for gain factors using is calculated for each vector base<sup>~</sup>communes = min { <sup>~</sup>gk, <sup>~</sup>gm, <sup>~</sup>gn}. Finally, the vector base corresponding to the highest value is used<sup>~</sup>municipalities. The resulting gain factors cannot be negative. Depending on the acoustics of the listening room, the gain factors can be normalized for energy conservation.
[0026] The following describes the Ambisonics format, which is an exemplary sound field format. Ambisonics mapping is a way of describing the sound field, using mathematical approximation of the sound field in one location. When using a spherical coordinate system, the pressure at r = (r, θ, φ) in space is described by means of a spherical Fourier transform
<img file="PL2553947T3_D0001.tif" />
where k is a wave number. Usually, n tends to a finite value of the order M. Coefficients A<sup>m</sup>n (k) of the series describe the sound field (assuming sources outside the validity area), jn (kr) is the spherical Bessel function of the first type a
Ύ ™ π (θ, φ) denote spherical components. Coefficients A<sup>m</sup>n (k) are considered in this context as Ambisonics coefficients. Spherical components Ym n (θ, φ) depend only on the elevation angle and the azimuth angle and describe the function on the unit sphere.
[0027] For sound field reproduction, plane waves are often accepted for simplicity. Ambisonics coefficients describing a flat wave as an acoustic source from the Ωs direction are
<img file="PL2553947T3_D0002.tif" />
[0028] Their dependence on the wave number k in this particular case falls to a pure directional relationship. For a limited order of M, the coefficients form a vector that can be written as
<img file="PL2553947T3_D0003.tif" />
containing O = (M + 1)<sup>2</sup> elements. The same system is used for spherical component coefficients, giving a vector
<img file="PL2553947T3_D0004.tif" />
The superscript H denotes the transposition of a complex conjugate.
[0029] To calculate the speaker signals from the sound field Ambisonics mapping, a mode matching method is commonly used. The main idea is to express the given A (iV) description of the Ambisonics sound field by the weighted sum of the A (Ωι) descriptions of the speaker sound fields
Λ = £ «γΑίΏ;) where Ωι are the directions of the speakers, wi are weights, and L is the number of speakers. When determining the pan function from equation (8), it is assumed that the incidence direction Ωs is known. If both the source and speaker sound field are flat, 4ni<sup>n</sup> (see equation (6)) can be omitted and equation (8) depends only on the conjugated numbers of spherical components, also called "modes" here. When using matrix notation it is written as
Υ (Ω<sub>8</sub>) * = Ψ w (O<sub>s</sub>) (9) where Ψ is the fashion matrix of the speaker configuration '' (10) with elements O x L. To obtain the desired weight vector w, various known strategies are used. If we choose M = 3, then Ψ is square and can be reversible. However, due to the irregular configuration of the speakers, the matrix is not properly scaled. In this case, the pseudo inverse matrix and □ = [ψ are often chosen<sup>Η</sup>ψ] ·<sup>1</sup>ψ<sup>Η</sup> (11) gives the DL x O decoding matrix. Finally, you can write in (Q<sub>s</sub>) = DY (Q<sub>S</sub>) * (12) where the weights in ((V) are the solution to equation (9) with minimum energy. The consequences of using pseudo inversion are described below.
[0030] The relationship between panning functions and decoding matrix in the Ambisonics system is described below. Starting with Ambisonics, the panning functions for individual speakers are calculated using equation (12). Let
<img file="PL2553947T3_D0005.tif" />
will be a fashion matrix of S directions (£ 'V) of the input signal, i.e. a spherical grid, with the elevation angle changing by one degree in the range of 1 ... 180 ° and the azimuth angle in the range of 1 ... 360 ° respectively. This fashion matrix has O x S elements. Using equation (12), the resulting matrix W has L x S elements, row l contains S panning weights for the correct loudspeaker:
<img file="PL2553947T3_D0006.tif" />
[0031] As a representative example, the panning function of a single speaker 2 is shown as the directional characteristic in Fig. 3. In this example, the decoding matrix D is of the order M = 3. As can be seen, the panning function values do not refer to the physical placement of the loudspeaker at all. This is due to the mathematically irregular speaker arrangement, which is not sufficient as a spatial sampling scheme for the selected row. The decoding matrix is therefore called an unordered fashion matrix. This problem can be overcome by ordering the loudspeaker Ψ matrix in equation (11). This solution is effective at the price of the spatial resolution of the decoding matrix, which in turn can be expressed as lower-order Ambisonics. In Fig. 4 an exemplary directional characteristic is presented, resulting from decoding using an ordered fashion matrix, and especially for ordering using the average of the eigenvalues of the fashion matrix. Compared with Fig. 3, the direction of the selected speaker is now clearly distinguished.
[0032] As outlined in the introduction, for reproducing Ambisonics signals when panning functions are already known, another way of determining the decoding matrix D is possible. Functions In panning are perceived as the desired signal, specified on the set of Ω directions of the virtual source, and the mode matrix Ξ of these directions serves as the input signal. Then, you can determine the decoding matrix using the formula
D = EC<sup>H</sup>[EE<sup>H</sup>]’<sup>1</sup>= WE<sup>+</sup> (15) where Ξ<sup>H</sup> [ΞΞ ^<sup>-1</sup> or simply Ξ + is a pseudo inverse fashion matrix Ξ. In the new method, the WB panning function is taken from VBAP and the decoding matrix in the Ambisonics system is determined from this.
[0033] The panning functions for W are taken as gain values g ^), calculated using equation (4), where Ω is selected according to equation (13). The resulting decoding matrix using equation (15) is the decoding matrix in the Ambisonics system, facilitating the VBAP panning function. Fig. 5 shows an example in which the directional characteristics resulting from decoding using a decoding matrix obtained from VBAP are shown. Preferably, the SL side flaps are much smaller than the SLreg side flaps of the resulting ordered mode matching of Fig. 4. In addition, the directional characteristics obtained from VBAP for individual speakers correspond to the geometry of the speaker configuration, since the VBAP panning functions depend on the vector base of the selected direction. As a result, the new method according to the invention gives better results in all directions of speaker configuration.
[0034] Source directions 103 can be determined quite freely. The condition for the number of S directions of the source is that they must be at least (N + 1)<sup>2</sup>. Thus, given the given order N of the SFc sound field signal, it is possible to determine S according to S> (N + 1)<sup>2</sup>, and distribution of the S directions of the source evenly on the unit sphere. As mentioned above, the result may be a spherical grid with an elevation angle θ, changing by a constant value of x (e.g. x = 1 ... 5 or x = 10, 20 etc.) degrees in the range of 1 ... 180 ° and angle azimuth φ from 1 ... 360 ° respectively, in which each source direction Ω = (θ, φ) can be determined by the azimuth angle φ and the elevation angle θ.
[0035] The beneficial effect was confirmed in the listening test. To assess the location of a single source, the virtual source is compared with the real source as a reference. The speaker in the desired position was used as the actual source. VBAP playback methods, decoding in the Ambisonics mod-matching system and the newly proposed decoding system in the Ambisonics system were used, using the VBAP panning function of the present invention. In the last two methods, for each test position and each input signal tested, a third order Ambisonics signal is generated. This synthetic Ambisonics signal is then decoded using the correct decoding matrix. The test signals used are broadband pink noise and the male speech signal. Test locations are located in the frontal area with directions
Ω1 = (76.1 °, -23.2 °), Ω2 = (63.3 °, -4.3 °) (16) [0036] The listening test was carried out in an acoustic room with an average reverberation time of approximately 0.2 s. In the listening test it took nine people. Test objects were asked to evaluate the quality of spatial reproduction performance for all reproduction methods compared to the reference. A single value had to be found representing the location of the virtual source and the change in the tone of the voice. The results of the listening test are shown in Fig. 5.
[0037] As the results show, disordered mode-matching decoding in the Ambisonics system is palpably rated worse than other methods tested. This result refers to Fig. 3. The fashion matching method of the Ambisonics system serves as a reference in this listening test. Another advantage is that the confidence intervals for the noise signal are larger for VBAP than for other methods. The average values show the highest values for decoding in the Ambisonics system using the VBAP panning function. Therefore, although the spatial resolution is lowered - due to the order of the Ambisonics system used - this method has advantages over the parametric VBAP method. Compared to VBAP, both Ambisonics decoding and VBAP panning features have the advantage that not only three speakers are used to create a virtual source. In VBAP, single speakers can be dominant if the position of the virtual source is close to one of the positions of the physical positions of the speakers. Most test objects reported smaller changes in voice color with VBAP Ambisonics control than with directly applied VBAP. The problem of changing the timbre of the voice for VBAP is already known from Pulkka's work. In contrast to VBAP, in the newly proposed method, more than three speakers are used to play a virtual source, but it surprisingly produces less timbre.
[0038] In summary, a new method for obtaining the decoding matrix from the VBAP panning function in the Ambisonics system has been disclosed. For different speaker configurations, this approach is beneficial compared to the matrix matching method. The properties and consequences of using these decoding matrices are discussed above. In summary, the newly proposed Ambisonics decoding with VBAP panning features avoids the typical problems of the well-known mod matching method. The listening test showed that Ambisonics decoding, obtained from VBAP, provides spatial reproduction of better quality than that obtained by direct application of VBAP. The proposed method requires only a description of the sound field, while VBAP requires a parametric description of created virtual sources.
[0039] Although the basic, novel properties of the present invention have been demonstrated, described and indicated for use in preferred embodiments, it is understood that without departing from the spirit of the present invention by those skilled in the art, various omissions and substitutions may be made and changes in the device and method described, in the shape and details of the devices disclosed and in their operation. It is expressly intended that all combinations of those elements that perform essentially the same function in substantially the same manner to achieve the same results fall within the scope of the invention. Substitutions of elements between the described embodiments are also fully intended and thought out. It is understood that modifications of details can be made without departing from the scope of the invention. Each property disclosed in the description and (where applicable) in the claims and the drawing may be entered separately or in any appropriate combination. The properties may, where appropriate, be implemented in hardware, in software, or in a combination of both. The reference numbers in the claims are for illustrative purposes only and do not limit the scope of the claims.
74 members in 12 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 10305316 | European Patent Office (EPO) | A | |
| 2011054644 | European Patent Office (EPO) | W |
Members74
| Document | Office | Kind | |
|---|---|---|---|
| WO2011117399A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2011231565A1 | Australia | A1 | |
| CN102823277A | China | A | |
| US2013010971A1 | United States of America | A1 | |
| EP2553947A1 | European Patent Office (EPO) | A1 | |
| KR20130031823A | Republic of Korea | A | |
| HK1174763A | Hong Kong, China | A | |
| HK1174763A1 | Hong Kong, China | A1 | |
| JP2013524564A | Japan | A | |
| EP2553947B1 | European Patent Office (EPO) | B1 | |
| PT2553947E | Portugal | E | |
| ES2472456T3 | Spain | T3 | |
| JP5559415B2 | Japan | B2 | |
| AU2011231565B2 | Australia | B2 | |
| PL2553947T3This record | Poland | T3 | |
| JP2014161122A | Japan | A | |
| JP5739041B2 | Japan | B2 | |
| CN102823277B | China | B | |
| US9100768B2 | United States of America | B2 | |
| JP2015159598A | Japan | A | |
| US2015294672A1 | United States of America | A1 | |
| BR112012024528A2 | Brazil | A2 | |
| US9460726B2 | United States of America | B2 | |
| JP6067773B2 | Japan | B2 | |
| US2017025127A1 | United States of America | A1 | |
| JP2017085620A | Japan | A | |
| KR101755531B1 | Republic of Korea | B1 | |
| KR20170084335A | Republic of Korea | A | |
| US9767813B2 | United States of America | B2 | |
| KR101795015B1 | Republic of Korea | B1 | |
| KR20170125138A | Republic of Korea | A | |
| BR112012024528A8 | Brazil | A8 | |
| US2017372709A1 | United States of America | A1 | |
| JP6336558B2 | Japan | B2 | |
| US10037762B2 | United States of America | B2 | |
| KR101890229B1 | Republic of Korea | B1 | |
| KR20180094144A | Republic of Korea | A | |
| JP2018137818A | Japan | A | |
| US2018308498A1 | United States of America | A1 | |
| US10134405B2 | United States of America | B2 | |
| KR101953279B1 | Republic of Korea | B1 | |
| KR20190022914A | Republic of Korea | A | |
| US2019139555A1 | United States of America | A1 | |
| KR102018824B1 | Republic of Korea | B1 | |
| KR20190104450A | Republic of Korea | A | |
| US2019341062A1 | United States of America | A1 | |
| JP6615936B2 | Japan | B2 | |
| US10522159B2 | United States of America | B2 | |
| JP2020039148A | Japan | A | |
| KR102093390B1 | Republic of Korea | B1 | |
| KR20200033997A | Republic of Korea | A | |
| US10629211B2 | United States of America | B2 | |
| US2020273470A1 | United States of America | A1 | |
| BR122020001822B1 | Brazil | B1 | |
| BR112012024528B1 | Brazil | B1 | |
| JP6918896B2 | Japan | B2 | |
| KR102294460B1 | Republic of Korea | B1 | |
| KR20210107165A | Republic of Korea | A | |
| JP2021184611A | Japan | A | |
| US11217258B2 | United States of America | B2 | |
| US2022189492A1 | United States of America | A1 | |
| JP7220749B2 | Japan | B2 | |
| JP2023052781A | Japan | A | |
| KR102622947B1 | Republic of Korea | B1 | |
| KR20240009530A | Republic of Korea | A | |
| US11948583B2 | United States of America | B2 | |
| US2024304195A1 | United States of America | A1 | |
| JP7551795B2 | Japan | B2 | |
| JP2024164284A | Japan | A | |
| US12283279B2 | United States of America | B2 | |
| KR102803833B1 | Republic of Korea | B1 | |
| KR20250061865A | Republic of Korea | A | |
| JP7725680B2 | Japan | B2 | |
| JP2025163200A | Japan | A |
Numbers
- Application
- 11709968
Titles2
- English
- METHOD AND DEVICE FOR DECODING AN AUDIO SOUNDFIELD REPRESENTATION FOR AUDIO PLAYBACK
- Polish
- Sposób i urządzenie do dekodowania odwzorowania pola dźwiękowego audio do odtwarzania audio
Classification
- CPC, 5
- G10L19/008
- H04S3/02
- H04S2420/11
- H04S7/308
- H04S2400/13
- IPC, 3
- H04S3 02
- G10L19 00
- G10L19 008