An apparatus for determining a converted spatial audio signal
Abstract
This record has no abstract on file.
Term
2.4 yearsto projected expiry
Projected expiry 2 February 2029, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
13 claims: 3 independent, 10 dependent
- 1Patent claims Zastrzeżenia patentowe 1. A device (100) for determining a converted spatial audio signal, the converted spatial audio signal comprising an omnidirectional audio component (W) and at least one directional audio component (X;Y;Z) from the input spatial audio signal, wherein the input spatial signal audio contains the input audio representation (P), time and frequency dependent diffuseness parameter (Ψ) and the input direction of arrival (eDOA), containing an estimator (110) for estimating a wave representation comprising a wave field measure ((e (k, n) P (k, n)) and a measure of the direction of arrival of the wave (eDOA, x, eDOA, y, eDOA, z), the estimator (110) is adapted to determine the wave representation based on the input audio representation (P), dispersion parameter (Ψ) and input direction of arrival (eDOA), wherein the estimator (110) is adapted to determine the wave field measure based on the fraction (β (λ> / 7)) of the input audio representation (P (k, n)), with the fraction (e (k, / 7)) and the input audio representation are time dependent and frequency dependent and where the fraction (e (k, / 7)) is calculated based on the spread parameter (Ψ (Α> / 7));and a processor (120) for processing wave field measures (e (k, n) P (k, n) and measures (eDOA, x, eDOA, y, eDOA, z) of the direction of arrival of the wave to obtain at least one directional component (X ;Y;Z), where the omnidirectional component (W) is equal to the input audio representation. 1. Urządzenie (100) do wyznaczania konwertowanego przestrzennego sygnału audio, przy czym konwertowany przestrzenny sygnał audio zawiera dookolną składową audio (W) i co najmniej jedną składową kierunkową audio (X;Y;Z), z wejściowego przestrzennego sygnału audio, przy czym wejściowy przestrzenny sygnał audio zawiera wejściową reprezentację audio (P), zależny od czasu i od częstotliwości parametr rozproszenia (Ψ) oraz wejściowy kierunek nadejścia (eDOA), zawierające estymator (110) do estymacji reprezentacji falowej zawierającej miarę pola falowego ((e(k,n)P(k,n)) i miarę kierunku nadejścia fali (eDOA,x, eDOA,y, eDOA,z), przy czym estymator (110) jest przystosowany do wyznaczania reprezentacji falowej w oparciu o wejściową reprezentację audio (P), parametr rozproszenia (Ψ) i wejściowy kierunek nadejścia (eDOA), przy czym estymator (110) jest przystosowany do wyznaczania miary pola falowego w oparciu o ułamek (β(λ>/7)) wejściowej reprezentacji audio (P(k,n)), przy czym ułamek (e(k,/7)) i wejściowa reprezentacja audio są zależne od czasu i zależne od częstotliwości i przy czym ułamek (e(k,/7)) jest obliczony w oparciu o parametr rozproszenia (Ψ(Α>/7));oraz procesor (120) do przetwarzania miary pola falowego (e(k,n)P(k,n) i miary (eDOA,x, eDOA,y, eDOA,z) kierunku nadejścia fali dla uzyskania co najmniej jednej składowej kierunkowej (X;Y;Z), przy czym składowa dookolna (W) jest równa wejściowej reprezentacji audio.
- 7A device (300) for determining a combined converted spatial audio signal, the converted spatial audio signal comprising at least a first combined component and a second combined component from the first and second input spatial audio signals, wherein the first input spatial audio signal comprises a first input audio representation, first direction of arrival and first diffuseness parameter depending on time and frequency, including:7. Urządzenie (300) do wyznaczania połączonego konwertowanego przestrzennego sygnału audio, przy czym konwertowany przestrzenny sygnał audio zawiera co najmniej pierwszą połączoną składową i drugą połączoną składową z pierwszego i drugiego wejściowego przestrzennego sygnału audio, przy czym pierwszy wejściowy przestrzenny sygnał audio zawiera pierwszą wejściową reprezentację audio, pierwszy kierunek nadejścia i pierwszy parametr rozproszenia zależny od czasu i zależny od częstotliwości, zawierające: a first device (101) as defined in one of claims 1 to 6 for providing a first converted signal comprising a first omnidirectional component from the first device and at least one directional component from the first device (101);pierwsze urządzenie (101) określone w jednym z zastrzeżeń od 1 do 6, do dostarczania pierwszego sygnału konwertowanego, zawierającego pierwszą składową dookolną, z pierwszego urządzenia i co najmniej jedną składową kierunkową z pierwszego urządzenia (101);a second device (102) as defined in one of claims 1 to 6, for providing a second converted signal comprising a second omnidirectional component from the second device and at least one directional component from the second device (102);drugie urządzenie (102) określone w jednym z zastrzeżeń od 1 do 6, do dostarczania drugiego konwertowanego sygnału, zawierającego drugą składowa dookolną z drugiego urządzenia i co najmniej jedną składowa kierunkową z drugiego urządzenia (102);an audio effects generator (301) for rendering the first omnidirectional component of the first device or the directional component of the first device (101) to obtain the first rendered component;generator (301) efektów audio do renderowania pierwszej składowej dookolnej z pierwszego urządzenia lub składowej kierunkowej z pierwszego urządzenia (101) dla uzyskania pierwszej składowej renderowanej;a first connecting module (311) for combining the first rendered component, the first omnidirectional component and the second omnidirectional component, or for combining the first rendered component, the directional component from the first device (101) and the directional component from the second device (102) to obtain the first combined component;and a second connecting module (312) for connecting the directional component from the first device (101) and the directional component from the second device (102), or for connecting the first omnidirectional component and the second omnidirectional component to obtain a second combined component. pierwszy moduł łączący (311) do łączenia pierwszej składowej renderowanej, pierwszej składowej dookolnej i drugiej składowej dookolnej, lub do łączenia pierwszej składowej renderowanej, składowej kierunkowej z pierwszego urządzenia (101) i składowej kierunkowej z drugiego urządzenia (102) dla uzyskania pierwszej składowej łączonej;oraz drugi moduł łączący (312) do łączenia składowej kierunkowej z pierwszego urządzenia (101) i składowej kierunkowej z drugiego urządzenia (102), lub do łączenia pierwszej składowej dookolnej i drugiej składowej dookolnej dla uzyskania drugiej składowej łączonej.
- 12A method of determining a converted spatial audio signal, the converted spatial audio signal comprising an omnidirectional audio component (W) and at least one directional audio component (X;Y;Z) from the input spatial audio signal, wherein the input spatial audio signal comprises the input audio representation (P) and the time and frequency dependent scatter parameter (Ψ) and the input direction of arrival (eDOA) including the steps of estimating the waveform representation containing the wave field measure (e (k, n) P ( k, n)) and a measure of the direction of arrival of the wave (eDOA, x, eDOA, y, eDOA, z), whereby the wave representation is estimated based on the input audio representation (P), scatter parameter (Ψ) and input direction of arrival (eDOA), where the wave field measure is determined based on the fraction (β (,)) of the input audio representation (P (k, n)), with the fraction (e (k, / 7)) and the input audio representation are time-dependent and frequency-dependent and where the fraction (e (k, / 7)) is calculated based on the spread parameter ^ (kn));and processing wave field measures (e (k, n) P (k, n)) and measures (eDOA, x, eDOA, y, eDOA, z) the direction of arrival of the wave to obtain at least one directional component (X;Y;Z ), where the omnidirectional component (W) is equal to the input audio representation. 12. Sposób wyznaczania konwertowanego przestrzennego sygnału audio, przy czym konwertowany przestrzenny sygnał audio zawiera dookolną składową audio (W) i co najmniej jedną składową kierunkową audio (X;Y;Z), z wejściowego przestrzennego sygnału audio, przy czym wejściowy przestrzenny sygnał audio zawiera wejściową reprezentacje audio (P) i zależny od czasu i od częstotliwości parametr rozproszenia (Ψ) oraz wejściowy kierunek nadejścia (eDOA) obejmujący etapy estymacji reprezentacji falowej zawierającej miarę pola falowego (e(k,n)P(k,n)) i miarę kierunku nadejścia fali (eDOA,x, eDOA,y, eDOA,z), przy czym reprezentacja falowa jest estymowana w oparciu o wejściową reprezentację audio (P), parametr rozproszenia (Ψ) i wejściowy kierunek nadejścia (eDOA), przy czym miara pola falowego jest wyznaczana w oparciu o ułamek (β(,)) wejściowej reprezentacji audio (P(k,n)), przy czym ułamek (e(k,/7)) i wejściowa reprezentacja audio są zależne od czasu i zależne od częstotliwości i przy czym ułamek (e(k,/7)) jest obliczany w oparciu o parametr rozproszenia ^(kn));oraz przetwarzania miary pola falowego (e(k,n)P(k,n)) i miary (eDOA,x, eDOA,y, eDOA,z) kierunku nadejścia fali dla uzyskania co najmniej jednej składowej kierunkowej (X;Y;Z), przy czym składowa dookolna (W) jest równa wejściowej reprezentacji audio.
Independent claims3
188 paragraphs in 2 sections, as filed
[0001] The present invention relates to the field of audio processing, in particular spatial audio processing and the conversion of various spatial formats.
[0002] DirAC (Directional Audio Coding) audio coding is a method for reproducing and processing spatial audio. Conventional systems use DirAC in two-dimensional and three-dimensional high-quality reproduction of recorded sound, in teleconferencing applications, in directional microphones and in stereo-to-surround mix, see. V. Pulkki and C. Faller, Directional audio coding: Filterbank and STFT-based design, 120th AES Convention, May 20-23, 2006, Paris, France May 2006, V. Pulkki and C. Faller, Directional audio coding in spatial sound reproduction and stereo upmixing, AES 28th International Conference, Pitea, Sweden, June 2006, V. Pulkki, Spatial sound reproduction with directional audio coding, Journal of the Audio Engineering Society, 55 (6): 503-516, June 2007, Jukka Ahonen, V. Pulkki i Tapio Lokki, Teleconference application and B-format microphone array for directional audio coding, 30th AES International Conference.
[0003] Other conventional applications using DirAC are, for example, the universal format for coding and noise cancellation. In the DirAC format, some directional sound properties are analyzed in frequency bands depending on time. The analysis data is transmitted along with the audio data and synthesized for various purposes. The analysis is usually performed using B-format signals, although theoretically DirAC is not limited to this format. B-format, see Michael Gerzon, Surround sound psychoacoustics, Wireless World, volume 80, pages 483-486, December 1974. It was developed in the work on Ambisonics, a system developed by British researchers in the 1970s to transfer surround sound from concert halls to home salons. The B-format consists of four signals, namely in (t), x (t), y (t), and z (t). The first corresponds to the pressure measured by the omni-directional microphone, while the last three are pressure readings of the microphones directed towards the three axes of the Cartesian coordinate system. The signals x (t), y (t) and z (t) are proportional to the components of the velocity vector of particles directed along the x, y and z axes, respectively.
[0004] The DirAC stream consists of 1-4 audio channels with directional meta data.
In teleconferencing and in some other cases, the stream consists of only one audio channel with metadata, called the DirAC stream. This is a very compact way of describing spatial audio, because only one audio channel must be sent along with additional information, which, for example, ensures good spatial separation between speakers. However, in these cases, certain types of sound, such as reflected sound and ambient sound, may be produced with limited quality. To ensure better quality in these cases, additional audio channels must be transmitted.
[0005] The conversion from B-format to DirAC is described in V. Pulkki, A method for reproducing natural or modified spatial impression in multichannel listening, Patent WO 2004/077884 A1, September 2004. Directional audio coding is an efficient approach to analysis and surround sound reproduction. DirAC uses parametric representation of sound fields based on properties that are important for the perception of surround sound, namely DOA (DOA, direction of arrival) and sound field dispersion in frequency subbands. In fact, DirAC assumes that interaural time differences (ITD) and interaural level differences (ITD) are perceived correctly when sound field DOA is reproduced correctly, while interaural coherence (IC) is received correctly if the dispersion is reproduced accurately. These parameters, i.e. DOA and diffuseness, represent additional information that accompanies the mono signal in the so-called mono DirAC stream.
[0006] Fig. 7 shows the DirAC encoder, which from the correct microphone signals calculates the mono audio channel and additional information, i.e. spread enie (,) and arrival direction eDOA (k, n). Fig. 7 shows a DirAC encoder 200 that is adapted to calculate the mono audio channel and additional information from valid microphone signals. In other words, Fig. 7 shows a DirAC encoder 200 for determining the spread and direction of arrival from valid microphone signals. FIG. 7 shows the DirAC 200 encoder comprising the P / U estimation unit 210, where P (k, n) represents the pressure signal and U (k, n) represents the particle velocity vector. The P / U estimation unit receives the microphone signals as information on which the P / U estimation is based. The energy analysis member 220 enables estimation of the arrival direction and the DirAC mono stream dispersion parameter.
[0007] DirAC parameters, e.g. W (k, n) mono audio representation, W (kn) spread parameter and arrival direction (DOA) eDOA (k, n) can be obtained from the frequency-time representation of microphone signals. Hence, the parameters depend on time and frequency. On the playback side, this information enables accurate spatial rendering. A multi-speaker system is required to reproduce surround sound in the desired listening position. However, its geometry can be any. In fact, the speaker channels can be set as a function of the DirAC parameters.
[0008] There are significant differences between the DirAC format and parametric multi-channel audio coding, such as MPEG surround, (see: Lars Villemoes,
Juergen Herre, Jeroen Breebaart, Gerard Hotho, Sascha Disch, Heiko Purnhagen, and Kristofer
Kjrlingm, MPEG surround: The forthcoming ISO standard for spatial audio coding, AES 28th International Conference, Pitea, Sweden, June 2006) although they share very similar processing structures. While MPEG surround is based on the time-frequency analysis of channels of different speakers, the DirAC format treats as input signals coincidental microphone channels that effectively describe the sound field of one point. So the DirAC codec also represents efficient recording technology for spatial audio.
[0009] Another system dealing with spatial sound is coding spatial objects SAOC (Spacial Audio Object Coding), see Jonas Engdegard,
Barbara Resch, Cornelia Falch, Oliver Hellmuth, Johannes Hilpert, Andreas Hoelzer, Leonid Ternetiev, Jeroen Breebaart, Jeroen Koppens, Erik Schuijer, and Werner Oomen, Spatial audio object coding (SAOC) the upcoming MPEG standard on parametric object based audio coding, 124th AES Convention, 17-20 May 2008, Amsterdam, The Netherlands), which are currently at the ISO / MPEG standardization stage. It uses the MPEG Surround rendering mechanism and treats various sound sources as objects. Audio coding offers very high performance in terms of bit rate and unprecedented freedom of interaction on the playback side. This approach promises new attractive features and functions in existing systems as well as a number of other innovative applications.
[0010] Document 2006/0045275 A1 discloses a method for processing audio data and an audio acquisition device implementing this method. The method consists of coding signals representing sound propagating in three-dimensional space and obtained from a source located at the first distance from the reference point, in order to obtain a sound representation through the components expressed, the spherical harmonic base and regarding the said component compensation in the field of the near-field effect.
[0011] In the publication "A Distributed System for the Creation and Delivery of Ambisonic Surround Sound Audio" by R. Foss and A. Smith, AES 16th International Conference, 1999, pages
116-125, a system for reproducing ambisonic surround sound compositions utilizing client-server architecture is disclosed. Monaural audio data and three-dimensional coordinates are converted to a B-format audio representation that is decoded into a series of loudspeakers for ambisonic surround sound.
[0012] US 6,259,795 B1 discloses a method and apparatus for spatialized audio processing, in which at least one head transfer function is performed for each spatial component of a sound field having positional spatial components to produce a series of transmission signals. Transmission signals are sent to multiple users, and the current user orientation is determined for each of the multiple users and a current orientation signal indicative of it is generated, which is then used to mix transmission signals to play to the user. The sound field signal may contain a B-format signal.
[0013] It is an object of the present invention to provide an improved concept for spatial processing.
[0014] The object is achieved by a device for determining a converted spatial audio signal according to claim 1 and the corresponding method according to claim 12.
[0015] The present invention is based on the finding that improved spatial audio can be obtained e.g. when converting a spatial audio signal encoded as a mono DirAC stream to a B-format signal. In embodiments, the converted B-format signal may be processed or rendered before it is added to other audio signals and encoded back into the DirAC stream. The embodiments may have various applications, e.g. mixing different types of DirAC and B-format streams, based on DirAC etc. Embodiments can introduce the reverse operation to disclosed in WO 2004/077884 A1, i.e. to convert from mono to B-format DirAC stream.
[0016] The present invention is based on the finding that improved processing can be obtained if audio signals are converted to directional components. In other words, it is a finding of the present invention that improved spatial processing can be obtained when the format of the spatial audio signal corresponds to directional components recorded, for example, by a directional B-format microphone. Furthermore, it is a finding of the present invention that directional and omnidirectional components from various sources can be processed together, and thus with greater efficiency. In other words, especially when processing spatial audio signals from multiple audio sources, processing can be performed more efficiently if the signals from multiple audio sources are available in the format of their omnidirectional and directional components because they can be processed together. In embodiments, audio effect generators or audio processors can therefore be used more efficiently by processing combined components from multiple sources.
[0017] In embodiments of the invention, spatial audio signals may be represented as a mono DirAC stream specifying a DirAC streaming technique in which only one audio channel is associated with the media data during transmission. This format can be converted, for example, to a B-format stream that has many directional components. Embodiments may allow improved spatial audio processing by converting spatial audio signals to directional components.
[0018] Variants of the invention may provide an advantage over decoding
DirAC mono, when only one audio channel is used to create all speaker signals, in that additional spatial processing is enabled based on the directional audio components that are determined before the speaker signals are created. Embodiments may offer the advantage that the problems associated with reverberated sound creation are reduced.
[0019] In embodiments of the invention, for example, the DirAC stream may use a stereo audio signal instead of a mono audio signal when the stereo channels L (left stereo channel) and R (right stereo channel) are present and are sent for use in DirAC decoding. Embodiments can achieve better quality of the reverberated sound and, for example, provide direct compatibility with the stereo speaker system.
[0020] Embodiments may offer the advantage that virtual microphone DirAC decoding is possible. Details of virtual microphone DirAC decoding can be found in V. Pulkki, Spatial sound reproduction with directional audio coding, Journal of the Audio Engineering Society, 55 (6): 503-516, June 2007. These embodiments allow you to obtain audio signals for the speakers by placing virtual microphones oriented towards the position of the speakers and having point sources of sound whose position is determined by the DirAC parameters. Embodiments may offer the advantage that convenient conversion of audio signals is enabled through conversion.
[0021] Embodiments of the present invention will be described in detail using the attached figures in which
Fig. 1a shows an embodiment of a device for determining a converted spatial audio signal;
Fig. 1b shows the pressure and components of a Gaussian plane velocity vector for a plane wave;
Fig. 2 shows another embodiment for converting a mono DirAC stream to a B-format signal;
Fig. 3 shows an embodiment for combining multiple converted spatial audio signals;
Figures 4a to 4d show embodiments for combining multiple DirAC-based spatial audio signals using various audio effects;
Fig. 5 shows an embodiment of an audio effect generator;
Fig. 6 shows an embodiment of an audio effect generator implementing multiple audio effects in directional components; and Fig. 7 shows a prior art DirAC encoder.
[0022] Fig. 1a shows an apparatus 100 for determining a converted spatial audio signal, the converted spatial audio signal comprising an omnidirectional component and at least one directional component (X; Y; Z) from the input spatial audio signal, the input spatial audio signal includes audio input representation (W) and input arrival direction (φ).
[0023] The apparatus 100 includes an estimator 110 for estimating a waveform representation comprising a wave field measure and a wave direction arrival measure based on the audio input representation (W) and the input arrival direction (φ). In addition, the device 100 includes a processor 120 for processing the wave field measure and the direction of arrival of the wave to obtain an omnidirectional component and at least one directional component. Estimator 110 may be adapted to estimate the wave representation as a plane wave representation.
[0024] In embodiments of the invention, the processor may be adapted to provide the input audio representation (W) as an omnidirectional audio component (W '). In other words, the omnidirectional component W 'is equal to the input audio representation W. Hence, according to the dotted lines in Fig. 1a, the input audio representation can bypass estimator 110, processor 120, or both. In other embodiments, the omnidirectional audio component W 'may be based on the wave intensity and the direction of arrival of the wave, processed by the processor 120 together with the input audio representation. In embodiments, multiple directional audio components (X; Y; Z) may be processed, such as e.g. the first (X), second (Y) and / or third (Z) directional audio component corresponding to different spatial directions. In embodiments, for example, three different directional audio components (X; Y; Z) may be obtained according to different directions of the Cartesian coordinate system.
[0025] The estimator 110 may be adapted to estimate the wave field measure within the wave field amplitude and wave field phase. In other words, in variants of the invention, the wave field measure can be estimated as a complex quantity. The wave field amplitude may correspond to the sound pressure module, and the wave field phase may correspond to the sound pressure phase, in some embodiments.
[0026] In embodiments of the invention, the measure of the direction of arrival of the wave may correspond to any direction-defining quantity, e.g. a vector, one or more angles etc., and may be obtained from any direction measure representing the audio component, e.g. intensity vector, velocity vector particles, etc. The field measure may correspond to any physical quantity describing the audio component, which may be real or complex, may correspond to pressure signal, amplitude or modulus of particle velocity, loudness, etc. In addition, these measures may be in the time domain and / or in the frequency domain.
[0027] Variants of the present invention may be based on the estimation of a plane wave representation for each of the input streams, which may take place in the estimator 110 shown in Fig. 1a. In other words, a wave field measure can be modeled using a plane wave representation. Generally speaking, there are a number of equivalent, exhaustive (i.e. full) descriptions of plane waves or plane waves in general. Below, a mathematical description will be introduced for calculating dispersion parameters and arrival directions or direction measures for various components. Although only a few descriptions relate directly to physical quantities, such as pressure, particle velocity, etc., there is potentially an infinite number of different ways to describe wave representations, one of which will be presented below as an example, but not intended to be limited in any way. embodiments of the present invention. Any combination may correspond to a wave field measure and a measure of wave arrival direction.
[0028] To further detail the various potential descriptions, two numbers a and b are considered. The information contained in waib can be sent by sending c and d when
<img file="PL2154677T3_D0001.tif" />
where Ω is known as a 2x2 matrix. The example applies only to linear combinations, but in general any combinations are possible, i.e. also non-linear.
[0029] Below, scalar sizes are represented by lower case letters a, b, c while column vectors are represented by bold lower case letters a, b, c. Superscript ()<sup>T</sup> indicates transposition, respectively, while (·) and (·) * represent the complex conjugate. The complex phasor notation differs from the time notation. For example, the pressure p (t), which is a real number, and from which a possible measure of the wave field can be obtained, can be expressed by means of the phasor P, which is an imaginary number, and from which another possible measure of the wave field can be obtained, by the formula
<img file="PL2154677T3_D0002.tif" />
where Re {·} is the real part and c> = 2nf is the angular frequency. In addition, capital letters used for physical quantities, in the following represent phasors. For the introductory example below, and to avoid ambiguities, it should be noted that all quantities considered below with the "PW" subscript refer to plane waves.
[0030] For an ideal monochrome plane wave, the vector velocity of UPW particles can be written as
<img file="PL2154677T3_D0003.tif" />
where the unit vector ed indicates the direction of wave propagation, e.g. by corresponding direction measure. You can prove that
<img file="PL2154677T3_D0004.tif" />
where Ia means active intensity, ρ0 means air density, c means speed of sound, E means sound field energy Ψ means dispersion.
[0031] It is noteworthy that since all ed components are real numbers, all UPW components are in phase with PPW. Fig. 1b shows examples of UPW and PPW in the Gaussian plane. As mentioned just now, all the components of the PCA have the same phase as the PPWs, specifically θ. On the other hand, their modules are linked as follows
<img file="PL2154677T3_D0005.tif" />
[0032] Embodiments of the present invention may provide a method for converting a mono DirAC stream to a B-format signal. The mono DirAC stream can be represented by a pressure signal captured, for example, by an omni-directional microphone and by additional information. Additional information may include time-frequency-dependent measures of dispersion and direction of arrival.
[0033] In embodiments, the input spatial audio signal may further include a spread parameter Ψ, and the estimator 110 may be adapted to estimate the wave field measure in addition based on the spread parameter Ψ.
[0034] The input direction of arrival and the measure of the direction of arrival of the wave may refer to a reference point corresponding to the recording location of the input spatial audio signal, i.e. in other words all directions may refer to the same reference point. The reference point can be the place where the microphone or multiple directional microphones are placed to record the sound field.
[0035] In embodiments, the converted spatial audio signal may include a first (X), a second (Y) and a third (Z) directional component. The processor 120 may be adapted to further process the wave field measure and the direction of arrival measure to obtain the first (X) and / or second (Y) and / or third (Z) directional component and / or omnidirectional audio components.
[0036] The notation and data model will be introduced below.
[0037] Let p (t) iu (t) = [ux (t), uy (t), uz (t)]<sup>T</sup> will be respectively the pressure and the vector of particle velocity, for a specific point in space, where [•]<sup>T</sup> means transposition. p (t) may correspond to the audio representation, au (t) = [ux (t), uy (t), uz (t)]<sup>T</sup> may correspond to directional components. These signals can be processed into the time-frequency domain using the appropriate filter bank or STFT (Short Time Fourier Transform) as suggested, for example, in publication V.
Pulkki and C. Faller, Directional audio coding: Filterbank and STFT-based design, 120th AES
Convention, May 20-23, 2006, Paris, France, May 2006.
[0038] Let P (k, n) and U (k, n) = [Ux (k, n), Uy (k, n), Uz (k, n)]<sup>T</sup> denote processed signals, where the kin are indexes for frequency (or frequency band) and time, respectively. The active intensity vector Ia (k, n) can be defined as
<img file="PL2154677T3_D0006.tif" />
where (·) * is the complex conjugate and Re {·} gets the real part. The active intensity vector can express the net energy flow characterizing the sound field, see FJ Fahy, Sound Intensity, Essex: Elsevier Science Publishers Ltd., 1989.
[0039] Let c denote the velocity of sound in the considered medium and E denote sound field energy defined by FJ Fahy
<img file="PL2154677T3_D0007.tif" />
where Μ calculates the 2nd norm. The compactness of the mono DirAC stream will be detailed below.
[0040] A mono DirAC stream may consist of a mono p (t) signal or an audio representation and additional information, e.g. an arrival direction measure. This additional information may include a time-frequency direction of arrival and a time-frequency dispersion measure. The first can be designated as eDOA (k, n), which is a unit vector directed in the direction from which the sound comes, i.e. it can model the direction of arrival. Second, dispersion, can be marked by
<img file="PL2154677T3_D0008.tif" />
[0041] In embodiments, the estimator 110 and / or processor 120 may be adapted to estimate / process the input DOA and / or measure the wave DOA in the eDOA unit vector range (k, n). The direction of arrival can be obtained as
<img file="PL2154677T3_D0009.tif" />
where the unit vector eI (k, n) indicates the direction indicated by the active intensity, i.e., respectively
<img file="PL2154677T3_D0010.tif" />
Alternatively, in the embodiments, the DOA or DOA measure can be expressed by azimuth and center angle in a spherical coordinate system. For example, if φ (,) and ^ (kn) are azimuth and elevation angles, respectively
<img file="PL2154677T3_D0011.tif" />
where eDOA, x (k, n) is a component of the eDOA (k, n) unit vector of the input direction of arrival along the x axis in the Cartesian coordinate system, eDOA, y (k, n) is the eDOA (k, n) component along the ya eDOA axis, z (k, n) is the eDOA component (k, n) along the z axis.
[0042] In embodiments, the estimator 110 may be adapted to estimate the wave field measure additionally based on the spread parameter Ψ, optionally also expressed by Ψ (k, n) in a time-frequency manner. Estimator 110 may be adapted for estimation based on a spread through parameter
<img file="PL2154677T3_D0012.tif" />
where <·> is the time average.
[0043] There are various strategies for obtaining P (k, n) and U (k, n) in practice. One option is to use a B-format microphone that provides 4 signals, namely in (t), x (t), y (t) and z (t). First, in (t) can correspond to the pressure reading of an omni-directional microphone. The next three may correspond to the microphone pressure readings directed in the directions of the three axes of the Cartesian coordinate system. These signals are also proportional to the speed of the particles. Hence, in some embodiments
<img file="PL2154677T3_D0013.tif" />
where W (k, n), X (k, n), Y (k, n) and Z (k, n) are processed B-format signals corresponding to the omnidirectional component W (k, n) and the three directional components X (k , n), Y (kn), Z (kn). It should be noted that the V2 factor in equation (6) comes from the convention used in the definition of B-format signals, see Michael Gerzon, Surround sound psychoacoustics, Wireless World, volume 80, pages 483-486, December 1974.
[0044] Alternatively, P (k, n) and U (k, n) can be estimated using a multidirectional microphone grid, as suggested in J. Merimaa, Applications of a 3-D microphone array, 112th AES Convention, Paper 5501, Munich, May 2002. The processing steps described above are also illustrated in Fig. 7.
[0045] Fig. 7 shows a DirAC encoder 200 that is adapted to calculate the mono audio channel and additional information from appropriate microphone signals. In other words, Fig. 7 illustrates the DirAC 200 encoder for determining the rozprosz (Αη) spread and the arrival direction of eDOA (k, n) from appropriate microphone signals. Fig. 7 shows a DirAC encoder 200 comprising a P / U estimation unit 210. The P / U estimation unit receives the microphone signals as input information on which the P / U estimation is based. Since all information is available, the P / U estimation is trivial, according to the above equations. The energy analysis member 220 allows estimation of the arrival direction and the dispersion parameter of the combined stream.
[0046] In embodiments, the estimator 110 may be adapted to determine the wave field measure or amplitude based on the fraction β (,) of the input audio representation P (k, n). Fig. 2 shows the processing steps of an embodiment for calculating B-format signals from a mono DirAC stream. All quantities depend on time and frequency indices (k, n) and are below, partially simplified for simplicity.
[0047] In other words, Fig. 2 shows another embodiment. According to equation (6), W (k, n) is equal to the pressure P (k, n). Therefore, the problem of B-format synthesis from the mono DirAC stream is limited to the estimation of the velocity vector of particles U (k, n), because its components are proportional to X (k, n), Y (k, n) and Z (k, n).
[0048] Embodiments may approach the estimation based on the assumption that the field consists of a plane wave summed to the scattering field. Hence, the pressure and speed of the particles can be expressed as
<img file="PL2154677T3_D0014.tif" />
where the subscripts "PW" and "diff" denote plane wave and scattering field, respectively.
[0049] The DirAC parameters contain information regarding active intensity only. Therefore, the vector Ukn) of the particle velocity is estimated as UFw (k, n), which is an estimator for the particle velocity only for the plane wave. Can be defined as
<img file="PL2154677T3_D0015.tif" />
where the real number β (k, n) is the proper weighting factor, which generally depends on the frequency and can display inverse proportionality to the dispersion Ψ (,). In fact, for low scattering, i.e. Ψ (k, n) close to 0, it can be assumed that the field consists of a single plane wave, so that
<img file="PL2154677T3_D0016.tif" />
implying that β (,) = 1.
[0050] Considering the above equation and equation (6), the omnidirectional and / or first and / or second and / or third directional components may be expressed as
Χ (Μ) = P ^ nj-e ^ k, ^ γ- (11)
Κ (Μ) = W ») 'P (k,«) · (k<sub>t</sub>«)
Z {k<sub>t</sub> η) = η} - P (k> η) e<sub>MA</sub>.<sub>t</sub> (ł. «) where eDOA, x (k, n) is a component of the eDOA (k, n) unit vector of the input direction of arrival along the x axis of the Cartesian coordinate system, eDOA, y (k, n) is the component of eDOA (k, n) along the y axis, eDOA, z (k, n) is a component of eDOA (k, n) along the z axis. In the embodiment shown in Fig. 2 the arrival direction measure estimated by the estimator 110 corresponds to eDOA, x (k, n), eDOA, y (k, n) and eDOA, z (k, n), and the wave field measure corresponds to e (k, n) P (k , n) The first directional component derived by the processor 120 may correspond to any of X (k, n), Y (k, n) or Z (k, n), and the second directional component of any remaining X (k, n), respectively, Y (k, n) and Z (k, n).
[0051] Two practical examples of the method for determining the coefficient e (kn) will be presented below.
[0052] The first embodiment aims first to estimate the plane wave pressure, i.e. PPW (k, n), and then to obtain a particle velocity vector therefrom.
[0053] By setting the air density ρ0 equal to 1 and ignoring the functional relationship (k, n) for simplicity, we can write
<img file="PL2154677T3_D0017.tif" />
Taking into account the statistical properties of scattering fields, you can enter an approximation
<img file="PL2154677T3_D0018.tif" />
where is Ediff is the energy of the scattering field. In this way, the estimator can be obtained by <| / '^ 1>. * <ΙΛ<sub>Ι</sub>τ |>, = 4ΐ<sup>Σ</sup>Ψ <| ί |>, · (<sup>14</sup>>
[0054] To calculate instant estimates, i.e. for each time-frequency wafer, expectation operators can be removed to obtain
<img file="PL2154677T3_D0019.tif" />
[0055] Using the plane wave assumption, particle velocity estimation can be obtained directly from
<img file="PL2154677T3_D0020.tif" />
(16} from which it follows that β ^ η) = φ-ν &, η) · (17) [0056] In other words, the estimator 110 can be adapted to estimate the fraction β (,) based on the scatter parameter Ψ (Α, η), according to
<img file="PL2154677T3_D0021.tif" />
and wave field measures according to
<img file="PL2154677T3_D0022.tif" />
wherein the processor 120 may be adapted to obtain a module of the first directional component X (k, n) and / or the second directional component Y (k, n) and / or the third directional component Z (k, n) and / or the omnidirectional component W (k , n) by = Τϊβ <Μ P (k> n) ^ <Μ)
K (L<sub>t</sub>n) - (Μ) '
Z (k<sub>t</sub> η) = 72β <7, η) - Ρ (Α, «) - (Μ) where the measure of the direction of arrival of the wave is represented by the unit vector [eDOA, x (k, n), eDOA, y (k, n), eDOA from (k, n)]<sup>T</sup>, where x, y and z are directions of the Cartesian coordinate system.
[0057] An alternative solution in the embodiments can be obtained by obtaining the coefficient β (,) directly from the expression for the dispersion Ψ (,). As already mentioned, the velocity of particles U (k, n) can be modeled
<img file="PL2154677T3_D0023.tif" />
Equation (18) can be replaced by equation (5) leading to
<img file="PL2154677T3_D0024.tif" />
[0058] In order to obtain instantaneous values, estimation operators can be removed, and the solution for e (k, n) brings
<img file="PL2154677T3_D0025.tif" />
[0059] In other words, in embodiments the estimator 110 may be adapted to estimate the fraction e (k, n) based on Ψ (Κ, π) according to
<img file="PL2154677T3_D0026.tif" />
[0060] In embodiments, the input spatial audio signal may correspond to a mono DirAC signal. The embodiments can be extended to process other streams. In the event that the surround audio stream or input does not contain an omnidirectional channel, embodiments may combine available channels to approximate the omnidirectional sensitivity pattern. For example, in the case of a stereo DirAC stream as an input spatial audio signal, the pressure signal P in Fig. 2 can be approximated by adding L and R channels.
[0061] The embodiment for Ψ = 1 will be explained below. Fig. 2 shows that if the dispersion is equal to one for both embodiments, the sound is only associated with the W channel because β is equal to zero, so that the X, Y and Z signals, i.e. the direction components are also equal to zero. If Ψ = 1 invariably in time, the mono audio channel can be bound to the W channel without any further calculations. The physical interpretation is that the audio signal is presented to the listener as a purely reactive field, because the particle velocity vector has a zero module.
[0062] Another case when Ψ = 1 occurs when considering the situation in which the audio signal is present only in one or any subset of dipole signals, and not in the W signal. In the DirAC scattering analysis this scenario is analyzed so that Ψ = 1 in equation 5, because the intensity vector is constantly zero, as the pressure P is zero in equation (1). The physical interpretation of this situation is also that the audio signal is presented to the listener as reactive, because this pressure signal is invariably zero over time, while the particle velocity vector is nonzero.
[0063] Due to the fact that the B-format is a natural representation independent of the speaker arrangement, embodiments may use the B-format as a common language of communication between different audio devices, which means that conversion from one format to another may be possible in embodiments by indirect conversion to B-format. For example, embodiments may combine DirAC streams from different recorded acoustic environments with other synthesized B-format audio environments. Combining mono DirAC streams with B-format streams can also be enabled by embodiments.
[0064] Embodiments may allow combining multi-channel audio signals in any surround format with a mono DirAC stream. In addition, embodiments may allow the combination of multi-channel audio signals in any surround format with a mono DirAC stream. In addition, embodiments may allow combining a mono DirAC stream with a B-format stream. In addition, embodiments may allow combining a mono DirAC stream with a B-format stream.
[0065] These embodiments may offer the advantage of, e.g., creating reverberation or introducing audio effects, as will be discussed in detail below. In music production, reverberators can be used as effects devices that perceptually place processed audio in virtual space. In virtual reality, reverberation synthesis may be needed when virtual sources are auralized within confined spaces, e.g. in rooms or concert halls.
[0066] When a signal for reverberation is available, such auralization can be accomplished by embodiments by using dry sound and reverberated sound in various DirAC streams. The embodiments may use different approaches as to how to process the reverberated signal in the context of DirAC, where the embodiments may produce the reverberated sound maximally dispersed around the listener.
[0067] Fig. 3 shows an embodiment of a device 300 for determining a combined converted spatial audio signal, wherein the combined converted spatial audio signal comprises at least a first combined component and a second combined component, wherein the converted spatial audio signal is determined from the first and second input a spatial audio signal comprising the first and second audio input representations and the first and second wave arrival directions.
[0068] The device 300 includes a first embodiment of a device 101 for determining a converted spatial audio signal as described above for providing a first converted signal comprising a first omnidirectional component and at least one directional component from the first device 101. In addition, the device 300 includes another embodiment of a device 102 for determining a converted spatial audio signal as described above for providing a second converted signal comprising a second omnidirectional component and at least one directional component from the second device 102.
[0069] In general, the embodiments are not limited to only devices 100, in general, the device 300 may include the many devices described above, e.g. the device 300 may be adapted to combine multiple DirAC signals.
[0070] According to Fig. 3, the device 300 further includes an audio effect generator 301 for rendering the first omnidirectional and first directional audio component from the first device 101 to obtain the first rendered component.
[0071] In addition, the device 300 includes a first connecting module 311 for connecting the first rendered component to the first and second omnidirectional components or for connecting the first rendered component with the directional components from the first device 101 and the second device 102 to obtain the first combined component. The device 300 further includes a second connecting module 312 for connecting the first and second omnidirectional components or directional components of the first or second devices 101 and 102 to obtain a second combined component.
In other words, the audio effect generator 301 may render the first omnidirectional component such that the first combiner 311 can then combine the rendered first omnidirectional component, the first omnidirectional component and the second omnidirectional component to obtain the first combined component. The first combined component may then correspond, for example, to the combined omnidirectional component. In this embodiment, the second connecting module 312 may combine the directional component from the first device 101 and the directional component from the second device to obtain a second combined component, for example corresponding to the first combined directional component.
[0073] In other embodiments, the audio effects generator 301 may render directional components. In these embodiments, the connecting module 311 may combine the directional component from the first device 101, the directional component from the second device 102 and the first rendered component to obtain the first combined component, in this case corresponding to the combined directional component. In this embodiment, the second connecting module 312 may combine the first and second omnidirectional components from the first device 101 and the second device 102 to obtain a second combined component, i.e., an combined omnidirectional component.
[0074] In accordance with the embodiments described above, each device may produce a plurality of directional components, e.g., X, Y and Z components. In variants of the invention, multiple audio effect generators may be used, as indicated in Fig. 3 by intermittent blocks 302, 303 and 304. These optional audio effect generators can generate corresponding rendered components based on omnidirectional and directional input signals. In one embodiment, the audio effects generator may render a directional component based on the omnidirectional component. In addition, the device 300 may include a plurality of connecting modules, i.e., modules 311, 312, 313 and 314 to combine the combined omnidirectional component and the plurality of combined directional components, e.g. for three spatial directions.
[0075] One advantage of the structure of the device 300 is that in general a maximum of four audio effect generators are needed to render an unlimited number of audio sources.
[0076] As indicated by the connecting modules 331, 332, 333 and 334 indicated by the dashed lines in Fig. 3, the audio effect generator can be adapted to render combinations of directional or omnidirectional components from devices 101 and 102. In one embodiment, the audio effect generator 301 may be adapted to render combinations of omnidirectional components of the first device 101 and the second device 102, or to render combinations of the directional components of the first device 101 and the second device 102 to obtain the first rendered component. As shown by the dashed tracks in Fig. 3, combinations of multiple components may be provided to various audio effect generators.
[0077] In one embodiment of the invention, all omnidirectional components from all sound sources, in Fig. 3 represented by the first device 101 and the second device 102, may be combined to generate multiple rendered components. In each of the four tracks shown in Fig. 3, each audio effect generator can generate a rendered component to be added to the respective directional or omnidirectional components from the sound sources.
[0078] In addition, as shown in Fig. 3, multiple delay and scaling members 321 and 322 may be used. In other words, each device 101 or 102 has one delay and scaling member 321 or 322 in its output path to delay one or more of its output components. In some embodiments, delay and scaling members can delay and scale only the corresponding omnidirectional components. In general, delay and scaling members can be used for omnidirectional and directional components.
[0079] In embodiments, the device 300 may include a plurality of devices 100 representing audio sources and correspondingly a plurality of audio effect generators, wherein the number of audio effect generators is less than the number of devices corresponding to the sound sources. As mentioned above, in one embodiment there can be up to four audio effect generators with a substantially unlimited number of sound sources. In embodiments, the audio effect generator may correspond to a reverberator.
[0080] Fig. 4a shows in more detail another embodiment of the device 300. Fig. 4a shows two devices 101 and 102, each of which outputs the omnidirectional audio component W and the three directional components X, Y, Z. According to the embodiment shown in Fig. . 4a, the omnidirectional components of each of the devices 101 and 102 are provided to two delay and scaling members 321 and 322, which output three delayed and scaled components, which are then added by connecting modules 331, 332, 333 and 334. Each of the combined signals is then rendered separately by one of the four audio effect generators 301, 302, 303 and 304, which are implemented as reverberators in Fig. 4a. As shown in Fig. 4a, each of the audio effect generators outputs one component, corresponding to one omnidirectional component and a total of three directional components. The connecting modules 311, 312, 313 and 314 are then used to connect the respective rendered components to the original components output by devices 101 and 102, wherein in Fig. 4a there may generally be multiple devices 100.
[0081] In other words, in the connecting module 311, a rendered version of the combined omnidirectional output signals of all devices can be combined with the original or non-rendered omnidirectional output components. Similar joining can be performed by other joining modules relative to directional components. In the embodiment shown in Fig. 4a, rendered directional components are created based on delayed and scaled versions of omnidirectional components.
[0082] In general, embodiments may efficiently apply an audio effect, such as reverberation, to one or more DirAC streams. For example, at least two DirAC streams are introduced into the embodiment of the device 300 as shown in Fig. 4a. In embodiments, these streams may be real DirAC streams or synthesized streams, for example, by taking the mono signal and adding additional information as direction and spread. As discussed above, devices 101, 102 can generate up to four signals for each stream, namely W, X, Y and Z. In general, embodiments of devices 101 or 102 can provide less than three directional components, for example, only X or X and Y, or any other combination thereof.
[0083] In some embodiments, omnidirectional components W may be provided to audio effect generators, e.g. reverbers, to create rendered components. In some embodiments, for each input DirAC stream, the signals may be copied to the four branches shown in Fig. 4a, which may be delayed independently, i.e. individually for each of the devices 101 or 102, four independently delayed, e.g. by Tw, Tx, Ty, Tz and scaled delays, e.g. by the scaling factorsyw, yx, yy, yz, versions can be combined before being delivered to the audio effects generator.
[0084] According to Figs. 3 and 4a, branches of different streams, i.e. the output signals of devices 101 and 102, can be combined to obtain four combined signals. The combined signals can then be rendered independently by audio generators, for example conventional mono converters. The resulting rendered signals can then be added to the W, X, Y and Z signals originally derived from various devices 101 and 102.
[0085] In variants of the invention, general B-format signals can be obtained, which can then be played, for example, using a B-format decoder, which is for example included in Ambisonics. In other embodiments, the B-format signals may be encoded, for example, with the DirAC encoder shown in Fig. 7 in such a way that the resulting DirAC stream can then be transmitted, further processed or decoded with a conventional mono DirAC decoder. The decoding step may correspond to calculating the speaker signals to be played.
[0086] Fig. 4b shows another embodiment of the device 300. Fig. 4b shows two devices 101 and 102 with the corresponding four output components. In the embodiment shown in Fig. 4b, only the omnidirectional components W are used for the initial individual delay and scaling in the delay members 321 and 322 before they are connected by the connecting module 331. The combined signal is then supplied to the audio effect generator 301, which FIG. 4b is again implemented as a reverberator. Rendered output from reverser 301 is then combined with the original omnidirectional components from devices 101 and 102 through the linking module 311. Other linking modules 312, 313 and 314 are used to connect the directional components X, Y and Z from devices 101 and 102 to obtain relevant combined directional components.
[0087] Compared to the embodiment shown in Fig. 4a, the embodiment shown in Fig. 4b corresponds to setting the scale factors for the branches X, Y and Z at 0. In this embodiment, only one audio effect generator or reverberator 301 is used.
[0088] In general, since devices 101, 102 and potentially N devices corresponding to N sound sources, potentially N members 321 delay and scaling can simulate the distances of sound sources, a smaller delay may correspond to the perception of the smaller distance of the virtual sound source from the listener. The spatial impression of the environment can then be created by appropriate audio effect generators or reverberators.
[0089] The variants of the invention shown in Figs. 3, 4a and 4b can be used in cases where mono DirAC decoding is used for N sound sources, which are then reverberated together. Since it can be assumed that the output signal from the reverser can be completely diffused, it can also be interpreted as an omnidirectional signal W. This signal can be combined with other Bformat synthesized signals, such as B-format signals from the N audio sources themselves, thus representing a direct path to the listener. When the resulting B-format signal is then encoded and decoded in the DirAC format, the reverberated audio can be made available by the embodiments.
[0090] In Fig. 4c, another embodiment of the device 300 is shown. In the embodiment of Fig. 4c, based on the output signals of omnidirectional devices 101 and 102, reverberated rendered directional components are created. Hence, based on the omnidirectional output signal, delay and scaling members 321 and 322 form individually delayed and scaled components that are combined in connecting modules 331, 332 and 333. Different reverberators 301, 302 and 303 are used for each of the combined signals, which generally correspond to different audio effect generators. As described above, the respective omnidirectional, directional and rendered components are combined by connecting modules 311, 312, 313 and 314 to provide a combined omnidirectional component and combined directional components.
[0091] In other words, W signals or omni-directional signals for each stream are provided to three audio effect generators, e.g. reverbers, as shown in the figures. In general, there can also be only two branches depending on whether a two-dimensional or three-dimensional audio signal is to be generated. Once the B-format signals are obtained, the streams can be decoded via the virtual microphone's DirAC decoder. The latter is described in detail in publication V. Pulkki, Spatial Sound
Reproduction With Directional Audio Coding, Journal of the Audio Engineering Society, 55 (6): 503-516.
[0092] Using this decoder, the Dp (k, n) loudspeaker signals can be obtained as a linear combination of the W, X, Y and Z signals, for example according to ~ + X (A, tt) cos (o<sub>;</sub>)something(/)<sub>f</sub>), + r (t, n> sin (<t,) cosQ3<sub>p</sub> ) + Z (*, n) sin ^)] where ap and βμ are the azimuth and mid angle of the p-th speaker. The word G (kn) is a panning gain depending on the direction of arrival and the configuration of the speakers.
In other words, the embodiment shown in Fig. 4c may provide loudspeaker audio signals corresponding to audio signals that can be obtained by placing virtual microphones oriented towards the location of the speakers and having point sound sources whose location is determined by the DirAC parameters. Virtual microphones can have a sensitivity pattern in the shape of a cardioid, dipole, or any first order directional pattern.
[0094] Reversed sounds can for example be efficiently used as X and Y in B-format summation. Such embodiments can be used for horizontal speaker spacing with any number of speakers, without creating the need for more reverberators.
[0095] As discussed earlier, mono DirAC decoding has restrictions on the quality of reverberation, while in embodiments the quality can be improved in the virtual microphone DirAC decoding, which also uses dipole signals in the B-format stream.
[0096] In embodiments, the correct creation of B-format signals may be performed to convert audio signal for virtual microphone DirAC decoding. A simple and effective concept that can be used in embodiments is to combine different audio channels with different dipole signals, for example with X and Y channels. Embodiments can accomplish this by using two reverbers producing non-coherent mono audio channels from the same input channel, treating their input signals as Bformat X and Y audio dipole channels, respectively, as shown in Fig. 4c for directional components. Because the signals do not apply to the W component, they will be analyzed as completely diffused in subsequent DirAC encoding. Also, with the virtual microphone's DirAC decoding, better reverberation quality can be obtained because the dipole channels contain reverberated sound in different ways. Embodiments can use them to create the impression of greater "width" and better "enveloping" of reverberation than with DirAC mono decoding. Thus, embodiments may use a maximum of two reverberators in horizontal speaker locations, and three in three-dimensional speaker locations in the described reverberation based on DirAC.
[0097] Embodiments may not be limited to reverberation signals, but may use any other audio effects that are aimed, e.g., at creating the impression of a complete sound dispersion. As in the above-described embodiments, the converted B-format signal can be added to other B-format synthesized signals in the embodiments, such as signals from N audio sources alone, thus representing a direct path to the listener.
[0098] Still another embodiment is shown in Fig. 4d. Fig. 4d shows a similar embodiment to Fig. 4a, however delay and scaling members 321 and 322 are not present, i.e. the individual signals in the branches are only reverberated. The embodiment shown in Fig. 4d can also be seen as similar to the embodiment shown in Fig. 4a with delays and scaling or gain before setting the reverberators to 0 and 1, respectively, however, in this embodiment it is not assumed that converters 301, 302, 303 and 304 are arbitrary and independent. In the embodiment shown in Fig. 4d, it is assumed that four audio effect generators are interdependent having a specific structure.
[0099] Each of the audio effect generators or reverbers can be implemented as a branch delay line, as will be explained in detail below with the help of Fig. 5. Delays and amplification or scaling can be chosen correctly in such a way that each branch models one separate echo whose direction, delay and power can be set freely.
[0100] In such an embodiment, the i-th echo may have a weighting coefficient, e.g. with respect to the DirAC sound p<sub>and</sub>, delays τ and arrival direction θί and φ<sub>and</sub> corresponding to a central angle and azimuth.
[0101] The reverberator parameters can be set as follows <sup>= =</sup> You = f<sub>FROM</sub> = Tj y * = p ,, for the W reverberator,
7x = A. <sup>,SOMETHING</sup>(AND)*<sup>SOMETHING</sup>(^) »For the X reverberator,
7r »A> sin (A) * cos (4), for the 7 reveberator, y<sub>2</sub> = p, · sin (0,), for the Z reverberator.
[0102] In some embodiments, the physical parameters of each echo can be obtained from random processes or taken from a room spatial impulse response. The latter can, for example, be measured or simulated using a ray tracking tool.
[0103] In general, embodiments may offer the advantage that the number of audio effects is independent of the number of sources.
[0104] Fig. 5 shows an embodiment using the conceptual scheme of a mono audio effect as for example used in an audio effect generator that is extended in the context of DirAC. For example, the reverberator can be implemented according to this scheme. Fig. 5 shows an embodiment of the reverser 500. Fig. 5 generally shows the structure of the FIR filter (Finite Impulse Response). Other embodiments may also use IIR (Infinite Impulse Response) filters. The input signal is delayed by K delay stages marked from 511 to 51K. K delayed copies of the signal for which delays are designated τ are τ<sub>κ</sub> it is then amplified by 521 to 52K amplifiers with γ γ gain factors<sub>κ</sub> before they are added together in summation 530.
[0105] Fig. 6 shows another embodiment with processing chain extension of Fig. 5 in the context of DirAC. The output from the processing block may be a B-format signal. Fig. 6 shows an embodiment in which a plurality of summing elements 560, 562 and 564 are used, resulting in three output signals W, X and Y. To establish other combinations, delayed copies of signals can be scaled differently before they are added in three different additions 560, 562 and 564. This is done by additional amplifiers from 531 to 53K and from 541 to 54K. In other words, embodiment 600 shown in Fig. 6 performs reverberation for various B-format signal components based on a mono DirAC stream. Three different reverberated copies of the signal are generated using three different FIR filters established using different filter ratios from pi to ρκ and ni to ηκ.
[0106] The following variant of the invention may be used in a reverberator or in audio effects that can be modeled as in Fig. 5. The input signal runs through a straight branched delay line in which multiple copies of it are added together. i-th of K branches is delayed and suppressed by Ti and γ, respectively.
[0107] The γ and τ factors can be obtained depending on the desired audio effect. In the case of the reverberator, these factors mimic the impulse response of the room to be simulated. In any case, their determination is not explained so it is assumed that they are given.
[0108] Fig. 6 shows a variant of the invention. The scheme of Fig. 5 is extended so that two additional layers are obtained. In embodiments, an approach angle θ obtained in a stochastic process may be associated with each branch.
For example, θ may be the implementation of a homogeneous distribution in the range [-π, π]. the i-th branch is multiplied by the coefficients n<sub>and</sub> ip<sub>and</sub>that can be defined as
Tl. = Sin (dJ (21}
<img file="PL2154677T3_D0027.tif" />
[0109] In embodiments, the i-th echo can be seen as coming from O ,. The extension to three-dimensionality is simple. In this case, one or more layers need to be added and the center angle must be taken into account. After generating the Bformat signal, i.e. W, X, Y and potentially Z, it can be combined with other B-format signals. Then it can be sent to the virtual microphone DirAC decoder, or after DirAC encoding, the mono DirAC stream can be sent to the mono DirAC decoder.
[0110] Embodiments may include a method of determining a converted spatial audio signal, the converted spatial audio signal comprising a first audio directional component and a second audio directional component from the input spatial audio signal, wherein the input spatial audio signal comprises an input audio representation and an input direction of arrival . The method comprises the step of estimating a wave representation comprising a wave field measure and a wave arrival direction measure based on the input audio representation and the wave arrival input direction. In addition, the method includes the step of processing the wave field measure and the arrival direction measure to obtain the first directional component and the second directional component.
[0111] In embodiments of the invention, the method of determining the converted spatial audio signal may include the step of obtaining a mono DirAC stream to be converted to B-format. Optionally, W can be obtained from P when available. If not, the W zoom step can be implemented as a linear combination of available audio signals. Then the step of calculating the β coefficient as a time-frequency weighting coefficient, inversely proportional to the dispersion can be implemented, e.g.
<img file="PL2154677T3_D0028.tif" />
or
<img file="PL2154677T3_D0029.tif" />
[0112] The method may further include the step of calculating the X, Y and Z signals of P, β and <sup>e</sup>DOA<sup>.</sup> [0113] In cases where Ψ = 1, the step of obtaining W from P can be replaced by obtaining W from P, respectively, when X, Y and Z are zero, obtaining at least one dipole signal X, Y or Z from P; respectively W is zero. Embodiments of the present invention can implement B-format domain signal processing, offering the advantage that advanced signal processing can be performed prior to generating speaker signals.
[0114] Depending on some require implementation of the methods of the invention, the methods of the present invention can be implemented in hardware or in software. Implementations may be made using digital memory media, in particular Flash memory, disk, DVD or CD containing electronically readable control signals that interact with the programmed computer system in such a way that the methods of the present invention may be implemented. In general, therefore, the present invention is a computer program code, where the program code is stored on a machine-readable medium, and the program code can operate to implement the methods of the present invention when the computer program is implemented on a computer or processor. In other words, the methods of the present invention are a computer program containing program code for implementing at least one method of the present invention when the computer program is implemented in a computer
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV, Germany Representative:
EP 2 154 677 B1
Z-10929/13
Contents2
29 members in 14 offices
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 8851308 | United States of America | P | |
| 9168208 | United States of America | P | |
| 09001398 | European Patent Office (EPO) | A | |
| EP20090001398 | – | – | – |
| US20080088513P | – | – | – |
| US20080091682P | – | – | – |
Members29
| Document | Office | Kind | |
|---|---|---|---|
| EP2154677A1 | European Patent Office (EPO) | A1 | |
| AU2009281367A1 | Australia | A1 | |
| CA2733904A1 | Canada | A1 | |
| WO2010017978A1 | World Intellectual Property Organization (WIPO) | A1 | |
| HK1141621A1 | Hong Kong, China | A1 | |
| EP2311026A1 | European Patent Office (EPO) | A1 | |
| KR20110052702A | Republic of Korea | A | |
| MX2011001657A | Mexico | A | |
| CN102124513A | China | A | |
| US2011222694A1 | United States of America | A1 | |
| JP2011530915A | Japan | A | |
| HK1155846A1 | Hong Kong, China | A1 | |
| RU2011106584A | Russian Federation | A | |
| AU2009281367B2 | Australia | B2 | |
| EP2154677B1 | European Patent Office (EPO) | B1 | |
| KR20130089277A | Republic of Korea | A | |
| ES2425814T3 | Spain | T3 | |
| RU2499301C2 | Russian Federation | C2 | |
| US8611550B2 | United States of America | B2 | |
| PL2154677T3This record | Poland | T3 | |
| CN102124513B | China | B | |
| JP5525527B2 | Japan | B2 | |
| EP2311026B1 | European Patent Office (EPO) | B1 | |
| CA2733904C | Canada | C | |
| ES2523793T3 | Spain | T3 | |
| KR101476496B1 | Republic of Korea | B1 | |
| PL2311026T3 | Poland | T3 | |
| BRPI0912451A2 | Brazil | A2 | |
| BRPI0912451B1 | Brazil | B1 |
Numbers
- Publication, DOCDB
- 2154677
- Publication, EPODOC
- PL2154677T
- Application
- 1398
- Application, DOCDB
- 09001398
- Application, EPODOC
- PL20090001398T
Titles2
- English
- An apparatus for determining a converted spatial audio signal
- Polish
- Urządzenie do wyznaczania konwertowanego przestrzennego sygnału audio
Classification
- CPC, 5
- H04S3/02
- G10L19/008
- H04S2400/15
- H04S2420/03
- H04S2420/11
- IPC, 3
- G10H1 00
- G10L19 008
- H04S3 02