Apparatus for merging spatial audio streams
Abstract
An apparatus (100) for merging a first spatial audio stream with a second spatial audio stream to obtain a merged audio stream comprising an estimator (120) for estimating a first wave representation comprising a first wave direction measure and a first wave field measure for the first spatial audio stream, the first spatial audio stream having a first audio representation and a first direction of arrival. The estimator (120) being adapted for estimating a second wave representation comprising a second wave direction measure and a second wave field measure for the second spatial audio stream, the second spatial audio stream having a second audio representation and a second direction of arrival. The apparatus (100) further comprising a processor (130) for processing the first wave representation and the second wave representation to obtain a merged wave representation comprising a merged wave field measure and a merged direction of arrival measure, and for processing the first audio representation and the second audio representation to obtain a merged audio representation, and for providing the merged audio stream comprising the merged audio representation and the merged direction of arrival measure.

Term
2.9 yearsto projected expiry
Projected expiry 11 August 2029, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
14 claims: 4 independent, 10 dependent
- 1Claims Zastrzeżenia patentowe 1. An apparatus (100) for combining a first spatial stream of an audio signal with a second spatial audio stream to obtain a combined audio stream, including an estimator (120) for estimating a first wavelength representation comprising a first wavelength measure (1) as the first wavelength;and first measure ' 1 in a wave field associated with a first wave module for the first spatial stream of the audio signal, which first stream of spatial audio signal comprises a first representation of an audio signal comprising a measure of pressure or a modulus of the first audio signal (P)(1)) and the first direction of arrival, eoox) · and to estimate a second wavelet representation comprising a second wavelength measure which is a forward wavelength () and a second wavelength measure () of the wavelength field associated with a second wavelength modulus for the second spatial flux of the acoustic signal;an acoustic signal representation containing a measure of pressure or a second acoustic signal module (P(2)) and a second approach direction * and a processor (130) for processing the first wave representation and the second wave representation to obtain a combined wave representation comprising a combined field measure (Λ), a combined measure of the direction of arrival, and and pABOUT}andFRONTny scattering parameter where the combined scattering parameter is based on the combined wave field (4), the first representation (P(1)) acoustic signal and second representation (P(2)) the acoustic signal, and where the combined wave field measure (Λ) is based on the first wave field, the second wave field, the first wave direction and the second wave direction ('and where the processor (130) is configured to process the first representation) P(1)) acoustic signal and second representation (P(2)) of an acoustic signal to obtain a combined representation (P) and to provide a combined audio signal stream including a combined representation of an audio signal (P), a combined measure of arrival, and a combined leakage parameter 1. Urządzenie (100) do łączenia pierwszego strumienia przestrzennego sygnału akustycznego z drugim strumieniem przestrzennego sygnału akustycznego do uzyskania połączonego strumienia sygnału akustycznego, zawierające estymator (120) do szacowania pierwszej reprezentacji falowej zawierającej pierwszą miarę ( )kierunku fali, będącą wielkością kierunkową pierwszej fali oraz pierwszą miarę ' 1 w 'pola falowego związaną z modułem pierwszej fali dla pierwszego strumienia przestrzennego sygnału akustycznego, który to pierwszy strumień przestrzennego sygnału akustycznego zawiera pierwszą reprezentację sygnału akustycznego zawierającą miarę ciśnienia lub modułu pierwszego sygnału akustycznego (P(1)) oraz pierwszy kierunek nadejścia, eoox) · oraz do szacowania drugiej reprezentacji falowej zawierającej drugą miarę kierunku fali, będącą wielkością kierunkową drugiej fali ( ) oraz drugą miarę ( )pola falowego związaną z modułem drugiej fali dla drugiego strumienia przestrzennego sygnału akustycznego, który to drugi strumień przestrzennego sygnału akustycznego zawiera drugą reprezentację sygnału akustycznego zawierającą miarę ciśnienia lub modułu drugiego sygnału akustycznego (P(2)) oraz drugi kierunek nadejścia *doa oraz procesor (130) do przetwarzania pierwszej reprezentacji falowej i drugiej reprezentacji falowej do uzyskania połączonej reprezentacji falowej, zawierającej połączoną miarę pola (Λ), połączoną miarę kierunku nadejścia doai oraz pO}ąCZOny parametr rozproszenia gdzie połączony parametr rozproszenia oparty jest na połączonej mierze pola falowego (4), pierwszej reprezentacji (P(1)) sygnału akustycznego i drugiej reprezentacji (P(2)) sygnału akustycznego, oraz gdzie połączona miara pola falowego (Λ) oparta jest na pierwszej mierze pola falowego, drugiej mierze pola falowego, pierwszej mierze kierunku fali i drugiej mierze kierunku fali ( ' oraz gdzie procesor (130) jest skonfigurowany do przetwarzania pierwszej reprezentacji (P(1)) sygnału akustycznego i drugiej reprezentacji (P(2)) sygnału akustycznego do uzyskania połączonej reprezentacji (P) oraz do dostarczania połączonego strumienia sygnału akustycznego zawierającego połączoną reprezentację sygnału akustycznego (P), połączoną miarę nadejścia oraz połączony parametr rozproszenia
- 1113. A method for combining a first spatial stream of an audio signal with a second spatial audio stream to obtain a combined audio stream, comprising estimating a first wavelength representation comprising a first measure (UPW, Wave blend, which is the directional size of the first wave and the first measure / »<ni ' 7 pw 'of the wave field associated with the first wave module for the first spatial stream of the acoustic signal, which first stream of spatial audio signal comprises a first representation of an acoustic signal comprising a measure of pressure or a modulus of the first audio signal (P)(1)) and the first direction (e(,) 1 · arrival,1 doa '· estimation of a second wavelet representation comprising a second wavelength measure which is the directional amount of the second wave () and a second measure () of the wave field associated with the second wave module for the second spatial stream of the audio signal, wherein the second spatial audio stream comprises a second signal representation acoustic containing a measure of the pressure or module of the second audio signal (P(2)) and second direction of arrival) · processing of the first wave representation and the second wave representation to obtain a combined wave representation comprising a combined field measure (4), a combined measure 13. Sposób łączenia pierwszego strumienia przestrzennego sygnału akustycznego z drugim strumieniem przestrzennego sygnału akustycznego do uzyskania połączonego strumienia sygnału akustycznego, zawierający szacowanie pierwszej reprezentacji falowej zawierającej pierwszą miarę ( UPW, Iłcierunku fali, będącą wielkością kierunkową pierwszej fali oraz pierwszą miarę / »<n i ' 7 pw 'pola falowego związaną z modułem pierwszej fali dla pierwszego strumienia przestrzennego sygnału akustycznego, który to pierwszy strumień przestrzennego sygnału akustycznego zawiera pierwszą reprezentację sygnału akustycznego zawierającą miarę ciśnienia lub modułu pierwszego sygnału akustycznego (P(1)) oraz pierwszy kierunek ( e(,) 1 · nadejścia,1 doa ' · szacowanie drugiej reprezentacji falowej zawierającej drugą miarę kierunku fali, będącą wielkością kierunkową drugiej fali ( ) oraz drugą miarę ( )pola falowego związaną z modułem drugiej fali dla drugiego strumienia przestrzennego sygnału akustycznego, który to drugi strumień przestrzennego sygnału akustycznego zawiera drugą reprezentację sygnału akustycznego zawierającą miarę ciśnienia lub modułu drugiego sygnału akustycznego (P(2)) oraz drugi kierunek nadejścia ) · przetwarzanie pierwszej reprezentacji falowej i drugiej reprezentacji falowej do uzyskania połączonej reprezentacji falowej, zawierającej połączoną miarę pola (4), połączoną miarę Λ approach direction (®DOa) and a combined spread parameter where the combined dispersion parameter is based on the combined wavefield range (Iand}, the first representation (P(1)) acoustic signal and second representation (P(2)) acoustic signal, and where the combined wave field measure (Λ) is based on the first wave field, the second on the wave field, the first wave direction measure (, jdrUgeand measure the direction of the wave ( ;processing of the first representation (Ρω) acoustic signal and second representation (P(2)) of an acoustic signal to obtain a combined representation (P) and providing a combined audio signal stream including a combined representation of an audio signal (P), a combined measure 1 DOA '1 direction of arrival Λ and combined parameter of dispersion (Ψ). Λ kierunku nadejścia (®DOa)oraz połączony parametr rozproszenia gdzie połączony parametr rozproszenia oparty jest na połączonej mierze pola falowego (Ia}, pierwszej reprezentacji (P(1)) sygnału akustycznego i drugiej reprezentacji (P(2)) sygnału akustycznego, oraz gdzie połączona miara pola falowego (Λ) oparta jest na pierwszej mierze pola falowego, drugiej mierze pola falowego, pierwszej mierze kierunku fali ( , j drUgiej mierze kierunku fali ( ;przetwarzanie pierwszej reprezentacji (Ρω) sygnału akustycznego i drugiej reprezentacji (P(2)) sygnału akustycznego do uzyskania połączonej reprezentacji (P) oraz dostarczanie połączonego strumienia sygnału akustycznego zawierającego połączoną reprezentację sygnału akustycznego (P), połączoną miarę 1 DOA'1 kierunku nadejścia Λ oraz połączony parametr rozproszenia (Ψ).
- 1315. A computer program containing a program code for carrying out the method specified in 15. Program komputerowy zawierający kod programu do realizacji sposobu określonego w
- 1416. Claim 14, when the program code is implemented on a computer or processor. 16. zastrzeżeniu 14, kiedy kod programu realizowany jest w komputerze lub w procesorze. Fraunhofer-Gesellschaft zur Fórderung der angewandten Forschung e.V., Niemcy Pełnomocnik:Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung eV, Germany Plenipotentiary: Jan Dobrzański Rzecznik Patentowy Jan Dobrzański, Patent Attorney EP 2 324 645 B1 EP 2 324 645 BI Z-9545 first representation Z-9545 pierwsza reprezentacja FIGURA IB FIGURE IB EP 2 324 645 B1 EP 2 324 645 BI Z-9545 Z-9545 FIGURA 2 FIGURE 2 ΕΡ 2 324 645 BI ΕΡ 2 324 645 BI Ζ-9545 Ζ-9545 CO WHAT CO co co What what what FIGURA 3 FIGURE 3 ΕΡ 2 324 645 Β1 ΕΡ 2 324 645 Β1 Ζ-9545 Ζ-9545 CZ LOT Ρ Μ, eD0A(k, n), Ψ '(k, n) Ρ Μ, eD0A(k,n), Ψ '(k,n) FIGURA 4 FIGURE 4 ΕΡ 2 324 645 BI ΕΡ 2 324 645 BI Ζ-9545 Ζ-9545 FIGURA 5 FIGURE 5 ΕΡ 2 324 645 BI ΕΡ 2 324 645 BI Ζ-9545 Ζ-9545 FIGURA 6 FIGURE 6
Independent claims4
147 paragraphs in 1 section, as filed
The present invention relates to the field of audio signal processing, particularly the spatial processing of an audio signal, and combining multiple streams of spatial audio signal.
[0002] DirAC coding, Directionai Audio Coding (see V. Pulkki and C. Faller, Directionai audio coding in spatiai sound reproduction and stereo upmixing, In AES 28th International Conference, Pitea, Sweden, June 2006 and K Pulkki, A method for reproduction natura! or modified spatiai impression in Muitichannei iistening, Patent WO 2004/077884 Al, September 2004} represents an effective approach to the analysis and reproduction of surround sound. In the DirAC format, the parametric representation of sound fields is used based on features relevant to the reception of spatial sensations , specifically the direction of arrival (DOA, Direction Of Arrivaf) and sound field diffusion in frequency sub-bands, in fact, in the DirAC format it is assumed that the intra-hearing time difference (ITD.Interaural Time Difference) and intra-auditory level difference (ILD, Interaurai Levei Differencd) are felt correctly when DOA sound fields are reproduced correctly, while intra-auditory coherence (IC, Interaural Coherence) is felt correctly when distraction is reproduced.
[0003] These parameters, namely DOA and dissipation, represent additional information that accompanies a mono signal in the so-called monophonic DirAC stream. The DirAC parameters are obtained from the time-frequency representation of the microphone signals. Therefore, these parameters depend on time and frequency. On the play side, this information allows for accurate spatial generation. To reproduce surround sound in a desired listening position, a multi-speaker system is needed. However, their geometry is a matter of choice. In fact, the signals for the speakers are determined in the function of the DirAC parameters.
[0004] There are significant differences between the DirAC format and parametric multi-channel audio coding, such as MPEG surround, although they are combined by very similar processing structures (cf. Lars ViHemoes, Juergen Herre, Jeroen Breebaart, Gerard Hotho, Sascha Disch, Heiko Purnhagen, and Kristofer Kjriingm, MPEG surround: The forthcoming ISO standard for spatiai audio coding, in AES 28th International Conference, Pitea, Sweden, June 2006). While MPEG surround is based on the time-frequency analysis of different speaker channels, the DirAC format, as input, treats coincidence microphone channels that effectively describe the sound field of one point. So the DirAC codec also represents an efficient spatial sound recording technique.
[0005] Another conventional system dealing with spatial sound is the encoding of SAOC spatial objects {Falls! Audio Object Coding, see Jonas Engdegard, Barbara Resch, Corneiia Faich, Oiiver Hellmuth, Johannes Hiipert, Andreas Hoe / zeros, Leonid Ternedev, Jeroen Breebaart, Jeroen Koppens, Erik Schuijer, and Werner Oomen, Spada and audio object coding (SAOC) the upcoming MPEG standard on parametric object based audio coding, in 124th AES Convention, 17-20 May 2008, Amsterdam, Netherlands), which are currently in the ISO / MPEG standardization stage.
[0006] It uses the MPEG Surround rendering engine and treats different sound sources as objects. Acoustic signal coding offers high performance in terms of bit rate and unprecedented freedom of interaction on the reconstruction side. This approach promises new attractive features and functions in existing systems as well as a number of other innovative applications.
[0007] The object of the present invention is to provide a recognized concept for combining spatial acoustic signals.
[0008] This object is achieved by a joining device according to claim 1 and by a joining method according to claim 13.
[0009] It should be noted that combining would be a trivial matter for the multichannel DirAC stream, i.e. if 4 audio channels of the B-format are available. In fact, signals from different sources can be directly added to obtain Bformat signals of the combined stream. However, if such channels are not available, direct connection is problematic.
[0010] The present invention is based on the finding that spatial acoustic signals can be represented by the sum of wave representations, e.g. flat wave representation and the representation of a scattered field. You can assign a direction to the first one. When combining multiple audio streams, variants of the invention may allow additional information on the combined stream, e.g. in the form of dispersion and direction. Variants of the invention may obtain this information from a wave representation as well as from input audio stream streams. When combining multiple audio signal streams that can all be modeled by a wave portion or representation and a splitting part or representation, the wave parts or components and the part or components of the dispersion can be combined separately. As a result of combining the wave parts, a combined wave part is created, for which a combined direction can be obtained based on the directions of the wave representations. Moreover, the part of the dispersion can also be combined separately, and a general dissipation parameter can be obtained from the combined part of the dispersion.
[0011] Variants of the present invention may provide a method for combining at least two spatial audio signals encoded as monophonic streams
DirAC. The resulting combined signal can also be represented as a mono DirAC stream. In embodiments of the present invention, mono DirAC encoding may be a compact way of describing surround sound, as only a single audio channel is required to be transmitted along with additional information.
[0012] In variants of the invention, the actual possible situation may be the situation of use in a teleconference in which more than two parties are involved. For example, let User A communicate with users B and C who generate two separate mono DirAC streams. In user A's location, the variant of the invention may allow the streams from users B and C to be combined into one monophonic DirAC stream that can be reproduced using conventional DirAC synthesis techniques. In a variant using a network topology in which the presence of a multipoint contro unit (MCU) is visible, the merging operation will be performed by the MCU itself, so that user A receives a single monophonic DirAC stream already containing the speech of both users B and C. Of course, DirAC streams can also be generated synthetically, which means that appropriate additional information can be added to the mono audio signal. In the aforementioned example, user A can receive two streams of audio signal from users B and C without any additional information. It is then possible to assign to each of the streams a certain direction and dispersion, thus adding the additional information needed to produce the DirAC streams, which can then be combined by an embodiment of the present invention. user A may receive two streams of audio signal from users B and C without any additional information. It is then possible to assign to each of the streams a certain direction and dispersion, thus adding the additional information needed to produce the DirAC streams, which can then be combined by an embodiment of the present invention. user A may receive two streams of audio signal from users B and C without any additional information. It is then possible to assign to each of the streams a certain direction and dispersion, thus adding the additional information needed to produce the DirAC streams, which can then be combined by an embodiment of the present invention.
[0013] Another possible scenario for the use of embodiments of the present invention may be multi-player multiplayer network games, and the use of virtual reality. In these cases, many streams are generated by players or virtual objects. Each stream is characterized by a certain direction of arrival with respect to the listener, so it can be expressed by the DirAC stream. Variants of the present invention may be used to combine different streams into a single DirAC stream, which is then reproduced at the listener location.
[0014] Variants of the present invention will be described in detail using the attached figures, in which:
Fig. 1a shows a variant of the connection device;
Fig. Lb shows the pressure and components of the particle velocity vector in the Gaussian plane for a plane wave;
Fig. 2 shows a variant of the DirAC encoder;
Fig. 3 shows the ideal combination of audio signal streams;
Fig. 4 shows the inputs and outputs of a variant of a general processing block for combining in the DirAC format;
Fig. 5 is a block diagram of an embodiment of the invention; and
Fig. 6 is a diagram of a variant of the connection method.
[0015] Fig. 1a illustrates a variant of the apparatus 100 for combining a first spatial stream of an audio signal with a second spatial audio signal stream to obtain a combined audio signal stream. The variant shown in Fig. 1a illustrates combining two audio signal streams, however, it is not limited to two audio signal streams because a plurality of spatial audio streams can be combined in a similar manner. The first stream of spatial audio signal and the second stream of spatial audio signal may, for example, correspond to the mono DirAC streams, and the combined audio signal stream may also correspond to a single monophonic DirAC audio stream. As will be discussed later, the mono DirAC stream may include a pressure signal, e.g. captured by a multi-directional microphone, and additional information. Additional information may include time and frequency dependent measure of distraction and direction of sound arrival.
wherein the second stream of spatial audio signal has a second representation of the audio signal and a second direction of arrival. In embodiments of the present invention, the first and / or second wave representation may correspond to a plane wave representation.
[0017] In the variant illustrated in Fig. La, the apparatus 100 further comprises a processor 130 for processing the first waveform and the second wave representation to obtain a combined waveform comprising a combined field measure and a combined measure of the arrival direction, and for processing the first representation of the audio signal and the second. the representation of the audio signal to obtain a combined representation of the audio signal, wherein the processor 130 is further adapted to provide a combined audio signal stream including a combined representation of the audio signal and a combined measure of the arrival direction.
[0018] Estimator 120 may be adapted to estimate the first wave field measure, in the form of the first wave field amplitude, to estimate the second wave field measure, in the form of the second wave field amplitude, and to estimate the phase difference between the first wave field measure and the second field measure wave. In embodiments of the present invention, the estimator may be adapted to estimate the first phase of the wave field and the second phase of the wave field. In variants of the invention, the estimator 120 may only estimate the phase shift, or the difference between the first and second wave representations respectively, or the first and second wave field measures.
[0019] In embodiments of the invention, the processor 130 may further be adapted to process the first waveform and the second wave representation to obtain the combined wave representation comprising the combined wave field measure, the combined arrival direction flight and the combined dispersion parameter, and to provide the combined signal stream. acoustic comprising a combined representation of the audio signal, a combined measure of the direction of arrival and a combined parameter of dispersion.
[0020] In other words, in embodiments of the present invention, the scattering parameter may be determined based on the waveforms of the combined audio signal stream. The scattering parameter can determine the measure of the spatial dispersion of the audio stream, i.e. the spatial distribution measure, such as, for example, an angular distribution around a certain direction. In a variant of the invention, it is possible to combine two mono synthetic signals containing only direction information.
The processor 130 may be adapted to process the first wave representation and the second wave representation to obtain a combined wave representation in which the combined dispersion parameter is based on the first waveform measure and the second wave direction. In the embodiments of the present invention, the first and second wave representations may have different arrival directions and the combined arrival direction may be located between them. In this variant, although the first and second flux of the spatial audio signal may not provide any scattering parameters, the combined scattering parameter may be determined based on the first and second wave representations, i.e. on the basis of the first wave direction measure and the second wave direction measure. E.g, when the two plane waves overlap from different directions, i.e. the first wave direction measure differs from the second wave direction measure, the combined acoustic signal representation may include a combined combined arrival direction with a non-zero merged leakage parameter to include the first wave direction and the second measure wave direction measures. In other words, although two directional streams of a spatial audio signal may not contain (or not provide) any dispersion, the combined audio signal stream may include a non-zero spread because it is based on the angular distribution determined by the first and second audio stream. the combined representation of the audio signal may include a combined, connected arrival direction with a non-zero combined dispersion parameter to account for the first wave direction measure and the second wave direction measure. In other words, although two directional streams of a spatial audio signal may not contain (or not provide) any dispersion, the combined audio signal stream may include a non-zero spread because it is based on the angular distribution determined by the first and second audio stream. the combined representation of the audio signal may include a combined, connected arrival direction with a non-zero combined dispersion parameter to account for the first wave direction measure and the second wave direction measure. In other words, although two directional streams of a spatial audio signal may not contain (or not provide) any dispersion, the combined audio signal stream may include a non-zero spread because it is based on the angular distribution determined by the first and second audio stream.
[0022] Variants can determine the scattering parameter Ψ, for example for the combined DirAC stream. In general, variants of the invention may in such cases set or assume scattering parameters of individual streams with a constant value ( e.g. 0 or 0.1) or with a variable value obtained from the analysis of the representation of the audio signal and / or the direction representation.
In other embodiments, the apparatus 100 for combining the first spatial stream of the audio signal and the second spatial stream of the audio signal to obtain the combined audio stream may include an estimator 120 for estimating the first waveform comprising a first measure of the wave direction and a first wave field measure for a first spatial stream of the audio signal, which first stream of spatial audio signal comprises a first representation of an audio signal, a first direction of arrival and a first scattering parameter. In other words, the first representation of an acoustic signal may correspond to an acoustic signal having a certain spatial width or to some extent dispersed. In one variant, this may correspond to the situation in the computer game. The first player may be in a situation in which the first representation of the audio signal represents an acoustic signal source, such as a passing train, producing to some extent a diffused sound field. In such a variant, the sound generated by the train itself can be diffused, and the sound of its warning signal, i.e. its corresponding frequency components, may not be distracted.
[0024] Estimator 120 may further be adapted to estimate a second wavelet representation comprising a second wavelength measure and a second wavelength measure for a second spatial stream of acoustic signal, the second spatial audio sound stream comprising a second acoustic signal representation, a second arrival direction, and a second parameter. distraction. In other words, the second representation of an acoustic signal may correspond to an acoustic signal having a certain spatial width or to some extent dispersed. Likewise, this may correspond to a situation in a computer game where the second sound source may be represented by a second audio stream, e.g. background noise from another passing train on another track. For the first player in a computer game,
In embodiments of the invention, the processor 130 may be adapted to process the first wave representation to obtain a combined wave representation comprising a combined wave field measure and a combined measure of the arrival direction, and to process the first representation of the audio signal and a second representation of the audio signal to obtain a combined the representation of an acoustic signal, and for providing a combined audio signal stream including a combined representation of the audio signal and a combined measure of the direction of arrival. In other words, the processor 130 may not determine a combined dispersion parameter. This may correspond to the sound field experienced by the other player in the above-mentioned computer game. The second player can be located further from the railway station,
In embodiments of the present invention, the apparatus 100 may further comprise means 110 for determining, for the first stream of spatial audio signal, a first representation of the audio signal and a first arrival direction, and, for a second stream of spatial audio signal of the second representation of the audio signal and the second direction of arrival. In embodiments, the means 110 for such determination may be provided along with a direct audio stream, i.e. the determination may simply refer to a reading of an audio signal representation in the form of e.g. a pressure signal and DOA and optionally also scatter parameters in the form of additional information.
Estimator 120 may be matched to estimate the first waveform of the first spatial audio signal having additionally a first scattering parameter and / or for estimating the second waveform from the second spatial stream of the audio signal having an additional second scattering parameter, the processor 130 may be matched processing a combined wave field measure, a first and a second representation of the audio signal, and a first and a second scattering parameter to obtain a combined spread parameter for the combined audio stream, and the processor 130 may be further adapted to provide an audio signal stream including a combined scatter parameter.The setting means 110 may be adapted to determine the first spread parameter for the first spatial stream of the audio signal and the second scatter parameter for the second spatial stream of the audio signal.
The processor 130 may be matched to a block (i.e. in the form of sample segments or values) of processing spatial streams of an audio signal, an acoustic signal representation, DOA and / or stray parameters. In some embodiments, the segment may include a fixed number of samples corresponding to the frequency representation of a frequency band at a certain time of the spatial stream of the audio signal. The segment may correspond to a mono representation and be associated with the DOA and the dispersion parameter.
[0029] In embodiments of the present invention, the setpoint means 110 may be adapted to determine the first and second audio signal representation, the first and second arrival direction and the first and second scatter parameters in a time and frequency dependent manner and / or the processor 130 may be adapted to converting the first and second waveforms, scattering parameters and / or DOA measures and / or for establishing a combined audio signal representation, a measure of the combined arrival direction and / or the combined scatter parameter in a time and frequency dependent manner.
[0030] In embodiments of the present invention, the first representation of the audio signal may correspond to a first monaural representation, the second representation of the audio signal may correspond to the second monophonic representation, and the combined audio signal representation may correspond to the combined mono representation. In other words, the representation of an audio signal may correspond to a single audio channel.
In embodiments of the present invention, the determination means 110 may be adapted to be determined and / or the processor may be adapted for processing the first and second monophonic representation, the first and second DAO and the first and second scattering parameters, and the processor 130 may provide a combined a mono representation, a combined DAO measure and / or a combined dispersion parameter in a time and frequency dependent manner. In embodiments of the present invention, the first stream of spatial audio signal may already be provided in the form of, for example, the DirAC representation, the determining means may be adapted to determine the first and second monophonic representations, the first and second DAO and the first and second dissipation parameters,
[0032] In the following, an embodiment of the present invention will be discussed in detail, with the notation and the data model first introduced. In embodiments of the present invention, the determining means 110 may be adapted to determine the first and second audio signal representations and / or the processor 130 may be adapted to provide a combined monophonic representation in the form of a pressure signal p (t) or a time-frequency processed pressure signal P ( k, n), where k is the frequency index and n is the time index.
[0033] In embodiments of the present invention, the first and second measures of the wave direction and the combined measure of the arrival direction may correspond to any directional dimension, e.g. a vector, angle, direction, etc., and may be obtained from any measure of the direction representing the acoustic signal component, such as intensity vector, particle velocity vector, etc. The first and second wave field measure as well as the combined wave field measure may correspond to any physical quantity describing the acoustic component, which may be real or imaginary, may correspond to a pressure signal, amplitude or modulus Particle velocities, loudness, etc. In addition, these measures can be in the time domain and / or frequency domain.
[0034] Variants of the present invention may be based on estimating the plane wave representation for the wave field measures of the wave representation of the input streams, which may take place in the estimator 120 shown in Fig. 1a. In other words, the wave field measure can be modeled using a plane wave representation. Generally speaking, there are a number of equivalent, comprehensive (ie full) descriptions of a plane wave or waves in general. Below, a mathematical description will be introduced for the calculation of scatter parameters and directions of arrival or for direction measures for various components. Although only a few descriptions refer directly to physical quantities, such as pressure, particle velocity, etc., there is potentially an infinite number of different ways to describe a wave representation,
[0035] In order to further present various potential descriptions, two numbers a and b are considered. The information contained in waib can be sent by sending a cid when
Γ ΊΙ
<img file="PL2324645T3_D0001.tif" />
where Ω is known as the 2x2 matrix. The example deals only with linear combinations, but generally any combination is possible, i.e. also non-linear.
[0036] Below, scalar quantities are represented by lowercase letters a, b, c while column vectors are represented by bold lowercase letters a, b, c. Upper index ()<sup>T </sup>means transpose accordingly, while () and (·) * denote the complex conjugate. The notation of a complex phasor differs from the instantaneous notation. For example, the pressure p (t), which is a real number, and from which a possible wave field measure can be obtained, can be expressed by means of a phasor P, which is an imaginary number, and from which another possible measure of the wave field can be obtained by means of a formula p (Z) = Re {/ V "}, where Re {·} is the real part, and a) = 2nf is the angular frequency Furthermore, the capital letters used for the physical quantities designation hereinafter represent the phasors. and in order to avoid ambiguities, it should be noted that all of the below-considered values with the subscript "PW" refer to flat waves.
[0037] For the ideal monochromatic plane wave, particle velocity vector U<sub>PW</sub> can be saved as u = L ™ L<sub>P</sub><sup>at</sup> PW
P <A where the unit vector e<sub>d</sub> indicates the direction of wave propagation, e.g. You can prove that
<img file="PL2324645T3_D0002.tif" />
where and<sub>and</sub> means active intensity, p<sub>0</sub> means the density of air, c means the speed of sound, E is the energy of the sound field and Ψ means the dissipation.
[0038] It is worth noting that because all the components e<sub>d</sub> are real numbers, all U components<sub>PW</sub> they are in phase with P<sub>PW</sub>. Fig. Ib is an example of U<sub>PW</sub> and P<sub>PW</sub> in the Gaussian plane. As mentioned above, all U components<sub>PW</sub> they have the same phase as P<sub>PW</sub>, specifically Θ. On the other hand, their modules are related in the following way. Even if many sound sources are present, the pressure and velocity of the particle can still be expressed as the sum of the individual components. Without losing the general look, we can highlight the case of two sources of sound. In fact, the extension to more sources of sound is simple. Let F ^ and F & be the pressures registered for the first and second sources, e.g. representing the first and second wave field measures. Similarly, Niech i ć /<sup>2)</sup> they will be composite particle velocity vectors. Considering the linear nature of the propagation phenomenon, when the sources emit simultaneously, the observed pressure. The time of the particle velocity i / are
P = + P @ u = u<sup>m</sup> + u<sup>m</sup> [0040] Thus, the active intensities are /?>
With <<sup>2</sup>> = ź-Re {p<sup>(2,</sup>£ Z<sup><2)</sup>} [0041] Hence /, = la<sup>}</sup> + la<sup>21</sup> + | Rejp<sup>(,)</sup>· V<sup>(2></sup>+ P<sup>m</sup>-σ ">}.
[0042] It should be noted that except in special cases
<img file="PL2324645T3_D0003.tif" />
<img file="PL2324645T3_D0004.tif" />
[0043] When two, e.g. flat, waves are perfectly consistent in phase (although they are traveling in different directions), where γ is a real number, then
With <<sup>2)</sup> Rel = {p<sup>(2)</sup>T ^}
ΙΑΉΠΜ and la = (1 + 7) ^ + (1 + -) / ^.
[0044] When the waves are in phase and move in the same direction, they can of course be interpreted as one wave.
[0045] For γ = -1 and any direction, the pressure disappears and there can be no energy flow, i.e. || I<sub>and</sub>|| = 0.
[0046] When the waves are ideally in a square, then p <2) = χ. £ Υ<sup>π</sup>'<sup>2</sup>/><sup>(,)</sup> <7<sup>(2)</sup> = / - e ' "<sup>/ 2</sup>tZ<sup><</sup>'<sup>) </sup>iZx<sup><2></sup>Xe =<sup>AND</sup>'<sup>2</sup>TFX<sup>(L></sup> .
CZ /<sup>2)</sup> tZ<sub>2</sub><sup>(2)</sup> where γ is a real number. It follows that
AND? iR =<sub>e</sub>{P < '>. ! / ( '>}
I <"= iRejp<sup>B</sup>-P "'),
Κ'ΙΙΉΊΙ /: "] and
I = /<sup>(1)</sup>+ /<sup><2)</sup> and <sup>Λ</sup> and <sup>Ύ</sup> a '[0047] By means of the above equations, it can be easily proved that for a plane wave, each of the exemplary quantities U, P and e<sub>d</sub> or P and I<sub>and</sub> they may represent an equivalent and exhaustive description, as all other physical quantities can be derived therefrom, i.e. in the embodiments of the present invention, any combination thereof may be used in place of a wave field measure or a measure of the wave direction. For example, in embodiments of the present invention, the 2norma of the active intensity vector can be used as a measure of the wave field.
[0048] A minimum description for carrying out the connection defined in the embodiments of the invention may be specified. The particle velocity and vectors for the i-th plane wave can be expressed as p = | p <> | e (/) (») u<sup>(and)</sup> = where Z / *<sup>0</sup> represents the phase f®. By expressing the combined intensity vector, i.e. the combined measure of the wave field and the combined measure of the direction of arrival with respect to these variables, you can write + -Re 2
2p<sub>about</sub>c
After<sup>c</sup> + -Re 2
pq<sup>c</sup> [0049] It should be noted that the first two components of the sum are i / J<sup>2)</sup>. The equations can be further simplified to
And a
<img file="PL2324645T3_D0005.tif" />
2p<sub>0</sub>c ·
<img file="PL2324645T3_D0006.tif" />
| P<sup>(2)</sup>MP<sup>and,)</sup>k '<sup>)</sup> cos (zp<sup>(2)</sup> - ZP<sup>(,)</sup>)
2p<sub>about</sub>c [0050] After introduction
Δ<sup>(1</sup>·<sup>2)</sup>= | ΖΡ<sup>(2)</sup>-ΖΡ<sup>(,)</sup>| we get
<img file="PL2324645T3_D0007.tif" />
[0051] This equation shows that the information needed to calculate I<sub>and</sub> it can be reduced to l ^ l, '| ZP<sup>(2)</sup> - ZP<sup>(1)</sup>|. In other words, the representation of each, e.g. flat, wave can be reduced to the amplitude of the wave and the direction of propagation. In addition, the relative phase difference between waves can also be taken into account. When more than two waves are combined, phase differences between all wave pairs can be considered. Of course, there are a number of other descriptions that contain the same information. For example, knowledge of intensity vectors and phase differences may be equivalent.
[0052] Generally, the energy description of the flat waves may not be sufficient to properly carry out the joining. Combining can be approximated by assuming that the waves are in a square. A comprehensive wave descriptor (i.e. all physical quantities are known) may be sufficient to connect, although it may not be necessary in all embodiments of the invention. In the variants of the invention performing the correct connection, the amplitude of each wave, the propagation direction of each wave, and the relative phase difference between each of the wave pairs to be combined can be taken into account.
[0053] The determining means 110 may be adapted to delivery and / or the processor 130 may be adapted to process the first and second arrival directions and / or to provide a combined measure the arrival direction as a common vector e<sub>DOWN</sub>A (k, n), with Gdoa (k, n) = - el (k, n) Qttz Ia (k, n} = \\ IJJ <, n) \\ 'ei (k, n} z | ZP ( 2) - ZP (1) |
<img file="PL2324645T3_D0008.tif" />
and denoting the time-frequency processed particle velocity vector u (t) = [u<sub>x</sub>(T), Uy (t)<sub>/</sub>at<sub>with</sub>(T)]<sup>T</sup>. In other words, let p (t) and u (t) = [u<sub>x</sub>(here<sub>s</sub>(T)<sub>/</sub>at<sub>with</sub>(T)]<sup>r</sup> will be respectively the pressure and the particle velocity vector for a specific point in space, where [']<sup>r</sup> means transposition. These signals can be processed into a time-frequency domain by means of a suitable filter bank, e.g. a short-term Fourier transform (STFT), as suggested by, e.g., V. Pulkki and C. Faller, Direction and audio coding: Filterbank and STFT-based design, in 120th AES Convention, May 20-23, 2006, Paris, France, May 2006.
[0054] Let P (k, n) and υ / ^^^ / υ ^, η ^ υ / ^ η ^ υ ^, η)]<sup>7</sup>Q-shaped signals are processed, where the kinases are the frequency (or frequency band) and time indexes, respectively. The active intensity vector Ijk, ri) can be defined as
AND<sub>and</sub>{k, n) = L ^ βp ^ k, rij-U \ k, n)} (1) where (·) * denotes the complex conjugate and Re {·} extracts the real part. The active intensity vector expresses the pure energy flow characteristic of the sound field (see FJ Fahy, Sound Intensity, Essex: Eisevier Science Pubiishers Ltd., 1989) and can therefore be used as a measure of the wave field.
[0055] Let c denote the sound speed in a given medium and E the energy of the sound field defined by FJ. Fahy
Α || σ (Α ») || ΥΥΑ (Ζ :,<sub>Π</sub>) |<sup>2</sup> , (2) where || · || calculates the 2-norm. The monophonic DirAC stream will be discussed in detail below.
[0056] A mono DirAC stream may consist of a monophonic pff signal) and additional information. Additional information may include time and frequency-dependent direction of arrival and a time-and-frequency-dependent measure of dispersion. The direction of arrival can be saved as e<sub>DOWN</sub>A (k, n), and is a unit vector indicating the direction from which the sound arrives. The measure of distraction, however, is written as<a name="caption1"></a>W, «) · [0057] In embodiments of the present invention, means 110 and / or processor 130 may be adapted to deliver / process the first and second DAO and / or combined DAO in the form of a unit vector e<sub>D0</sub>A (k, n). The direction of arrival can be obtained as «DOA ^<sup>k</sup>, nj = -e, {k, ri), where the unit vector e<sub>r</sub>(k, n) indicates the direction that is indicated by the active intensity, specifically
Λ (*, η) = || /<sub>ο</sub> (k, η) || · E<sub>f</sub> (k, η), e, (k, η) = Ι<sub>α</sub> (k, η) / || /<sub>β</sub> (ł, λ) || . (3) [0058] Alternatively in embodiments of the present invention, the DAO may be expressed in the form of an azimuth (azimuthal length) and elevation angles in a spherical coordinate system. For example, if φ and θ are respectively azimuth and elevation angles, then<sup>e</sup>DOA <sup>Λ</sup>) - [cos (^) 'cos (i9), sin (^) · cos («9), sin (i9) P. (4) [0059] In embodiments of the present invention, the determining means 110 and / or the processor 130 may be adapted to deliver / process the first and second scattering parameters and / or the combined scattering parameter through Ψ (Ą n) in a time dependent manner. and frequency. The set-up means 110 may be adapted to provide the first and / or second stray parameter and / or the processor 130 may be adapted to provide a combined dispersion parameter in the form (5) where it is the mean over time.
[0060] In practice, there are different approaches to obtaining P (k, n) and U (k, n). One of the possibilities is to use the B-format microphone, which provides 4 signals, namely in {£), KJj y (f) and Ąt). The first one, corresponds to the pressure reading of the omnidirectional microphone. The next three signals are readings of microphones directed in the directions of the three axes of the Cartesian coordinate system. These signals are also proportional to the speed of the particle. Therefore, in some embodiments of the present invention
P (k, n) = W (k, n)! / (*, «) = -7 =! - [* (*,«), Z (fc, «), Ζ (ί,") Γ <sup>(6)</sup> where WJęn), Ąk<sub>f</sub>n), y [k, n) and Ąk, n) are processed B-format signals. It should be noted that the V2 coefficient in equation (6) derives from the convention used in the definition of * B-format signals (see Michaei Gerzon, Surround sound psychoacoustics, In Wire / ess World, voiume 80, pages 483-486, December 1974).
[0061] Alternatively, P (k, n) and U (k, n) can be estimated using a number of multidirectional microphones, as suggested in J. Merima, Applications of 3-D microphone array, in 112th AES Convention, Paper 5501, Munich, May 2002. The processing steps described above are also illustrated in Fig. 2.
[0062] Fig. 2 shows a DirAC encoder 200 that is adapted to calculate a mono audio channel and additional information from respective input signals, e.g., microphone signals. In other words, Fig. 2 shows the encoder 200
DirAC to determine the dispersion and direction of arrival from the correct microphone signals. Fig. 2 shows a DirAC encoder 200 comprising a P / U estimator module 210. The P! U estimator module 210 receives the microphone signals as the input information on which the Ρ / U estimation is based. Since all information is available, the P / iZ estimation is simple, according to the above equations. Step 220 of the energy analysis allows estimating the direction of arrival and the dispersion parameter of the combined stream.
[0063] In embodiments of the present invention, other audio signal streams may be combined than a mono DirAC audio stream. In other words, in embodiments of the present invention, the determining means 110 may be adapted to process any other stream of audio signals into a first and a second stream of an audio signal, such as, for example, stereo or surround audio data. In the event that the variants of the invention combine DirAC streams other than mono, they may differ from case to case. If the DirAC stream includes B-format signals, as acoustic signals, the particle velocity vector will be known and combining will be trivial, as will be discussed in more detail below. If the DirAC stream contains signals other than B-format or a monaural omnidirectional signal, the determining means 110 may be adapted to process the two mono DirAC streams first, after which the variants of the invention may respectively combine the processed streams. In embodiments of the present invention, the first and second spatial audio streams can thus represent processed mono DirAC streams.
[0064] Variants of the present invention may combine available audio channels to approximate a multidirectional reception system. For example, in the case of a stereo DirAC stream, this can be achieved by summing the left L channel and the right R channel.
[0065] In the following, physical phenomena will be explained in the field produced by many sound sources. When dealing with many sound sources, it is still possible to present the pressure and velocity of the particle as the sum of the individual components.
[0066] Let A (k, n) and lfl (k, n) show the pressure and velocity of a particle that would be registered for the i-th source, if it were the only source. Assuming the linearity of the propagation phenomenon, when N sources are active at the same time, the observed pressure P (k, ri) and the velocity of the particle L {k, ri) are (7) i = l and respectively
V {k, n) * j / Xk, n). (8) fs = l [0067] The previous equations show that if both the pressure and the velocity of the particle are known, obtaining a combined monophonic DirAC stream is simple. Such a situation is shown in Fig. 3. Fig. 3 shows a variant of the invention implementing an optimized or possibly perfect combination of multiple audio streams. In the example of Fig. 3, it is assumed that all vectors of pressure and particle velocity are known. Unfortunately, such trivial merging is not possible with monophonic DirAC streams for which the particle velocity of iPfcn) is unknown.
[0068] Fig. 3 shows N streams, and for each of them in blocks 301, 302-30N an estimation of PI U is performed. The result of the actions in the P / U estimation blocks are the respective time-frequency representations of individual signals ^ \ k, n) il / \ k, n), which can then be combined according to the above equations (7) and (8), as illustrated in the two summaries 310 and 311. After obtaining the combined P (k, ri) and U (k, n), Stage 320 of the energy analysis can determine the dispersion parameter Ψ (k, ri) and the direction of arrival e<sub>D0</sub>A (k, n) \ n a simple way.
[0069] Fig. 4 shows an embodiment of the invention for combining multiple mono DirAC streams. As described above, the N streams are to be connected by a variant of the apparatus 100 shown in Fig. 4. As shown in Fig. 4, each of the input N-streams can be represented by a time and frequency dependent mono-frequency representation of 't', k) , from the direction of arrival<sup>e</sup>DOA ^^) <sub>and</sub> where <sup>(1)</sup> represents the first stream. A corresponding representation is also shown in Fig. 4 for combined streams.
The task of combining at least two mono DirAC streams is shown in Fig. 4. Since the pressure P (k, ri) can be obtained simply by summing the known values f ^ '\ k, n} as in equation (7), the problem combining at least two monophonic DirAC streams is reduced to determine e<sub>D0</sub>A (k, n} and Ψ (λ; / 7) The following variant of the invention is based on the assumption that the field of each source consists of a plane wave summed to the scattering field, hence the pressure and velocity of the particle for the i-th source can be expressed as
P "\ k, n) = P ^ (k, n) + P $ (k, n) lP '\ k,<sub>n</sub>) = U ^ (k, n) + U% (k, "), (9) (10) where the subscripts" PW "and" diff "mean flat wave and stray field respectively. A variant of the present invention is provided below, including a strategy for estimating the direction of sound arrival and dispersion. The corresponding processing steps are shown in Fig. 5.
[0071] Fi 5 shows another device 500 for combining multiple audio streams, which will be discussed in detail below. Fig. 5 illustrates the processing of the first spatial stream of the audio signal in the form of a first monophonic representation of the first incoming direction ^% οα and the first scattering parameter Ψ<sup>(1)</sup>. As shown in Fig. 5, the first stream of spatial audio signal is divided into an approximate representation f<sub>and</sub>a flat stream as well as a second stream of spatial audio signal and potentially other streams of spatial audio signal, respectively
- Ratings are marked with a visor over the appropriate representation in the pattern.
[0072] Estimator 120 may be adapted to estimate N wave representations
Ppw (fi * <sup>in</sup>)<sub>ABOUT</sub>once the representation of the scatter field as the approximations ^ (ννι) for N spatial streams of acoustic signal, where l <i <N. The processor 130 may be adapted to determine the combined arrival direction based on estimating ^ DOA where
¡(Ł, n) = 2 Re {p<sub>in</sub> (k, n) · U '<sub>PW</sub> (K ")} <a name="caption2"></a>Aw (M) = £ Ą<sup>(</sup>£ (M), / «1
Ppi (k, n) = a<sup>{,</sup>\ k, n) · P<sup>{</sup>'\ k, ri),
V<sub>n</sub>(k, n) = ^ (k, n), il </ & (*, «) = - p- \ k, n) Ρ<sup>ω</sup><Μ e ^ k.ri),
Pr \ C with real values a<sup>in</sup>(K, n) 3<sup>in</sup>(K, n) e {0.,. L}.
[0073] Fig. 5 shows a dashed line 120 and a processor 130. In the variant 20 shown in Fig. 5, the determining means 110 are not present because the first stream of spatial audio signal and the second stream of spatial audio signal are assumed to be as also potentially other acoustic signal streams are provided in the mono DirAC representation, i.e. mono representations, DOA and scatter parameters are just separated from the stream. As shown in Fig. 5, the processor 130 may be adapted to determine the combined DOA based on the estimate.
[0074] The direction of the arrival of the sound, i.e. the direction measure, can be estimated by which it is calculated as (11)<sub>and</sub>(k> n).
|| λ (Μ) 1 where X (k, n) is an estimate of the active intensity for the combined stream. It can be obtained as follows / "(*,«) = j (k, n) U '<sub>PW</sub> (k, π)}, (12)
Λ where <sub>and</sub> AT<sub>PW</sub>{k, n) estimates of the pressure and velocity of the particle corresponding only to flat waves, e.g. as a wave field measure. They can be defined as follows
P "(k, n) = ^ (k, n), (13) / = 1
P &&, «) = a<sup>(and</sup>\ k, ri) P<sup>{and</sup>\ k, n), (14)
AT<sub>PW</sub>{kA = YU<sup>{</sup>A {k, n), (15) i = 1
U ^ (km) ^ - fi <\ k, n) P<sup>in</sup>(k, n) · «& / *.«). (16)
After<sup>c</sup> [0075] The ratios i.c.<sub>f</sub>ri) and fi'Xk, ri) are generally frequency-dependent and may have inverse proportionality to distraction Ψ<sup>(1)</sup>(Α // 7). Actually, when the dispersion Ψ<sup>(1)</sup>(Α; ζ7) has a value close to zero, it can be assumed that the field consists of a single flat wave, so that
P $ (k, n) «P (k, n) (17) and ΰρ'ιί- (k, n) a - 1- P<sup>(,)</sup> (k, n) · (k, n), (18)
After<sup>c</sup> implying that a<sup>(L)</sup>(K, n) = 3<sup>(L)</sup>(K, n) = l.
[0076] Two variants of the present invention will be set out below which determine a<sup>0)</sup>(k, n) ip<sup>in</sup>(K, n). First, the energy considerations of the scatter fields are considered. In these variants, it can be assumed that the field consists of a plane wave summed to the ideal field of dissipation.
[0077] In embodiments of the present invention, the estimator 120 may be adapted to determine a<sup>in</sup>(k, n) and 3<sup>(L)</sup>(k, n) according to (19) a<sup>(,</sup>\ k<sub>9</sub>n) - β<sup>ω</sup>(/ ε, η) (k, η) = y / lF<sup>(and</sup>\ k, η) 'by setting the air density p<sub>0</sub> equal to 1, and eliminating the functional relationship (k, ri) for simplicity, you can save
<img file="PL2324645T3_D0009.tif" />
(20) [0078] In embodiments of the present invention, the processor 130 may be adapted to approximate scattered fields based on their statistical properties, which approximation may be obtained in the following manner>, + 2c '(21) where F ^ is the energy of the scattered field . Variants can therefore estimate pw \, = 7ι-ψ<sup>(0</sup> (22) [0079] To calculate immediate estimates (i.e. for each time-frequency plate, "tHć<sup>1</sup>} variants of the invention may remove the expectation operators, obtaining (k, ri) = ri) (23) [0080] By using the plane wave assay, the particle velocity estimation can be obtained directly - P ^ (k, n) e ^ \ k, ri). (24) [0081]
In embodiments of the present invention, simplified particle speed modeling may be used. In these variants, the estimator 120 may be adapted to approximate the coefficients a (i) (k, n) and 3 (i) (k, n) based on simplified modeling. In the variants of the invention an alternative approach can be used, which can be achieved by introducing simplified modeling of the particle speed p<sup>l</sup>'\ k, ri) (25)
Ι-! Γ<sup>ω</sup>(λ, «) [0082] The approach is shown below. The particle velocity ΐΧ (λ> / 7) is modeled as -X '' (*, «). (26)
Pci [0083] Coefficient β<sup>(0</sup>(/ ^ ζ7) can be obtained by replacing equation (25) with equation (5), which leads to the form - <| ^ <'> (fc, n) - Ρ<sup>(0</sup>(Λ, zi) f - /> (k, nj>, | P<sup>{,</sup>\ k, nj = \ -Fń-_ (27) · | / Ά / 1) | (Y (Λ, «) +!)>,
2p ^ c '[0084] In order to obtain immediate values, expectation operators can be removed and solved for β<sup>ω</sup>(A> / 7) getting
P<sup>v</sup>\ k, nj =
(28) [0085] It should be noted that this approach leads to similar directions of sound arrival, as in the case of equation (19), but with a lower complexity of calculations, taking into account the fact that the coefficient a<sup>(/)</sup>(Ąz7) is one.
[0086] In embodiments of the present invention, the processor 130 may be adapted to estimate the spread, i.e. to estimate the combined dispersion parameter. The dissociation of the combined streams, written as Ψ (λ> / 7) can be estimated directly from known quantities Ψ<sup>ω</sup>(λ> / 7) and κ \ k, n) and from the estimation of I kk, r i}, obtained in the manner described above. After considering the energy factors introduced in the previous section, the variants of the invention may use the estimator <II // *, / i) j + X ς<sup>η</sup>Ί | χά Y>, ι = 1 (29) ,, ρ (') rz (0 [Know] <sup>J</sup>/ '»' And allows the use of alternative representations obtained in equation (b) variants. In fact, the wave direction can be obtained by U<sub>PW</sub> during p (O when <sup>r</sup>pw gives the amplitude and phase of the i-th wave. From the latter, it is easy to calculate all phase differences Δ<sup>(Ϋ)</sup>. The DirAC parameters of the combined stream can then be calculated by replacing equation (b) with equation (a), (3) and (5).
[0088] Fig. 6 shows variants of the method of combining at least two DirAC streams. These variants may provide a method for combining a first spatial stream of an audio signal and a second spatial stream of an audio signal to obtain a combined audio signal stream. In variations, the method may comprise the steps of determining the first spatial stream of the audio signal, the first representation of the audio signal and the first DOA, as well as the second spatial stream of the audio signal, the second representation of the audio signal, and the second DOA. In the method variants, DirAC representations of the spatial audio streams may be available, so the determining step is in this case a simple reading of the respective representations from the audio streams. In Fig.
[0089] In variations, the method may include the step of estimating a first wavelet representation comprising a first wavelength measure and a first wavelength field measure for a first spatial stream of the audio signal based on the first audio signal representation, the first DOA and optionally the first scatter parameter. Accordingly, the method may comprise the step of estimating a second wavelet representation comprising a second wavelength measure and a second wavelength field measure for the second spatial stream of the audio signal based on the second audio signal representation, the second DOA and optionally the second scatter parameter.
The method may further comprise the step of combining the first wave representation and the second wave representation to obtain a combined waveform comprising a combined field measure and a combined measure DOA and a step of combining the first representation of the audio signal and the second representation of the audio signal to obtain the combined representation of the audio signal, as shown in Fig. 6, as step 620 for monaural audio channels. The variant shown in Fig. 6 includes the step of calculating a<sup>in</sup>(A> z7) and β<sup>(Ι)</sup>(1<sub>ζ</sub>(7) according to equations (19) and (25), allowing the vectors of pressure and particle velocity to be estimated for the plane wave representation at step 640. In other words, the steps of estimating the first and second plane wave representations occur in steps 630 and 640 in Fig.
in the form of flat wave representations.
The step of combining the first and second plane wave representations occurs at step 650 where the vectors of pressure and particle velocity of all streams can be summed.
[0092] In step 660 in Fig. 6, the calculation of the active intensity vector and the estimate of DOA are made based on the combined representation of the plane wave.
[0093] Method variants may include the step of combining or processing a combined field measure, a first and a second monophonic representation, and a first and a second scattering parameter to obtain a combined spread parameter. In the method variants depicted in Fig. 6, the scattering calculation takes place at step 670, for example, on the basis of equation (29).
[0094] Variants of the method may provide the advantage that combining the spatial streams of the audio signal may take place with high quality and moderate complexity.
[0095] Depending on certain requirements for implementing the methods of the invention, the methods of the present invention may be implemented in hardware or in software. Implementations may be made using digital storage media, in particular Flash, disk, DVD or CD memories containing electronically readable control signals that interact with the programmed computer system in a manner that the methods of the present invention may be implemented. Generally, the present invention is a computer program code, where the program code is stored on a machine-readable carrier, and the program code may be operable to implement the methods of the present invention when the computer program is implemented in a computer or processor. In other words, the methods of the present invention,
Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung eV, Germany
Proxy:
Jan Dobrzański, Patent Attorney
EP 2 324 645 B1 Z-9545
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
25 members in 15 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 8852008 | United States of America | P | |
| 8852008 | United States of America | P | |
| 09001397 | European Patent Office (EPO) | A | |
| 09001397 | European Patent Office (EPO) | A | |
| 09806392 | European Patent Office (EPO) | A | |
| 2009005827 | European Patent Office (EPO) | W | |
| 2009005827 | European Patent Office (EPO) | W | |
| EP20090001397 | – | – | – |
| EP20090806392 | – | – | – |
| US20080088520P | – | – | – |
| WO2009EP05827 | – | – | – |
Members25
| Document | Office | Kind | |
|---|---|---|---|
| EP2154910A1 | European Patent Office (EPO) | A1 | |
| AU2009281355A1 | Australia | A1 | |
| CA2734096A1 | Canada | A1 | |
| WO2010017966A1 | World Intellectual Property Organization (WIPO) | A1 | |
| MX2011001653A | Mexico | A | |
| EP2324645A1 | European Patent Office (EPO) | A1 | |
| KR20110055622A | Republic of Korea | A | |
| CN102138342A | China | A | |
| US2011216908A1 | United States of America | A1 | |
| JP2011530720A | Japan | A | |
| EP2324645B1 | European Patent Office (EPO) | B1 | |
| ATE546964T1 | Austria | T1 | |
| ES2382986T3 | Spain | T3 | |
| HK1157986A1 | Hong Kong, China | A1 | |
| PL2324645T3This record | Poland | T3 | |
| RU2011106582A | Russian Federation | A | |
| KR101235543B1 | Republic of Korea | B1 | |
| AU2009281355B2 | Australia | B2 | |
| RU2504918C2 | Russian Federation | C2 | |
| CN102138342B | China | B | |
| US8712059B2 | United States of America | B2 | |
| JP5490118B2 | Japan | B2 | |
| CA2734096C | Canada | C | |
| BRPI0912453A2 | Brazil | A2 | |
| BRPI0912453B1 | Brazil | B1 |
Numbers
- Publication, DOCDB
- 2324645
- Publication, EPODOC
- PL2324645T
- Application
- 806392
- Application, DOCDB
- 09806392
- Application, EPODOC
- PL20090806392T
Titles2
- English
- APPARATUS FOR MERGING SPATIAL AUDIO STREAMS
- Polish
- Urządzenie do łączenia strumieni przestrzennego sygnału akustycznego
Classification
- CPC, 5
- H04S3/008
- H04S3/00
- G10L19/008
- H04S2420/03
- H04S2420/11
- IPC, 2
- H04S3 00
- G10L19 008