Method and apparatus for enhancement of audio reconstruction
Abstract
antagonists of trpv1 and uses thereof the invention relates to compounds of the formula i and pharmaceutically acceptable derivatives thereof, compositions comprising an effective amount of a compound of the formula i or a pharmaceutically acceptable derivative thereof, and methods for the treatment or prevention of a condition, for example, pain, iu, ulcers, dii and sii, comprising administering an effective amount of a compound of the formula to an animal in need of said compound or a pharmaceutically acceptable derivative thereof.
Term
1.4 yearsleft in the term
Expires 1 February 2028.
- Priority
- Filed
- Granted
- Today
- Expires
14 claims: 5 independent, 9 dependent
- 1Method for the reconstruction of an audio signal having at least one audio channel and associated direction parameters indicating an origin direction of a portion of the audio channel in relation to the recording position, the method comprising:selecting a direction set from origin in relation to the recording position;and modifying the portion of the audio channel to obtain a reconstructed portion of the reconstructed audio signal, characterized by the fact that the modification involves increasing the intensity of the portion of the audio channel, having direction parameters indicating a direction of origin next to a set of origin direction in relation to another portion of the audio channel having direction parameters indicating a direction of origin more distant from the set of origin direction. 1. Método para a reconstrução de um sinal de áudio tendo pelo menos um canal de áudio e parâmetros de direção associados indicando uma direção de origem de uma porção do canal de áudio em relação à posição de gravação, o método compreendendo: selecionar um conjunto de direção de origem em relação à posição de gravação;e modificar a porção do canal de áudio para a obtenção de uma porção reconstruída do sinal reconstruído de áudio, caracterizado pelo fato de que a modificação compreende o aumento de uma intensidade da porção do canal de áudio, tendo parâmetros de direção indicando uma direção de origem próxima a um conjunto de direção de origem em relação a outra porção do canal de áudio tendo parâmetros de direção indicando uma direção de origem mais distante do conjunto de direção de origem.
- 1112. Method, according to any one of the preceding claims, characterized by the fact that it also comprises:panoramicization of the amplitude of the channel portions so that a perceived direction of origin of the reconstructed channel portions corresponds to the direction of origin when reproduced using a predetermined installation of speakers. 12. Método, de acordo com qualquer uma das reivindicações anteriores, caracterizado pelo fato de que compreende ainda: panoramização da amplitude das porções de canais de maneira que uma direção de origem percebida das porções de canais reconstruidas corresponda à direção de origem quando reproduzidas usando uma instalação predeterminada de altofalantes .
- 1213. Method for enhancing the directional perception of an audio signal, the method comprises:obtaining at least one audio channel and associated direction parameters indicating an origin direction of a portion of the audio channel in relation to 13. Método para o realce da percepção direcional de um sinal de áudio, o método compreende: obter pelo menos um canal de áudio e parâmetros de direção associados indicando uma direção de origem de uma porção do canal de áudio em relação à Petition 870190102982, of 10/14/2019, p. 39/42 Petição 870190102982, de 14/10/2019, pág. 39/42 4/5 posição de gravação;selecionar um conjunto de direção de origem em relação à posição de gravação;e modificar uma porção do canal de áudio para obter uma porção de um sinal de áudio realçado, caracterizado pelo fato de que a modificação compreende o aumento da intensidade de uma porção do canal de áudio tendo parâmetros de direção indicando uma direção de origem próxima a um conjunto de direção de origem em relação a uma outra porção do canal de áudio tendo parâmetros de direção indicando uma direção de origem mais distante do conjunto de direção de origem. 4/5 recording position;select a set of origin direction in relation to the recording position;and modifying a portion of the audio channel to obtain a portion of an enhanced audio signal, characterized by the fact that the modification comprises increasing the intensity of a portion of the audio channel having direction parameters indicating a direction of origin close to a set of origin direction in relation to another portion of the audio channel having direction parameters indicating a direction of origin more distant from the set of origin direction.
- 1314. Audio decodifloader for the reconstruction of an audio signal having at least one audio channel and associated direction parameters indicating an origin direction of a portion of the audio channel in relation to a recording position, comprising:an adapted direction selector to select a set of home direction in relation to the recording position;and an audio portion modifier to modify the audio channel portion to obtain a reconstructed portion of the reconstructed audio signal, characterized by the fact that the modification comprises increasing the intensity of the audio channel portion having direction parameters indicating a direction of origin close to a set of direction of origin in relation to a 14. Decodifloader de áudio para a reconstrução de um sinal de áudio tendo pelo menos um canal de áudio e parâmetros de direção associados indicando uma direção de origem de uma porção do canal de áudio em relação a uma posição de gravação, compreendendo: um seletor de direção adaptado para selecionar um conjunto de direção de origem em relação à posição de gravação;e um modificador da porção de áudio para modificar a porção do canal de áudio para a obtenção de uma porção reconstruída do sinal reconstruído de áudio, caracterizado pelo fato de que a modificação compreende o aumento da intensidade da porção do canal de áudio tendo parâmetros de direção indicando uma direção de origem próxima a um conjunto de direção de origem em relação a uma directional perception of an audio signal, the audio encoder comprising: a signal generator to obtain at least one audio channel and associated direction parameters indicating a direction percepção direcional de um sinal de áudio, o codificador de áudio compreendendo: um gerador de sinais para obter pelo menos um canal de áudio e parâmetros de direção associados indicando uma direção Petition 870190102982, of 10/14/2019, p. 40/42 Petição 870190102982, de 14/10/2019, pág. 40/42 5/5 de origem de uma porção do canal de áudio em relação a uma posição de gravação;um seletor de direção adaptado para selecionar um conjunto de direção de origem em relação à posição de gravação;e um modificador de sinais para modificar a porção do canal de áudio na obtenção de uma porção de um sinal de áudio realçado, caracterizado pelo fato de que a modificação compreende o aumento da intensidade de uma porção do canal de áudio tendo parâmetros de direção indicando uma direção de origem próxima a um conjunto de direção de origem em relação a uma outra porção do canal de áudio tendo parâmetros de direção indicando uma direção de origem mais distante do conjunto de direção de origem. 5/5 origin of a portion of the audio channel in relation to a recording position;a direction selector adapted to select a set of original direction in relation to the recording position;and a signal modifier to modify the portion of the audio channel in obtaining a portion of an enhanced audio signal, characterized by the fact that the modification comprises increasing the intensity of a portion of the audio channel having direction parameters indicating a direction of origin close to a set of direction of origin in relation to another portion of the audio channel having direction parameters indicating a direction of origin more distant from the set of direction of origin.
- 1416. System for enhancing a reconstructed audio signal, the system comprising:an audio encoder for obtaining an audio signal having at least one audio channel and associated direction parameters indicating an origin direction for a portion of the audio channel audio in relation to a recording position;a direction selector adapted to select a set of original direction in relation to the recording position;and an audio decoder having an audio portion modifier to modify the audio channel portion in obtaining a reconstructed portion of the reconstructed audio signal, characterized by the fact that the modification involves increasing the intensity of the portion of the audio channel having direction parameters indicating a direction of origin close to a set of direction of origin in relation to another portion of the audio channel having direction parameters indicating a direction of origin furthest from the origin direction set. 16. Sistema para o realce de um sinal reconstruído de áudio, o sistema compreendendo: um codificador de áudio para a obtenção de um sinal de áudio tendo pelo menos um canal de áudio e parâmetros de direção associados indicando uma direção de origem de uma porção do canal de áudio em relação a uma posição de gravação;um seletor de direção adaptado para selecionar um conjunto de direção de origem em relação à posição de gravação;e um decodificador de áudio tendo um modificador da porção de áudio para modificar a porção do canal de áudio na obtenção de uma porção reconstruída do sinal reconstruído de áudio, caracterizado pelo fato de que a modificação compreende o aumento da intensidade da porção do canal de áudio tendo parâmetros de direção indicando uma direção de origem próxima a um conjunto de direção de origem em relação a uma outra porção do canal de áudio tendo parâmetros de direção indicando uma direção de origem mais distante do conjunto de direção de origem.
Independent claims5
147 paragraphs, as filed
AUDIO METHOD AND DECODER FOR THE RECONSTRUCTION OF AN AUDIO SIGNAL, AUDIO METHOD AND ENCODER FOR ENHANCING THE DIRECTIONAL PERCEPTION OF AN AUDIO SIGNAL AND SYSTEM FOR ENHANCING A RECONSTRUCTED AUDIO SIGNAL
FIELD OF THE INVENTION [0001] The present invention relates to techniques on how to improve the perception of a direction of origin of a reconstructed audio signal. In particular, the present invention proposes an equipment and method for reproducing recorded audio signals in such a way that a selectable direction of audio sources can be emphasized or overweight in relation to audio signals from other directions.
BACKGROUND OF THE INVENTION AND PREVIOUS TECHNIQUE [0002] In general, in multi-channel reproduction and listening, the listener is surrounded by multiple speakers. There are several methods for capturing audio signals for specific installations. A general objective in reproduction is to reproduce the spatial composition of the originally recorded signal, that is, the origin of the individual audio source, like the place of a trumpet in the orchestra. Various speaker installations are quite common, and different spatial impressions can be created. Without the use of special post-production techniques, commonly known two-channel stereo installations can only recreate auditorium events on a line between the two speakers. This is commonly done by the so-called '' amplitude-panning '', in which the amplitude of the signal associated with an audio source is distributed between the two speakers, depending on the position of the audio source in relation to the speakers. loudspeakers. This is
Petition 870190102982, of 10/14/2019, p. 7/42
2/30 normally done during recording or subsequent mixing, that is, an audio source from the far left in relation to the position of the listener will be mainly played through the left speaker, whereas an audio source in front of the position of the listener listener will be played with identical amplitude (level) by both speakers. However, sound from other directions cannot be played.
[0003] As a consequence, using more speakers that are positioned around the listener, more directions can be covered, creating a more natural spatial impression. The probably most well-known multi-channel speaker layout is the 5.1 standard (ITU-R775-1), which consists of 5 speakers, whose azimuth angles to the listener's position are predetermined as 0<sup>O</sup>, ± 30 ° and ± 110 °, that is, during recording or mixing, the signal is configured for that specific speaker configuration, and deviations from the standard of a reproduction installation will result in reduced reproduction quality.
[0004] Several other systems have been proposed with several numbers of speakers located in different directions. Professional and special systems, especially in theaters and sound installations, also include speakers at different heights.
[0005] According to different reproduction facilities, several different recording methods have been designed and proposed for the aforementioned speaker systems, in order to record and reproduce the spatial impression in the hearing situation as land was perceived in the
Petition 870190102982, of 10/14/2019, p. 8/42
3/30 recording environment. A theoretically ideal way to record the spatial sound of a chosen multichannel speaker system would be to use the same number of microphones and speakers. In this case, the directivity standards of the microphones should also correspond to the speaker layout, so that the sound from any single direction would only be registered with a small number of microphones (1, 2 or more). Each microphone is associated with a specific speaker. The more speakers that are used for playback, the narrower the microphone's directivity standards should be. However, narrow directional microphones are quite expensive and typically have a non-flat frequency response, which reduces the quality of the recorded sound undesirably. In addition, the use of multiple microphones with very broad directivity patterns as input for multichannel reproduction results in a blurred and colorful perception of hearing due to the fact that sound from a single direction would always be reproduced with more speakers than necessary, as it would be registered with microphones associated with the different speakers. In general, the microphones currently available are best suited for recordings and reproductions on two channels, that is, they are designed without the objective of reproducing a spatial surround type impression.
[0006] From the point of view of the microphone project, several approaches were discussed to adapt the directivity standards of the microphones to the demands in audio-spatial reproduction. In general, all microphones capture sound differently, depending on the direction of sound arrival at the microphone, that is, microphones have different sensitivities, depending on the
Petition 870190102982, of 10/14/2019, p. 9/42
4/30 direction of arrival of the recorded sound. In some microphones, this effect is reduced, as they capture the sound almost independently of the direction. These microphones are generally referred to as omnidirectional microphones. In a typical microphone design, a circular diaphragm is attached to a small, air-tight wrap. If the diaphragm is not attached to the casing and the sound reaches it equally from each side, its directional pattern has two lobes, that is, this microphone captures the sound with equal sensitivity both from the front and the rear of the diaphragm , although with reverse polarities. This microphone does not capture sound from the direction coinciding with the diaphragm plane, that is, perpendicular to the direction of maximum sensitivity. This directional pattern is called a dipole, or figure of eight.
[0007] Omnidirectional microphones can also be changed to directional microphones, through a non-airproof wrap for the microphone. The envelope is specially constructed so that the sound waves can propagate through the envelope and reach the diaphragm, where some propagation directions are preferred, so that the directional pattern of this microphone becomes a pattern between the omnidirectional and the dipole. These patterns can, for example, have two lobes. However, the lobes can have different strengths. Some commonly known microphones have patterns that have only a single lobe. The most important example is the cardioid pattern, where the directional function D can be expressed as D = 1 + cos (θ), Θ being the direction of arrival of the sound. Therefore, the directional function quantifies what fraction of the amplitude of the incoming sound is
Petition 870190102982, of 10/14/2019, p. 10/42
5/30 captured, depending on direction.
[0008] The omnidirectional patterns discussed above are also called zero order patterns and the other previously mentioned patterns (dipole and cardioid) are called first order patterns. All microphone designs discussed above do not allow arbitrary conformation of directivity standards, since their directivity standards are totally determined by their mechanical constructions.
[0009] To partially solve this problem, some specialized acoustic structures have been designed, which can be used to create directional patterns narrower than those of the first order microphones. For example, when a tube with holes is attached to an omnidirectional microphone, a microphone with a narrow directional pattern can be created. These microphones are called shotgun or rifle microphones. However, they may not typically have a flat frequency response, that is, the directivity standard is narrowed at the cost of the recorded sound quality. In addition, the directivity pattern is predetermined by the geometric construction and, therefore, the directivity pattern of a recording made with this type of microphone cannot be controlled after recording.
[00010] Therefore, other methods have been proposed to allow partial change of the directivity pattern after the actual recording. In general, this is based on the essential idea of recording sound with a set of omnidirectional or directional microphones and then applying signal processing. Several of these techniques have been proposed recently. A good example
Petition 870190102982, of 10/14/2019, p. 11/42
6/30 simple is to record the sound with two omnidirectional microphones, which are placed close together, and subtract both signals from each other. This creates a virtual microphone signal having a directional pattern equivalent to a dipole.
[00011] In another method, more sophisticated microphone signal schemes can also be delayed or filtered before being added together. Using beam forming, a technique also known from the wireless LAN, a signal corresponding to a narrow beam is formed by filtering each microphone signal with a specially designed filter and by summing the signals after filtering. (filter sum beam formation). However, these techniques are blind to the signal itself, that is, they do not know the direction of arrival of the sound. Thus, a predetermined directional pattern must be defined, which is independent of the actual presence of a sound source in the predetermined direction. In general, estimating the direction of arrival of the sound is already a task in itself.
[00012] In general, several different spatial directional characteristics can be formed with the above techniques. However, the formation of spatially arbitrary selective sensitivity patterns (that is, the formation of narrow directional patterns) requires a large number of microphones.
[00013] An alternative way to create multichannel records is to locate a microphone close to each sound source (eg, an instrument) to be recorded and recreate a spatial impression by controlling the levels of the close-up microphone signals in the mix Final. However, this system requires a large number of microphones and a lot of user interaction to create the downPetition 870190102982, of 10/14/2019, p. 12/42
7/30 final mix.
[00014] A method for overcoming the above problem has recently been proposed, being called directional audio coding (DirAC), which can be used with different microphone systems and can record sound for reproduction with arbitrary speaker installations . The purpose of DirAC is to reproduce a spatial impression of an existing acoustic environment as precisely as possible, using a multichannel speaker system with an arbitrary geometric installation. Within the recording environment, the environment responses (which can be continuous recorded sound or impulse responses) are measured with an omnidirectional microphone (W) and a set of microphones that measure the direction of arrival of the sound and the diffusibility of the sound. sound. In the following paragraphs and within the application, the term diffusibility should be understood as a measure of the non-directivity of the sound, that is, the sound that reaches the listening or recording position has equal resistance in all directions, being diffused to the maximum . A common way to quantify the diffusion is to use the diffusibility values in the range [0, ..., 1], where the value 1 describes the sound with maximum diffusion and a value 0 describes a perfectly directional sound, that is, the sound that comes only from a clearly distinguishable direction. A commonly known method of measuring the direction of arrival of the sound is to apply 3 figure eight microphones (XYZ) aligned with the coordinated Cartesian axes. Special microphones have been designed, called SoundField microphones, which directly produce all the desired responses. However, as mentioned above, the W, X, Y and Z signals can also be
Petition 870190102982, of 10/14/2019, p. 13/42
8/30 computed from the set of discrete omnidirectional microphones.
[00015] In DirAC analysis, a recorded sound signal is divided into frequency channels, which correspond to the frequency selectivity of human hearing perception, that is, the signal, for example, is processed by a filter bank or a Transform Fourier to divide the signal into several frequency channels, having a bandwidth adapted to the frequency selectivity of human hearing. Then, the frequency band signals are analyzed to determine the direction of origin of the sound and a diffusibility value for each frequency channel with a predetermined time resolution. This time resolution does not need to be fixed and can, of course, be adapted to the recording environment. In DirAC, one or more audio channels are recorded or transmitted, together with the analyzed direction and the diffusibility data.
[00016] In synthesis or decoding, the audio channels finally applied to the speakers can be based on the omnidirectional channel W (registered with a high quality due to the standard of omni-directional directivity of the microphone used), or the sound of each speaker it can be computed as a weighted sum of W, X, Y and Z, thus forming a signal that has a certain directional characteristic for each speaker. Corresponding to the encoding, each audio channel is divided into frequency channels, which are optionally, furthermore, divided into diffuse and non-diffuse streams, depending on the analyzed diffusibility. If the diffusibility has been measured as high, a diffuse flow can be reproduced using a technique that produces a diffuse perception of sound, such as
Petition 870190102982, of 10/14/2019, p. 14/42
9/30 also used in Binaural Cue Coding. The non-diffuse sound is reproduced using a technique that aims to produce a point-type virtual audio source, located in the direction indicated by the direction data found in the analysis, that is, the generation of the Dirac signal, that is, the special reproduction is not dimensioned for an ideal specific speaker installation, as in the prior art (eg 5.1). This is particularly the case when the source of the sound is determined as direction parameters (that is, described by a vector) using knowledge about the directivity patterns in the microphones used in the recording. As already discussed, the origin of sound in three-dimensional space is parameterized in a frequency selective manner. Thus, directional printing can be reproduced with high quality for arbitrary speaker installations, as long as the geometry of the speaker installation is known. DirAC is therefore not restricted to special speaker geometries and, in general, allows for more flexible spatial reproduction of sound.
[00017] Although several techniques have been developed for the reproduction of multichannel audio recordings and the recording of the appropriate signals for later multichannel reproduction, none of the previous techniques allows to influence an already recorded signal, in a way that an emphasis can be emphasized.
<td>direction</td><td>source of signals</td><td>audio during</td><td>reproduction</td><td>in</td>
<td>way</td><td>that, for example,</td><td>intelligibility of</td><td>sign of</td><td>an</td>
<td>direction</td><td>distinct desired</td><td>be highlighted.</td><td></td><td></td>
<td></td><td colspan="2">SUMMARY OF THE INVENTION</td><td></td><td></td>
<td> [00018]</td><td colspan="2">According to a configuration</td><td colspan="2">of this</td>
invention, an audio signal can be reconstructed having at least
Petition 870190102982, of 10/14/2019, p. 15/42
10/30 an audio channel and associated direction parameters indicating the origin direction of a part of the audio channel in relation to a recording position, allowing an enhancement of the signal's perceptiveness coming from a different direction or from numerous different directions.
[00019] This means that, during playback, a desired direction of origin can be selected in relation to the recording position. While receiving a reconstructed portion of the reconstructed audio signal, the portion of the audio channel is modified so that the intensity of the portions of the audio channel is increased by having direction parameters indicating a direction of origin close to the desired direction of origin in relation to to the other portions of the audio channel having direction parameters indicating a direction of origin furthest from the desired direction of origin. The directions of origin of the portions of an audio channel or of a multichannel signal can be emphasized, in order to allow a better perception of the audio objects, which were located in the selected direction during the recording.
[00020] According to another configuration of the present invention, the user can choose, during reconstruction, which direction or which directions should be emphasized so that the portions of the audio channel or portions of multiple audio channels, which are associated with that chosen direction are emphasized, that is, so that their intensities or amplitudes are increased in relation to the remaining portions. According to a configuration, emphasis or attenuation of sound can be given from a specific direction with a more precise spatial resolution than with systems that do not implement the parameters of
Petition 870190102982, of 10/14/2019, p. 16/42
11/30 direction. According to another embodiment of the present invention, arbitrary spatial weighting functions, which cannot be obtained with ordinary microphones, can be specified. In addition, the weighting functions can vary in time and frequency, so that other configurations of the present invention can be used with great flexibility. In addition, weighting functions are extremely easy to implement and update, as they should only be loaded into the system instead of replacing hardware (for example, microphones).
[00021] According to another configuration of the present invention, audio signals having an associated diffusibility parameter, the diffusibility parameter indicating the diffusibility of the audio channel portion, are reconstructed so that the intensity of a portion of the audio channel with high diffusibility it is reduced in relation to another portion of the audio channel, having associated less diffusibility.
[00022] Thus, in the reconstruction of an audio signal, the diffusibility of the individual portions of the audio signal can be taken into account to further increase the directional perception of the reconstructed signal. Also, this can increase the redistribution of audio sources compared to techniques using only portions of diffused sound to increase the overall diffusibility of the signal instead of making use of the diffusibility information for a better redistribution of audio sources. Note that the present invention also allows, in contrast, to emphasize portions of the recorded sound that are of diffuse origin, such as ambient signals.
[00023] According to another configuration, at least one audio channel is subjected to upmixing on multiple channels of
Petition 870190102982, of 10/14/2019, p. 17/42
12/30 audio. The multiple audio channels can correspond to the number of speakers available for playback. Arbitrary speaker installations can be used to enhance the redistribution of audio sources, and it can be ensured that the direction of the audio source is always reproduced in the best way with existing equipment, regardless of the number of speakers available.
[00024] According to another configuration of the present invention, reproductions can even be made using a monophonic speaker. It is clear that the direction of origin of the signal will, in this case, be the physical location of the speaker. However, by selecting a desired direction of signal origin in relation to the recording position, the audibility of the signal from the selected direction can be significantly increased, when compared to the playback of a simple down-mix.
[00025] According to another configuration of the present invention, the direction of origin of the signal can be precisely reproduced, when one or more audio channels are subjected to upmixing to the number of channels corresponding to the speakers. The original direction can be reconstructed in the best way using, for example, amplitude panning techniques. To further increase the quality of perception, other phase changes can be introduced, which are also dependent on the selected direction.
[00026] Certain configurations of the present invention can also reduce the cost of microphone capsules for recording the audio signal without seriously affecting the audio quality, since at least the microphone used to determine the
Petition 870190102982, of 10/14/2019, p. 18/42
13/30 direction / diffusion estimate should not necessarily have a flat frequency response.
BRIEF DESCRIPTION OF THE DRAWINGS [00027] Various configurations of the present invention will be described below with reference to the accompanying drawings.
<td> [00028]</td><td>THE</td><td>Fig</td><td> . 1</td><td>show</td><td>a configuration of a</td><td>method</td>
<td colspan="2">for reconstruction</td><td>in</td><td>one s</td><td>after</td><td>audio;</td><td></td>
<td> [00029]</td><td>THE</td><td>Fig</td><td> . 2</td><td>show</td><td>a block diagram</td><td>on one</td>
<td>equipment for</td><td>The</td><td colspan="3">reconstruction of</td><td>an audio signal; and</td><td></td>
<td> [00030]</td><td>THE</td><td>Fig.</td><td> . 3</td><td>show</td><td>a block diagram of</td><td>another</td>
<td>configuration;</td><td></td><td></td><td></td><td></td><td></td><td></td>
<td> [00031]</td><td>THE</td><td>Fig</td><td> . 4</td><td>show</td><td>an application example</td><td>on one</td>
inventive method or inventive equipment in a teleconference setting;
[00032] Fig. 5 shows a configuration of a method for enhancing the directional perception of an audio signal;
[00033] Fig. 6 shows a decodifloader configuration for the reconstruction of an audio signal; and [00034] Fig. 7 shows a system configuration for enhancing the directional perception of an audio signal.
DETAILED DESCRIPTION OF THE PREFERRED CONFIGURATIONS [00035] Fig. 1 shows a configuration of a method for the reconstruction of an audio signal having at least one audio channel and associated direction parameters indicating an origin direction of a portion of the audio channel in relation to a recording position. In a selection step 10, a desired direction of origin is selected in relation to the recording position for a reconstructed portion of the reconstructed audio signal, in which the
Petition 870190102982, of 10/14/2019, p. 19/42
14/30 reconstructed portion corresponds to the portion of the audio channel, that is, for a portion of the signal to be processed, a desired direction of origin is selected, from which portions of signals will be clearly audible after reconstruction. The selection can be made directly by user input or automatically, as detailed below.
[00036] The portion may be a portion of time, a portion of frequency or a portion of time from a given frequency range of an audio channel. In a modification step 12, the portion of the audio channel is modified to obtain the reconstructed portion of the reconstructed audio signal, where the modification comprises increasing an intensity of a portion of the audio channel having direction parameters indicating a direction of origin close to the desired direction of origin in relation to another portion of the audio channel having direction parameters indicating a direction of origin more distant from the desired direction of origin, that is, these portions of the audio channel are emphasized by increasing their intensities or levels, which can, for example, be implemented by multiplying a scale factor of the portion of the audio channel. According to a configuration, portions originating from a direction close to the selected (desired) direction are multiplied by large scale factors, to emphasize these portions of signals in the reconstruction and to improve the audibility of these recorded audio objects, in which the listener is interested . In general, in the context of this application, increasing the strength of a signal or a channel will be understood as any measure that makes the signal better audible. This may, for example, be an increase in the breadth of the
Petition 870190102982, of 10/14/2019, p. 20/42
15/30 signal, the energy carried by the signal or by multiplying the signal by a scale factor greater than the unit. Alternatively, the volume of competitive signals can be reduced to achieve the effect.
[00037] The selection of the desired direction can be made directly through the user interface at the hearing location. However, according to alternative configurations, the selection can be made automatically, for example, by analyzing the directional parameters, so that the portion of frequencies having approximately the same origin is emphasized, considering that the remaining portions of the audio channel are suppressed. . Thus, the signal can be automatically focused on the predominant audio sources, without requiring additional user input at the listening tip.
[00038] According to other configurations, the selection step is omitted, since a direction of origin has been established, that is, the intensity of a portion of the audio channel is increased having direction parameters indicating a direction of origin close to the established direction. The established direction can, for example, be physically connected, that is, the direction can be predetermined. If, for example, only the central party is interested in a teleconference scenario, this can be implemented using a predetermined established direction. Other configurations can read the established direction from a memory that may also have stored some alternative directions to be used as established directions. One of these can, for example, be read when an equipment of the invention is connected.
Petition 870190102982, of 10/14/2019, p. 21/42
16/30 [00039] According to an alternative configuration, the selection of the desired direction can also be made on the encoder side, that is, when recording the signal, so that other parameters are transmitted with the audio signal, indicating the desired direction for playback. Thus, a spatial perception of the reconstructed signal in the encoder can already be selected without knowledge about the specific installation of the speaker used for reproduction.
[00040] Since the method for reconstructing an audio signal is independent of the specific speaker installation that is to reproduce the reconstructed audio signal, the method can be applied to monophonic as well as stereo or multichannel speaker configurations , that is, according to another configuration, the spatial impression of a reproduced environment is post-processed to enhance the perceptibility of the signal.
[00041] When used for monophonic playback, the effect can be interpreted as recording the signal with a new type of microphone capable of forming arbitrary directional patterns. However, this effect can be fully achieved at the receiving end, that is, during signal playback, without changes in the recording installation.
[00042] Fig. 2 shows a device configuration (decodifreader) for the reconstruction of an audio signal, that is, a configuration of a decoder 20 for the reconstruction of an audio signal. The decoder 20 comprises a direction selector 22 and an audio portion modifier 24. According to the configuration in Fig. 2, a multichannel audio input 26 registered by several microphones is analyzed by means of
Petition 870190102982, of 10/14/2019, p. 22/42
17/30 a direction analyzer 28 which obtains direction parameters indicating a direction of origin of a portion of the audio channels, that is, the direction of origin of the portion of the analyzed signal. According to a configuration of the present invention, the direction from which most of the energy is incident on the microphone is chosen. The recording position is determined for each specific signal portion. This can, for example, also be done using the DirAC microphone techniques previously described. Of course, another method of directional analysis based on the recorded audio information can be used to implement the analysis. As a result, the direction analyzer 28 obtains direction parameters 30, indicating the origin direction of a portion of an audio channel or the multichannel signal 26. In addition, directional analyzer 28 can operate to obtain a diffusibility parameter 32 for each signal portion (for example, for each frequency range or for each signal time period).
[00043] Direction parameter 30 and, optionally, diffusibility parameter 32 are transmitted to direction selector 22 which is implemented to select the desired origin direction in relation to a recording position for the reconstructed portion of the reconstructed signal from audio. Information about the desired direction is transmitted to the audio portion modifier 24. The audio portion modifier 24 receives at least one audio channel 34, having a portion, for which the direction parameters have been obtained. The at least one channel modified by the audio portion modifier can, for example, be a downmixing of the multichannel signal 26, generated by the algorithms
Petition 870190102982, of 10/14/2019, p. 23/42
Conventional 18/30 multichannel downmixing. An extremely simple case would be the direct sum of the signals from the multichannel audio input 26. However, since the configurations of the invention are not limited to the number of input channels, in an alternative configuration, all the audio input channels 26 can be processed simultaneously by the audio decoder 20.
[00044] The audio portion modifier 24 modifies the audio portion to obtain the reconstructed portion of the reconstructed audio signal, wherein the modification comprises increasing the intensity of a portion of the audio channel having direction parameters indicating a direction of origin close to the desired origin direction in relation to another portion of the audio channel having direction parameters indicating a direction of origin more distant from the desired direction of origin. In the example of Fig. 2, the modification is made by multiplying the scale factor 36 (q) by the portion of the audio channel to be modified, that is, if the portion of the audio channel is analyzed as originating from a direction close to the selected desired direction, a large scale factor 36 is multiplied by the audio portion. Thus, at its output 38, the audio portion modifier sends a reconstructed portion of the reconstructed audio signal corresponding to the portion of the existing audio channel at its input. As also indicated by the dashed lines at the output 38 of the audio portion modifier 24, this cannot only be done for a mono-output signal, but also for multichannel output signals, for which the number of output channels is not fixed or predetermined.
[00045] In other words, the configuration of the audio decoder 20 takes its input from this analysis
Petition 870190102982, of 10/14/2019, p. 24/42
19/30 directional as, for example, used in DirAC. The audio signals 26 from a set of microphones can be divided into frequency bands according to the frequency resolution of the human auditory system. The direction of the sound and, optionally, the diffusibility of the sound depending on the time in each frequency channel are analyzed. These attributes are also provided as, for example, azimuth (azi) and elevation (he) direction angles, and as Psi diffusibility index, which varies between zero and one.
[00046] Then, the desired or selected directional characteristic is imposed on the acquired signals using a weighting operation, which depends on the direction angles (azi and / or it) and, optionally, on the diffusibility (Psi). Of course, this weighting can be specified differently for different frequency bands and, in general, will vary over time.
[00047] Fig. 3 shows another configuration of the present invention, based on the DirAC synthesis. Thus, the configuration in Fig. 3 can be interpreted as enhancing DirAC reproduction, which allows controlling the sound level, depending on the direction analyzed. This makes it possible to emphasize sound from one or multiple directions, or to suppress sound from one or multiple directions. When applied to multichannel reproduction, post-processing of the reproduced sound image is obtained. If only one channel is used as an output, the effect is equivalent to using a directional microphone with arbitrary directional patterns when recording the signal. In the configuration shown in Fig. 3, the derivation of the direction parameters is shown, as well as the derivation of a transmitted audio channel. Analysis is based on channels
Petition 870190102982, of 10/14/2019, p. 25/42
20/30
W, X, Y and Z of B-format microphones, as, for example, recorded by a sound field microphone.
[00048] Processing is done by frames. Therefore, continuous audio signals are divided into frames, which are scaled by a windowing function to avoid discontinuities at the edges of the frame. The windowed signal frames are subjected to a Fourier transform in a Fourier transform block 40, dividing the microphone signals into N frequency bands. For simplicity, the processing of an arbitrary frequency band will be described in the following paragraphs, since the remaining frequency bands are processed in an equivalent manner. The Fourier transform block 40 produces coefficients that describe the resistance of the frequency components present in each of the W, X, Y and Z channels of B-shaped microphones within the analyzed window frame. These frequency parameters 42 are sent to the audio encoder 4 4 to obtain an audio channel and associated direction parameters. In the configuration shown in Fig. 3, the transmitted audio channel is chosen as the omnidirectional channel 46 having information about the signal from all directions. Based on the coefficients 42 of the omnidirectional and directional portions of the B-shaped microphone channels, a directional and diffusibility analysis is performed by a direction analysis block 48.
[00049] The direction of origin of the sound of the analyzed portion of the audio channel 46 is transmitted to an audio decoder 50 for the reconstruction of the audio signal together with the omnidirectional channel 46. When the diffusibility parameters 52 are
Petition 870190102982, of 10/14/2019, p. 26/42
21/30 present, the signal path is divided into a non-diffuse path 54a and a diffuse path 54b. The non-diffuse path 54a is scaled according to the diffusibility parameter, so that when the diffusibility Ψ is high, most of the energy or amplitude will remain in the non-diffusive path. Otherwise, when the diffusibility is high, most of the energy will be diverted to the diffuse path 54b. In the diffuse path 54b, the signal is either correlated or diffused using either 56a or 56b propagators. Correlation can be done using conventionally known techniques, such as convolution with a white noise signal, in which the white noise signal may differ from frequency channel to frequency channel. As long as the delay preserves energy, the final output can be regenerated by simply adding the signals from the non-diffuse signal path 54a and the diffuse signal path 54b at the output, since the signals in the signal paths have already been scaled, as indicated by diffusibility parameter Ψ. The diffuse signal path 54b can be scaled, depending on the number of speakers, using an appropriate scaling rule. For example, signals on the diffuse path can be scaled by i / 4n, where N is the number of speakers.
[00050] When the reconstruction is done for a multi-channel installation, the direct signal path 54a and the diffuse signal path 54b are divided into a number of subpaths corresponding to the individual speaker signals (in the split positions 58a and 58b). For this, the division in positions 58a and 58b can be interpreted as equivalent to an upmixing of at least one audio channel for multiple playback channels.
Petition 870190102982, of 10/14/2019, p. 27/42
22/30 by the multi-speaker speaker system. Therefore, each of the multiple channels has a channel portion of the audio channel 46. The origin direction of the individual audio portions is reconstructed by the redirect block 60 which further increases or reduces the intensity or amplitude of the channel portions corresponding to the speakers used for playback. For this purpose, the redirection block 60 in general requires knowledge about the installation of speakers used for playback. Actual redistribution (redirection) and derivation of associated weighting factors can, for example, be implemented using vector-based amplitude panning techniques. By providing different geometric speaker installations to the redistribution block 60, arbitrary playback speaker configurations can be used to implement the concept of the invention, without loss of reproduction quality. After processing, multiple inverse Fourier transforms are made in the frequency domain signals per block of inverse Fourier transforms 62, in order to obtain a signal in the time domain, which can be reproduced by the individual speakers. Before playback, an overlapping and addition technique must be performed by the 6 4 sum units to concatenate the individual audio frames so that continuous time-domain signals are obtained, ready to be played by the speakers.
[00051] According to the configuration of the invention shown in Fig. 3, the processing of Dir-AC signals is altered so that a modifier of the audio portion 66 is introduced to modify the portion of the audio channel actually processed and
Petition 870190102982, of 10/14/2019, p. 28/42
23/30 that allows to increase the intensity of a portion of the audio channel having direction parameters indicating a direction of origin close to the desired direction. This is achieved by applying an additional weighting factor to the direct signal path, that is, if the processed frequency portion originates from the desired direction, the signal is emphasized by applying an additional gain to this specific signal portion. The gain can be applied before the split point 58a, since the effect should contribute equally to all portions of channels.
[00052] The application of the additional weighting factor can, in an alternative configuration, also be implemented within the redistribution block 60 which, in this case, applies redistribution gain factors increased or reduced by the additional weighting factor.
[00053] When using directional enhancement to reconstruct a multichannel signal, playback can, for example, be done in the style of a DirAC presentation, as shown in Fig. 3. The audio channel to be played is divided into bands of frequency equal to those used in directional analysis. These frequency bands are then divided into streams, a diffuse and a non-diffuse stream. The diffuse flow is reproduced, for example, applying the sound to each speaker after convolution with large bursts of noise of 30ms. The noise bursts are different for each speaker. The non-diffuse flow is applied in the direction from directional analysis which is clearly time dependent. To obtain directional perception in multichannel speaker systems, simple amplitude panning in pairs or triplets can be used. In addition, each frequency channel is
Petition 870190102982, of 10/14/2019, p. 29/42
24/30 multiplied by a gain factor or scale factor, which depends on the direction analyzed. In general terms, a function can be specified by defining a desired directional pattern for reproduction. For example, it may be that only one direction should be emphasized. However, arbitrary directional patterns are easily implementable with a configuration in Fig. 3.
[00054] In the following approach, another embodiment of the present invention is described in the form of a list of processing steps. The list is based on the assumption that the sound is recorded with a B-format microphone, and is then processed for listening with multichannel or monophonic speakers using a DirAC-style presentation or presentation of a supply of directional parameters, indicating the direction source of portions of the audio channel. Processing is as follows:
1. Divide the microphone signals into frequency bands and analyze the direction and, optionally, the diffusibility in each band, depending on the frequency. As an example, the direction can be parameterized by azimuth and elevation angle (azi, ele).
2. Specify an F function, which describes the desired directional pattern. The function can have an arbitrary format. It typically depends on the direction. In addition, it can also depend on diffusibility, if diffusibility information exists. The function can be different for different frequencies and can also be changed
Petition 870190102982, of 10/14/2019, p. 30/42
25/30 depending on the weather. In each frequency band, obtain a directional factor q of the F function for each instant of time, which is used for the subsequent weighting (scaling) of the audio signal.
3. Multiply the values of the audio sample by the q values of the directional factors corresponding to each time and frequency portion to form the output signal. This can be done in a representation in the time domain and / or in the frequency domain. In addition, this processing can, for example, be implemented as part of a DirAC presentation for any number of desired output channels.
[00055] As previously described, the result can be heard using a multichannel or monophonic speaker system.
[00056] Fig. 4 shows an illustration of how the equipment and methods of the invention can be used to greatly increase the perception of a participant within a teleconference scenario. On the recording side 100, four interlocutors 102a-102d are illustrated with different orientations in relation to the recording position 104, that is, an audio signal originating from the interlocutor 102c has a fixed origin direction in relation to the recording position 104 . Supposing that the audio signal recorded at the recording position 10 4 has a contribution from the caller 102c and some background noise that originates, for example, from a discussion between the speakers
Petition 870190102982, of 10/14/2019, p. 31/42
26/30
102a and 102b, a broadband signal recorded and transmitted to a listening location 110 will comprise both signal components.
[00057] As an example, an installation of interlocutors having six speakers 112a-112f is outlined, which surround the listener located at the position of listener 114. Therefore, in principle, the sound emanating from almost arbitrary positions around the listener 114 can be reproduced in the installation shown in Fig.
4. Conventional multichannel systems would reproduce sound using these six speakers 112a-112f to reconstruct the spatial perception experienced at recording position 104 during recording, as closely as possible. Therefore, when the sound is reproduced using conventional techniques, the contribution of speaker 102c as background of the participating speakers 102a and 102b would also be clearly audible, reducing the intelligibility of the signal of speaker 102c.
[00058] According to a configuration of the present invention, a direction selector can be used to select the desired direction of origin in relation to the recording position that is used for a reconstructed version of a reconstructed audio signal that must be reproduced through speakers 112a-112f. Therefore, listener 114 can select the desired direction 116, corresponding to the position of speaker 102c. Thus, the audio portion modifier can modify the audio channel portion to obtain the reconstructed portion of the reconstructed audio signal, so that the intensity of the portions of the audio channel that originate from a direction close to the selected direction is emphasized. 116. The listener can, at the receiving end, decide which direction of origin will be played. Having made this selection, only
Petition 870190102982, of 10/14/2019, p. 32/42
27/30 emphasized those portions of signals that originate from the direction of the speaker 102c and, thus, the participating interlocutors 102a and 102b will become less disturbing. In addition to emphasizing the signal of the selected direction, the direction can be reproduced by panning through amplitude, as indicated symbolically by the waveforms 120a and 120b. As the 102c callers would be located closer to the 112d speaker than to the 112c speaker, amplitude panning will lead to a reproduction of the signal emphasized by the 112c and 112d speakers, whereas the remaining speakers will be almost muted (eventually reproducing diffuse portions of signals). Amplitude panning will increase the level of speaker 112d in relation to speaker 112c, since speaker 102c is located closer to speaker 112d.
[00059] Fig. 5 illustrates a block diagram of a method configuration for enhancing the directional perception of an audio signal. In a first analysis step 150, at least one audio channel and associated direction parameters are obtained indicating an origin direction of a portion of the audio channel in relation to a recording position.
[00060] In a selection step 152, the desired direction of origin is selected in relation to the recording position for a reconstructed portion of the reconstructed audio signal, the reconstructed portion corresponding to a portion of the audio channel.
[00061] In a 154 modification step, the audio channel portion is modified to obtain the reconstructed portion of the reconstructed audio signal, where the modification comprises increasing the intensity of a portion of the audio channel having
Petition 870190102982, of 10/14/2019, p. 33/42
28/30 direction parameters indicating a direction of origin close to the desired direction of origin in relation to another portion of the audio channel, having direction parameters indicating a direction of origin furthest from the desired direction of origin.
[00062] Fig. 6 illustrates an audio decoder configuration for the reconstruction of an audio signal having at least one audio channel 160 and associated direction parameters 162 indicating an origin direction of a portion of the audio channel. in relation to a recording position.
[00063] The audio decoder 158 comprises a direction selector 164 for selecting the desired origin direction in relation to the recording position of a reconstructed portion of the reconstructed audio signal, the reconstructed portion corresponding to a portion of the audio channel. The decoder 158 further comprises an audio portion modifier 166 to modify the audio channel portion in obtaining the reconstructed portion of the reconstructed audio signal, where the modification comprises increasing the intensity of a portion of the audio channel having direction parameters indicating an origin direction close to the desired origin direction in relation to another portion of the audio channel, having direction parameters indicating a direction of origin furthest from the desired direction of origin.
[00064] As indicated in Fig. 6, a single reconstructed portion 168 can be obtained or multiple reconstructed portions 170 can be obtained simultaneously, when the decoder is used in a multi-channel reproduction facility. The configuration of a system to enhance a directional perception of an audio signal 180, as shown in
Petition 870190102982, of 10/14/2019, p. 34/42
29/30
Fig. 7 is based on the decoder 158 of Fig. 6. Therefore, in the following, only the elements additionally introduced will be described. The system for enhancing a directional perception of an audio signal 180 receives an audio signal 182 as input, which may be a monophonic signal or a multichannel signal recorded by multiple microphones. An audio encoder 184 obtains an audio signal having at least one audio channel 160 and associated direction parameters 162 indicating an origin direction of a portion of the audio channel in relation to the recording position. The at least one audio channel and the associated direction parameters are further processed as already described for the audio decoder.
<td>Fig.</td><td>6, to</td><td>get a signal</td><td>in</td><td>highlighted output</td>
<td>perceptually</td><td> 170 .</td><td></td><td></td><td></td>
<td> [00065]</td><td>although</td><td>of the invention</td><td>Tue</td><td>been described</td>
<td>mainly</td><td>in the field</td><td>playback</td><td colspan="2">multichannel audio,</td>
different fields of application can have benefits with the methods and equipment of the invention. As an example, the concept of the invention can be used to target (by enlarging or attenuating) specific individuals speaking in a teleconference setting. It can also be used to reject (or amplify) ambient components, as well as for reverberation or reverb enhancement. Other possible application scenarios include noise cancellation of ambient noise signals. Another possible use could be the directional enhancement of signals with hearing aids.
[00066] Depending on certain implementation requirements of the methods of the invention, the methods of the invention can be implemented in hardware or in software. Implementation can
Petition 870190102982, of 10/14/2019, p. 35/42
30/30 be made using a digital storage medium, in particular a disc, DVD or CD having stored control signals with
<td colspan="2">electronic reading, which cooperates with</td><td>one</td><td>system</td><td>in</td><td colspan="2">computer</td>
<td>programmable, for</td><td>that are carried out</td><td>the</td><td>methods</td><td>of</td><td>invention.</td><td>In</td>
<td>general, this</td><td>invention is therefore</td><td>one</td><td>product</td><td>in</td><td>program</td><td>in</td>
computer with a program code stored in a machine-readable vehicle, the program code operating to carry out the methods of the invention when the computer program product operates on a computer. In other words, the methods of the invention are, therefore, a computer program having a program code for carrying out at least one of the methods of the invention when the computer program product operates on a computer.
[00067] Although the above has been shown and described particularly with reference to its particular configurations, it will be understood by the technicians in the subject that several other changes of form and details can be made without abandoning its spirit and scope. It should be understood that several changes can be made to adapt different configurations without abandoning the broader concepts revealed in the present and encompassed by the following claims.
11 priority claims, no other members on record
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 60896184 | United States of America | – | |
| 89618407 | United States of America | P | |
| 11742488 | United States of America | – | |
| 74248807 | United States of America | A | |
| 2008000829 | European Patent Office (EPO) | W | |
| 11742488 | – | – | – |
| 60896184 | – | – | – |
| PCTEP2008000829 | – | – | – |
| US20070742488 | – | – | – |
| US20070896184P | – | – | – |
| WO2008EP00829 | – | – | – |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent or certificate of addition of invention grantedGrantedB16A | B16A | |
| Decision: intention to grantB09A | B09A | |
| Others concerning applications: alteration of classificationB15K | B15K | |
| Notification to applicant to reply to the report for non-patentability or inadequacy of the application according art. 36 industrial patent lawB06A | B06A | |
| Objections, documents and/or translations needed after an examination request according art. 34 industrial property lawB06F | B06F |
Numbers
- Publication
- PI0808225
- Publication, DOCDB
- PI0808225
- Publication, EPODOC
- BRPI0808225
- Application
- 8225
- Application, DOCDB
- PI0808225
- Application, EPODOC
- BR2008PI08225
Titles2
- Portuguese
- MÉTODO E DECODIFICADOR DE AUDIO PARA A RECONSTRUÇÃO DE UM SINAL DE ÁUDIO, MÉTODO E CODIFICADOR DE AUDIO PARA O REALCE DA PERCEPÇÃO DIRECIONAL DE UM SINAL DE ÁUDIO E SISTEMA PARA O REALCE DE UM SINAL RECONSTRUÍDO DE ÁUDIO
- English
- AUDIO METHOD AND DECODER FOR THE RECONSTRUCTION OF AN AUDIO SIGNAL, AUDIO METHOD AND ENCODER FOR ENHANCING THE DIRECTIONAL PERCEPTION OF AN AUDIO SIGNAL AND SYSTEM FOR ENHANCING A RECONSTRUCTED AUDIO SIGNAL
Classification
- CPC, 5
- H04S7/302
- H04S2400/11
- H04S2400/13
- H04S2400/15
- H04S2420/11