A spatial audio processing method, a program product, an electronic device and a system
Abstract
A method comprising the following steps: - the reception of a first audio signal (S1); and - the generation of a digital representation (S1 ') of the first audio signal (S1) by applying a head-related transfer function (HRTF) in a first sound reproduction position (r1); characterized in that: the method further comprises the step of changing the first sound reproduction position (r1) to a second sound reproduction position (r3) in response to the reception of a second audio signal (S2) or a precursor signal for a second audio signal (S2).
Term
Term ended
Projected expiry passed 27 June 2025, 1.2 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
13 claims: 8 independent, 5 dependent
- 1ES 2 584 869 T3 REIVINDICACIONES 1. Un método que comprende los siguientes pasos:- la recepción de una primera señal de audio (S1);y - la generación de una representación digital (S1') de la primera señal de audio (S1) mediante la aplicación de una función de transferencia relacionada con la cabeza (HRTF) en una primera posición de reproducción de sonido (r1);que se caracteriza porque: el método comprende además el paso de cambiar la primera posición de reproducción de sonido (r1) a una segunda posición de reproducción de sonido (r3) como respuesta a la recepción de una segunda señal de audio (S2) o una señal precursora para una segunda señal de audio (S2).
- 2Un método, de conformidad con la reivindicación 1, en el que:- la segunda señal de audio (S2) es una señal de aviso (Sr) o una señal de voz (Sp) recibidas;o - la señal precursora para una segunda señal de audio (S2) es un mensaje para el establecimiento de una llamada telefónica o un mensaje activado por una llamada telefónica que va a ser establecida.
- 3Un método, de conformidad con las reivindicaciones 1 o 2, que además comprende el paso de:definir la primera posición de reproducción de sonido (r1) y la segunda posición de reproducción de sonido (r3) para la primera señal de audio (S1).
- 4Un método, de conformidad con la reivindicación 3, que además comprende el paso de:visualizar las mencionadas posiciones de reproducción de sonido.
- 5Un método, de conformidad con cualquiera de las reivindicaciones anteriores, que además comprende el paso de:generar una representación digital (S2') de la segunda señal de audio (S2) mediante la aplicación de una función de transferencia relacionada con la cabeza (HRTF) en una tercera posición de reproducción de sonido (r2);y en el que la mencionada tercera posición de reproducción de sonido (r2) está más cerca del centro de la cabeza del usuario que la mencionada segunda posición de reproducción de sonido, es decir | r2 | < | r3 |.
- 6Un producto de programa (51) que comprende:- medios para recibir una primera señal de audio (S1);y - medios para generar una representación digital (S1') de la primera señal de audio (S1) mediante la aplicación de una función de transferencia relacionada con la cabeza (HRTF) en una primera posición de reproducción de sonido (r1);que se caracteriza porque: el producto de programa (51) comprende además medios para cambiar la primera posición de reproducción de sonido (r1) a una segunda posición de reproducción de sonido (r3) como respuesta a la recepción de una segunda señal de audio (S2) o una señal precursora para una segunda señal de audio (S2).
- 7Un producto de programa (51), de conformidad con la reivindicación 6, en el que:- la segunda señal de audio (S2) es una señal de aviso (Sr) o una señal de voz (Sp) recibidas desde una red de comunicaciones (39);o - la señal precursora para una segunda señal de audio (S2) es un mensaje para el establecimiento de una llamada telefónica en el dispositivo electrónico (30) o un mensaje activado por una llamada telefónica que va a ser establecida.
- 8Un producto de programa (51), de conformidad con las reivindicaciones 6 o 7, que además comprende:medios para definir la primera posición de reproducción de sonido (r1) y la segunda posición de reproducción de sonido (r3) para la primera señal (S1).
- 9Un producto de programa (51), de conformidad con la reivindicación 7, que además comprende:medios para visualizar las mencionadas posiciones de reproducción de sonido.
- 10Un producto de programa (51), de conformidad con cualquiera de las reivindicaciones anteriores comprendidas entre la 6 y la 9, que además comprende:medios para generar una representación digital (S2') de la segunda señal de audio (S2) mediante la aplicación de una función de transferencia relacionada con la cabeza (HRTF) en una tercera posición de reproducción de sonido (r2);y en el que: la mencionada tercera posición de reproducción de sonido (r2) está más cerca del centro de la cabeza del usuario que la mencionada segunda posición de reproducción de sonido, es decir | r2 | < | r3 |.
- 11Un dispositivo electrónico (30) que se caracteriza porque:el dispositivo electrónico (30): - está adaptado para: llevar a cabo un método, de conformidad con cualquiera de las reivindicaciones ES 2 584 869 T3 comprendidas entre la 1 y la 5;o - comprende: un producto de programa (51), de conformidad con cualquiera de las reivindicaciones comprendidas entre la 6 y la 10.
- 12Un sistema que comprende:un dispositivo electrónico (30), de conformidad con la reivindicación 11;y auriculares que poseen al menos dos canales (100).
- 13Un método, un producto de programa (51), un dispositivo electrónico (30) o un sistema, de conformidad con cualquiera de las reivindicaciones anteriores, en el que:- la primera señal de audio (S1) comprende una señal para un canal izquierdo y un canal derecho;- la representación digital (S1') comprende una señal de audio para el canal izquierdo y el canal derecho, el canal izquierdo de la misma comprende una combinación de canales izquierdos obtenida mediante la aplicación de la función de transferencia relacionada con la cabeza (HRTF) en la primera o la segunda posición de reproducción de sonido (r1 o r3) a los canales izquierdo y derecho de la primera señal de audio (S1), y el canal derecho de la misma comprende una combinación de canales derechos obtenida mediante la aplicación de la función de transferencia relacionada con la cabeza (HRTF) en la misma posición de reproducción de sonido (r1 o r3) a los canales izquierdo y derecho de la primera señal de audio (S1).
Independent claims13
67 paragraphs in 9 sections, as filed
ES 2 584 869 T3
DESCRIPTION
A spatial audio processing method, a program product, an electronic device, and a system
BACKGROUND OF THE TECHNIQUE
Advances in computer science and acoustic field theory have revealed interesting possibilities in sound technology. As a practical example of new technologies, a relatively new tool on the market is a software product that can be used to create an impression of the position of an audio signal source when a user hears a representation of the audio signal through headphones that have at least two channels.
In practice, when such a tool is run on a processor in the form of a software product, the audio signal passes through a Head-Related Transfer Function (HRTF) in order to generate, for a user wearing headphones with at least two channels (for example, in stereo), the psychoacoustic impression that the audio signal comes from a predefined position.
The mechanism by which this psychoacoustic impression is created can be illustrated by means of an example. As we know from observing everyday life, a person can observe the position “r” (bold here denotes a vector that can be expressed with “r”, “Φ” and “θ” in spherical coordinates) of a source sound with fairly good precision. Therefore, if the sound is emitted by a sound source located near the left ear (r = 30 cm, Φ = 3π / 2, θ = 0), it is received first by the left ear and only a fraction of a second later. through the right ear. Now, if an audio signal is played through headphones first in the left ear and a fraction of a second later in the right ear through those headphones, this can be done by filtering the signal through a respective transfer function related to the head, the listener has the impression that the sound source is located near the left ear.
A more detailed discussion of the different properties of HRTF and how it can be obtained can be found, for example, in published US Patent Application No. 2004/0136538 A1, and the references mentioned therein.
SUMMARY OF THE INVENTION
The human capacity to receive information through the ear is quite limited. Especially the ability to follow a sound source can be greatly affected when another sound source is present. Accordingly, an object of the invention is to create a method, a program product, an electronic device and a system by which the perception of an audio signal from a first sound source can be improved when a signal is received from audio from another sound source simultaneously with the signal from the first source. This objective can be achieved in accordance with any of the independent patent claims.
The dependent patent claims describe various advantageous embodiments of the invention.
ADVANTAGES OF THE INVENTION
If the first position in which a head-related transfer function is applied to a first audio signal is shifted to a second sound reproduction position in response to the reception of a second audio signal or a precursor signal for a second audio signal, the user can be in a better position to better distinguish between the first and second signals.
Also, the transfer of the first audio signal from the first sound reproduction position to the second sound reproduction position can be automated.
By switching in response to the reception of a precursor signal, the transfer can be carried out before the second audio signal begins to play, which improves user comfort as the position of the first signal can be transferred before starting to play the second audio signal.
If the second audio signal is a warning signal or a voice signal, it may be easier for the user to focus on the second audio signal while listening to the first audio signal. For example, if a phone call is played as the second audio signal, the user can continue to listen to the first audio signal, such as radio or music from an MP3 or CD, while continuing a phone conversation at the same time. .
On the other hand, transfer from the second sound reproducing position to the first sound reproducing position can be performed in response to the fact that the second audio signal is no longer received. For example, after hanging up on a phone call, the first sound playback position can be used automatically.
ES 2 584 869 T3
If the precursor signal is a message for the establishment of a telephone call or a message triggered by a telephone call to be established, the convenience of the user when receiving the telephone call can be improved. The beginning of a phone call is generally of the utmost importance, as the caller and / or called person usually proceed to identify themselves.
Consequently, the user might find it disturbing that the first audio signal is transferred only when a call has already been established. This way, he or she can have some time to prepare for an upcoming phone call.
If the second sound reproduction position is further away than the first sound reproduction position, the user's ability to differentiate the signals can be improved.
Furthermore, if a head-related transfer function, preferably the same head-related transfer function as that corresponding to the first audio signal, is applied to the second audio signal in a third sound reproduction position, being the third sound reproduction position closer to the user's head than the second sound reproduction position, the user's concentration on the second audio signal may not be affected as much by the disturbance caused by the first audio signal.
LIST OF DRAWINGS
The invention is described in greater detail below with reference to the examples shown in the attached drawings in Figures 1 to 5B, with respect to which:
An example of a sound source location in head coordinates is illustrated in Figure 1A;
A user wearing headphones is illustrated in Figure 1B;
Figure 2 illustrates how the sound reproduction position can be changed;
In Figure 3 some functional blocks of an electronic device are shown;
Figure 4 is a flow chart illustrating signal processing in the example of Figure 2;
Signal processing in the case of a signal source is illustrated in Figure 5A; Y
Signal processing for two signal sources is illustrated in Figure 5B.
The same reference symbols are used to refer to similar features throughout the drawings.
DETAILED DESCRIPTION
Certain current development activities of the applicant are aimed at creating an electronic device that can be used by a user wearing headphones possessing at least two channels (eg in stereo). The electronic device is adapted to transmit a signal of at least two channels (eg a stereo signal) to the headphones, preferably via a wireless link.
An example of head coordinates in a plane is shown in Figure 1A. A sound source (13) is located at point "r" (at a distance "r" and at an angle "Φ"), as seen from the center of the person's head (11). The acoustic conditions of the room are denoted by “e”, mainly as a consequence of echo and background noise.
Figure 1B illustrates the head (11) of a user of an electronic device (30) wearing headphones with at least two channels (for example, in stereo) (100) that are adapted to receive a representation S '' ' of an audio signal S from the electronic device (30) through its receiving means (101). The headphones (100) comprise at least two acoustic transducers (for example, loudspeakers) (104 and 105), one for the right ear (14) and one for the left ear (15). The headphones (100) are adapted to reproduce the sound of the received representation S '' 'for at least two channels (that is, at least the left and the right). The electronic device (30) is described in more detail below with reference to Figure 3.
As is known in the prior art, by properly selecting a head-related transfer function (HRTF) that causes adequate phase differences and attenuation, possibly in a frequency-dependent manner, and applying it to an audio signal S in the processing unit (34) for at least two channels (at least the left and the right), thus generating a digital representation S ', which can then be controlled in the electronic device (30) and, finally, passed to the headphones (100) as representation S '' ', a user, when listening to its reproduction, has the impression that the sound source (13 ) this
ES 2 584 869 T3 located at a defined position (sound reproduction position "r"). The easiest way to express the sound reproduction position "r" is as a point in polar or spherical coordinates, but it can also be expressed in any other coordinate system.
The location of the sound source (13), as shown in Figure 1A, can be chosen almost deliberately in the electronic device (30), for example in its processing unit (34), by selecting a playback position “r” sound that is used by the HRTF to modify its filtering characteristics. Alternatively, separate HRTFs (one for each "r" sound reproduction position) can be used, and then proceed to change the HRTF to be used when the "r" sound reproduction position changes.
On the one hand, a HRTF as described in application No. 2004/0136538 A1 can be used in order to carry out the present invention if high quality 3D printing is desired. If this approach were taken, the HRTF could be stored in the electronic device (30). Since an electronic device can have several users (eg, members of a family), the electronic device (30) can therefore comprise a greater number of HRTFs, one for each user. The selection of the HRTF to be used can be selected, for example based on a code that the user enters into the electronic device (30). Alternatively, if users prefer to use their personal headphones, the selection can be based on an identifier that identifies the headphones (100).
On the other hand, a simpler method of defining the HRTF will be appropriate, especially if the reproduction of the 2D sound image is sufficient. This operation is becoming increasingly simple, as appropriate software modules are already available on the market.
A general HRTF can also be used for all users. An especially suitable HRTF of this type is one that has been recorded using a head and torso simulator. The HRTF is then preferably stored for a large selection of angles around the head. In order to obtain 2 degree resolution, 180 HRTF positions must be stored. In order to obtain 5 degree resolution, 72 HRTF positions must be stored for 2D playback of the sound source. To control the distance, additional HRTF positions are preferably needed.
In the case of "sound source 2D reproduction", the position of the sound source (13) would be located approximately at one level, preferably at the level of the user's ears. In the case of "sound source 3D reproduction", the sound source (13) may also be located below or above this level.
Figure 2 illustrates how the position of sound reproduction can be changed - that is, the position from which the user who listens to a representation reproduction Yes "observes the sound source (13) that is being located - of a signal audio signal S1 from the first sound reproduction position π to a second sound reproduction position re, in accordance with one aspect of the invention.
First, an audio signal S1 from a sound source (13) is received or reproduced by the electronic device (30). Next, the audio signal S1 is controlled by the electronic device (30) by applying an HRTF with a first sound reproduction position n. The controlled signal, after being converted to an analog signal and after amplification, gives the impression that the sound source (13) is located in the π position when listening through headphones having at least two channels (100).
In response to the reception of a second audio signal S2 from a second sound source (13B), or a precursor signal for a second audio signal S2, the first sound reproduction position π of the HRTF is replaced by a second sound reproduction position r3, so that the representation S1 '' 'of the audio signal S1 provides the impression, when listening through headphones having at least two channels (100), that the sound source (13) is located in the re position.
Also, the HRTF can be applied to the second audio signal S2 with a third sound reproduction position r2. Next, the representation S2 '' 'of the audio signal S2 provides the impression, when listening through headphones having at least two channels, that the second sound source (13B) is located at position r2 .
The transition from position r1 to position r3 can be made without complications, that is, step by step. This gives the impression that the sound source (13) is moving.
In Figure 3 some functional blocks of the electronic device (30) are shown.
The electronic device (30) preferably comprises means (35) for receiving and transmitting data to / from a communication network (39), especially a radio receiver and a radio transmitter. Data transmission between the electronic device (30) and the communication network (39) can occur through a wireless interface or an electrical interface. An example of the first is the air interface of a cellular communications network, especially
ES 2 584 869 T3 is a GSM network, and an example of the second is the traditional interface between a telephone device and a Public Switched Telephone Network (PSTN).
The electronic device (30) also comprises input / output means (32) for the operation of the electronic device (30). The input / output means (32) may comprise a keyboard and / or a joystick which is preferably suitable for dialing a number or selecting a destination address or name from a telephone directory stored in memory (36) ; the keyboard preferably further comprises a dial switch and an answer button. The input / output means (32) may also comprise a screen.
An electronic device (30), according to the invention, comprises means (31) for passing an S '' 'representation of an S audio signal to the headphones (100). The means (31) may comprise a wireless transmitter.
The electronic device (30) further comprises a processing unit (34), such as a microprocessor, and a memory (36). The processing unit (34) is adapted to read software as executable code and then execute it. The software is normally stored in memory (36). The memory (36) also houses the HRTF, making it accessible to the processing unit (34).
The electronic device (30) may further comprise one or more sound sources (13 and 13B). The sound sources (13 and 13B) can be FM or digital radio receivers, or music players (in particular, MP3 or CD players). The sound sources (13 and 13B) can also be located externally to the electronic device (30), which means that a corresponding audio signal is received through means (35) to receive data from a communications network (39 ), especially through a radio receiver, a generic receiver (for example, Bluetooth) or a dedicated receiver. Next, the audio signal received from an external sound source (13 and 13B) is controlled similarly to an audio signal received from an internal sound source. Therefore, the audio signal S can be any audio signal generated in the electronic device (30), reproduced from a music file (especially an MP3 file), received from the communication network (39) or from FM radio. or digital. The representation S '' 'can be passed to the headphones (100) through the use of a wireless connection, such as Bluetooth, or through a cable.
Other components (37) may exist between the processing unit (34) and the means (31) for passing an S '' 'representation of an S audio signal to the headphones (100). They are, to some extent, necessary to change a digital representation S 'from the processing unit (34) to a signal S' 'appropriate for the media (31) to pass a representation S' '' of an audio signal. S to headphones (100). These components (37) may comprise a digital-to-analog converter, an amplifier, and filters. However, a more detailed description of them is omitted here, as it is irrelevant to understanding the nature of this invention, and because these components are well known in the prior art.
Figure 4 is a flow chart illustrating signal processing in the example of Figure 2. The flow chart is explained in conjunction with Figures 5A and 5B, which illustrate signal processing in the case of one and two signal sources, respectively.
The processing unit (34) executes an audio program module (51) stored in memory (36). Initially, the audio program module (51) can be installed in the electronic device (30) by using input / output means (32), an interchangeable memory medium such as a USB memory device, or You can download from a communications network (39) or from a remote device. Before installation, the audio program module 51 preferably takes the form of a program product that can be sold to customers.
The audio program module (51) comprises the HRTF, which can be defined by the user, so that each user can have their own HRTF in order to improve the acoustic quality. However, for basic range purposes, a simple HRTF will suffice.
The audio program module (51) is started in step 401, as soon as the sound source (13) producing the audio signal S1 is activated. Typically, the audio signal S1 is controlled by the audio program module (51) using a first sound reproduction position r1 selected in step 403. If the second sound source (13B) is inactive, ie there is no other active sound (13B) present (which is detected in step 405), the audio signal S1 in step 407 passes through the HRTF. The audio program module (51) generates a digital representation S1 'by applying the HRTF with the first sound reproduction position π to the audio signal S1. This operation is repeated until the sound source (13) is inactive.
The audio signal S1 can comprise a signal for more than one channel. For example, if the audio signal S1 is a stereo signal (for example, from an MP3 player as the signal source (13)), it will already comprise a signal for two channels (left and right). HRTF can be applied with the first sound reproduction position π to the left and right channel separately. The resulting four digital representations in total can then be combined in order to have only one signal for the left and right channels.
ES 2 584 869 T3
More than two sound sources are supported, for example a stereo MP3 signal (as sound source (13)) comprises two sound sources, the audio signals of which need to be placed in different positions. The other sound source (13B) could preferably be an audio signal of an incoming call or an audio signal (eg, a ringtone) generated to alert the user.
If in step 405 it is detected that a second sound source (13B) is active, in step 421 the sound reproduction position r3 is selected for the sound source (13) and the sound reproduction position r2 is selected for the other sound source (13B). Next, in step 423 a digital representation S 'is generated by applying the HRTF with the second sound reproduction position r3 to the audio signal S1, and optionally by applying the HRTF with the third reproduction position sound r2 to the second audio signal S2. This operation is repeated until one of the sound sources (13 or 13B) becomes inactive or the audio program module (51) stops receiving a corresponding audio signal (S1 and S2) (tested in steps 427 and 425 , respectively).
If the sound source (13) becomes inactive or the audio signal S1 is not received at the audio program module (51), in step 429 the audio signal S1, possibly received by the audio program module ( 51), is ignored in step 429.
If the sound source 13B becomes inactive or the audio signal S2 is not received at the audio program module 51, step 425 returns execution control to step 403.
Accordingly, the audio program module (51) can generate in step 423, when executed in the processing unit (34), a digital representation signal S2 'of the second audio signal S2 for at least two channels of sound (LEFT and RIGHT) by applying the HRTF in a third position of sound reproduction r2. The digital representation signal S2 'is adapted to provide the impression, after being converted from digital to analog, amplified and filtered, when listening through headphones having at least two channels (100), that the second signal of S2 audio comes from the third sound playback position r2.
The HRTF is applied in the processing unit (34), preferably separately, for the audio signals S1 and S2, both with different sound reproduction positions (ie r3 and r2). Then the digital representations S1 'and S2' can be combined into a combined digital representation S '= S1' + S2 '. Since both digital representations S1 'and S2' comprise information for at least two channels (left and right), it may also be advantageous to perform channel synchronization when digital representations S1 'and S2' are combined.
In other words, if a sound source 13 is adapted to provide a stereo signal such as the S1 audio signal, each channel of the S1 audio signal passes separately through the HRTF, with the playback position of r3 (or r3) sound. The four resulting signals are then added (two by two) in order to generate the digital representation S1 '. The same applies if the other sound source (13B) is adapted to provide a stereo signal like the audio signal S2, but now with r2 as the sound reproduction position (r2).
If the third sound reproduction position r2 is closer to the middle of the user's head than the second sound reproduction position r3, ie | r2 | <| r3 |, the user can be in a better position to follow the second sound source (13B), that is, the disturbance caused by the sound source (13) can be reduced.
The second audio signal (S2) can be a warning signal or a voice signal received from the communication network (39).
The precursor signal for a second audio signal (S2) may be a message from the communication network (39) for the establishment of a telephone call or a message triggered by a telephone call to be established.
The user can preferably define, using the input means (32), the first sound reproduction position π and / or the second sound reproduction position r3 for the first audio signal (S1). By using output means (32), the aforementioned sound reproduction positions can be displayed, for example on the screen of the electronic device. This should make it easier to define the addresses.
Although the invention has been described above with reference to the examples shown in the accompanying drawings, it is clear that the invention is not limited thereto, but can be modified by those skilled in the art without leaving the scope of the invention.
For example, in addition to the sound reproduction positions n, r2 and r3, a parameter, sometimes referred to as the "room parameter", can also be defined, which is fed to the audio program module (51). The room parameter describes the effect of the “surrounding room”, for example, the possible echo that can bounce off the walls of an artificial room. The room parameter and consequently the effect of the surrounding room can be changed at the same time when
ES 2 584 869 T3 change the sound reproduction position from ri to re. In this way, the user can hear, for example, a change from a smaller room to a larger room, or the opposite. For example, if | r3 | is greater than | π |, such that π would be close to or beyond the wall of the “surrounding room”, it may be appropriate to increase the size of the room.
Contents9
9 members in 6 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 04026708 | European Patent Office (EPO) | A | |
| 04026708 | European Patent Office (EPO) | – | |
| 2005052997 | European Patent Office (EPO) | W |
Members9
| Document | Office | Kind | |
|---|---|---|---|
| EP1657961A1 | European Patent Office (EPO) | A1 | |
| WO2006051001A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200629962A | Taiwan Province of China | A | |
| US2007291967A1 | United States of America | A1 | |
| EP1902597A1 | European Patent Office (EPO) | A1 | |
| US8488820B2 | United States of America | B2 | |
| EP1902597B1 | European Patent Office (EPO) | B1 | |
| ES2584869T3This record | Spain | T3 | |
| HUE029900T2 | Hungary | T2 |
Numbers
- Publication
- 2584869
- Application
- 5760883
Titles2
- Spanish
- Un método de procesamiento de audio espacial, un producto de programa, un dispositivo electrónico y un sistema
- English
- A spatial audio processing method, a program product, an electronic device and a system
Classification
- CPC, 5
- H04S1/002
- H04S1/005
- H04S3/002
- H04S2400/11
- H04S2420/01
- IPC, 2
- H04S3 00
- H04S1 00