Method for generating an overall room noise for passing to a real endpoint, use of said method and teleconferencing system
Abstract
The invention relates to a method for generating a to a real endpoint (EA, EB, EC) to be transmitted overall sound (GA, GB, GC), where the real endpoint (EA, EB, EC) an audio signal (AA, AB, AC), with the method steps: a. Set a default recipient position (PS) within a virtual space (V2, V3), b. Assigning the real endpoint (EA, EB, EC) with associated audio signal (AA, AB, AC) to a defined position (PA, PB, PC) within the virtual space (V2, V3c. for at least one real endpoint (EA, EB, EC) Generating a surround sound (RA, RB, RC) based on the assigned audio signal (AA, AB, AC), the assigned defined position (PA, PB, PC) and the standard receiver position (PS), d. Generating at least one overall sound (GA, GB, GC) for at least one real endpoint (EA, EB, EC) from at least a part of previously generated spatial sounds (RA, RB, RC) with the exception of the surround sound (RA, RB, RC) of the at least one real endpoint (EA, EB, EC) and e. Transmitting the at least one overall sound (GA, GB, GC) to the at least one real endpoint (EA, EB, EC).

Term
9.8 yearsto projected expiry
Projected expiry 8 July 2036, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
11 claims: 11 independent, 0 dependent
- 1Method for generating a to a real endpoint (EA, EB, EC) to be transmitted overall sound (GA, GB, GC), where the real endpoint (EA, EB, EC) an audio signal (AA, AB, AC), with the method steps:a. Set a default recipient position (PS) within a virtual space (V2, V3) b. Assigning the real endpoint (EA, EB, EC) with associated audio signal (AA, AB, AC) to a defined position (PA, PB, PC) within the virtual space (V2, V3) c. for at least one real endpoint (EA, EB, EC) Generating a surround sound (RA, RB, RC) based on the assigned audio signal (AA, AB, AC), the assigned defined position (PA, PB, PC) and the standard receiver position (PS) d. Generating at least one overall sound (GA, GB, GC) for at least one real endpoint (EA, EB, EC) from at least a part of previously generated spatial sounds (RA, RB, RC) with the exception of the surround sound (RA, RB, RC) of the at least one real endpoint (EA, EB, EC) and e. Transmitting the at least one overall sound (GA, GB, GC) to the at least one real endpoint (EA, EB, EC). Verfahren zur Erzeugung eines an einen realen Endpunkt (EA, EB, EC) zu übermittelnden Gesamtraumklangs (GA, GB, GC), wobei dem realen Endpunkt (EA, EB, EC) ein Audiosignal (AA, AB, AC) zugeordnet ist, mit den Verfahrensschritten: a. Festlegen einer Standard-Empfängerposition (PS) innerhalb eines virtuellen Raums (V2, V3), b. Zuordnen des realen Endpunktes (EA, EB, EC) mit zugehörigem Audiosignal (AA, AB, AC) zu einer definierten Position (PA, PB, PC) innerhalb des virtuellen Raums (V2, V3), c. für zumindest einen realen Endpunkt (EA, EB, EC) Erzeugen eines Raumklangs (RA, RB, RC) auf Basis des zugeordneten Audiosignals (AA, AB, AC), der zugeordneten definierten Position (PA, PB, PC) und der Standard-Empfängerposition (PS), d. Erzeugen zumindest eines Gesamtraumklangs (GA, GB, GC) für zumindest einen realen Endpunkt (EA, EB, EC) aus zumindest einem Teil zuvor erzeugter Raumklänge (RA, RB, RC) mit Ausnahme des Raumklangs (RA, RB, RC) des zumindest einen realen Endpunkts (EA, EB, EC) und e. Übermitteln des zumindest einen Gesamtraumklangs (GA, GB, GC) an den zumindest einen realen Endpunkt (EA, EB, EC).
- 2Method according to one of the preceding claims, characterized, that audio signals (AA, AB, AC) can be preselected overall and / or in packages. Verfahren nach einem der vorhergehenden Ansprüche, dadurch gekennzeichnet, dass Audiosignale (AA, AB, AC) gesamthaft und/oder paketweise vorselektiert werden.
- 3Method according to one of the preceding claims, characterized, that a surround sound (RA, RB, RC) is generated by means of an interaural transit time difference and / or an interaural level difference. Verfahren nach einem der vorhergehenden Ansprüche, dadurch gekennzeichnet, dass ein Raumklang (RA, RB, RC) mittels einer interauralen Laufzeitdifferenz und/oder einer interauralen Pegeldifferenz erzeugt wird.
- 4Method according to one of the preceding claims, characterized, that an audio signal (AA, AB, AC) and / or a surround sound (RA, RB, RC) and / or a total sound (GA, GB, GC) are filtered by means of a head-related transfer function. Verfahren nach einem der vorhergehenden Ansprüche, dadurch gekennzeichnet, dass ein Audiosignal (AA, AB, AC) und/oder ein Raumklang (RA, RB, RC) und/oder ein Gesamtraumklang (GA, GB, GC) mittels einer Head-Related Transfer Function gefiltert werden.
- 5Method according to one of the preceding claims, characterized, that an audio signal (AA, AB, AC), a surround sound (RA, RB, RC) and / or a total-space sound (GA, GB, GC) an echo and / or reverberation effect, preferably as a function of at least one defined position (PA, PB, PC), to be added. Verfahren nach einem der vorhergehenden Ansprüche, dadurch gekennzeichnet, dass einem Audiosignal (AA, AB, AC), einem Raumklang (RA, RB, RC) und/oder einem Gesamtraumklang (GA, GB, GC) ein Echo- und/oder ein Nachhall-Effekt, vorzugsweise in Abhängigkeit wenigstens einer definierten Position (PA, PB, PC), hinzugefügt werden.
- 6Method according to one of the preceding claims, characterized, that the echo and / or the reverberation effect, for example, for imitation of different wall or ceiling materials of the virtual space (V2, V3). Verfahren nach einem der vorhergehenden Ansprüche, dadurch gekennzeichnet, dass der Echo- und/oder der Nachhall-Effekt, beispielsweise zur Imitation unterschiedlicher Wand- oder Deckenmaterialien des virtuellen Raums (V2, V3), eingestellt werden.
- 7Use of the method according to one of the preceding claims in a multimedia system for at least two users, in particular a teleconferencing system (100) or an Internet-based multi-player system. Verwendung des Verfahrens nach einem der vorhergehenden Ansprüche in einem Multimediasystem für wenigstens zwei Nutzer, insbesondere einem Telekonferenzsystem (100) oder einem internetbasierten Viel-Spieler-System.
- 8Teleconferencing system (100) with at least three teleconferencing endpoints (A, B, C), each of the teleconferencing endpoints (A, B, C) comprising an audio playback unit (101) and an audio recording unit (102), wherein the teleconferencing system (100), a first and a second audio signal (AA, AB, AC), for example, a first and a second participant (TA, TB, TC), via the audio recording units (102) of a first and a second teleconferencing endpoint (A, B, C), characterized, that the teleconferencing system (100), a total sound (GA, GB, GC) according to a method according to one of the preceding claims 1 to 6 from the first and the second audio signal (AA, AB, AC) to a third teleconferencing endpoint (A, B, C), for example a third party (TA, TB, TC) and via its audio playback unit (101), each teleconferencing endpoint (A, B, C) corresponding to a real endpoint (EA, EB, EC) in the sense of the procedure. Telekonferenzsystem (100) mit wenigstens drei Telekonferenzendpunkten (A, B, C), wobei jeder der Telekonferenzendpunkte (A, B, C) eine Audio-Wiedergabeeinheit (101) und eine Audio-Aufnahmeeinheit (102) aufweist, wobei das Telekonferenzsystem (100) eingerichtet ist, ein erstes und ein zweites Audiosignal (AA, AB, AC), beispielsweise eines ersten und eines zweiten Teilnehmers (TA, TB, TC), über die Audio-Aufnahmeeinheiten (102) eines ersten und eines zweiten Telekonferenzendpunkts (A, B, C) aufzunehmen, dadurch gekennzeichnet, dass das Telekonferenzsystem (100) eingerichtet ist, einen Gesamtraumklang (GA, GB, GC) gemäß einem Verfahren nach einem der vorhergehenden Ansprüche 1 bis 6 aus dem ersten und dem zweiten Audiosignal (AA, AB, AC) zu erzeugen, an einen dritten Telekonferenzendpunkt (A, B, C), beispielsweise eines dritten Teilnehmers (TA, TB, TC), zu übermitteln und über dessen Audio-Wiedergabeeinheit (101) auszugeben, wobei jeder Telekonferenzendpunkt (A, B, C) einem realen Endpunkt (EA, EB, EC) im Sinne des Verfahrens entspricht.
- 9Teleconferencing system according to claim 8, characterized, that the teleconferencing system (100) a moderation unit (106), which is arranged, at least one audio signal (AA, AB, AC), a surround sound (RA, RB, RC) and / or overall sound (GA, GB, GC) to activate, deactivate, amplify and / or mitigate. Telekonferenzsystem nach Anspruch 8, dadurch gekennzeichnet, dass das Telekonferenzsystem (100) eine Moderationseinheit (106) aufweist, die eingerichtet ist, wenigstens ein Audiosignal (AA, AB, AC), einen Raumklang (RA, RB, RC) und/oder einen Gesamtraumklang (GA, GB, GC) zu aktivieren, zu deaktivieren, zu verstärken und/oder abzuschwächen.
- 10Teleconferencing system (100) according to one of claims 8 or 9, characterized, that the teleconferencing system (100) is designed as a video conference system, preferably as a web conferencing system, wherein preferably each teleconferencing end point (A, B, C) is a picture display unit (105) having. Telekonferenzsystem (100) nach einem der Ansprüche 8 oder 9, dadurch gekennzeichnet, dass das Telekonferenzsystem (100) als Videokonferenzsystem, bevorzugt als Webkonferenzsystem, ausgebildet ist, wobei bevorzugt jeder Telekonferenzendpunkt (A, B, C) eine Bildwiedergabeeinheit (105) aufweist.
- 11Teleconferencing system according to one of the preceding claims 8 to 10, characterized, that the teleconferencing system (100) is set up by means of the image display unit (105) of the third teleconferencing endpoint (A, B, C) the virtual space (V2, V3), wherein in each case one of the first and the second participant (TA. TB, TC) associated first and second virtual object at the defined positions (PA, PB, PC) within the virtual space (V2, V3) associated with the first and second teleconferencing end points (A, B, C), respectively, and a third subscriber (TA, TB, TC) associated third virtual object at the standard receiver position (PS) is pictured. Telekonferenzsystem nach einem der vorhergehenden Ansprüche 8 bis 10, dadurch gekennzeichnet, dass das Telekonferenzsystem (100) eingerichtet ist, mittels der Bildwiedergabeeinheit (105) des dritten Telekonferenzendpunkts (A, B, C) den virtuellen Raum (V2, V3) darzustellen, wobei jeweils ein dem ersten und dem zweiten Teilnehmer (TA, TB, TC) zugeordnetes erstes und zweites virtuelles Objekt an den definierten Positionen (PA, PB, PC) innerhalb des virtuellen Raums (V2, V3) dargestellt werden, die dem ersten bzw. zweiten Telekonferenzendpunkt (A, B, C) zugeordnet sind, und ein dem dritten Teilnehmer (TA, TB, TC) zugeordnetes drittes virtuelles Objekt an der Standard-Empfängerposition (PS) dargestellt wird.
Independent claims11
88 paragraphs in 1 section, as filed
The invention relates to a method for generating a total spatial sound to be transmitted to a real endpoint, wherein an audio signal is assigned to the real endpoint. Furthermore, the invention relates to the use of the method and a teleconferencing system.
As a rule, several participants participate in a teleconference and communicate with each other. For this purpose, audio signals, in this case speech signals of the participants, must be processed and transmitted between the participants. In particular, one participant in each case must receive the audio signals of the other participants in the form of a total signal.
In advanced teleconferencing systems, a virtual (conference) room is simulated, which is virtually entered by all participants for purposes of the teleconference, or appears to communicate in it.
Such a teleconferencing system is for example from the <patcit><text>US 2010/0316232 A1</text></patcit> known. The teleconferencing system has teleconferencing end points each with an audio recording and playback unit, for example in the form of a headset with a headphone and a microphone. The teleconferencing endpoints allow participants to participate in a teleconference.
The document specifically describes that each of the participants is assigned a virtual position in a virtual space, in particular on a virtual conference table.
In order to provide the subscribers of the teleconferencing system with a realistic or lifelike hearing impression, room sounds are generated from audio signals recorded by respective other teleconferencing end points via their audio recording units. For this purpose, a receiving teleconferencing endpoint or listening subscriber is given the impression that an audio signal of a respective other teleconferencing endpoint or subscriber originates from its assigned virtual position within the virtual space and is received at its own position on the virtual conference table. would be heard.
For each (receiving) teleconferencing endpoint or subscriber, a total room sound from all room sounds of all other teleconferencing endpoints or subscribers is then mixed and presented to the associated subscriber via the associated audio playback unit, so that each subscriber ultimately receives an auditory impression as if he were actually would sit at the (virtual) conference table and communicate with the other participants.
The disadvantage here, however, is that the number of room sounds to be generated increases approximately quadratically with the number of participants. Thus, demands on available computing power and the like increase disproportionately with the number of participants. As a result, the number of participants in such a teleconference is usually limited and / or it requires large server farms to cope with the resulting computational effort.
From the <patcit><text>WO 199053673 A1</text></patcit> is another teleconferencing system known with similar disadvantages.
Another way to enable teleconferencing for larger subscriber numbers is in the <patcit><text>US Pat. No. 8,627,213 B1</text></patcit> proposed. It is envisaged to worsen the quality of the generated room sounds with increasing number of participants and thus reduce the amount of computation required per generated surround sound. This, however, suffers the ease of use of such a teleconferencing system.
The problem outlined above is further exacerbated by the fact that the spatial sounds ultimately represent multi-channel, in particular stereophonic, signals. Thus, at least two individual signals must be calculated for each surround sound to be generated. This further increases the computational effort actually required.
It is therefore an object of the present invention to provide a method and a teleconferencing device with which the computational outlay for the generation of overall room sounds, in particular for a large number of subscribers, can be reduced.
The object is achieved by a method for generating a total space sound to be transmitted to a real end point, wherein an audio signal is assigned to the real end point, with the method steps: <ul list-style="bullet"><li>a. Set a default recipient location within a virtual space</li><li>b. Assigning the real endpoint with associated audio signal to a defined position within the virtual space,</li><li>c. for at least one real endpoint generating a surround sound based on the associated audio signal, the associated defined position and the default receiver position,</li><li>d. Generating at least one total spatial sound for at least one real endpoint from at least a part of previously generated spatial sounds withException of the spatial sound of the at least one real endpoint and</li><li>e. Transmitting the at least one overall sound to the at least one real endpoint.</li></ul>
Thus, audio signals from a real endpoint, such as a first endpoint of a teleconferencing system, to another real endpoint, such as a second endpoint of the teleconferencing system, can be communicated by generating and transmitting a total surround sound for each participant based on the audio generated by other participants.
The audio signals can be recorded, for example via microphones at the real endpoints and used according to the method. The audio signals may preferably be voice signals, in particular monophonic voice signals.
According to the invention, a respectively "listening" participant or a real endpoint assigned to it for receiving audio signals of other participants or endpoints is positioned virtually in a virtual space at a position identical to all participants or real endpoints, the standard receiver position. The standard receiver position may, for example, correspond to a head page of a virtual conference table arranged in virtual space. This definition can already take place in the construction of a system applying the method according to the invention. It can also be done at the beginning of each use of the method.
In addition, a defined, preferably different from the standard receiver position, position within the virtual space is assigned to each real endpoint. For example, the defined positions may correspond to virtual seating positions on side areas of the virtual conference table.
Based on the audio signal of at least one real endpoint as well as its defined position and the standard receiver position in the virtual space, a surround sound is now generated. In particular, the surround sound may be generated as if the audio signal were being radiated from the defined position and received at the standard receiver position. For example, a direction and a distance from the position of the defined position and the standard receiver position to each other can be detected and a surround sound can be generated from the audio signal according to the direction and distance.
Thus, the impression is created as if audio signals of their respective associated real endpoints were each emitted from the respective defined position.
To generate the surround sound, it can be simulated that a receiver pair with at least two receivers spaced apart from one another is arranged at the standard receiver position. Then, for example, additionally a standard receiver direction, thus a "gaze" or "hearing direction", can be taken into account, with which the relative orientation of the receiver pair to the defined position can be described. Thus, further spatial sound influencing properties, for example predefinable emission characteristics or reception characteristics, can be taken into account.
For at least one real endpoint, spatial sounds of one or more other real endpoints are then combined to form a total sound sound. This overall sound is then transmitted to the at least one real endpoint.
Thus, the at least one real endpoint receives surround sounds based on the respective audio signals of other real endpoints. In particular, a real endpoint "hears" the overall sound associated with it. Its overall sound provides a surround sound impression, after which the other real endpoints would send their audio signals from the respective defined positions to its apparently own receiver position within the virtual space. Since the at least one real endpoint is only transmitted in each case the total space sound assigned to it, there is no direct possibility for this or its associated subscriber to recognize that the apparently own receiver position corresponds in each case to the standard receiver position and identical for all real endpoints is.
According to the invention, it suffices thereby to generate at most one surround sound and one overall sound for each real endpoint.
Thus, the number of room sounds required no longer increases approximately quadratically, but only approximately linearly with the number of real endpoints. The inventive method thus proves to be particularly efficient.
Also, an additional real endpoint with associated audio signal may simply be added by assigning it a defined position and creating surround sound for it, preferably before full-scale sounds are generated.
The method can be used or requires in particular for large numbers of participants Given number of participants only low computing power and transmission capacity. In other words, compared to the prior art, a substantial increase in efficiency is achieved without there being noticeable quality losses with regard to the spatial sound impression.
A further increase in efficiency can be achieved if audio signals are preselected overall and / or packet by packet. For example, only the audio signals of the real endpoints that exceed a predefinable level or performance level can be taken into account for the generation of surround sounds and / or overall room sounds. This can take place overall, that is to say for a complete audio signal, or else packet by packet, that is to say for segments of the audio signal. Other audio signals can thus be hidden. Thus, for example, unnecessary background noise can be eliminated. Also, the amount of signals or signal data to be processed can be reduced.
A surround sound or effect can be generated particularly easily if a surround sound is generated by means of an interaural transit time difference (ILD) and / or an interaural level difference (IPD). ILD and IPD, for example, make it possible to convert a monophonic audio signal into a (stereophonic) surround sound and to memorize the surround sound in an apparent direction. For this purpose, the audio signal can be fed into different channels of the room sound with a time delay to produce an ILD. To produce an IPD, the audio signal can be fed with different volume levels or different gains and / or attenuations into different channels of the surround sound.
The realism of generated spatial sounds or total room sounds can be further increased if an audio signal and / or a surround sound and / or a total room sound are filtered by means of a head-related transfer function.
Also, an echo and / or a reverberation effect, preferably as a function of at least one defined position, can be added to an audio signal, a surround sound and / or an overall sound.
It is particularly advantageous if the echo and / or the reverberation effect, for example, for imitation of different wall or ceiling materials of the virtual space to be set.
For example, the number of simulated reflections of audio signals on walls and ceilings can be adjusted. For example, it can be set whether only direct connections between a defined position and the standard receiver position for spatial sound generation are taken into account and / or indirect connections, in particular two, three or more virtual reflections, are taken into account. It is also possible to simulate materials in virtual space arranged virtual objects or space limitations, for example, in terms of their reflection, absorption and / or sound conduction properties. Thus, further, realism-enhancing effects can be implemented in the generated total sounds. This has a particularly advantageous effect,
The scope of the invention also includes a use of the method according to the invention in a multimedia system for at least two users, in particular a teleconferencing system or an Internet-based multi-player system. Multi-user multimedia systems may be designed, for example, as web-based audiovisual systems. In these, it is often provided that an audio signal is assigned to a real endpoint, for example a computer unit equipped with an audio playback unit, in particular loudspeakers. The audio signal may originate from a sound recording of an audio recording unit of the computer unit or in the computer unit and / or in a central unit of the multimedia system, for. As a central Internet server, be generated as the real endpoint associated audio signal.
Then, the inventive method can be used to generate based on the audio signals overall sounds with a high degree of realism and to transmit this efficiently to other real endpoints of the multimedia system.
The scope of the invention further includes a teleconferencing system having at least three teleconferencing endpoints, each of the teleconferencing endpoints having an audio playback unit and an audio recording unit, wherein the teleconferencing system is configured, a first and a second audio signal, such as a first and a second subscriber to record via the audio recording units of a first and a second teleconferencing endpoint, wherein the teleconferencing system is set to, a total room sound according to the inventive method from the first and the second audio signal to generate to transmit to a third teleconferencing endpoint, such as a third party, and output via the audio playback unit, each teleconferencing endpoint corresponding to a real endpoint in the sense of the method.
Such a teleconferencing system is suitable for applying the method according to the invention and thus drastically reducing the computation and / or transmission capacities required for a teleconference, but at the same time satisfying high quality requirements with regard to the overall room sounds formed.
It is particularly preferred if the teleconferencing system has a moderation unit which is set up to activate, deactivate, amplify and / or attenuate at least one audio signal, surround sound and / or overall sound. Thus, a guided via the teleconferencing teleconference can be moderated by means of one or more moderation units. For example, subscribers of the teleconference can thus be switched on or hidden and / or subsets of subscribers can be formed within a teleconference, so that one subgroup can be temporarily separated acoustically from another subgroup, for example.
In particular, it can be provided that the teleconferencing system is designed as a video conferencing system, preferably as a web conferencing system, wherein preferably each teleconferencing endpoint has an image display unit. The web conferencing system can be designed in particular as a browser-based web conferencing system. A web conference system makes it possible, for example, to be able to participate in a web conference produced by means of the web conferencing system via a browser which is executively installed on the respective computer unit. In particular, it can be provided that the teleconferencing system has a computer program component that is downloadable and / or executable by means of a browser.
It is particularly preferred if the teleconferencing system is set up to display the virtual space by means of the image display unit of the third teleconference endpoint, wherein in each case a first and second virtual object assigned to the first and the second participant are displayed at the defined positions within the virtual space corresponding to the first are assigned to the first and second teleconferencing endpoint and a third virtual object associated with the third participant is displayed at the standard receiver position.
Thus, for a subscriber or user of a teleconferencing endpoint or receiver of the associated overall sound, the created acoustic impression can be enriched by a visual impression matching thereto. Thus, an impression of a virtual reality can be further increased.
Further features and advantages of the invention will become apparent from the following detailed description of variants of the method according to the invention and of embodiments of the subjects of the invention, with reference to the figures of the drawing, which shows details essential to the invention, and from the claims. The features shown there are not necessarily to scale and presented in such a way that the features of the invention can be made clearly visible. The various features may be implemented individually for themselves or for a plurality of combinations in variants of the invention.
In the schematic drawing variants and embodiments of the invention are shown and are explained in more detail in the following description.
Show it:
<figref>1</figref> a teleconference with three subscribers according to the prior art;
<figref>2</figref> a teleconference in which the method according to the invention is used;
<figref>3</figref> an inventive teleconferencing system;
<figref>4a</figref> and <figref>4b</figref> Teleconference endpoints of the teleconferencing system of <figref>3</figref> two different participants.
<figref>1</figref> schematically illustrates the operation of a teleconferencing system in the prior art.
By way of example, three participants T take part in a teleconference<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub> part. The participants T<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub> each use teleconferencing endpoints A ', B', C '.
The teleconferencing endpoints A ', B', C 'are in a virtual space V<sub>1</sub> each positions P<sub>A '</sub>, P<sub>B '</sub>, P<sub>C '</sub> assigned. To make a teleconference, first take theTeleconferencing endpoints A ', B', C 'each associated with them audio signals A<sub>A</sub>, A<sub>B</sub>, A<sub>C</sub>, For example, monophonic speech signals, from the teleconference endpoints A ', B', C 'using participants T<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub>on microphones, for example. Subsequently, suitable room sounds R<sub>1</sub> to R<sub>6</sub> and transmitted to the respective other (receiving) teleconferencing end points A ', B', C '. There the room sounds R<sub>1</sub> to R<sub>6</sub> the participants T<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub>, for example via stereo headphones, presented. The room sounds R<sub>1</sub> to R<sub>6</sub> be formed, the participants T<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub> to give a surround sound impression, as if the audio signals A<sub>A</sub>, A<sub>B</sub>, A<sub>C</sub> apparently in virtual space V<sub>1</sub> from the positions P assigned to the respective other teleconferencing end points A ', B', C '<sub>A '</sub>, P<sub>B '</sub>, P<sub>C '</sub> to their own position P<sub>A '</sub>, P<sub>B '</sub> or P<sub>C '</sub> would arrive.
Again <figref>1</figref> It is not difficult to deduce that in the prior art, only three subscribers T<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub> the total of 6 room sounds R<sub>1</sub> to R<sub>6</sub> are generated, each of the room sounds R<sub>1</sub> to R<sub>6</sub> each comprises at least one left and one right channel.
Thus, already with only three participants T<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub> Teleconference endpoints A ', B', C 'a total of 12 surround channels or individual signals are calculated and transmitted. The computational effort increases approximately quadratically with the number of teleconferencing endpoints.
This high calculation and transmission costs can be drastically reduced by using the method according to the invention. This should now be based on the<figref>2</figref> be explained in more detail using the example of a teleconference to be carried out.
In the <figref>2</figref> schematically a teleconference is shown, at which the three participants T<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub> over 3 teleconferencing endpoints A, B, C participate. The teleconferencing endpoints A, B, C represent real endpoints E<sub>A</sub>, E<sub>B</sub>, E<sub>C</sub> in the sense of the method according to the invention. In turn, the teleconferencing end points A, B, C are each the audio signals A<sub>A</sub>, A<sub>B</sub>, A<sub>C</sub> assigned.
First, a standard receiver position P<sub>S</sub> within a virtual space V<sub>2</sub> established.
Subsequently, the telephoto endpoints A, B, C defined positions P<sub>A</sub>, P<sub>B</sub>, P<sub>C</sub> in virtual space V<sub>2</sub> assigned.
In the <figref>2</figref> the case is shown in which the teleconferencing end point A the audio signals A<sub>B</sub> and A<sub>C</sub> in the form of a total room sound G<sub>A</sub> from the other two teleconferencing B, C should receive. In the case illustrated here, therefore, the total space sound G is<sub>A</sub> in such a way and to transmit to the teleconferencing end point A, that a Raumklangeindruck arises as if the audio signals A<sub>B</sub> and A<sub>C</sub> from the defined positions P<sub>B</sub>, P<sub>C</sub> to the standard receiver position P<sub>S</sub> sent and received there. Analogous to the procedure illustrated for this case, total space sounds G<sub>B</sub> and G<sub>C</sub> for the teleconferencing endpoints B, C formed by the defined positions P<sub>A</sub>, P<sub>C</sub> (for teleconferencing endpoint B) or P<sub>A</sub>, P<sub>B</sub> (for teleconferencing endpoint C) in each case in connection with the standard receiver position P<sub>S</sub> be based on.
For this purpose, in a further method step, at least for the teleconferencing end points B and C, room sounds R<sub>B</sub>, R<sub>C</sub> on the basis of the respective assigned audio signal A.<sub>B</sub> or A<sub>C</sub>, the assigned defined positions P<sub>B</sub>, P<sub>C</sub> and the standard receiver position P<sub>S</sub> generated.
If the total room sounds G<sub>B</sub> and / or G<sub>C</sub> are generated to increase efficiency analogous to a surround sound R<sub>A</sub> generated and the room sounds R<sub>C</sub> (for total room sound G<sub>B</sub>) and R<sub>B</sub> (for total room sound G<sub>C</sub>) reused, that is reused without recalculation. If already created, the surround sound R<sub>A</sub>, reused.
The example of the room sound R<sub>C</sub> let the generation of the spatial sounds R<sub>A</sub>, R<sub>B</sub> or R<sub>C</sub> be explained in more detail.
First of all, the audio signal A becomes<sub>C</sub> pre-selected in packages. For this purpose, the audio signal A<sub>C</sub> disassembled into individual packages. Within each individual package it is checked whether a given minimum sound level is reached. If this is not the case, this individual package is replaced by a zero signal of equal duration, otherwise the individual package is processed unchanged. Thus, unnecessary background noise can be minimized and at the same time the required later processing or computational effort can be further reduced.
In the <figref>2</figref> It can now be seen that the standard receiver position P<sub>S</sub> within the virtual space V<sub>2</sub> two more positions P<sub>L</sub>, or P<sub>R</sub> assigned or previously defined. These are spaced from each other by a distance d.
Thus, within the virtual space different distances d<sub>L</sub> and d<sub>R</sub> of positions P<sub>L</sub> and P<sub>R</sub> to the teleconference end point C associated position P<sub>C</sub>,
A virtually from the position P<sub>C</sub> propagating audio signal would thus correspond to a suitably selected speed of sound with an interaural transit time difference ILD, in other words proportional to the difference of the distances d<sub>L</sub> and d<sub>R</sub>, at positions P<sub>L</sub> or P<sub>R</sub> arrive.
Also, due to propagation at different levels of sound, the audio signal would be at positions P<sub>L</sub> or P<sub>R</sub> Thus, this would also result in an interaural level difference IPD.
Therefore, the surround sound R<sub>C</sub> formed in this process variant by the audio signal A<sub>C</sub> offset in time by the resulting interaural transit time difference ILD and with different sound levels corresponding to the resulting interaural level difference IPD in a left and a right channel of the spatial sound R to be generated<sub>C</sub> fed or calculated for the two channels.
In addition, both channels of the room sound R<sub>C</sub> by means of a position P<sub>L</sub>, P<sub>R</sub> and P<sub>C</sub> appropriately selected Head-Related Transfer Function filtered.
In this variant of the method, the surround sound R is used to improve a spatial sound impression<sub>C</sub> added or added further effects.
In particular, echo and reverberation effects become the surround sound R<sub>C</sub> according to a suitably chosen embodiment of the virtual space V<sub>2</sub> added to even apparent non-direct acoustic connections between the positions P<sub>L</sub>, P<sub>R</sub> and P<sub>C</sub> to account for surround sound generation. In particular, there are limitations in advance<figref>21</figref>. <figref>22</figref>. <figref>23</figref>. <figref>24</figref> of the virtual space V<sub>2</sub>For example, virtual walls, ceilings and floors, as well as virtual objects, such as a virtual conference table <figref>25</figref>, predefined in dependence on a desired embodiment and within the virtual space V<sub>2</sub> positioned. In addition, beforehand the limits<figref>21</figref>. <figref>22</figref>. <figref>23</figref>. <figref>24</figref> and sound properties associated with the virtual objects, in particular acoustic reflection and absorption properties, according to which the echo and reverberation effects are subsequently generated or adjusted.
In a further method step, the total space sound G<sub>A</sub> for the teleconferencing endpoint A from the generated surround sounds R<sub>B</sub>, R<sub>C</sub>that is, from all room sounds R<sub>A</sub>, R<sub>B</sub>, R<sub>C</sub> with the exception of the "own" room sound R<sub>A</sub>, generated by channelwise addition of the two spatial sounds.
In a last method step, the total space sound G<sub>A</sub> transmitted to the teleconferencing endpoint A and stereo headphones of the teleconferencing endpoint A the subscriber T.<sub>A</sub> presents.
Since it is thus only the generation of three room sounds R<sub>A</sub>, R<sub>B</sub>, R<sub>C</sub> requires, is saved in accordance with the method in this way, compared to the prior art significantly computational capacity. In addition, the computational need increases only approximately linearly with the number of teleconferencing endpoints.
Furthermore, the <figref>2</figref> to deduce the idea of the inventive method that the apparent radiation of the audio signals A<sub>A</sub>, A<sub>B</sub>, A<sub>C</sub> or the room sounds R<sub>A</sub>, R<sub>B</sub>, R<sub>C</sub> is spatially separated from the apparent reception of the total space sounds G<sub>A</sub>, G<sub>B</sub>, G<sub>C</sub>, which according to the invention at the standard recipient position P<sub>S</sub> is located.
In the <figref>3</figref> is now an inventive teleconferencing system <figref>100</figref> shown.
Of the <figref>3</figref> it can be seen that in turn the participants T<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub> Teleconferencing endpoints A, B, C are assigned. The teleconferencing end points A, B, C correspond to real endpoints E<sub>A</sub>, E<sub>B</sub>, E<sub>C</sub> of the inventive method. Other participants or teleconferencing endpoints can the teleconferencing system<figref>100</figref> to be added at will; the number of 3 participants or teleconferencing endpoints is chosen here only as an example.
The teleconferencing endpoints A, B, C each have an audio playback unit <figref>101</figref> and an audio recording unit <figref>102</figref> on. The audio playback unit<figref>101</figref> and the audio recording unit <figref>102</figref> are in this embodiment as a headset, that is designed as headphones and microphones. The audio playback units<figref>101</figref> or the audio recording unit <figref>102</figref> are each with a computer unit <figref>103</figref> the respective teleconferencing endpoint A, B or C connected. About the audio recording units<figref>102</figref> become audio signals A<sub>A</sub>, A<sub>B</sub>, A<sub>C</sub>, in particular speech signals of the participants T<sub>A</sub>, T<sub>B</sub>, T<sub>C</sub>, added.
Furthermore, the teleconferencing endpoints A, B, C are by means of their computer units <figref>103</figref> with a central unit designed as a computing server <figref>108</figref> in terms of data, in particular via a communications network, in this case the Internet.
The central unit <figref>108</figref> receives from the teleconferencing end points A, B, C the audio signals A<sub>A</sub>, A<sub>B</sub>, A<sub>C</sub>, generates the total space sounds G from these<sub>A</sub>, G<sub>B</sub>, G<sub>C</sub> and transmits them to the respective teleconferencing end points A, B, C for playback or presentation via the respective audio playback units <figref>101</figref>, The processing of the audio signals A<sub>A</sub>. A<sub>B</sub>, A<sub>C</sub> takes place by means of a central computer program component <figref>109</figref> on that in the central unit <figref>108</figref> executable installed.
The teleconferencing system <figref>100</figref> is especially designed as a web conferencing system. These are on the computer units<figref>103</figref> the teleconferencing endpoints B and C respectively include instances of an endpoint computer program component <figref>104</figref> executable installed.
On the teleconference endpoint A is an instance of an endpoint computer program component <figref>104 '</figref> executable installed. The endpoint computer program component<figref>104 '</figref> corresponds to the endpoint computer program component <figref>104</figref> with the difference that they additionally have a moderation unit <figref>106</figref> having.
In particular, these instances may be an endpoint computer program component <figref>104 '</figref> by means of one on the computer units <figref>103</figref> executable installed browser. The endpoint computer program components<figref>104 '</figref> are also by the computer units <figref>103</figref> downloadable on the central unit <figref>108</figref> installed, allowing use of the teleconferencing system <figref>100</figref> is easily enabled by the endpoint computer program components <figref>104 '</figref> from the central unit <figref>108</figref> by means of a browser on computer units <figref>103</figref> downloaded and executed there.
The moderation unit <figref>106</figref> is designed to provide settings with which a participant each one of the audio signals A<sub>A</sub>, A<sub>B</sub>, A<sub>C</sub> amplify and generally be able to activate or deactivate. For this, the respectively selected settings are changed by the moderation unit<figref>106</figref> by means of the respective computer units <figref>103</figref> to the central unit <figref>108</figref> which transmits the audio signals A received from it<sub>A</sub>, A<sub>B</sub>, A<sub>C</sub> appropriately reinforced, attenuated and / or selected before using the method according to the invention.
Thus, for example, temporary subgroups of subscribers of the teleconference can be formed via these setting options, between which at least temporarily no audio signals A are formed<sub>A</sub>, A<sub>B</sub>, A<sub>C</sub> be replaced. In this embodiment, therefore, the subscriber has T<sub>A</sub> Access to these settings of the moderation unit installed on its assigned teleconference point A. <figref>106</figref>, Thus participant T represents<sub>A</sub> a presenter of the teleconference.
Furthermore, the <figref>3</figref> Image display units <figref>105</figref> to be taken with the respective computer units <figref>103</figref> the teleconferencing endpoints A, B, C are connected.
By way of example, the show <figref>4a</figref> and the <figref>4b</figref> Screen issues, as they participants T<sub>A</sub> (<figref>4a</figref>) or T<sub>B</sub> (<figref>4b</figref>) are presented via their teleconferencing endpoints A and B, respectively. A virtual space V can be seen in each case<sub>3</sub>in which virtual objects A ', B', C ', in this case avatars, at the defined positions P<sub>A</sub>, P<sub>B</sub>, P<sub>C</sub> or the standard receiver position P<sub>S</sub> corresponding graphic positions within the virtual space V<sub>3</sub> are arranged. In each case, this is the respective subscriber T<sub>A</sub> or T<sub>B</sub> associated virtual object A 'and B' at the standard receiver position P<sub>S</sub> corresponding graphic position. The other virtual objects B ', C' or A ', C' are on the other hand at the positions P<sub>B</sub> and P<sub>C</sub> or P<sub>A</sub> and P<sub>C</sub>, Thus, a visual impression corresponding to the auditory auditory impression by the screen output of the image display units<figref>105</figref> generated for the respective teleconferencing end points A, B, C.
QUOTES INCLUED IN THE DESCRIPTION
This list of the documents listed by the applicant has been generated automatically and is included solely for the better information of the reader. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions.
Cited patent literature
<ul list-style="bullet"><li>US 2010/0316232 A1 <b>[0004]</b></li><li>WO 199053673 A1 <b>[0009]</b></li><li>US 8627213 B1 <b>[0010]</b></li></ul>
5 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US2010316232A1 | Cites | United States of America | Applicant |
| US2015373477A1 | Cites | United States of America | Search report |
| US6850496B1 | Cites | United States of America | Search report |
| US8627213B1 | Cites | United States of America | Applicant |
| WO9053673A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 102016112609 | Germany | A | |
| DE201610112609 | – | – | – |
9 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent grant now finalGrantedR020 | R020 | |
| Change of applicant/patenteeR081 | R081 | |
| Change of representativeR082 | R082 | |
| Grant decision by examination section/examining divisionR018 | R018 | |
| Change of applicant/patenteeR081 | R081 | |
| Change of representativeR082 | R082 | |
| Change of representativeR082 | R082 | |
| Response to examination communicationR016 | R016 | |
| Request for examination validly filedR012 | R012 |
Numbers
- Publication
- 102016112609
- Publication, DOCDB
- 102016112609
- Publication, EPODOC
- DE102016112609
- Application
- 10112609
- Application, DOCDB
- 102016112609
- Application, EPODOC
- DE201610112609
Titles2
- English
- Method for generating a total room sound to be transmitted to a real endpoint, use of the method and teleconferencing system
- German
- Verfahren zur Erzeugung eines an einen realen Endpunkt zu übermittelnden Gesamtraumklangs, Verwendung des Verfahrens sowie Telekonferenzsystem
Classification
- CPC, 5
- H04L12/1827
- H04R27/00
- H04S7/302
- H04S7/305
- H04S2420/01