Spatial audio rendering for beamforming loudspeaker array
Abstract
Problem to be solved.To provide a method for obtaining an immersive reproduction by using a very limited number of loudspeaker cabinets (for example, only one). A method of reproducing sound using a loudspeaker array housed in a loudspeaker cabinet 2 is to select several sound rendering modes and to change in one or both of sensor data and user interface selection. Includes changing the selected sound rendering mode based on. This sound rendering mode includes several central modes and at least one direct perimeter mode. [Selection diagram] Fig. 1

Term
Projected expiry 15 August 2037.
- Priority
- Filed
- Published
- Today
- Projected expiry
20 claims: 4 independent, 16 dependent
- 1ラウドスピーカアレイを有するオーディオシステムであって、 複数のラウドスピーカドライバを内蔵したラウドスピーカキャビネットと、 出力が前記複数のラウドスピーカドライバの入力に連結された複数のオーディオ増幅器と、 前記ラウドスピーカドライバによってサウンドに変換されるサウンドプログラムコンテンツについての複数の入力オーディオチャンネルを受信するレンダリングプロセッサであって、前記複数のオーディオ増幅器の入力に連結された出力を有し、a)複数の第1のモードと、b)第2のモードとを含む複数のサウンドレンダリング動作モードを有する、レンダリングプロセッサと、 決定論理入力としてセンサデータ及びユーザインタフェース選択のうちの一方又は双方を受信する決定論理であって、前記決定論理入力の各々は、i)部屋の特徴又はii)聴取場所のうちの1つを示す、決定論理と、 を備え、前記レンダリングプロセッサの前記複数の第1のモードの各々において、前記レンダリングプロセッサの前記出力は、前記複数のラウドスピーカドライバに、i)前記複数の入力オーディオチャンネルのうちの2つ以上の合計を含む全方向性パターンを有するサウンドビームを生成させ、前記全方向性パターンは、ii)複数のローブを有する指向性パターンであって、各ローブは前記2つ以上の入力オーディオチャンネルの差分を含む、指向性パターンと重ね合わされており、 前記レンダリングプロセッサの前記第2のモードにおいて、前記レンダリングプロセッサの前記出力は、前記複数のラウドスピーカドライバに、i)前記聴取場所に照準が向けられた直接コンテンツパターンを有するサウンドビームを生成させ、前記直接コンテンツパターンは、ii)前記聴取場所から離れて照準が向けられた周囲コンテンツパターンと重ね合わされており、 前記決定論理は、前記レンダリングプロセッサの前記複数のサウンドレンダリングモードのうちの1つのレンダリングモードの選択を行い、これに従って、前記レンダリングプロセッサは、前記サウンドプログラムコンテンツの再生中に前記複数のラウドスピーカドライバを駆動するように構成され、前記決定論理は、前記決定論理入力の変化に基づいて、前記レンダリングモードの選択を変更する、システム。
- 2前記ラウドスピーカキャビネット内の前記複数のドライバによって、500Hzを超える全てのコンテンツがサウンドに変換される、請求項1に記載のシステム。
- 3前記ラウドスピーカキャビネット内の前記複数のドライバは、前記サウンドプログラムコンテンツについての前記複数の入力オーディオチャンネルよりも数が多い、請求項2に記載のシステム。
- 4前記レンダリングプロセッサの前記複数の第1のモードの各々において、前記指向性パターンの前記複数のローブの各ローブは、前記2つ以上の入力チャンネルの差分を含み、前記複数のローブのうちの隣接するローブは互いに逆極性である、請求項2に記載のシステム。
- 5前記レンダリングプロセッサの前記複数の第1のモードの各々において、前記指向性パターンの前記複数のローブの各ローブは、前記2つ以上の入力チャンネルの差分を含み、前記複数のローブのうちの隣接するローブは互いに逆極性である、請求項1に記載のシステム。
- 6前記複数の第1のモードは、低次の第1のモードと高次の第1のモードとを含み、前記高次の第1のモードは、前記低次の第1のモードよりも大きな指向性指数又はより多数のローブを有したビームパターンを有する、請求項1に記載のシステム。
- 7前記決定論理は、前記複数の入力オーディオチャンネルを解析して相関コンテンツ及び無相関コンテンツを発見するものであり、前記相関コンテンツは、次に、前記直接コンテンツパターンでレンダリングされる一方、前記無相関コンテンツは、前記周囲コンテンツパターンでレンダリングされる、請求項1に記載のシステム。
- 8前記サウンドプログラムコンテンツは、映画フィルムの前記サウンドトラックであり、前記複数のオーディオチャンネルは、前記サウンドトラックの前記オーディオチャンネルの全てである、請求項1に記載のシステム。
- 9ラウドスピーカキャビネットに収容されたラウドスピーカアレイを用いてサウンドを再生するための方法であって、 ラウドスピーカキャビネット内に収容されたラウドスピーカアレイによってサウンドに変換されるサウンドプログラムコンテンツについての複数の入力オーディオチャンネルを受信することと、 決定入力としてセンサデータ及びユーザインタフェース選択のうちの一方又は双方を受信することであって、前記決定入力の各々は、i)部屋の特徴又はii)聴取場所のうちの1つを示す、ことと、 前記サウンドプログラムコンテンツの再生が前記ラウドスピーカアレイによって行われる複数のサウンドレンダリングモードのうちの1つを選択することと、前記決定入力の変化に基づいて前記選択されたサウンドレンダリングモードを変更することと、 を含み、前記複数のサウンドレンダリングモードは、a)複数の第1のモードと、b)第2のモードとを含み、 前記複数の第1のモードの各々において、前記ラウドスピーカアレイは、i)前記複数の入力オーディオチャンネルのうちの2つ以上の合計を含む全方向性パターンを有するサウンドビームを生成し、前記全方向性パターンは、ii)複数のローブを有する指向性パターンであって、各ローブは前記2つ以上の入力オーディオチャンネルの差分を含む指向性パターンと重ね合わされており、 前記第2のモードにおいて、前記ラウドスピーカアレイは、i)前記聴取場所に照準が向けられた直接コンテンツパターンを有するサウンドビームを生成し、前記直接コンテンツパターンは、ii)前記聴取場所から離れて照準が向けられた周囲コンテンツパターンと重ね合わされている、方法。
- 10前記サウンドレンダリングモードのうちの1つを選択することは、前記サウンドプログラムコンテンツを解析することに基づき、 低次の指向性パターンを有する前記複数の第1のモードのうちの1つは、前記サウンドプログラムコンテンツが主に周囲音又は拡散音である場合に選択され、 高次の指向性パターンを有する前記複数の第1のモードのうちの1つは、前記サウンドプログラムコンテンツがパンされたサウンドを含む場合に選択される、請求項9に記載の方法。
- 11前記サウンドプログラムコンテンツを解析することは、前記複数の入力オーディオチャンネルを解析して相関コンテンツ及び無相関コンテンツを発見することを含み、前記第2のモードにおいて、前記相関コンテンツは、前記周囲コンテンツパターンではなく前記直接コンテンツパターンでレンダリングされる一方で、前記無相関コンテンツは、前記直接コンテンツパターンではなく前記周囲コンテンツパターンでレンダリングされる、請求項10に記載の方法。
- 12前記サウンドプログラムコンテンツについての前記複数の入力オーディオチャンネルの全てにおいて、500Hz未満の周波数を超える全てのコンテンツが、前記ラウドスピーカキャビネットに収容された前記ラウドスピーカアレイによってサウンドに変換される、請求項9に記載の方法。
- 13前記サウンドプログラムコンテンツをサウンドに変換するために使用される前記ラウドスピーカアレイ内のドライバの数は、前記サウンドプログラムコンテンツについての前記複数の入力オーディオチャンネルよりも数が多い、請求項12に記載の方法。
- 14前記複数の第1のモードの各々において、前記指向性パターンにおける前記複数のローブの各ローブが前記2つ以上の入力チャンネルの差分を含み、前記複数のローブのうちの隣接するローブは、互いに反対の極性である、請求項9に記載の方法。
- 15前記複数の第1のモードは、低次の第1のモードと高次の第1のモードとを含み、前記高次の第1のモードは、前記低次の第1のモードよりも大きな指向性指数又はより多数のローブを有するビームパターンを有する、請求項9に記載の方法。
- 16ラウドスピーカアレイを有するオーディオシステムであって、 ラウドスピーカキャビネットに収容されたラウドスピーカアレイによってサウンドに変換されるサウンドプログラムコンテンツについての複数の入力オーディオチャンネルを受信する手段と、 部屋の音響又は聴取者の位置のうちの1つを示すセンサデータとユーザインタフェース選択とのうちの一方又は双方を受信する手段と、 前記サウンドプログラムコンテンツに対してコンテンツ解析を実行する手段と、 前記サウンドプログラムコンテンツの再生が前記ラウドスピーカアレイによって行われる複数のサウンドレンダリングモードのうちの1つを選択する手段と、 前記聴取者の位置、部屋の音響又はコンテンツ解析のうちの1つ以上の変化に基づいて前記選択されたサウンドレンダリングモードを変更する手段と、 を備え、前記複数のサウンドレンダリングモードは、a)複数の第1のモードと、b)第2のモードとを含み、 前記複数の第1のモードにおいて、前記ラウドスピーカアレイは、増加する次数の複数のサウンドビームパターンを別々に生成し、 前記第2のモードにおいて、前記ラウドスピーカアレイは、i)前記聴取者の位置に照準が向けられた直接コンテンツパターンを有するサウンドビームを生成し、前記直接コンテンツパターンは、ii)前記聴取者の位置から離れて照準が向けられた周囲コンテンツパターンと重ね合わされている、オーディオシステム。
- 17前記複数のサウンドビームパターンは、それぞれ、増加するステレオ密度を有し、前記複数のサウンドビームパターンの各々は、360度にわたる複数の隣接するステレオセクタを含み、各ステレオセクタは、左チャンネル領域と右チャンネル領域とに挟まれた中央チャンネル領域から構成されている、請求項16に記載のオーディオシステム。
- 18前記サウンドプログラムコンテンツについてのコンテンツ解析に基づいて、前記サウンドレンダリングモードのうちの1つを選択する場合、 低次の指向性パターンを有する前記複数の第1のモードのうちの1つは、前記サウンドプログラムコンテンツが主に周囲音又は拡散音である場合に選択され、 高次の指向性パターンを有する前記複数の第1のモードのうちの1つは、前記サウンドプログラムコンテンツがパンされたサウンドを含む場合に選択される、請求項16に記載のオーディオシステム。
- 19前記サウンドプログラムコンテンツをコンテンツ解析することは、前記複数の入力オーディオチャンネルを解析して相関コンテンツ及び無相関コンテンツを発見することを含み、前記第2のモードにおいて、前記相関コンテンツは前記直接コンテンツパターンでレンダリングされる一方で、前記無相関コンテンツは前記周囲コンテンツパターンでレンダリングされる、請求項16に記載のオーディオシステム。
- 20前記サウンドプログラムコンテンツについての前記複数の入力オーディオチャンネルの全てにおいて、500Hz未満の周波数を超える全てのコンテンツは、前記ラウドスピーカキャビネットに収容された前記ラウドスピーカアレイによってサウンドに変換され、前記サウンドプログラムコンテンツをサウンドに変換するために使用される前記ラウドスピーカアレイ内のドライバの数は、前記サウンドプログラムコンテンツについての前記複数の入力オーディオチャンネルよりも数が多い、請求項16に記載のオーディオシステム。
Independent claims20
36 paragraphs, as filed
0001This application claims the benefit of the earlier filing date of the co-pending US Patent Provisional Application No. 62 / 402,836 filed on September 30, 2016. One embodiment of the present invention relates to spatially selective rendering of audio by a loudspeaker array for playing back stereo recordings indoors. Other embodiments are also described.
0002Much effort has been put into developing techniques aimed at playing sound recordings with improved quality, resulting in sounds as natural as the original recording environment. This method creates a sound field around the listener whose spatial distribution is closer to the spatial distribution of the original recording environment. Early experiments in this area have shown that, for example, a loudspeaker in front of the listener plays a music signal, and a loudspeaker behind the listener produces a slightly delayed version of the same signal. Playback revealed that it gave the listener the impression that the music was being played in front of them. Add another loudspeaker to the left of the listener and another loudspeaker to the right to provide the same signal to these side speakers with a delay different from the delay between the front and rear loudspeakers. By doing so, the above configuration can be improved.
0003Stereo recording captures the sound environment by simultaneously recording from at least two microphones strategically placed with respect to the sound source. While playing these (at least two) input audio channels through their respective loudspeakers, the listener roughly derives the location of the sound source (using the small perceived timing and volume differences), thereby. , Enjoy a sense of space. In one approach, two signals, the central signal containing the central information, and the lateral signal that starts at essentially zero for the centrally located sound source and then increases with angular deviation (ie, obtains "side" information). And, you may choose the configuration of the microphone to generate. Such reproduction of the central and side signals can be performed by their respective loudspeaker cabinets that are adjacent to each other and oriented vertically to each other, which are substantially directional enough to replicate the recording with the microphone arrangement. Can have.
0004Loudspeaker arrays such as line arrays have been used in large venues such as outdoor music festivals to generate spatially selective sound (beams) directed at the audience. Line arrays are also used in large closed spaces such as chapels, sports arenas, and malls.
0005One embodiment of the present invention uses a loudspeaker array to render audio that has both clarity and immersive or spatial sensation in a room or other confined space. The system has a loudspeaker cabinet with a large number of built-in drivers and a large number of audio amplifiers connected to the driver inputs. The rendering processor receives several input audio channels (eg, left and right of a stereo recording) for sound program content such as music that should be converted to sound by the driver. The rendering processor has an output coupled to the input of the amplifier on a digital audio communication link. The rendering processor also has several sound rendering operating modes that generate individual signals for driver input. The decision logic (decision processor) receives one or both of the sensor data and the user interface selection as the decision logic input. The decision logic input represents or is defined by the characteristics of the room (eg, where the loudspeaker cabinet is located) and / or the listening location (eg, the position of the listener in the room or with respect to the loudspeaker cabinet). can do. Content analysis may be performed by decision theory on the input audio channel. Using content analysis, room features (eg, room acoustics), and one or more of the listener's location or listening location, decision logic then makes the selection of rendering mode for the rendering processor. , Accordingly, drive the loudspeaker during playback of the sound program content. The choice of rendering mode can be changed automatically during playback, for example, based on changes in decision logic inputs.
0006The sound rendering mode includes several first modes (eg, center mode) and one or more second modes (eg, direct ambient mode). The rendering processor can be configured in any one of the first modes or in the second mode. In one embodiment, in each of the central modes, the loudspeaker driver (acting collectively as a beam forming array) is a predominantly omnidirectional beam (or beam pattern) superimposed on a directional beam (or beam pattern). ) To generate a sound beam.
0007In ambient direct mode, the loudspeaker driver i) produces a sound beam with a direct content pattern, but the direct content pattern is aimed at the listener's position and ii) away from the listener's position. It is superimposed on the surrounding content pattern that is aimed at. The direct content pattern includes a direct sound segment taken from the input audio channel (eg, a segment containing direct audio, dialogue or commentary, which should be perceived by the listener to come from a particular direction). The ambient content pattern is a segment of ambient or diffuse sound taken from the input audio channel (eg, rainfall or crowd noise that should be perceived by the listener as being around or completely surrounded by the listener. Segments that contain). In one embodiment, the surrounding content pattern is more directional than the direct content pattern, while in other embodiments the opposite is true.
0008By being able to change between multiple first and second modes, the audio system uses a beam forming array, eg, in a single loudspeaker cabinet, to play music (eg, 500Hz). Not only can it be clearly rendered (due to the high directivity of the audio content above the low cutoff frequencies that can be below), but the room is "probably with a low or negative directivity index for ambient content playback". Can "satisfy". Thus, in one example, a single loudspeaker cabinet, for example, for all content above the lower cutoff frequency that is part, but not all, or all of the input audio channels. Can be used to render audio with both clarity and immersiveness.
0009In one embodiment, content analysis is performed on the input audio channel to discover correlated and uncorrelated content, eg, using time correlation / window correlation. Beamformers can be used to render correlated content directly with a content beam pattern, while uncorrelated content is rendered simultaneously with one or more ambient content beams. Knowing the acoustic interaction between the loudspeaker cabinet and the room (which can be partially based on the decision logic inputs that describe the room) can help render the surrounding content. For example, if it is determined that the loudspeaker cabinet is located near an acoustically reflective surface, then the acoustic knowledge of such room can be used to render the sound program content (rather than in one of the central modes). ) Surrounding direct mode can be selected.
0010In other cases of listener position and room acoustics, such as when the loudspeaker cabinet is placed away from any acoustic reflective surface, select one of the central modes to render the sound program content. can do. Each of these can be described as an "extended" omnidirectional mode in which the audio is played consistently over 360 degrees while retaining some spatial quality. Beamformers capable of producing higher order beam patterns, such as dipoles and quadrupoles, can be used, and uncorrelated content (eg, derived from the difference between the left and right input channels) is mainly monaural. It is added to or superposed on the beam (essentially an omnidirectional beam with the sum of the left and right input channels).
0011The above summary does not include an exhaustive list of all aspects of the invention. The present invention includes all feasible systems and methods from all suitable combinations of the various aspects summarized above, as well as those disclosed in the "forms for carrying out the invention" below, in particular. It is considered that those shown in the scope of claims filed with this application are included. Such a combination has certain advantages not specifically described in the above overview.
0012Embodiments of the present invention are shown in the drawings of the accompanying drawings as an example, but not as a limitation, in which the same reference numerals indicate similar elements. It should be noted that references to "some" or "one" embodiments of the invention in the present disclosure do not necessarily refer to the same embodiment, but they refer to at least one embodiment. Also, in order to simplify the drawings and reduce the total number of drawings, a given figure may be used to characterize a plurality of embodiments of the present invention, with all elements of the figure in the given embodiment. Not needed.
0013<figref num="1">FIG. 3 is a block diagram of an audio system with a beam forming loudspeaker array.</figref><figref num="2A">It is an elevation view of the sound beam generated in the center side rendering mode.</figref><figref num="2B">The spatial variation of the rendered audio content is shown in the horizontal plane as a superposition of the sound beams in Figure 2A.</figref><figref num="3A">It is an elevation view of the sound beam pattern generated by the higher-order center side rendering mode.</figref><figref num="3B">The rendered beam content in the embodiment of FIG. 3A is shown for the two input audio channels available to form the beam.</figref><figref num="3C">The spatial changes in the horizontal plane of FIGS. 3A and 3B of the rendered content resulting from the superposition of the beams are shown.</figref><figref num="4">An elevation view of an example of a sound beam pattern generated in direct ambient mode is shown.</figref><figref num="5">It is a downward view with respect to the horizontal plane of the room in which this audio system is operating.</figref>
0014Some embodiments of the present invention will be described below with reference to the accompanying drawings. Whenever the shape, relative position, and other aspects of the parts described in the embodiments are not explicitly defined, the scope of the invention is not limited to the parts shown, the parts shown are merely for illustration purposes. Means that it is for. Also, although many details have been described, it is understood that some embodiments of the invention can be practiced without these details. In other cases, well-known circuits, structures, and techniques are not shown in detail so as not to interfere with the understanding of this description.
0015FIG. 1 is a block diagram of an audio system with a beam forming loudspeaker array used to play sound program content within a large number of input audio channels. Loudspeaker cabinet 2 (also known as an enclosure) contains a large number of loudspeaker drivers 3 (at least three, in most cases more than the number of input audio channels). In one embodiment, the cabinet 2 may have a substantially cylindrical shape, as shown, for example, in FIG. 2A and in the top view of FIG. 5, and the driver 3 has a central vertical axis. They are arranged side by side in the circumferential direction around 9. Other arrangements are possible for driver 3. Further, the cabinet 2 can have other general shapes such as a substantially spherical or substantially ellipsoidal shape in which the driver 3 can be distributed substantially uniformly over the surface of the sphere. The driver 3 may be an electrodynamic driver and may include some specially designed for different frequency bands, including, for example, any suitable combination of tweeter and midrange drivers.
0016The loudspeaker cabinet 2 of this example also includes a number of power audio amplifiers 4, each of which has an output coupled to the drive signal input of the corresponding loudspeaker driver 3. Each amplifier 4 receives an analog input from the corresponding digital-to-analog converter (DAC) 5, where the latter receives its input digital audio signal via the audio communication link 6. DAC5 and amplifier 4 are shown as separate blocks, but in one embodiment it provides more efficient digital / analog conversion and amplification of individual driver signals (eg, using class D amplifier technology). Therefore, the components of these electronic circuits can be combined not only for each driver but also for a plurality of drivers.
0017Individual digital audio signals for each of the drivers 3 are supplied by the rendering processor 7 through the audio communication link 6. The rendering processor 7 is mounted in an enclosure separate from the loudspeaker cabinet 2 (eg, as part of a computing device 18 (see Figure 5), which may be a smartphone, laptop computer, or desktop computer). be able to. In such cases, the audio communication link 6 is more likely to be a wireless digital communication link such as a BLUETOOTH link or a wireless local area network link. However, in other cases, the audio communication link 6 may be via a physical cable such as a digital optical audio cable (eg, TOSLINK connection) or a high definition multimedia interface (HDMI) cable. In another embodiment, the rendering processor 7 and the decision logic 8 are both mounted within the outer enclosure of the loudspeaker cabinet 2.
0018The rendering processor 7 receives two input audio channels for stereo recording, namely, left (L) channel and right (R) channel, for some input audio channels for the sound program content shown in the example in Figure 1. is there. For example, the left and right input audio channels may be channels of music recorded as only two channels. Alternatively, there may be three or more input audio channels, such as, for example, the entire 5.1 surround format audio soundtrack for a movie, or a movie intended for a large public theater setting. These are those in which the rendering processor converts these input channels into individual input drive signals for driver 3 in one of several sound rendering modes of operation and then converts them into sound by driver 3. Is. The rendering processor 7 may be implemented entirely as a programmed digital microprocessor or as a combination of a programmed processor and dedicated hard-wired digital circuits such as digital filter blocks and state machines. The rendering processor 7 can include a beamformer that can be configured to generate the individual drive signals of the driver 3 and multiple simultaneous desired beams emitted by the driver 3 as a beam forming loudspeaker array. As, the audio content of the input audio channel can be "rendered". The beam may be shaped and guided by a beamformer according to several preset rendering modes (as described further below).
0019Rendering mode selection is made by decision theory 8. Decision logic 8 may be implemented as a programmed processor, for example by sharing a rendering processor 7 or by programming a different processor, as to which sound rendering mode to use based on a particular input. The rendering processor 7 drives the loudspeaker driver 3 according to the determined mode (the desired beam during playback of the sound program content). To generate). More generally, the selected sound rendering mode is automatic during playback, based on one or more listener positions, room acoustics, and, as described below, content analysis performed by decision logic 8. Can be changed.
0020Decision logic 8 may automatically change the rendering mode selection during playback (ie, does not require immediate input from the user or listener of the audio system) based on changes in its decision logic input. it can. In one embodiment, the decision logic input includes one or both of the sensor data and the user interface selection. The sensor data can include, for example, measurements taken by an imaging camera such as a proximity sensor, a depth camera, or a directional sound pickup system (eg, one that uses a microphone array). The process of decision logic 8 allows the sensor data and optionally user interface selection (which allows, for example, the listener to manually depict the boundaries of the room, as well as the size and position of furniture or other objects in the room. Can be used to calculate the position of the listener, eg, the radial position given by the angle of the loudspeaker cabinet 2 with respect to the forward or forward axis. Selecting a user interface can indicate the distance from a room feature, eg, a loudspeaker cabinet 2 to an adjacent wall, ceiling, window, or indoor object such as furniture. Sensor data can also be used to measure, for example, acoustic reflection or sound absorption values for a room or some feature in a room. More generally, the decision logic 8 can have the ability to evaluate the interaction between the individual loudspeaker drivers 3 and the room (including digital signal processing algorithms), for example, the loudspeaker cabinet 2. It is possible to determine when it was placed near the acoustic reflecting surface. In such cases, and as described below, the ambient beam (in direct ambient rendering mode) can be placed at different angles to promote the desired stereo effect enhancement or immersive effect.
0021Rendering Processor 7 has several sound rendering operating modes, including two or more central modes and at least one peripheral direct mode. Therefore, since the rendering processor 7 is preset in such an operation mode or has a function of performing beam formation in such an operation mode, the current operation mode is determined by the decision logic 8 during the reproduction of the sound program content. It can be selected and changed in real time. These modes allow the system to select an input audio channel (eg, for example, based on what is expected to give the best or greatest effect to the listener in a particular room and for the particular content being played. It is considered to be a remarkable improvement in stereo effect for L and R). Therefore, an improved stereo effect or indoor immersion can be achieved. Each of the different modes is based not only on the listener's position and room acoustics, but also on the content analysis of the particular sound program content (providing a more immersive stereo effect for the listener). It can be expected to have remarkable advantages (in terms of points). Further, in one embodiment of the present invention, all of the content exceeding the lower cutoff frequency in all the available input audio channels for the sound program content is determined only by the driver 3 of the loudspeaker cabinet 2. It may be selected based on the understanding that it will be converted to sound. The driver is treated as a loudspeaker array by a beamformer that calculates each individual driver signal based on the knowledge of each driver's physical position with respect to the other driver. In other words, except for the woofer and subwoofer content (eg, less than 300Hz), the original audio content in the input audio channel is not sent to another loudspeaker in the system. This is a single loudspeaker cabinet 2 (beam for all content above the lower cutoff frequency)
0022In each of the central modes of the render processor 7, the output of the render processor 7 produces a sound beam with multiple loudspeaker drivers 3 (i) an omnidirectional pattern that includes the sum of two or more input audio channels. This omnidirectional pattern may be ii) a directional pattern with multiple lobes, each lobe superimposed on a directional pattern containing the difference between two or more input audio channels. There is. As an example, FIG. 2A shows a sound beam produced in such a mode for two input audio channels L and R (stereo input). The loudspeaker cabinet 2 produces an omnidirectional beam 10 (having an omnidirectional pattern, as shown) superimposed on the dipole beam 11. The omnidirectional beam 10 can be regarded as a stereo (L, R) original monaural downmix. The dipole beam 11 is an example of a stronger directional pattern, in which case each lobe has two primary lobes that contain the difference between the two input channels L, R but have opposite polarities. In other words, the content output to the right-pointing lobe of the figure is LR, while the content output to the left-pointing lobe of the dipole is-(LR) = RL. To generate such a combination of beams, the rendering processor 7 makes a suitable linear combination of several pre-defined orthogonal modes to generate a superposition of the omnidirectional beam 10 and the dipole beam 11. It can have a beamformer to generate. This combination of beams delivers content within the sector of the overall circle, as shown in FIG. 2B, over the horizontal plane of FIG. 2A where the omnidirectional beam 10 and the dipole beam 11 are drawn. It is a downward view.
0023The resulting or synthesized sound beam pattern shown in Figure 2B is shown here (in the horizontal plane of loudspeaker cabinet 2 and around the central vertical axis 9) over 360 degrees of adjacent stereo sectors shown. It can be said that it has a "stereo density" determined by a number. Each stereo sector is composed of a central region C sandwiched between a left region L and a right region R. Therefore, in the case of the central mode shown in FIG. 2B, the stereo density there is defined only by two adjacent stereo sectors, each with a separate opposite central region C, a single located opposite to each other. It shares the left region L and a single right region R. Each of these stereo sectors, or the content of each of these stereo sectors, is the result of an overlay of the omnidirectional beam 10 and the dipole beam 11, as shown in FIG. 2A. For example, the left region L is obtained as the sum of the LR content in the rightward lobe of the dipole beam 11 and the L + R content of the omnidirectional beam 10, where the quantity L + R is also called C.
0024Another way to look at the dipole beam 11 shown in Figure 2A is to have a lower-order center with only two main lobes or main lobes in the directional pattern, each lobe containing the difference between the same two or three input channels. An example of a side rendering mode, adjacent ones of these main lobes are understood to have opposite polarities. This generalization also covers the specific embodiments shown in FIGS. 3A-3C in which the dipole beam 11 is replaced by a quadrupole beam 13 with four primary lobes in the directional pattern. This is a higher-order beam pattern as compared to the lower-order beam patterns of FIGS. 2A and 2B. In this case, each lobe contains the difference between two or more input channels (in this case only L and R as shown in FIG. 3B), and adjacent ones of the primary lobes have opposite polarities. Thus, looking at FIG. 3B, the forward lobe whose content is RL is adjacent to both the left primary lobe with the reverse polarity LR and the right primary lobe with the opposite polarity LR as well. Similarly, the backward lobe (shown hidden behind the loudspeaker cabinet 2) has a content RL of opposite polarity to its two adjacent lobes (the same left and right lobes with content LR). ..
0025The higher central mode shown in Figures 3A and 3B produces the combination or superposition sound beam pattern shown in Figure 3C, with four adjacent stereo sectors (the central vertical axis in the horizontal plane). There are 360 degrees around 9). As described above, each stereo sector is composed of a central region C sandwiched between the left channel region L and the right channel region R. Similar to FIG. 2B, there is overlap between adjacent sectors and the L region is shared by two adjacent stereo sectors as in the R region. Therefore, in FIG. 3C, there are four sectors corresponding to four adjacent central regions C sandwiched between the L region and the R region, respectively.
0026The above discussion is made by giving examples of the lower central mode (dipole beam 11) of FIGS. 2A and 2B and the higher central mode (quadrupole beam 13) of FIGS. 3A-3C. Expanded to the central mode of Rendering Processor 7. The higher central mode can be considered to have a beam pattern with a higher directivity index or to have more primary lobes than the lower central mode. In other words, the various central modes available in Rendering Processor 7 each produce an increasing order of sound beam patterns.
0027As mentioned above, the selection of rendering mode can be a function of content analysis of the input audio channel as well as the current listener position and room acoustics. For example, if this selection is based on content analysis for sound program content, then the selection of lower or higher directional patterns (one of the available central modes) is ambient or diffuse. It depends on the spectral and / or spatial characteristics of the input audio channel signal, such as the amount of sound (reverberation), the presence of isolated sound sources hard-panned (to the left or right), or the protrusion of vocal content. Such content analysis can be performed during playback, for example at predetermined intervals of 1 second or 2 seconds, for example, by processing the audio signal of the input audio channel. In addition, content analysis can also be performed by evaluating the metadata associated with the sound program content.
0028Note that certain types of diffuse content benefit from being played in a lower central mode that emphasizes the spatial separation of uncorrelated content (indoors). Other types of content, including already strong spatial isolation, such as hard-panned isolated sources, can benefit from higher-order central modes, resulting in a more uniform stereo experience around loudspeakers. In extreme cases, the lowest-order central mode may be a mode in which essentially only the omnidirectional beam 10 is generated without any directional beam such as the dipole beam 11, and the sound content is pure. Can be appropriate if it is monaural. An example of this is when calculating the difference RL (or LR) between two input channels yields essentially zero or very small signal components.
0029With reference to FIG. 4, this figure shows an elevation view of the sound beam pattern generated in an example of the perimeter direct rendering mode. Here, the output of the beamformer in the rendering processor 7 (see Figure 1) causes the loudspeaker driver 3 in the array to generate a sound beam with (i) a direct content pattern (direct beam 15), and this direct content pattern. (Ii) is superimposed on the surrounding content pattern (here, the peripheral right beam 16 and the peripheral left beam 17), which are more directional than the direct content pattern. The direct beam 15 is aimed at a predetermined listener axis 14, whereas the ambient beams 16 and 17 are aimed away from the listener axis 14. The listener axis 14 represents the listener's current position (with respect to the loudspeaker cabinet 2), or the current listening location. The listener's position is calculated by decision logic 8 as an angle to the front axis (not shown) of loudspeaker cabinet 2 using, for example, any suitable combination of its inputs, including sensor data and user interface selection. You may. Note that the direct beam 15 does not have to be omnidirectional and may be directional (similar to each of the ambient beams 16 and 17). Also, certain parameters of ambient direct mode may be variable (eg, beamwidth and angle) depending on the audio content, room acoustics, and loudspeaker placement.
0030Decision logic 8 analyzes the input audio channel, for example, using time window correlation to discover correlated and uncorrelated (or uncorrelated) content within it. For example, the L and R input audio channels can be analyzed to determine how the spacing or segments in the two channels (audio signals) correlate with each other. Such an analysis can reveal that a particular audio segment that effectively appears on both input audio channels is a real "dry" center image, where the dry left channel and the dry right channel are mutually exclusive. It is in phase. In contrast, another segment that is considered more "peripheral" may be detected, and from a correlation analysis point of view, one perimeter segment is less transient than a dry central image and is diff. It also appears in the calculation LR (or RL). As a result, this peripheral segment must be rendered as diffuse sound by the audio system by reproducing such segment only within the directional pattern of the peripheral right beam 16 and the peripheral left beam 17. , 17 are aimed away from the listener so that the audio content in them (called surrounding content or diffuse content) bounces off the walls of the room (see also Figure 1). In other words, the correlated content is rendered on the direct beam 15 (with the direct content pattern), while the uncorrelated content is rendered on, for example, the perimeter right beam 16 and the perimeter left beam 17 (with the perimeter content pattern). ..
0031Another example of ambient content is the reverberation of recorded audio. In this case, decision logic 8 detects a direct audio segment in the input audio channel and then signals the rendering processor 7 to render that segment directly into the beam 15. The decision logic 8 can also detect the reverberation of its direct audio segment, and the segment containing that reverberation is also extracted from the input audio channel, and in one embodiment, the surrounding right beam 16 and the peripheral left beam 17 are then Is rendered only by side-firing (more directional and aimed away from the listener's axis 14). In this way, the direct audio reverberation reaches the listener via an indirect route, thereby providing a more immersive experience for the listener. In other words, the direct beam 15 in that case should not contain the extracted reverberation, but only the direct audio segment, whereas the reverberation is the right beam around the more directional side firing. It is attributed to 16 and the surrounding left beam 17.
0032In summary, one embodiment of the invention is intended to improve playback or playback in a particular room, taking into account the acoustics of the room, the location of the listener, and the direct-to-environmental nature of the content in the original recording. It is a technology that attempts to repackage the original audio recording. The function of decision logic 8 is to determine instructions stored in a machine-readable medium in terms of content analysis, listener position or listening location determination, and room acoustic determination, as well as beamformer performance in the rendering processor 7. It can be achieved by the processor running. A machine-readable medium (eg, any form of solid-state digital memory), along with a processor, may be housed in a separately provided computing device 18 (see room shown in Figure 5), or an audio system. May be housed in the loudspeaker cabinet 2 (see further Figure 1). A processor so programmed receives the input audio channel of the sound program content, for example, by streaming a music or movie file from a remote server over the Internet. In addition, the processor receives one or both of the sensor data and user interface selection that indicates or represents (eg, represents or is defined by) the acoustics of the room or the location of the listener. The processor also performs content analysis on sound program content. For example, one of several sound rendering modes is selected based on the current listener position and room acoustic combination, which causes the loudspeaker array to play the sound program content. This rendering mode can be changed automatically based on changes in the listener's position, room acoustics, or content analysis. This sound rendering mode can include several central modes and at least one peripheral direct mode. In this central mode, the loudspeaker array is each Generates a sound beam pattern of increasing order. In direct ambient mode, the loudspeaker array produces a sound beam with a superposition of the direct content pattern (direct beam) and the ambient content pattern (one or more ambient beams). Content analysis extracts correlated and uncorrelated content from the original recording (input audio channel).
0033In one embodiment, when the rendering processor is configured in ambient direct operating mode, the correlated content is rendered only in the direct beam direct content pattern, while the uncorrelated content is in one or more ambient beams. Rendered only in the surrounding content pattern.
0034When the rendering processor is configured in one of its central operating modes, the low-order directional pattern is selected while the sound program content is primarily surrounding or diffuse, while Higher-order directional patterns are selected when the sound program content primarily contains panned sounds. This selection between different central modes, whether it is a piece of music or an audiovisual work such as a movie film, can occur dynamically during playback of sound program content.
0035The techniques described above are particularly useful when the audio system relies primarily on a single loudspeaker cabinet (accommodating speaker arrays), in which case 500 Hz or less on all input audio channels for sound program content. All content above the cutoff frequency, such as (eg 300Hz), is only converted to sound by the loudspeaker cabinet. This allows a very limited number of loudspeaker cabinets (eg only one) to provide a sophisticated solution to the problem of how to get immersive playback. Is especially desirable for use in small rooms (as opposed to public cinemas and other large acoustic venues).
0036Although some embodiments have been described and illustrated in the accompanying drawings, such embodiments are merely exemplary and not limiting of the invention, and various other modifications It should be understood that the invention is not limited to the particular configurations and arrangements illustrated and described, as can be recalled by those skilled in the art. For example, FIG. 5 shows an audio system as a combination of a computing device 18 and a loudspeaker cabinet 2 in the same room, as well as some furniture and listeners. In this case, there is only one case of loudspeaker cabinet 2 communicating with computing device 18, but in other cases additional loudspeaker cabinets communicating with computing device 18 during playback (eg, for example). There may be woofers and subwoofers) receiving audio content below the cutoff frequency below the loudspeaker array. Therefore, this description should be regarded as exemplary rather than limiting.
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| JP2005198249A | Cites | Japan | Y | Search report | 1-20 |
| WO2016048381A1 | Cites | World Intellectual Property Organization (WIPO) | Y | Search report | 1-20 |
| JPH05153698A | Cites | Japan | Y | Search report | 1-20 |
18 members in 6 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 62402836 | United States of America | – | |
| 201662402836 | United States of America | P | |
| 15593887 | United States of America | – | |
| 201715593887 | United States of America | A |
Members18
| Document | Office | Kind | |
|---|---|---|---|
| EP3301947A1 | European Patent Office (EPO) | A1 | |
| US2018098171A1 | United States of America | A1 | |
| US2018098172A1 | United States of America | A1 | |
| CN107889033A | China | A | |
| KR20180036524A | Republic of Korea | A | |
| US9942686B1 | United States of America | B1 | |
| JP2018061237AThis record | Japan | A | |
| AU2017216541A1 | Australia | A1 | |
| AU2017216541B2 | Australia | B2 | |
| AU2019204177A1 | Australia | A1 | |
| JP6563449B2 | Japan | B2 | |
| US10405125B2 | United States of America | B2 | |
| KR102078605B1 | Republic of Korea | B1 | |
| KR20200018537A | Republic of Korea | A | |
| EP3301947B1 | European Patent Office (EPO) | B1 | |
| CN107889033B | China | B | |
| KR102182526B1 | Republic of Korea | B1 | |
| AU2019204177B2 | Australia | B2 |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 |
Numbers
- Publication
- 2018061237
- Application
- 156885
Titles2
- Japanese
- ビーム形成ラウドスピーカアレイに関する空間オーディオレンダリング
- English
- Spatial audio rendering for beam-forming loudspeaker arrays
Classification
- CPC, 13
- H04R9/06
- H04S7/303
- H04S7/30
- H04R9/02
- H04R2400/11
- H04R1/403
- H04R5/02
- H04R5/04
- H04S3/008
- H04S7/305
- H04S2400/01
- H04S2420/03
- H04S2420/13
- IPC, 2
- H04S7 00
- H04R3 12