Hybrid, priority-based rendering system and method for adaptive audio
Abstract
The present invention relates to a hybrid priority-based rendering system and method for adaptive audio. The embodiment is directed to a method of rendering adaptive audio through the following steps: receiving input audio including channel-based audio, audio objects, and dynamic objects, where the dynamic objects are classified into a group of low-priority dynamic objects and a group of high-priority dynamic objects. High-level dynamic objects; rendering channel-based audio, audio objects, and low-priority dynamic objects in the first rendering processor of the audio processing system; and rendering high-priority dynamic objects in the second rendering processor of the audio processing system. The rendered audio then undergoes virtualization and post-processing steps for playback through sound bars and other similar speakers with limited height capabilities.

Term
9.4 yearsleft in the term
Expires 4 February 2036.
- Priority
- Filed
- Granted
- Today
- Expires
18 claims: 2 independent, 16 dependent
- 1一种渲染自适应音频的方法,包括: 接收输入音频,所述输入音频包括静态的基于声道的音频和至少一个动态对象,其中, 所述动态对象基于优先度值被分类为低优先度动态对象或高优先度动态对象,其中,所述 输入音频根据包括音频内容和渲染元数据的基于对象音频的数字比特流格式进行格式化;以及 使用第一渲染处理渲染低优先度动态对象,并使用第二渲染处理渲染高优先度动态对 象, 其中,基于为第一渲染处理和第二渲染处理中的每一个提供的相应的处理能力,第一 渲染处理不同于第二渲染处理, 其中,所述渲染包括基于所述优先度值与优先度阈值的比较来将动态对象分类为低优 先度动态对象或高优先度动态对象,并且其中,所述渲染包括基于所述分类来选择第一渲 染处理或第二渲染处理,并且与所述分类独立地渲染所述基于声道的音频。
- 2如权利要求1所述的方法,其中,所述基于声道的音频包括环绕声音频床,并且所述 输入音频还包括符合中间空间格式的音频对象,并且所述基于声道的音频是使用第一渲染 处理来渲染的。
- 3如权利要求1所述的方法,还包括对渲染的音频进行后处理以便传输到扬声器系统。
- 4如权利要求3所述的方法,其中,后处理步骤包括以下中的至少一个:上混、音量控 制、均衡化、和低音管理。
- 5如权利要求4所述的方法,其中,所述后处理步骤还包括虚拟化步骤,从而促进所述 输入音频中存在的高度提示的渲染以便通过扬声器系统回放。
- 6如权利要求2所述的方法,其中,第一渲染处理是在第一渲染处理器中执行的,所述 第一渲染处理器被优化以渲染基于声道的音频和静态对象;并且 第二渲染处理是在第二渲染处理器中执行的,所述第二渲染处理器被优化以通过第二 渲染处理器相对于第一渲染处理器的提高的性能能力、提高的存储器带宽以及提高的传输 带宽中的至少一个来渲染高优先度动态对象。
- 7如权利要求6所述的方法,其中,第一渲染处理器和第二渲染处理器被实现为通过传 输链路相互耦接的分开的渲染数字信号处理器DSP。
- 8如权利要求1所述的方法,其中,所述优先度阈值由以下中的一个定义:预先设置的 值、用户选择的值、以及自动化处理。
- 9如权利要求1所述的方法,其中,高优先度动态对象能够通过它们各自在对象音频元 数据OAMD比特流中的位置确定。
- 10一种包含指令的非暂时性计算机可读存储介质,所述指令当被处理器执行时执行 根据权利要求1所述的方法。
- 11一种用于渲染自适应音频的系统,包括: 接口,该接口接收比特流中的输入音频,所述比特流具有音频内容以及相关联的元数 据,所述音频内容包括动态对象,其中,所述动态对象基于优先度值被分类为低优先度动态 对象和高优先度动态对象,其中,所述输入音频根据包括音频内容和渲染元数据的基于对 象音频的数字比特流格式进行格式化; 渲染处理器,所述渲染处理器耦接到所述接口并且被配置为渲染所述动态对象,其中, 使用第一渲染处理来渲染低优先度动态对象,并且使用第二渲染处理来渲染高优先度动态 对象, 其中,基于为第一渲染处理和第二渲染处理中的每一个提供的相应的处理能力,第一 渲染处理不同于第二渲染处理, 其中,所述渲染包括基于所述优先度值与优先度阈值的比较来将动态对象分类为低优 先度动态对象或高优先度动态对象,并且其中,所述渲染包括基于所述分类来选择第一渲 染处理或第二渲染处理。
- 12如权利要求11所述的系统,还包括接收基于声道的音频,所述基于声道的音频包括 环绕声音频床,并且所述音频对象符合中间空间格式,并且还包括使用第一渲染处理来渲 染所述基于声道的音频。
- 13如权利要求11所述的系统,其中,所述处理器还被配置为对渲染的音频进行后处理 以便传输到扬声器系统。
- 14如权利要求13所述的系统,其中,所述后处理包括以下中的至少一个:上混、音量控 制、均衡化、和低音管理。
- 15如权利要求14所述的系统,其中,所述后处理还包括虚拟化步骤,从而促进所述输 入音频中存在的高度提示的渲染以便通过扬声器系统回放。
- 16如权利要求11所述的系统,所述渲染处理器包括用于处理第一优先度类型的音频 分量的第一渲染处理器和用于处理第二优先度类型的音频分量的第二渲染处理器,其中所 述第一渲染处理器被优化以渲染低优先度动态对象、基于声道的音频和静态对象,并且 所述第二渲染处理器被优化以通过第二渲染处理器相对于第一渲染处理器的提高的 性能能力、提高的存储器带宽以及提高的传输带宽中的至少一个来渲染高优先度动态对 象。
- 17如权利要求16所述的系统,其中,第一渲染处理器和第二渲染处理器被实现为通过 传输链路相互耦接的分开的渲染数字信号处理器DSP。
- 18如权利要求11所述的系统,其中,所述优先度阈值由以下中的一个定义:预先设置 的值、用户选择的值、以及自动化处理。 CN 111586552 Β
Independent claims18
114 paragraphs in 1 section, as filed
Hybrid priority-based rendering system and method for adaptive audio
[0001] This application is a divisional application of an invention patent application with an application number of 201680007206.4, an application date of February 4, 2016, and an invention title of "hybrid priority-based rendering system and method for adaptive audio" .
[0002] Cross-references to related applications
[0003] This application claims priority to U.S. Provisional Patent Application No. 62/113268 filed on February 6, 2015, which is incorporated herein by reference in its entirety.
Technical field
[0004] One or more implementations relate generally to audio signal processing, and more specifically to a hybrid priority-based rendering strategy for adaptive audio content.
Background technique
[0005] The introduction of digital cinema and the development of real three-dimensional ("3D") or virtual 3D content have created new sound standards, such as the merging of multiple channels of audio to allow content creators to be more creative and enjoy the audience. The auditory experience is more enveloping and more realistic. As a means for distributing spatial audio, it is critical to extend beyond traditional speaker feeds and channel-based audio, and there has been considerable interest in model-based audio descriptions. Model-based audio descriptions allow listeners to choose what they want Playback configuration to render audio specifically for the configuration they choose. The spatial presentation of sound utilizes audio objects, which are audio signals with related parametric source descriptions of apparent source position (for example, 3D coordinates), apparent source width, and other parameters. Further developments include the development of a next-generation spatial audio (also known as "adaptive audio") format that includes a mix of audio objects and traditional channel-based speaker feeds, together with the location metadata of the audio objects . In a spatial audio decoder, the channels are directly transmitted to their associated speakers, or downmixed to an existing speaker group, and the audio objects are rendered by the decoder in a flexible (adaptive) manner. A parametric source description associated with each object (such as a position trajectory in 3D space) is taken as input along with the number and position of speakers connected to the decoder. The renderer then uses certain algorithms (such as the translation law) to The audio associated with each object is distributed on a set of speakers. The creative spatial intention of each object is therefore best presented on the specific speaker configuration present in the listening room.
[0006] The advent of advanced object-based audio has significantly increased the nature of the audio content transmitted to various different speaker arrays and the complexity of the rendering process. For example, a theater soundtrack may include many different sound elements corresponding to images, dialogue, noise, and sound effects emitted from different places on the screen, and combined with background music and environmental effects to create an overall auditory experience. Accurate playback requires that the sound be reproduced in a way that corresponds as closely as possible to the display content on the screen in terms of sound source position, intensity, movement, and depth.
[0007] Although advanced 3D audio systems (such as Dolby® AtmosTM systems) are mostly designed and deployed for cinema applications, consumer-level systems are being developed to bring cinema-level, adaptive audio experiences to the home environment And office environment. Compared with theaters, these environments are subject to obvious constraints in terms of venue size, acoustic characteristics, system power, and speaker configuration. Current professional-grade spatial audio systems therefore need to be suitable for rendering high-level object audio content to listening environments characterized by different speaker configurations and playback capabilities. To this end, certain virtualization technologies have been developed to extend the capabilities of traditional stereo or surround sound speaker arrays, thereby using complex rendering algorithms and technologies (various
Such as content-related rendering algorithms, reflected sound transmission, etc.) to reconstruct spatial sound prompts. Such rendering techniques have led to the development of DSP-based renderers and circuits optimized for rendering different types of adaptive audio content, such as Object Audio Metadata Content (OAMD) Bed and ISF (Intermediate Space Format) objects. Different DSP circuits have been developed to take advantage of the different characteristics of adaptive audio for rendering specific OAMD content. However, such a multi-processor system needs to be optimized for the memory bandwidth and processing capacity of each processor.
[0008] Therefore, there is a need for a system that provides a scalable processor load for two or more processors in a multi-processor rendering system for adaptive audio.
[0009] The increasing use of surround sound and cinema-based audio in the home has also led to the development of different types and configurations of speakers that exceed the standard two-way or three-way upright or bookshelf speakers. Different speakers have been developed to play back specific content, such as soundbar speakers that are part of a 5.1 or 7.1 system. The sound bar refers to a type of speaker in which two or more drivers are juxtaposed in a single housing (speaker cabinet) and are typically arranged along a single axis. For example, a popular sound bar typically includes 4-6 speakers lined up in a rectangular box designed to be mounted on top, below, or directly in front of a television or computer monitor. Transfer out the screen directly. Due to the configuration of the soundbar, certain virtualization technologies may be difficult to implement compared to speakers that provide height cues through physical placement (for example, height drivers) or other technologies.
[0010] Therefore, there is a further need for a system that optimizes adaptive audio virtualization technology for playback through a soundbar speaker system.
[0011] The topic discussed in the background section should not be assumed to be prior art just because it is mentioned in the background section. Similarly, the issues mentioned in the background section or issues associated with the subject of the background section should not be assumed to have been previously recognized in the prior art. The topics in the background section only represent different methods, which can themselves be inventions. Dolby, Dolby TrueHD and Atmos are trademarks of Dolby Laboratories Licensing Corporation.
Summary of the invention
[0012] An embodiment of a method for rendering adaptive audio through the following steps is described: receiving input audio including channel-based audio, audio objects, and dynamic objects, where the dynamic objects are classified as low-priority dynamics A collection of objects and a collection of high-priority dynamic objects; channel-based audio, audio objects, and low-priority dynamic objects are rendered in the first rendering processor of the audio processing system; and the second rendering processor of the audio processing system Medium rendering of high-priority dynamic objects. Input audio can be formatted according to an object audio-based digital bitstream format that includes audio content and rendering metadata. Channel-based audio includes surround sound audio beds, and audio objects include objects that conform to an intermediate space format. Low-priority dynamic objects and high-priority dynamic objects are distinguished by a priority threshold. The priority threshold can be defined by one of the following: the creator of the audio content including the input audio, the value selected by the user, and the automation performed by the audio processing system deal with. In the embodiment, the priority threshold is encoded in the target audio metadata bitstream. The relative priority of the audio objects of the low-priority audio objects and the high-priority audio objects may be determined by their respective positions in the object audio metadata bitstream.
[0013] In an embodiment, the method further includes: during or after channel-based audio, audio objects, and low-priority dynamic objects are rendered in the first rendering processor to generate rendered audio, passing through the first The rendering processor passes the high-priority audio object to the second rendering processor; and post-processes the rendered audio for transmission to the speaker system. The post-processing step includes at least one of the following: upmixing, volume control, equalization, bass management, and a virtualization step for facilitating the rendering of the high prompts present in the input audio for playback through the speaker system.
[0014] In an embodiment, the speaker system includes a soundbar speaker having a plurality of juxtaposed drivers that transmit sound along a single axis, and the first rendering processor and the second rendering processor are embodied in In separate digital signal processing circuits coupled together through transmission links. The priority threshold is determined by at least one of the following: the relative processing capabilities of the first rendering processor and the second rendering processor, and the memory associated with each of the first rendering processor and the second rendering processor Bandwidth and the transmission bandwidth of the transmission link.
[0015] The embodiment is further directed to a method for rendering adaptive audio by the following steps: receiving an input audio bitstream including audio components and associated metadata, each of the audio components having an audio type selected from: Channel audio, audio objects, and dynamic objects; the decoder format of each audio component is determined based on the respective audio type; the priority field of each audio component is determined according to the priority field in the metadata associated with each audio component Priority; rendering the audio component of the first priority type in the first rendering processor; and rendering the audio component of the second priority type in the second rendering processor. The first rendering processor and the second rendering processor are implemented as separate rendering digital signal processors (DSP) coupled to each other through a transmission link. The audio components of the first priority type include low-priority dynamic objects, and the audio components of the second priority type include high-priority dynamic objects. The method further includes rendering channel-based audio and audio in the first rendering processor. Object. In an embodiment, the channel-based audio includes a surround sound audio bed, the audio object includes an object conforming to the Intermediate Space Format (ISF), and the low-priority dynamic object and the high-priority dynamic object include conforming object audio metadata (OAMD) Format object. The decoder format of each audio component produces at least one of the following: OAMD formatted dynamic objects, surround sound audio beds, and ISF objects. The method can further Including at least applying virtualization processing to high-priority dynamic objects to facilitate the rendering of high prompts present in the input audio for playback through a speaker system, and the speaker system may include a bar with multiple juxtaposed drivers that transmit sound along a single axis Speaker speaker.
[0016] The embodiments are further directed to a digital signal processing system that implements the foregoing methods and/or a speaker system that includes circuits that implement at least some of the foregoing methods.
[0017] Incorporation by reference
[0018] Each publication, patent, and/or patent application mentioned in this article is incorporated herein by reference in its entirety, as if each publication and/or patent application were explicitly and individually indicated by reference Incorporate the same degree.
Description of the drawings
[0019] In the following drawings, the same reference numerals are used to refer to the same elements. Although the following figures depict various examples, one or more implementations are not limited to the examples depicted in the figures.
[0020] FIG. 1 illustrates an exemplary speaker placement in a surround system (eg, 9.1 surround) that provides height speakers for playback height channels.
[0021] FIG. 2 illustrates combining channel-based data and object-based data to generate an adaptive audio mix under one embodiment.
[0022] FIG. 3 is a table illustrating the types of audio content processed in a hybrid priority-based system under one embodiment.
[0023] FIG. 4 is a block diagram of a multi-processor rendering system for implementing a hybrid priority-based rendering strategy under one embodiment.
[0024] FIG. 5 is a more detailed block diagram of the multi-processor rendering system of FIG. 4 under one embodiment.
[0025] FIG. 6 illustrates a method of implementing priority-based rendering in order to play back adaptive audio content through a soundbar under one embodiment.
[0026] FIG. 7 illustrates a soundbar speaker that can be used with an embodiment of a hybrid priority-based rendering system.
[0027] FIG. 8 illustrates the use of an adaptive audio rendering system based on priority in an exemplary consumer use case for televisions and sound bars.
[0028] FIG. 9 illustrates the use of an adaptive audio rendering system based on priority in an exemplary full surround sound home environment.
[0029] FIG. 10 is a table illustrating some exemplary metadata definitions in an adaptive audio system that utilizes priority-based rendering for a soundbar, under one embodiment.
[0030] FIG. 11 illustrates an intermediate space format for use with a rendering system under some embodiments.
[0031] FIG. 12 illustrates the arrangement of rings in a stacked ring format translation space for use with an intermediate space format under one embodiment.
[0032] FIG. 13 illustrates the speaker arc of the angle used in the ISF processing system when the audio object is translated under one embodiment.
[0033] FIGS. 14A-C illustrate the decoding of the intermediate space format of the stacked loop under different embodiments.
Detailed ways
[0034] A system and method for a hybrid priority-based rendering strategy are described, in which an object audio metadata (OAMD) bed or an intermediate space format (ISF) object is used for the temporal object audio on the first DSP component The renderer (OAR) component renders, and the OAMD dynamic object is rendered by the virtual renderer in the post-processing chain on the second DSP component. The output audio can be optimized by one or more post-processing and virtualization technologies for playback through the soundbar speakers. The aspects of one or more embodiments described herein may be implemented in an audio or audiovisual system that processes source audio information in a mixing, rendering, and playback system that includes one or more computers or processing devices that execute software instructions . Any of the described embodiments can be used alone or in any combination with each other. Although various embodiments may have been inspired by various deficiencies of the prior art that may be discussed or implied in one or more places in this document, the embodiments may not necessarily solve any of these deficiencies. In other words, different embodiments can solve different deficiencies that may be discussed in this document. Some embodiments may only partially solve some of the defects that may be discussed in this document or only one defect, and some embodiments may not solve any of these defects.
[0035] For the purpose of this description, the following terms have associated meanings: the term "channel" means an audio signal plus metadata, in which the position is encoded as a channel identifier, for example, front left or upper right Surround; "channel-based audio" is audio formatted for playback through a predefined set of speaker zones with associated nominal locations (for example, 5.1, 7.1, etc.); the term "object" or "object-based audio "Means one or more audio channels with parameterized source descriptions such as apparent source position (for example, 3D coordinates), apparent source width, etc.; "adaptive audio" means channel-based and/ Or object-based audio signal plus metadata, which is based on the playback environment, using the audio stream plus metadata in which the position is encoded as a 3D position in space to render the audio signal; and "listening environment" means any open, Partially enclosed or completely enclosed areas, such as rooms that can be used to play back audio content alone or play back audio content and video or other content, and can be embodied in homes, theaters, theaters, auditoriums, studios, game consoles, etc. Such an area may have one or more surfaces disposed therein, such as a wall or baffle that can directly or indirectly reflect sound waves.
[0036] Adaptive audio format and system
[0037] In an embodiment, the interconnection system is implemented as part of an audio system configured to work with a sound format and processing system. The sound format and processing system may be referred to as a "spatial audio system" or an "adaptive audio system". system". Such systems are based on audio formats and rendering technologies to allow for enhanced audience immersion, better artistic control, and system flexibility and scalability. The entire adaptive audio system generally includes an audio encoding, distribution, and decoding system that is configured to generate one or more bits containing both conventional channel-based audio elements and audio object encoding elements flow. Compared with the channel-based method or the object-based method separately, such a combined method provides better coding efficiency and rendering flexibility.
[0038] An exemplary implementation of an adaptive audio system and related audio format is the Dolby® AtmosTM platform. This system contains height (up/down) dimensions that can be implemented as a 9.1 surround system or similar surround sound configuration. Figure 1 illustrates speaker placement in a current surround system (for example, 9.1 surround) that provides height speakers for playback of height channels. 9.1 The speaker configuration of the system 100 consists of five speakers 102 in the floor plane and four speakers 104 in the height plane. Generally speaking, these speakers can be used to generate sounds that are designed to be more or less accurately emitted from any location in the room. Pre-defined speaker configurations (such as those shown in Figure 1) can naturally limit the ability to accurately represent the location of a given sound source. For example, the sound source cannot be panned to be more left than the left speaker itself. This applies to each speaker, thus forming one-dimensional (for example, left and right), two-dimensional (for example, front-rear), or three-dimensional (for example, left and right, front and back, up and down) where downmixing is constrained Geometric shapes. A variety of different speaker configurations and types can be used in such speaker configurations. For example, some enhanced audio systems may use speakers with 9.1, 11.1, 13.1, 19.4, or other configurations. Speaker types can include full-range direct speakers, speaker arrays, surround speakers, subwoofers, tweeters, and other types of speakers.
[0039] Audio objects can be considered as groups of sound elements that can be perceived as being emitted from a specific physical location or multiple physical locations in the listening environment. Such objects can be static (stationary) or dynamic (moving). Audio objects are controlled by metadata that defines the location of the sound at a given point in time and other functions. When objects are played back, they are rendered using existing speakers and based on location metadata, and not necessarily output to a predefined physical channel. Tracks in the session can be audio objects, and standard pan data is similar to position metadata. In this way, content placed on the screen can be effectively panned in the same way as channel-based content, but if necessary, content placed around can be rendered to individual speakers. Although the use of audio objects provides the desired control over discrete effects, other aspects of the soundtrack can work effectively in a channel-based environment. For example, many environmental effects or reverberation actually benefit from being fed to the speaker array. Although these can be seen as objects with enough width to fill the array, it is beneficial to retain some channel-based functions.
[0040] The adaptive audio system is configured to support audio beds in addition to audio objects, where the beds are effectively channel-based sub-mixes or stems. Depending on the intention of the content creator, these can either be delivered separately for final playback (rendering) or combined into a single bed. These beds can be created in different channel-based configurations (such as 5.1, 7.1, and 9.1) and arrays including overhead speakers (such as shown in Figure 1). Figure 2 illustrates combining channel-based data and object-based data to generate an adaptive audio mix under one embodiment. As shown in the process 200, the channel-based data 202 (for example, 5.1 or 7.1 surround sound data provided in the form of pulse code modulation (PCM) data) is combined with the audio object data 204 to generate an adaptive audio mix 208. The audio object data 204 is generated by combining elements of the original channel-based data with associated metadata, which specifies certain parameters related to the location of the audio object. As shown conceptually in Figure 2, the authoring tool provides
The ability of an audio program that is a combination of the device channel group and the target channel. For example, an audio program may contain one or more speaker channels, optionally organized into groups (or tracks, for example, stereo or 5.1 tracks), descriptive metadata for one or more speaker channels, one or more The target channel, and descriptive metadata for one or more target channels.
[0041] In an embodiment, the bed audio component and the object audio component of FIG. 2 may include content that meets a specific formatting standard. Fig. 3 illustrates the types of audio content processed in a hybrid priority-based rendering system under one embodiment. As shown in the table 300 of FIG. 3, there are two main types of content, channel-based content that is relatively static in terms of trajectory and dynamic content that moves between speakers or drivers in the system. Channel-based content can be embodied in the OAMD bed, and the dynamic content is prioritized as OAMD objects of at least two priority levels (low priority and high priority). Dynamic objects can be formatted according to certain object formatting parameters, and are classified into certain types of objects, such as ISF objects. The ISF format is described in more detail later in this description.
[0042] The priority of a dynamic object reflects certain characteristics of the object, such as content type (for example, dialogue vs. effect vs. environmental sound), processing requirements, memory requirements (for example, high bandwidth vs. low bandwidth), and other similar characteristic. In an embodiment, the priority of each object is defined along the scale and is encoded in the priority field, which is included as a part of the bitstream that encapsulates the audio object. The priority can be set as a scalar value, such as an integer value from 1 (lowest) to 10 (highest), or as a binary flag (0 low/1 high) or other similar encoding priority setting mechanism. The priority level is generally set once for each object by the content creator, and the content creator can determine the priority of each object based on one or more of the above-mentioned characteristics.
[0043] In alternative embodiments, the priority level of at least some objects may be set by the user, or may be based on certain runtime criteria such as dynamic processor load, object loudness, environmental changes, system failures, user preferences, Acoustic customization, etc.) to modify the automatic dynamic processing of the default priority level of the object to set.
[0044] In the embodiment, the priority level of the dynamic object determines the processing of the object in the multi-processor rendering system. The encoded priority level of each object is decoded to determine which processor (DSP) of the dual DSP or multi-DSP system will be used to render the specific object. This enables the use of priority-based rendering strategies when rendering adaptive audio content. Fig. 4 is a block diagram of a multi-processor rendering system for implementing a hybrid priority-based rendering strategy under one embodiment. FIG. 4 shows a multi-processor rendering system 400 including two DSP components 406 and 410. These two DSPs are contained in two separate rendering subsystems (decoding/rendering component 404 and rendering/post-processing component 408). These rendering subsystems generally include processing blocks that perform traditional object and channel audio decoding, object rendering, channel remapping, and signal processing before the audio is sent to further post-processing and/or amplification stages and speaker stages.
[0045] The system 400 is configured to render and play back audio content produced by one or more capture components, preprocessing components, authoring components, and encoding components that encode input audio into a digital bitstream 402. The adaptive audio component can be used to analyze the input audio by checking factors such as source interval and content type to automatically generate appropriate metadata. For example, location metadata can be derived from multi-channel recordings by analyzing the relative levels of related inputs between channel pairs. The detection of content types (such as voice or music) can be achieved, for example, through feature extraction and classification. Certain authoring tools allow the creation of audio programs by optimizing the input and organization of the sound engineer's creation intentions, so that he can create a final audio mix optimized for playback in almost any playback environment at once. This can be achieved by using audio objects and location metadata associated with and encoded with the original audio content. Once the adaptive audio content has been authored and encoded in the appropriate codec device, it is decoded and rendered for playback through speakers 414.
[0046] As shown in FIG. 4, object audio including object metadata and channel audio including channel metadata are used as input
The audio bitstream is input to one or more decoder circuits within the decoding/rendering subsystem 404. The input audio bitstream 402 contains data related to various audio components (such as those shown in FIG. 3), including OAMD beds, low-priority dynamic objects, and high-priority dynamic objects. The priority assigned to each audio object determines which of the two DSPs 406 or 410 performs rendering processing for that specific object. OAMD beds and low-priority objects are rendered in DSP 406 (DSP1), while high-priority objects are passed through rendering subsystem 404 for rendering in DSP 410 (DSP 2). The rendered bed, low-priority objects, and high-priority objects are then input to the post-processing component 412 in the subsystem 408 to generate an output audio signal 413, which is transmitted for playback through the speaker 414.
[0047] In the embodiment, the priority level that distinguishes low-priority objects and high-priority objects is set within the priority of the bitstream that encodes the metadata of each associated object. The cutoff value or threshold between low priority and high priority can be set to a value along the priority range, such as a value of 5 or 7 along the priority scale 1 to 10, or for binary priority flag 0 Or 1 simple detector. The priority level of each object can be decoded in the priority determination component within the decoding subsystem 402 to route each object to the appropriate DSP (DPS1 or DSP2) for rendering.
[0048] The multi-processing architecture of FIG. 4 facilitates efficient processing of different types of adaptive audio beds and objects based on the specific configuration and capabilities of the DSP and the bandwidth/processing capabilities of the network and processor components. In the embodiment, DSP1 is optimized to render OAMD bed and ISF objects, but may not be configured to best render OAMD dynamic objects, while DSP2 is optimized to render OAMD dynamic objects. For this application, OAMD dynamic objects in the input audio are assigned a high priority level so that they are passed to DPS2 for rendering, while the bed and ISF objects are rendered in DSP1. This allows the appropriate DSP to render the audio component or audio components that it can render best.
[0049] In addition to or instead of the type of audio component being rendered (eg, bed/ISF object vs. OAMD dynamic object), routing and distributed rendering of audio components can be performed based on certain performance-related metrics, such as based on two The relative processing power of one DSP and/or the bandwidth of the transmission network between two DSPs. Therefore, if one DSP is significantly more powerful than another DSP, and the network bandwidth is sufficient to transmit unrendered audio data, the priority level can be set such that the more powerful DSP is required to render more audio components among the audio components. For example, if DSP2 is much more powerful than DPS1, it can be configured to render all OAMD dynamic objects, or render all objects regardless of the format, assuming it can render these other types of objects.
[0050] In an embodiment, certain application-specific parameters (such as room configuration information, user selections, processing/network constraints, etc.) may be fed back to the object rendering system to allow the object priority level to be dynamically changed. Prior to being output for playback through speakers 414, the priority-ranked audio data is then processed through one or more signal processing stages such as equalizers and limiters.
[0051] It should be noted that the system 400 represents an example of a playback system for adaptive audio, and other configurations, components, and interconnections are also possible. For example, FIG. 3 illustrates two rendering DSPs for processing dynamic objects classified into two types of priority. In order to increase the processing power and priority levels, an additional number of DSPs can also be included. Therefore, N DSPs can be used for N different priority distinctions, such as three DSPs for high, medium, low priority, and so on.
[0052] In an embodiment, the DSPs 406 and 410 shown in FIG. 4 are implemented as separate devices coupled together through a physical transmission interface or a network. Each DSP may be contained in a separate component or subsystem (such as the subsystems 404 and 408 shown), or they may be separate components contained in the same subsystem (such as an integrated decoder/renderer component) . Alternatively, the DSPs 406 and 410 may be separate processing components within a monolithic integrated circuit device.
[0053] Exemplary Implementation
[0054] As mentioned above, the initial implementation of the adaptive audio format is to include content capture (object and channel) in digital video
In the context of the Institute, the content capture is created using novel creative tools, encapsulated using adaptive audio cinema encoders, and using PCM or using the existing Digital Cinema Initiative (DCI) distribution mechanism. There are lossless codecs distributed. In this case, the audio content is intended to be decoded and rendered in a digital cinema to create an immersive spatial audio cinema experience. However, it is now imperative to deliver the enhanced user experience provided through adaptive audio formats directly to consumers at home. This requires certain characteristics of the format and system to be suitable for use in a more restricted listening environment. For the purpose of description, the term "consumer-based environment" is intended to include any non-theatre environment, including listening environments for ordinary consumers or professionals, such as houses, studios, rooms, console areas, auditoriums, etc.
[0055] Current authoring and distribution systems for consumer audio create and deliver audio intended to be reproduced to a predefined and fixed speaker location, while the essence of the audio (ie, the actual audio played back by the consumer reproduction system) ) Has limited knowledge of the type of content conveyed. However, the adaptive audio system provides a new hybrid method for audio creation, which includes audio specific to a fixed speaker location (left channel, right channel, etc.) and based on generalized 3D spatial information including position, size, and speed. Options for both the audio elements of the object. This hybrid method provides a method of rendering (generalized audio objects) fidelity (provided by fixed speaker locations) and flexibility. The system also provides additional useful information about audio content via new metadata that is paired with the audio essence by the content creator at the time of content creation/creation. This information provides detailed information about the properties of the audio that can be used during rendering. Such attributes may include content type (for example, dialogue, music, effects, dubbing, background/environment, etc.) as well as audio object information and useful rendering information such as spatial attributes (for example, 3D position, object size, speed, etc.) (For example, alignment to speaker location, channel weight, gain, bass management information, etc.). Audio content and reproduction intent metadata can be created either manually by the content creator or by using automatic media intelligence algorithms, which can be created during the creation process. The time runs in the background and can be reviewed by the content creator during the final quality control phase, if needed.
[0056] FIG. 5 is a block diagram of a priority-based rendering system for rendering different types of channel-based components and object-based components, and is a more detailed illustration of the system shown in FIG. 4 according to an embodiment. As shown in FIG. 5, the system 500 processes an encoded input bitstream 506 that carries both the mixed object stream(s) and the channel-based audio stream(s). The bitstream is processed by rendering/signal processing blocks as indicated by 502 and 504, both of which are represented or implemented as separate DSP devices. The rendering functions executed in these processing blocks implement various rendering algorithms for adaptive audio and certain post-processing algorithms (such as upmixing).
[0057] The priority-based rendering system 500 includes two main components, a decoding/rendering stage 502 and a rendering/post-processing stage 504. The input bitstream 506 is provided to the decoding/rendering stage through HDMI (High Definition Multimedia Interface), but other interfaces are also possible. The bitstream detection component 508 parses the bitstream and directs different audio components to an appropriate decoder, such as a Dolby Digital Plus (Dolby Digital Plus) decoder, a MAT 2.0 decoder, a TrueHD decoder, and so on. The decoder generates various formatted audio signals, such as OAMD bed signals and ISF or OAMD dynamic objects.
[0058] The decoding/rendering stage 502 includes an OAR (Object Audio Renderer) interface 510, and the 04-port interface 510 includes an OAMD processing component 512, an OAR component 514, and a dynamic object extraction component 516. The dynamic object extraction component 516 obtains outputs from all decoders, and separates the bed, the ISF object, and any low-priority dynamic objects and high-priority dynamic objects. The bed, ISF object, and low-priority dynamic object are sent to the OAR component 514. For the example embodiment shown, the OAR component 514 represents the core of the processor (for example, DSP) circuit of the decoding/rendering stage 502, and renders to a fixed 5.1.2 channel output format (for example, the standard 5.1+2 Height channel), but other surround sound plus height configurations are also possible, such as 7.1.4. The rendering output 513 of the OAR component 514 is then transmitted to the digital audio processor (DAP) component of the rendering/post-processing stage 504. Should
The stage performs functions such as the following: upmixing, rendering/virtualization, volume control, equalization, bass management, and other possible functions. In an example embodiment, the output 522 of the rendering/post-processing stage 504 includes a 5.1.2 speaker feed. The rendering/post-processing stage 504 may be implemented as any suitable processing circuit, such as a processor, DSP, or similar device.
[0059] In an embodiment, the output signal 522 is transmitted to the sound bar or sound bar array. For specific use case examples such as those shown in Figure 5, the soundbar also utilizes a priority-based rendering strategy to support use cases with 31.1 object MAT 2.0 inputs without overlapping the memory bandwidth between the two stages 502 and 504 . In an exemplary implementation, the memory bandwidth allows up to 32 audio channels to be read and written from external memory at 48 kHz. Because 8 channels are required for the 5.1.2-channel rendering output 513 of the OAR component 514, a maximum of 24 OAMD dynamic objects can be rendered by the virtual renderer in the rendering/post-processing stage 504. If there are more than 24 OAMD dynamic objects in the input bitstream 506, the additional lowest priority objects must be rendered by the OAR component 514 on the decoding/rendering stage 502. The priority of dynamic objects is determined based on their position in the OAMD stream (for example, the highest priority object comes first, and the lowest priority object last).
[0060] Although the embodiments of FIGS. 4 and 5 are described with respect to beds and objects that conform to OAMD and ISF formats, it should be understood that priority-based rendering schemes using a multi-processor rendering system may be compatible with channel-based rendering schemes. Audio is used with any type of adaptive audio content of two or more types of audio objects, where the object types can be distinguished based on relative priority levels. An appropriate rendering processor (eg, DSP) may be configured to optimally render all types or only one type of audio object types and/or channel-based audio components.
[0061] The system 500 of FIG. 5 illustrates a rendering system that adapts the OAMD audio format to work with specific rendering applications that involve channel-based beds, ISF objects, and OAMD dynamic objects, and are targeted at bars. The playback of the speakers is rendered. The system implements a priority-based rendering strategy, which solves some implementation complexity problems of reconstructing adaptive audio content through a soundbar or similar co-located speaker system. FIG. 6 is a flowchart illustrating a method of implementing priority-based rendering in order to playback adaptive audio content through a soundbar under one embodiment. The process 600 of FIG. 6 generally represents the method steps performed in the priority-based rendering system 500 of FIG. 5. After receiving the input audio bitstream, audio components including channel-based beds and audio objects in different formats are input to an appropriate decoder circuit for decoding, 602. Audio objects include dynamic objects that can be formatted using different formatting schemes, and can be distinguished based on the relative priority of encoding with each object, 604. The process determines the priority level of each dynamic audio object compared to the defined priority threshold by reading the appropriate metadata field in the bitstream for each dynamic audio object. The priority threshold that distinguishes low-priority objects from high-priority objects can be programmed into the system as a hard-wired value set by the content creator, or it can be through user input, automated means, or other adaptive It should be set dynamically by the mechanism. The channel-based bed and low priority dynamic objects are then rendered in the first DSP of the system, 606 along with any objects optimized to be rendered in the first DSP of the system. High priority dynamic objects are passed along to the second DSP, where they are then rendered, 608. The rendered audio components are then transmitted through certain optional post-processing steps for playback through a soundbar or soundbar array, 610.
[0062] Implementation of sound bar
[0063] As shown in FIG. 4, the prioritized rendered audio output generated by the two DSPs is transmitted to the sound bar for playback to the user. Considering the popularity of flat screen TVs, soundbar speakers have become more and more popular. Such TV sets have become very thin and relatively light to optimize portability and installation options, despite providing ever-increasing screen sizes at affordable prices. However, considering space, power, and cost constraints, the sound quality of these TV sets is usually very poor. Sound bars are usually stylish power-on speakers that are placed under a flat-screen TV to improve the quality of the TV's audio, and can be used alone or as part of a surround sound speaker setup. Figure 7
Illustrated is a soundbar speaker that can be used with an embodiment of a hybrid priority-based rendering system. As shown in the system 700, the sound bar speaker includes a cabinet 701 that accommodates a number of drivers 703, and the drivers 703 are arranged along a horizontal (or vertical) axis to drive sound directly out of the front of the cabinet. Any actual number of drives 703 can be used according to size and system constraints, and a typical number is in the range of 2-6 drives. The drivers can be the same size and shape, or they can be an array of different drivers, such as a larger central driver for lower frequency sounds. The HDMI input interface 702 may be provided to allow direct interface with a high-definition audio system.
[0064] The soundbar system 700 may be a passive speaker system without onboard power and amplification and with minimal passive circuitry. It can also be a power-on system in which one or more components are installed in a cabinet or tightly coupled via external components. Such functions and components include power supply and amplification 704, audio processing (eg, EQ, bass control, etc.) 706, A/V surround sound processor 708, and adaptive audio virtualization 710. For the purpose of description, the term "driver" means a single electroacoustic transducer that generates sound in response to an electric audio input signal. The driver can be implemented in any suitable type, geometry, and size, and can include horns, paper cones, ribbon transducers, and the like. The term "speaker" means one or more drivers within an integral housing.
[0065] The virtualization function provided in the component 710 of the soundbar 700 or as a component of the rendering/post-processing stage 504 allows for adaptation in local applications such as televisions, computers, game consoles or similar devices An audio system, and allows spatial playback of the audio through speakers arranged in a plane corresponding to the viewing screen or monitor surface. Figure 8 illustrates the use of a priority-based adaptive rendering system in an exemplary consumer use case for televisions and sound bars. Generally speaking, based on speaker locations/configurations that may be limited in terms of spatial resolution (ie, no surround or rear speakers) and the generally reduced quality of equipment (TV speakers, soundbar speakers, etc.), TV use cases provide The challenge of creating an immersive consumer experience. The system 800 of FIG. 8 includes speakers (TV-L and TV-R) at the left and right locations of a standard television, and possibly optional left-up excitation drivers and right-up excitation drivers (TV-LH and TV-RH). ). The system also includes a sound bar 700 as shown in FIG. 7. As mentioned earlier, compared to standalone or home theater speakers, the size and quality of TV speakers are reduced due to cost constraints and design choices. However, the combined use of dynamic virtualization and soundbar 700 can help overcome these shortcomings. The sound bar 700 of FIG. 8 is shown as having forward excitation Generator drivers and possibly lateral excitation drivers, all of which are arranged along the horizontal axis of the soundbar cabinet. In FIG. 8, the dynamic virtualization effect is exemplified for a soundbar speaker so that a person at a specific listening position 804 will hear horizontal elements associated with appropriate audio objects individually rendered in the horizontal plane. The height element associated with the appropriate audio object can be rendered by dynamic control of the speaker virtualization algorithm parameters based on the object spatial information provided by the adaptive audio content in order to provide at least a partial immersive user experience. For co-located speakers of a soundbar, this dynamic virtualization can be used to create the perception of objects moving along the side of the room or other horizontal plane sound trajectory effects. This allows the soundbar to provide spatial cues that would otherwise not exist due to the lack of surround or rear speakers.
[0066] In an embodiment, the soundbar 700 may include non-collocated drivers, such as upward excitation drivers that utilize sound reflections to allow virtualization algorithms to provide a high degree of prompting. Certain drivers may be configured to radiate sound to other drivers in different directions, for example, one or more drivers may implement steerable sound beams with individually controlled sound zones.
[0067] In an embodiment, the sound bar 700 may be used as part of a full surround sound system with height speakers or floor mounted speakers that enable height. Such an implementation will allow the virtualized soundbar to expand the immersive sound provided by the surround speaker array. Figure 9 illustrates the use of an adaptive audio rendering system based on priority in an exemplary full surround sound home
Use in a courtyard environment. As shown in system 900, a soundbar 700 associated with a television or monitor 802 is used in conjunction with a surround sound array of speakers 904, such as in the 5.1.2 configuration shown. For this case, the soundbar 700 may include an A/V surround sound processor 708 to drive the surround speakers and provide at least part of the rendering and virtualization processing. The system of Figure 9 only illustrates a possible set of components and functions that can be provided by an adaptive audio system, and certain aspects can be reduced or removed based on the needs of the user, while still providing an enhanced experience.
[0068] FIG. 9 illustrates the use of dynamic speaker virtualization to provide an immersive user experience in addition to the immersive user experience provided by the soundbar in the listening environment. A separate virtualizer can be used for each related object, and the combined signal can be sent to the L speaker and R speaker to create a multi-object virtualization effect. As an example, dynamic virtualization effects are shown for L speakers and R speakers. These speakers can be used to create a diffuse or point source near-field audio experience along with audio object size and location information. A similar virtualization effect can also be applied to any or all of the other speakers in the system.
[0069] In an embodiment, the adaptive audio system includes a component that generates metadata from the original spatial audio format. The methods and components of system 500 include an audio rendering system configured to process one or more bitstreams containing both conventional channel-based audio elements and audio object encoding elements. A new extension layer containing audio object coding elements is defined and added to any one of the channel-based audio codec bitstream or audio object bitstream. This method enables the implementation of a bitstream that includes an extension layer that will be processed by the renderer for use in existing speaker and driver designs or next-generation speakers defined with individually addressable drivers and drivers. The spatial audio content from the spatial audio processor includes audio objects, channels, and location metadata. When an object is rendered, it is assigned to one or more drivers of the soundbar or soundbar array based on the location metadata and the location of the playback speakers. Metadata is generated in the audio workstation in response to engineers mixing input to provide rendering queues that control spatial parameters (eg, position, speed, intensity, timbre, etc.) and specify which driver(s) or speakers in the listening environment Play their respective sounds during the presentation. The metadata is associated with the respective audio data in the workstation for packaging and transportation of the spatial audio processor. Figure 10 is an example of the use of soundbars based on advantages in one embodiment A table of some exemplary metadata definitions used in the first rendered adaptive audio system. As shown in the table 1000 of FIG. 10, some metadata may include elements that define audio content types (for example, dialogue, music, etc.) and certain audio characteristics (for example, directness, diffusion, etc.). For a priority-based rendering system played through a soundbar, the driver definition included in the metadata can include the playback soundbar and other speakers that can be used with the soundbar (for example, other surround speakers or virtualization-enabled Speaker) configuration information (for example, drive type, size, power, built-in A/V, virtualization, etc.). Referring to Figure 5, metadata may also include fields and data that define the decoder type (for example, Digital+, TrueHD, etc.). From these fields and data, channel-based audio and dynamic objects (for example, OAMD bed, ISF object) can be derived , Dynamic OAMD objects, etc.). Alternatively, the format of each object can be clearly defined by specific associated metadata elements. The metadata also includes a priority field for dynamic objects, and the associated metadata can be expressed as a scalar value (for example, 1 to 10) or a binary priority flag (high/low). The metadata elements shown in FIG. 10 are intended to illustrate only some possible metadata elements that are encoded in the bitstream that transmits the adaptive audio signal, and many other metadata elements and formats are also possible.
[0070] Intermediate Space Format
[0071] As described above for one or more embodiments, certain objects processed by the system are ISF objects. ISF is a format that optimizes the operation of the audio object panner by dividing the pan operation into the following two parts: a time-varying part and a static part. Generally speaking, the audio object shifter moves the monophonic object (for example, Objecti) to N
CN 111586552 Β
Each speaker is operated, therefore, the translation gain is determined according to the function of the speaker location (Xl, yrZi),..., (Xνν, zν) and the target location XYZj (t). These gain values will continuously change over time because the location of the object will be time-varying. The goal of the intermediate space format is only to divide the translation operation into two parts. The first part (which will be time-varying) uses the object location. The second part (which uses a fixed matrix) will only be configured based on the speaker location. Figure 11 illustrates an intermediate space format for use with a rendering system under some embodiments. As shown in the diagram 1100, the spatial translator 1102 receives the object and speaker location information for the speaker decoder 1106 to decode. Between these two processing blocks 1102 and 1106, the audio object scene is represented by K-channel Intermediate Space Format (ISF) 1104. Multiple audio objects (l<=i<=Nj can be processed by a separate spatial shifter, and the outputs of the spatial shifter are added together to form an ISF signal 1104, so that a K-channel ISF signal set can contain X objects In some embodiments, the encoder can also be given information about the height of the speaker through elevation restriction data, so that detailed knowledge of the elevation of the playback speaker can be used by the spatial translator 1102.
[0072] In an embodiment, the spatial translator 1102 is not given detailed information about the location of the playback speaker. However, it is assumed that the locations of a series of "virtual speakers" are limited to several levels or layers and the distribution within each level or layer is approximate. Therefore, although the spatial translator is not given detailed information about the location of the playback speakers, some reasonable assumptions can usually be made about the approximate number of speakers and the approximate distribution of these speakers.
[0073] The quality of the resulting playback experience (ie, how close it matches the audio object translator of FIG. 11) can be achieved either by increasing the number of channels K, or by gathering more knowledge about the most likely playback speaker placement To improve. Specifically, in the embodiment, as shown in FIG. 12, the height of the speaker is divided into several planes. The desired composition sound field can be thought of as a series of sounding events emitted from any direction around the listener. The location of the vocal event can be considered to be defined on the surface of the sphere 1202 centered on the listener. The sound field format (such as High Order Ambisonics) is defined in a way that allows the sound field to be further rendered on (quite) arbitrary speaker arrays. However, in the sense that the height of the speaker is fixed in 3 planes (ear height plane, ceiling plane, and ground), the envisaged typical playback system may be constrained. Therefore, the concept of an ideal spherical sound field can be modified, where the sound field is composed of sound-producing objects in rings located at various heights on the surface of a sphere around the listener. For example, one such arrangement 1200 is illustrated in FIG. 12, which has a vertex ring, an upper ring, a middle ring, and a lower ring. If necessary, for completeness purposes, you can also include an additional ring at the bottom of the sphere (the bottom point, strictly speaking, it is also a point rather than a ring). In addition, there may be more or fewer loops in other embodiments.
[0074] In the embodiment, the stacked ring format is named BH9.5.0.1, where four numbers indicate the number of channels in the middle ring, upper ring, lower ring, and vertex ring, respectively. The total number of channels in the multi-channel bundle will be equal to the sum of these four numbers (so, the BH9.5.0.1 format contains 15 channels). Another example format using all four rings is BH15.9.5.1. For this format, the channel naming and ordering will be as follows: [Ml, M2, One M15, 111, 112-1191112, -25, 21], where the channels are arranged in a ring (in order of l1), and in each Within the ring, they are simply numbered in ascending cardinal order. Each ring can be thought of as being filled with a set of nominal speakers spread evenly around the ring. Therefore, the channels in each ring will correspond to a specific decoding angle, starting from channel 1 (which will correspond to 0. Azimuth (front)) and enumerating in counterclockwise order (so from the listener From the perspective, channel 2 will be on the left of the center). Therefore, the azimuth of channel n will be equal to X360. (Where n is the number of channels in the ring, and n is in the range from 1 to n).
[0075] Regarding certain use cases of object_priority related to ISF, OAMD generally allows each ring in the ISF to have an object.priority value. In the embodiment, these priority values are used in multiple ways to perform additional processing. head
First, the height loop and the lower plane loop are rendered by the smallest/suboptimal renderer, while the important listener plane loop can be rendered by a more complex/higher precision high-quality renderer. Similarly, in the encoding format, more bits (ie, higher quality encoding) can be used for the listener plane ring, and fewer bits can be used for the height ring and the terrestrial ring. This is possible in ISF because it uses rings, which is generally not possible in traditional high-end high-fidelity stereo formats, because each of the different channels interact in a way that compromises the overall audio quality Polar-pattern. Generally speaking, a slight reduction in the rendering quality of the height ring or the ground ring is not excessively harmful, because the content in these rings usually contains only atmospheric content. [0076] In an embodiment, the rendering and sound processing system uses two or more rings to encode a spatial audio scene, where different rings represent different spatially separated components of the sound field. The audio object is translated within the ring according to a translation curve that can be used for conversion, and the audio object is translated between the rings using a translation curve that cannot be used for conversion. The different spatially separated components are separated based on their vertical axis (ie, as vertically stacked rings). The sound field elements are transmitted in the form of "nominal speakers" in each ring; and the sound field elements in each ring are transmitted in the form of spatial frequency components. For each ring, a decoding matrix is generated by concatenating pre-calculated sub-matrices representing the segments of the ring. If there are no speakers in the first ring, the sound from one ring to the other can be redirected.
[0077] In the ISF processing system, the location of each speaker in the playback array can be expressed by coordinate (x, y, z) coordinates (which is the location of each speaker relative to the candidate listening position near the center of the array). In addition, the (x, y, z) vector can be converted to a unit vector to effectively project each speaker location onto the surface of the unit sphere:
[0078] Speaker location: dagger=% to" (1)
I! <sub>=</sub> . x V
[0079] Speaker unit vector:-Yasi
[0080] FIG. 13 illustrates a speaker arc where the audio object is translated to the angle used in the ISF processing system under one embodiment. Diagram 1300 illustrates a scenario where the audio object (0) is sequentially translated through several speakers 1302 so that the listener 1304 experiences the illusion that the audio object is moving through a trajectory sequentially passing through each speaker. Without loss of generality, it is assumed that the unit vectors of these speakers 1302 are arranged along a ring in the horizontal plane, so that the location of the audio object can be defined as a function of its azimuth angle Φ. In Figure 13, the audio object passes through speakers A, B, and C at an angle Φ (where these speakers are placed at azimuth angles 6a, 6b, and 6C, respectively). An audio object translator (eg, translator 1102 in FIG. 11) will typically use speaker gain to translate the audio object to each speaker, where speaker gain is a function of angle 6. The audio object translator can use a translation curve with the following properties: (1) When the audio object is translated to a position that coincides with the physical speaker location, the coincident speaker is used to exclude all other speakers; (2) When the audio object is shifted When panning to the angle Φ between the two speaker locations, only these two speakers are working, thus providing the least amount of "spreading" of the audio signal on the speaker array; (3) The panning curve can show a high level of "discrete" sex", "Discreteness" refers to the part where the energy of the translation curve is constrained in the area between a loudspeaker and its nearest neighbor. Therefore, referring to Figure 13, for speaker B:
CN 111586552 B
[0081] Discreteness:<sup>d</sup>B = get: one". (3)
[0082] Therefore, cIbWI, and when clb=1, this implies that the translation curve for speaker B is only in the area between 6A and 6c (the angular positions of speakers A and C, respectively) (in space (Above) is completely constrained to be non-zero. On the contrary, translation curves that do not exhibit the above-mentioned "discreteness" property (ie, c!b<1) can exhibit one other important property: the translation curves are spatially smoothed, so that they are constrained in the spatial frequency. , In order to satisfy the Nyquist sampling theorem.
[0083] Any translation curve that is limited in space cannot be compact in its spatial support. In other words, these translation curves will spread over a wide range of angles. The term "stopband fluctuation" refers to the (undesirable) non-zero gain that appears in the translation curve. By satisfying the Nyquist sampling theorem, these translation curves have the problem of not being too "discrete". By being properly "Nyquist sampling", these translation curves can be moved to alternative speaker locations. This means that a set of loudspeaker signals that have been created for a specific arrangement of ν loudspeakers (the loudspeakers are evenly spaced in the circle) can be remixed to an alternative set of N loudspeakers at different angular locations (recreated with an nXΝ matrix) Hybrid); that is, the speaker array can be rotated to a new set of angular speaker locations, and the original N speaker signals can be converted into the new set of N speakers. Generally speaking, this "re-useable" property allows the system to remap N speaker signals to S speakers through the SXN matrix, provided that for the case of S>N, the new speaker feed is no longer more than the original n Channel "discrete" is acceptable.
[0084] In the embodiment, the intermediate space format of the overlapping ring represents each object according to the (time-varying) (x, y, z) location of each object through the following steps:
[0085] 1. Place the object i at X, yj, zj, and assume that the location is in the cube (so | <sub>Yi</sub> K
1 and T zj W1) or within the unit sphere (+ yj + <= 1 ).
[0086] 2. Use the vertical location (zj) to translate the audio signal of the object i to each of several (R) spatial regions according to a translation curve that is not convertible.
[0087] 3. Represent each spatial region (ie region r: lWrWR) in the form of * nominal speaker signals (according to the figure
4. It represents the audio components located in the annular area of the space), the X nominal speaker signals are created using a transformable translation curve, and the transformable translation curve is the azimuth (6) of the object i function.
[0088] Note that for the special case of a ring with a size of zero (according to Figure 12, the vertex ring), the above step 3 is unnecessary, because the ring will contain at most one channel.
[0089] As shown in FIG. 11, the ISF signal 1104 for K channels is decoded in the speaker decoder 1106. Figures 14A-C illustrate the decoding of the intermediate space format of the overlapping ring under different embodiments. Figure 14A illustrates that the stacked ring format is decoded into individual rings. Fig. 14B illustrates a stacked loop format decoded without a vertex speaker. Figure 14c illustrates the stacked loop format decoded without vertex speakers or ceiling speakers.
[0090] Although the embodiments are described above with respect to the ISF object as a type of object in comparison with the dynamic 0AMD object, it should be noted that audio objects that are formatted in different formats but can be distinguished from the dynamic 0AMD object can also be used. [0091] The aspects of the audio environment described herein represent the playback of audio or audio/visual content through appropriate speakers and playback devices, and can represent any environment in which the listener is experiencing the playback of captured content, such as a theater , Concert hall, amphitheatre, home or room, listening booth, car, game console, earphone or headset system, public place
Address (PA) system or any other playback environment. Although the embodiments have been described mainly in terms of examples and implementations in a home theater environment in which spatial audio content is associated with TV content, it should be noted that the embodiments can also be implemented in other consumer-based systems, such as games, screenings System and any other monitor-based A/V system. Spatial audio content including object-based audio and channel-based audio can be used in combination with any related content (associated audio, video, graphics, etc.), or it can constitute independent audio content. The playback environment can be any suitable listening environment from earphones or near-field monitors to small or large rooms, cars, amphitheaters, concert halls, etc.
[0092] The various aspects of the system described herein can be implemented in a suitable computer-based processing network environment for processing digital or digitized audio files. The various parts of the adaptive audio system may include one or more networks including any desired number of individual machines, including one or more routers (not shown) for buffering and routing data transmitted between the computers. Such a network can be built on a variety of different network protocols, and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof. In embodiments where the network includes the Internet, one or more machines may be configured to access the Internet through a web browser program.
[0093] One or more of the components, blocks, processing, or other functional components may be implemented by a computer program that controls the execution of a processor-based computing device of the system. It should also be noted that with regard to the behaviors, register transfers, logic components, and/or other characteristics of the various functions disclosed herein, these functions can use hardware, firmware, and/or be included in various machine-readable or computer-readable Read any combination of data and/or instructions in the medium to describe. The computer-readable media that may contain such formatted data and/or instructions include, but are not limited to, various forms of physical (non-transitory) non-volatile storage media, such as optical, magnetic, or semiconductor storage media.
[0094] Unless the context clearly requires otherwise, throughout and in the claims, the words "including", "including", etc., shall be interpreted in an inclusive sense that is completely different from the exclusive or exhaustive sense; that is Said, in the sense of "including but not limited to" to explain. Words using the singular or plural number also include the plural or singular number respectively. In addition, the words "herein", "in the following", "above", "below" and words of similar meaning refer to the entire application, rather than to any specific part of the application. When the word "or" is used when referring to a list of two or more items, the word encompasses all the following interpretations of the word: any item in the list, all items in the list, and the words in the list Any combination of items.
[0095] The term "one embodiment", "some embodiments" or "embodiments" throughout this specification means that a specific feature, structure, or characteristic described in combination with the embodiment is included in the disclosed system (one or more A) and method(s) in at least one embodiment. Thus, the appearances of the phrases "in one embodiment", "in some embodiments" or "in an embodiment" in various places throughout this specification may refer to the same embodiment, or may not necessarily refer to the same implementation example. In addition, the specific features, structures or characteristics can be combined in any suitable manner understood by those of ordinary skill in the art.
[0096] Although one or more implementations have been described with respect to specific embodiments by way of example, it is understood that one or more implementations are not limited to the disclosed embodiments. On the contrary, the intention is to cover various modifications and similar arrangements that are apparent to those skilled in the art. Therefore, the scope of the appended claims should be given the broadest interpretation so as to encompass all such modifications and similar arrangements.
1 sheet
Sheet 1
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| CN102576533A | Cites | China | Y | Search report | 1-19 |
| WO2015009748A1 | Cites | World Intellectual Property Organization (WIPO) | A | Search report | 1-19 |
| US2005093839A1 | Cites | United States of America | A | Search report | 1-19 |
| WO2013112564A1 | Cites | World Intellectual Property Organization (WIPO) | A | Search report | 1-19 |
| EP1724684A1 | Cites | European Patent Office (EPO) | A | Search report | 1-19 |
| CN101223778A | Cites | China | A | Search report | 1-19 |
| CN102549655A | Cites | China | A | Search report | 1-19 |
| 《Audio-pro with multiple DSPs and dynamic load distribution》;B Vercoe;《BT Technology Journal》;20041004;第1,5-6部分,附图3 | Non-patent | – | – | Search report | – |
30 members in 5 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 201562113268 | United States of America | P | |
| 201562113268 | United States of America | P | |
| 62113268 | United States of America | – | |
| 201680007206 | China | A | |
| 201680007206 | China | A | |
| 2016800072064 | – | – | – |
| 62113268 | – | – | – |
| CN20168007206 | – | – | – |
| CN2016807206 | – | – | – |
| US201562113268P | – | – | – |
Members30
| Document | Office | Kind | |
|---|---|---|---|
| WO2016126907A1 | World Intellectual Property Organization (WIPO) | A1 | |
| CN107211227A | China | A | |
| EP3254476A1 | European Patent Office (EPO) | A1 | |
| US2017374484A1 | United States of America | A1 | |
| JP2018510532A | Japan | A | |
| US10225676B2 | United States of America | B2 | |
| US2019191258A1 | United States of America | A1 | |
| US10659899B2 | United States of America | B2 | |
| CN107211227B | China | B | |
| JP6732764B2 | Japan | B2 | |
| CN111556426A | China | A | |
| CN111586552A | China | A | |
| JP2020174383A | Japan | A | |
| EP3254476B1 | European Patent Office (EPO) | B1 | |
| US2021112358A1 | United States of America | A1 | |
| EP3893522A1 | European Patent Office (EPO) | A1 | |
| CN111586552BThis record | China | B | |
| US11190893B2 | United States of America | B2 | |
| JP7033170B2 | Japan | B2 | |
| CN111556426B | China | B | |
| CN114374925A | China | A | |
| JP2022065179A | Japan | A | |
| US2022159394A1 | United States of America | A1 | |
| CN114554386A | China | A | |
| CN114554387A | China | A | |
| EP3893522B1 | European Patent Office (EPO) | B1 | |
| US11765535B2 | United States of America | B2 | |
| JP7362807B2 | Japan | B2 | |
| CN114374925B | China | B | |
| CN114554386B | China | B |
4 legal events, as 2 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Patent grantGrantedGR01 | GR01 | CN | |
| Requests to designate patent in hong kongDE | DE | HK | |
| Entry into force of request for substantive examinationSE01 | SE01 | CN | |
| PublicationPB01 | PB01 | CN |
Numbers
- Publication
- 111586552
- Publication, DOCDB
- 111586552
- Publication, EPODOC
- CN111586552B
- Application
- 2020104531452
- Application, DOCDB
- 202010453145
- Application, EPODOC
- CN202010453145
Titles2
- Chinese
- 用于自适应音频的混合型基于优先度的渲染系统和方法
- English
- Hybrid priority-based rendering system and method for adaptive audio
Classification
- CPC, 11
- H04S3/008
- H04R1/403
- H04R5/02
- H04S7/302
- G10L19/008
- G10L19/167
- G10L19/20
- H04S2420/03
- H04R27/00
- H04S2400/11
- H04R2499/13
- IPC, 7
- H04S3 00
- H04S7 00
- G10L19 008
- G10L19 16
- G10L19 20
- H04R1 40
- H04R5 02