Audio gain control using specific-loudness-based auditory event detection
Abstract
This record has no abstract on file.
Term
0.5 yearsto projected expiry
Projected expiry 30 March 2027, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
13 claims: 6 independent, 7 dependent
- 1Zastrzeżenia patentowe 1. Sposób modyfikowania parametru przetwarzania dynamiki dźwięku, obejmujący:wykrywanie zmian w charakterystykach widmowych względem czasu w sygnale dźwiękowym, identyfikowanie zmian granic zdarzenia słuchowego większych niż wartość progowa w charakterystykach widmowych względem czasu we wspomnianym sygnale dźwiękowym, gdzie zdarzenie słuchowe stanowi segment dźwięku pomiędzy kolejnymi granicami, generowanie sygnału sterującego modyfikującego parametr na podstawie wspomnianych zidentyfikowanych granic zdarzenia, a także modyfikowanie parametru przetwarzania dynamiki dźwięku w funkcji sygnału sterującego.
- 2Sposób według zastrz. 1, w którym parametrem jest z:czasu natarcia, czasu zwolnienia i stosunku.
- 3Sposób według zastrz. 1, w którym modyfikowanym parametrem jest stała czasowa wygładzania wzmocnienia.
- 4Sposób według zastrz. 3, w którym stała czasowa wygładzania wzmocnienia jest stałą czasową natarcia wygładzania wzmocnienia.
- 5Sposób według zastrz. 3, w którym stała czasowa wygładzania wzmocnienia jest stałą czasową zwolnienia wygładzania wzmocnienia.
- 6Sposób według dowolnego z zastrz. 1 - 5, w którym sygnał sterujący modyfikujący parametr jest oparty na położeniu wspomnianych zidentyfikowanych granic zdarzenia słuchowego oraz stopniu zmiany w charakterystykach widmowych stowarzyszonych z każdą ze wspomnianych granic zdarzenia słuchowego.
- 7Sposób według zastrz. 6, w którym generowanie sygnału sterującego modyfikującego parametr obejmuje:dostarczenie impulsu na każdej z granic zdarzenia słuchowego, przy czym każdy taki impuls ma amplitudę proporcjonalną do stopnia wspomnianych zmian w charakterystykach widmowych, a także wygładzenie w dziedzinie czasu tego impulsu, tak aby jego amplituda płynnie opadała do zera, otrzymując w ten sposób sygnał sterujący modyfikujący parametr.
- 8Sposób według dowolnego z zastrz. 1 - 7, w którym zmiany w charakterystykach widmowych względem czasu są wykrywane poprzez porównanie różnic w głośności właściwej.
- 9Sposób według zastrz. 8, w którym wspomniany sygnał dźwiękowy jest reprezentowany przez dyskretną sekwencję czasową x[n], która została spróbkowana ze źródła dźwięku z częstotliwością próbkowania fs, zaś zmiany w charakterystykach widmowych względem czasu są obliczane poprzez porównanie różnicy w głośności właściwej N[b,t] w pasmach częstotliwości b pomiędzy kolejnymi blokami czasowymi t.
- 10Sposób według zastrz. 9, w którym różnica w zawartości widmowej pomiędzy kolejnymi blokami czasowymi sygnału dźwiękowego jest obliczana zgodnie z wyrażeniem:-ΐ]| b gdzie
- 11Sposób według zastrz. 9, w którym różnica w zawartości widmowej pomiędzy kolejnymi blokami czasowymi sygnału dźwiękowego jest obliczana zgodnie z wyrażeniem:gdzie H/łORAt [A *]
- 12Urządzenie zawierające środki przystosowane do wykonywania sposobu według dowolnego z zastrzeżeń 1 - 11.
- 13Program komputerowy przechowywany na odczytywanym komputerowo nośniku, przeznaczony do spowodowania, aby komputer wykonywał sposób według dowolnego z zastrzeżeń 1 - 11. Uprawniony:Dolby Laboratories Licensing Corporation Pełnomocnik: mgr inż. Irena Rachubik Rzecznik patentowy Dźwięk I Identyfikuj zdarzenia słuchowe .. ł ΣΖ~ Identyfikuj właściwości zdarzeń słuchowych (opcjonalne) Modyfikuj parametry dynamiki FIG. 5 □ 100 200 300 400 500 600 700 800 900 100O Indeks paczki k FIG. 6 Amplituda Wzmocnienie (dB) Sita zdarzenia Amplituda Wzmocnienie (dB) Amplituda _I-1-1-10.5 1 1.5 2 2.5 3 • Czas (próbki) χ FIG. 9 a) Sygnał oryginalny b) Wzmocnienie DRC c) Sygnał zmodyfikowany wzmocnieniem DRC d) Sygnał kontrolowany zdarzeniowo e) Wzmocnienie DRC kontrolowane zdarzeniowo f) Sygnał modyfikowany przez wzmocnienie DRC kontrolowane zdarzeniowo a) Sygnał oryginalny b) Wzmocnienie DRC c) Sygnał zmodyfikowany wzmocnieniem DRC d) Sygnał kontrolowany zdarzę niowo e) Wzmocnienie DRC kontrolowane zdarzeniowo f) Sygnał modyfikowany przez wzmocnienie DRC kontrolowane zdarzeniowo FIG. 10 Czas (próbki) x 10 Czas (próbki) x 10 wo 200 3CD 400 500 600 CO Czas (bloki) u -10 100 200 300 40D 500 500 CO Czas (bloki) Czas (próbki) x 10 100 200 30D ullD b.nn 600 700 Czas (bloki) DOKUMENTY PRZEDSTAWIONE W OPISIE Ta lista dokumentów przedstawionych przez Zgłaszającego została przyjęta jedynie dla informacji czytającego i nie jest częścią składową europejskiego opisu patentowego. Została ona utworzona z dużą starannością;Europejski Urząd Patentowy nie ponosi jednak żadnej odpowiedzialności za ewentualne błędy i braki. Dokumenty patentowe przedstawione w opisie • WO 2006047600 A1 [0003] [0032] • US 6002776 A, Bhadkamkar [0004] • US 2005038579 W [0032] [0060] [0090] • US SN10474387 A [0084] • US 20040122662 A1 [0084] • US SN10478398 A [0085] • US 20040148159 A1 [0085] • US SN10478538 A [0086] US 20040165730 A1 [0086] US SN10478397 A [0087] US 20040172240 A1 [0087] US 0524630 W [0088] WO 2006026161 A [0088] US 2004016964 W [0089] WO 2004111994 A2 [0089] WO 2006047600 A [0090] Dokumenty niepatentowe cytowane w opisie • Albert S. Bregman. Auditory Scene Analysis--The Perceptual Organization of Sound. Second MIT Press, 1991 [0004] • Limiters and Compressors. Alan Tutton. Audio Engineer's Reference Book.Focal Press,Reed Educational and Professional Publishing, Ltd, [0083] • Brett Crockett ;Michael Smithers. A Method for Characterizing and IdentifyingAudioBased on Auditory SceneAnalysis.Audio EngineeringSociety Convention Paper 6416, 118th Convention, 28 May 2005 [0091] • Brett Crockett. High Quality Multichannel Time Scaling and Pitch-Shifting using Auditory Scene Analysis. Audio Engineering Society Convention Paper 5948, October 2003 [0092] • Alan Seefeldt et al. A New Objective Measure of Perceived Loudness. Audio Engineering Society Convention Paper 6236, 28 October 2004 [0093] • Handbook for Sound Engineers, The New Audio Cyclopedia. Focal Press, 1998, 850-851 [0094] • Limiters and Compressors. Alan Tutton. Audio Engineer's Reference Book. Focal Press, Reed Educational and Professional Publishing, Ltd, 1999, 2.149-2.165 [0095]
Independent claims13
156 paragraphs, as filed
Technical field [0001] The present invention relates to methods and apparatus for dynamic control of the sound range in which the sound processing device analyzes the sound signal and changes the level, gain or dynamic range of the sound, all or some of the parameters of the sound amplification and dynamic processing being generated as auditory event function. The invention also relates to computer programs designed to implement such methods or control such apparatus.
[0002] The present invention also relates to methods and apparatus employing detection based on the specific loudness of auditory events. The invention also relates to computer programs designed to implement such methods or control such apparatus.
State of the art
Audio dynamics processing [0003] Automatic gain control (AGC) and dynamic range control (DRC) techniques are well known and are a common element in many audio signal paths. In short, both techniques involve measuring the level of the audio signal in a certain way and then modifying the signal in terms of gain by a quantity that is a function of the level being measured. In a 1: 1 linear dynamics processing system, the input audio signal is not processed, and the output audio signal ideally matches the input audio signal. In addition, in the case of a sound dynamics processing system that automatically measures the characteristics of the input signal and uses this measurement to control the output signal, if the input signal increases by 6 dB, and the output signal is processed such that it only increases by 3 dB, then it means that the output signal has been compressed in a ratio of 2: 1 relative to the input signal. In international publication number WO 2006/047600 A1 ("Calcula2 ting and Adjusting the Perceived Loudness and / or the Perceived Spectral Balance of an Audio Signal" by Alan Jeffrey Seefeldt) a detailed overview of five basic types of sound processing is given: compression, limitation, automatic gain control (AGC), expansion and gating.
Auditory events and detection of auditory events [0004] The division of sounds into units or segments perceived as separate and different is sometimes referred to as "auditory scene analysis" or "auditory scene analysis" ("ASA") and these segments are sometimes referred to as "auditory events" or "sound events." An extended discussion on the subject of auditory scene analysis was presented by Albert S. Bregman in his book Auditory Scene Analysis - The Perceptual Organization of Sound, Massachusetts Institute of Technology, 1991, fourth edition, 2001, Second MIT Press, paper cover). In addition, US Patent Document No. 6002776 for Bhadkamkar et al. Of December 14, 1999 cites publications dating back to 1976 as "prior art associated with sound separation by auditory scene analysis". However, the patent of Bhadkamkar and others advises against the practical application of auditory scene analysis, arriving at the conclusion that "techniques associated with auditory scene analysis, despite being scientifically interesting as models of processing human hearing, are now too demanding in terms of computing and specialized so that they can be considered practical techniques for sound separation until fundamental progress is made. "
[0005] A suitable method for identifying auditory events has been described by Crockett and Crocket and others in various patent applications and articles listed below under the heading "Attached by Reference". According to these documents, the sound signal is divided into auditory events, each of which tends to be perceived as separate and different, by detecting changes in the spectral composition (amplitude as a function of frequency) versus time. This can be done, for example, by calculating the spectral content of successive sound signal time blocks, calculating the difference in spectral content between successive sound signal time blocks, and also identifying the auditory event boundary when the difference in spectral content between successive time blocks exceeds the threshold value. Alternatively, changes in amplitude versus time can be calculated instead of or in addition to changes in spectral composition over time.
[0006] In its least computationally demanding implementation, this process divides the audio signal into time segments by analyzing the entire frequency band (full-range audio technique) or essentially the entire frequency band (in practical implementations, filtering limiting the band at the ends of the spectrum is often used) and transmission of greatest importance for the loudest sound signal components. This approach uses the advantage of the psychoacoustic phenomenon, in which, on smaller time scales (20 milliseconds (ms) and less), the ear may tend to focus on a single auditory event at a time. This implies that although many events may occur at the same time, one component tends to be the most distinctive in terms of perception and can be processed individually as if it were the only one that takes place. The use of this effect also allows you to scale the detection of the auditory event along with the complexity of the processed sound. For example, if the processed audio input signal is the sound of a single instrument, auditory events that are identified will likely be the individual notes that are played. Similarly for the input voice signal, it is the individual elements of speech, vowels and consonants that are likely to be identified as individual elements of sound. As the complexity of sound increases, such as music with a drum set or numerous instruments and voice, auditory event detection identifies the "most expressive" (i.e. the loudest) element of sound at any given time.
[0007] At the expense of greater computational complexity in the process, changes in the spectral composition versus time can also be taken into account in discrete frequency subbands (fixed or dynamically determined subbands or both dynamically determined and fixed) instead of the full bandwidth. In this alternative approach, more than one audio stream is considered on different frequency subbands, and it is not assumed that only a single stream can be seen at a time.
[0008] Auditory event detection can be accomplished by dividing the time domain waveform into time intervals or blocks, and then converting the data in each block to the frequency domain using either a filter bank or time domain conversion into a frequency domain, such as for example fast Fourier FFT transform. The amplitude of the spectral content of each block can be normalized to eliminate or reduce the effect of amplitude changes.
Each resulting representation in the frequency domain gives information about the spectral content of sound in a given block. The spectral content of subsequent blocks is compared and changes greater than the threshold may be considered as an indication of the instantaneous or temporary termination of the auditory event.
[0009] Preferably, data in the frequency domain is normalized as described below. The extent to which frequency-domain data must be normalized gives information about the amplitude. Hence, if the change exceeds the assumed threshold to this extent, it can also be used as information about the event boundary. Event start and end points resulting from spectrum changes and amplitude changes can be logically added to each other so that event boundaries resulting from both types of changes can be identified.
[0010] Although the techniques described in the Crockett and Crockett and other applications and articles are particularly useful in connection with embodiments of the present invention, other techniques for identifying auditory events and event boundaries can be used in the embodiments of the present invention.
Disclosure of the Invention [0011] Conventional sound dynamics processing techniques known in the art are associated with multiplication of sound by a time-varying control signal that adjusts the gain of the sound, giving the desired result. "Gain" is a scaling factor that scales the amplitude of the sound. This control signal can be generated in a continuous mode or based on blocks of audio data, but is generally obtained by some form of measurement of the processed sound, and the speed of its change is determined by smoothing filters, sometimes with constant characteristics and sometimes with characteristics that change with with sound dynamics. For example, response times can be adjusted depending on changes in volume or sound power. Existing methods, such as Automatic Gain Control (AGC) and Dynamic Range Compression (DRC), for example, do not specify in any psychoacoustically justified way the intervals during which gain changes can be seen as deterioration and when they can be used without introducing audible artifacts. Thus, conventional sound dynamics processes can often introduce audible artifacts, i.e. the effects of dynamic processing can introduce unwanted perceptible changes in sound.
[0012] Auditory scene analysis identifies perceptively discrete sound events, with each event occurring between two successive boundaries of the sound event. Audible deterioration caused by a change in gain can be significantly reduced by ensuring that within a sound event the gain is more constant and by limiting a large portion of the change to the vicinity of the event boundary. In the context of compressors or expanders, the response to an increase in sound level (often called an attack) may be fast, comparable or shorter than the minimum duration of sound events, but the appropriate response to a decrease (release or regeneration) may be slower, so that the sounds that produce permanent or gradually disappearing, can be audibly disturbed. Under certain conditions, it is very beneficial to delay gain recovery until the next limit or to slow down the rate of change of gain during an event. For automatic gain control applications where the medium to long-term sound level or volume is normalized, and both attack and release times can be long compared to the minimum auditory event duration, it is beneficial to delay changes or slow down gain change rates during events. to the next event boundary both for increasing gain and reducing gain.
[0013] According to one embodiment of the present invention, the audio processing system receives the audio signal and performs analysis and changes the gain and / or characteristics of the dynamic range of the sound. The modification of the dynamic range of the sound is often controlled by the parameters of the dynamics processing system (attack and release times, compression ratio and others), which have a significant impact on the perceived artifacts introduced by dynamics processing. Changes in signal characteristics relative to time in the sound signal are detected and identified as sound event boundaries, so that the sound segment between successive boundaries creates a sound event in the sound signal. The characteristics of the audio events of interest may include event properties such as perceived strength or duration. Some of said dynamics processing parameters are generated at least partially in response to sound events and / or the degree of change in signal characteristics associated with the boundaries of the sound event.
[0014] Typically, an audio event is a segment of an audio signal that tends to be perceived as separate or different. One useful measure of signal characteristics includes a measure of the spectral content of an audio signal, for example as described in the cited documents of Crockett and Crockett and others.
All or part of the dynamics processing parameters may be generated at least partially in response to the presence or absence and characteristics of one or more audio events. The boundary of an audio event can be identified as a change in signal properties over time that exceeds the threshold. Alternatively, all or some of the parameters may be generated at least partially in response to a continued measure of the degree of change in signal properties associated with the boundaries of the sound event. Although, in principle, embodiments of the invention may be implemented in the analog and / or digital domain, practical implementations will likely be implemented in the digital domain, in which each of the audio signals is represented by individual samples or samples within data blocks. In this case, the signal properties may be the spectral content of the sound within the block, the detection of changes in signal properties over time may be the detection of changes in the spectral content of the sound from block to block, and the instantaneous start and end limits of the sound event agree with the data block boundary. It should be noted that for a more traditional case of performing dynamic gain changes in sample-by-sample mode, the described audio scene analysis can be performed in block mode and the resulting information about the sound event is used to perform dynamic gain changes that are used in sample-by-sample mode. .
[0015] By controlling key parameters of sound dynamics processing using the results of the sound stage analysis, a radical reduction of audible artifacts introduced by dynamics processing can be obtained.
[0016] The present invention presents two ways of performing sound stage analysis. The first involves performing spectral analysis and identifying the location of perceived audio events that are used to control dynamic gain parameters by identifying changes in spectral content. The second way is to transform the sound into the domain of perceived loudness (which can give information more related to psychoacoustics than it is in the case of the first way) and identifies the location of sound events, which are then used to control dynamic gain parameters. It should be noted that the second way requires that sound processing should take into account absolute levels of acoustic reproduction, which may not be possible in some implementations. The presentation of both methods of sound stage analysis allows for dynamic modification of the gain controlled on the basis of ASA auditory stage analysis using processes or devices that may or may not be calibrated to take into account absolute reproduction levels.
[0017] This document describes aspects of the present invention in a sound dynamics processing environment that includes embodiments of other inventions. Other such inventions are described in various current U.S. and international patent applications of Dolby Laboratories Licensing Corporation, which owns this application, which applications are referred to herein.
Description of the drawings [0018]
Fig. 1 is a flowchart showing an example of processing steps for performing sound stage analysis.
Fig. 2 shows an example of block processing, windowing, and performing the DFT procedure on sound while performing sound stage analysis. Fig. 3 is a flowchart or block diagram showing parallel processing in which sound is used to identify sound events and to identify sound event characteristics, so that events and their characteristics are used to modify dynamics processing parameters.
Fig. 4 is a flowchart or block diagram showing processing in which sound is only used to identify sound events, and event characteristics are determined based on the detection of the sound event, so that these events and their properties are used to modify dynamics processing parameters.
Fig. 5 is a flowchart or block diagram showing processing in which sound is only used to identify sound events, and event properties are determined based on the detection of the sound event and so that only the properties of the sound event are used to modify the dynamics processing parameters .
Fig. 6 is a set of idealized auditory filter response responses that approximate critical banding on the ERB scale. The horizontal scale is the frequency scale expressed in hertz, while the vertical scale represents the level in decibels.
Fig. 7 shows the contours of equal loudness according to ISO 226. The horizontal scale represents the frequency in hertz (logarithmic scale at base 10), and the vertical scale represents the sound pressure level expressed in decibels.
Fig. 8a - c show the idealized input / output characteristics and the input characteristics of the dynamic range compressor input.
Fig. 9a-f shows an example of the use of audio events to control release time in the digital implementation of the traditional Dynamic Range Controller (DRC), in which gain control is obtained based on the mean square signal power.
Fig. 10a-f shows an example of using audio events to control release time in the digital implementation of a conventional Dynamic Range Controller (DRC) in which gain control is obtained based on the mean square power (RMS) of the signal for an alternative signal to that used in Fig. 9.
Fig. 11 shows the corresponding set of idealized AGC and DRC curves for applying AGC followed by DRC in a loudness dynamics processing system. The purpose of this combination is to ensure that all processed sound has approximately the same perceived volume, while still maintaining at least some of the original sound dynamics.
The best way to implement the invention
Analysis of the sound stage (original method not in the field of loudness) [0019] According to an embodiment of one embodiment of the present invention, the auditory stage analysis may be composed of four general stages as shown in part of Figure 1. In the first stage 1-1 ("Perform Spectral Analysis"), a time domain audio signal is taken, divided into blocks, and the spectral profile or spectral content calculated for each block. Spectral analysis transforms the audio signal into a short-term frequency domain. It can be made using any filter bank, or based on transformers or bandpass filter banks, as well as in either linear or curved frequency space (such as the Barka scale or critical band, which better approximates the characteristics of the human ear). There is a trade-off between time and frequency for each filter bank. Higher resolution in the time domain, and thus shorter time intervals, leads to lower frequency resolution. Higher frequency resolution, and thus narrower subbands, leads to longer time intervals.
[0020] In the first step conceptually illustrated in Fig. 1, the spectral content of successive time segments of the audio signal is calculated. In a practical embodiment, the ASA block size may be from any number of samples of the input audio signal, although 512 samples are a good compromise between time and frequency resolution. In the second stage 1-2 the differences in spectral content between successive blocks are determined ("Measure the differences in spectral content"). In the second stage, the difference in spectral content is calculated between successive segments of the audio signal. As stated above, a strong indicator of the beginning or end of a perceived sound event is believed to be a change in spectral content. In the third stage 1-3 ("Identify the location of the boundaries of the sound event"), when the difference in spectral content between one block of the spectral profile and the next is greater than the threshold value, the block boundary is considered the boundary of the sound event. The sound segment between successive boundaries creates a sound event. Thus, in the third stage, the boundary of the sound event is determined between successive time segments when the difference in the spectral content profile exceeds the threshold value, thus defining the sound events. In this embodiment, the sound event boundaries define the sound events having a length that is an integer multiple of the spectral profile blocks with a minimum length of one spectral profile block (512 samples in this example). In fact, event boundaries don't have to be so limited. In an alternative embodiment of the practical embodiments discussed herein, the size of the input block may vary, for example, to have substantially the size of the audio event.
[0021] After identification of the audio event boundaries, the key properties of the audio event are recognized as shown in step 1-4. [0022] Overlapping or non-overlapping audio segments may be windowed and used to calculate the spectral profiles of the input audio signal. Overlapping results in better resolution with respect to the location of the sound events, and also makes it less likely to miss an event such as, for example, a short transient. However, overlapping also increases the complexity of the calculations. Application can therefore be omitted. In fig. 2 presents a conceptual representation of non-overlapping blocks of N samples that are windowed and transformed into the frequency domain by a discrete Fourier transform (DFT).
Each block can be windowed and converted to the frequency domain, for example using the DFT procedure, preferably implemented as a Fast Fourier Transform (FFT) to increase speed.
[0023] The following variables can be used to calculate the spectral profile of the input block:
M = number of windowed samples in the block used to calculate the spectral profile
P = number of spectral samples of the computational tab [0024] In general, any integer can be used as the above variables.
However, implementation will be more efficient if the number M is equal to the power of 2, so that standard FFT procedures for spectral profile calculations can be used. In the practical implementation of the auditory scene analysis process, the listed parameters can be set to values:
M = 512 samples (or 11.6 ms at 44.1 kHz)
P = 0 samples (no bookmark) [0025] The above-mentioned values were determined experimentally and it turned out that they identify with sufficient accuracy the location and duration of auditory events. However, it turned out that setting the P value to 256 samples (50% tab) rather than zero samples (no tab) is useful for identifying some hard-to-detect events. Although many different types of windows can be used to minimize spectral artifacts resulting from windowing, the window used to calculate the spectral profile is a M-point Hanning, Kaiser-Bessel window, or other suitable, preferably non-rectangular window. The above-mentioned values and the Hanning window were selected after extensive experimental analysis, as they showed that they give excellent results in a wide range of sound material. Non-rectangular windowing is preferred for processing audio signals with predominantly low-frequency content. Rectangular windowing produces spectral artifacts that can cause incorrect detection of events. Unlike some coder / decoder (codec) applications, where the general application / adding process must give a certain level of stability, this type of restriction does not take place in this case and the window can be selected for characteristics such as its time / frequency resolution and damming band suppression.
[0026] In step 1-1 (Fig. 1) the spectrum of each M-sample block can be calculated by windowing data using the M-point Hanning window, KaiserBessel or other appropriate window, by converting to the frequency domain using the M-point fast Fourier transform, as well as calculating the size of complex coefficients FFT The resulting data is normalized so that the highest value is reduced to unity, while the normalized array of M-numbers is converted to the logarithmic domain. Data can also be normalized using some other metric, such as the average size value or the average data power value. This table does not have to be converted to the logarithmic domain, but such conversion simplifies the calculation of the differential measure in step 1-2. In addition, the logarithmic domain is better suited to the nature of the human hearing system. The resulting logarithmic domain values range from minus infinity to zero. In practical implementation, a lower limit can be imposed on the range of values. This limit may be constant, for example -60 dB, or it may be frequency dependent to reflect the lower audibility of quiet sounds at low and very high frequencies. (Note that it would be possible to reduce the array size to M / 2 because the fast Fourier FFT transform represents both negative and positive frequencies.)
[0027] In step 1-2, a measure of the difference between the spectra of adjacent blocks is calculated. For each block, each of the M (log) spectrum coefficients from step 1-1 is subtracted from the corresponding coefficient for the preceding block, after which the difference is calculated (the sign is omitted). These M differences are then added to one number. This differential measure can also be expressed as the average difference per spectral coefficient by dividing this differential measure by the number of spectral coefficients used in total (in this case M coefficients).
[0028] In step 1-3, the locations of the auditory event boundaries are identified by applying a threshold value to the table of differential measures from step 1-2 with the threshold value. When the differential measure exceeds the threshold, the change in spectrum is considered sufficient to signal a new event and the block number of that change is recorded as the event boundary. For the above M and P values and logarithmic values (in step 1-1) expressed in dB decibels, the threshold value can be set to 2500 if the whole FFT value is compared (including the mirror part) or 1250 if it compares half the FFT (as noted above, FFT represents both negative and positive frequencies - for the FFT size one is a mirror image of the other). This value was chosen experimentally and gives good detection of auditory event boundaries. This parameter value can be changed to reduce (increase the threshold) or increase (decrease the threshold) event detection.
[0029] The process of Fig. 1 can be represented more generally by the equivalent solutions of Figs. 3, 4 and 5. Fig. 3 shows the audio signal given in parallel to the "Identify auditory events" function or step 3-1 that divides this an audible signal for auditory events, each of which tends to be perceived as separate and different, and for the function "Identify auditory event properties" or step 32. The process of Fig. 1 it can be used to divide the sound signal into auditory events and their properties identified, or some other suitable process can be used. Auditory event information, which can be an identification of an auditory event boundary, determined by a function or step 3-1, is then used to modify sound dynamics processing parameters (e.g., attack, release, factor and the like), as needed, by a function or step 3-3 "Modify dynamics parameters". The optional function or step 3-3 "Identify properties" also receives information about the auditory event. The "Identify properties" function or step 3-3 may characterize some or all of the auditory events by one or more properties. Such properties may include identifying the dominant auditory event subband as described in connection with the process of Fig. 1. The property may also contain one or more sound properties, including, for example, a measure of auditory event power, a measure of auditory event amplitude, a measure of spectral flatness of an auditory event, and whether the auditory event is essentially silent or other properties that help in modifying dynamics parameters such as negative audible processing artifacts are reduced or removed. Properties may also include other properties, such as whether the auditory event contains a transient course. [0030] Different solutions compared to the solution of Fig. 3 are shown in Fig. 4 and
5. Fig. 4 shows an input audio signal that is not fed directly to a function or step 4-3 "Identify properties, but it receives information from a function or step 4-1" Identify auditory events ". The solution of Fig. 1 is a specific example of this type of solution. In Fig. 5, the functions or steps 5-1, 5-2 and 5-3 are arranged in series.
[0031] The details of this practical embodiment are not critical. Other methods of calculating the spectral content of subsequent time segments, calculating the differences between successive time segments, as well as setting auditory event boundaries within appropriate boundaries between successive time segments can also be used when the difference in the spectral content profile between successive time segments exceeds the threshold value.
Auditory scene analysis (new way in the field of loudness) [0032] In an international application filed on October 25, 2005 under the Patent Cooperation Treaty under the number PCT / US2005 / 038579, published as the International Publication No. WO 2006/047600 A1, under the title "Calculating and Adjusting the Perceived Loudness and / or the Perceived Spectral Balance of an Audio Signal" by Alan Jeffrey Seefeldt revealed, among others, an objective measure of perceived loudness based on the psychoacoustic model. Said application is hereby attached by reference in its entirety. As described in this application, the excitation signal E [b, t] is calculated based on the sound signal x [n], which approximates the energy distribution along the basal membrane of the inner ear in the critical band b during the time block t. This stimulation can be calculated from the short-time discrete Fourier transform (STDFT) of the audio signal as follows:
£ [*. Ί] = \ Eib, t -1) - (- 0 * (1) where X [k, t] represents the STDFT transformation of the audio signal x [n] in the time block ti pack k. Note that in equation 1 t, it represents time in discrete units of a transform block as opposed to a continuous unit, such as a second, for example. T [k] represents the frequency response of the filter simulating sound transmission through the outer and middle ear, and Cb [k] represents the frequency response of the basal membrane at a location corresponding to the critical band b. Figure 6 shows the corresponding set of critical band filter responses in which bands are distributed evenly along the scale of the Equivalent Rectangular Bandwidth (ERB) as defined by Moore and Glasberg. Each filter shape is described by a rounded exponential function and the bands are spread using a gap of 1 ERB. The final smoothing time constant 1b in equation 1 can be advantageously selected in proportion to the integration time of human loudness perception within band b.
[0033] Using equal loudness contours, such as those shown in Fig. 7, excitation in each band is transformed to an excitation level that would generate the same perceived loudness at 1 kHz. Then the specific loudness is calculated, a measure of the perceived loudness distributed in frequency and time based on the transformed excitation E1kHz [b, t] by the compression non-linearity. One such function used to calculate the specific loudness N [b, t] is given by the expression:
<img file="PL2011234T3_D0001.tif" />
(2) where TQ1kHz is the silence threshold at 1 kHz, and the white constants are chosen to match the increase in loudness data collected from listening experiments. In short, this transformation based on excitation to specific loudness can be represented by the function y {} so that:
<img file="PL2011234T3_D0002.tif" />
[0034] Finally, the overall loudness L [t], represented in sons, is calculated by adding up the specific loudness in the bands:
(3) [0035] The specific loudness N [b, t] is a spectral representation that is intended to simulate the way in which man perceives sound as a function of frequency and time. It captures changes in sensitivity to different frequencies, changes in level sensitivity as well as changes in frequency resolution. As such, it is a spectral representation well suited to the detection of auditory events. Although computationally more complex, comparing the difference N [b, t] in the bands between successive time blocks can in many cases result in perceptually more accurate detection of auditory events as compared to the direct application of the subsequent FFT spectra described above.
[0036] The said patent application discloses several applications for modifying sound based on this psychoacoustic loudness model. Among them are sharp dynamics processing algorithms such as AGC and DRC. These disclosed algorithms can use auditory events to control various associated parameters. Due to the fact that the specific loudness has already been calculated, it is easily available for determining these events. Details of the preferred embodiment are discussed below.
Control of sound dynamics processing parameters with auditory events [0037] Two embodiments of the present invention will now be presented.
The first describes the use of auditory events to control slow-down time in the digital implementation of the dynamic range controller (DRC), in which gain control is obtained based on the mean square value (RMS) of the signal strength. The second embodiment describes the use of auditory events to control certain forms of the more advanced combination of AGC and DRC implemented in the context of the psychoacoustic loudness model described above. These two embodiments are intended to serve only as examples of the invention and it should be understood that the use of auditory events to control the parameters of the dynamics processing algorithm is not limited to the properties described below.
Dynamic range control [0038] The described digital implementation of DRC segment the audio signal x [n] into windowed, overlapping half blocks, and gain modification is calculated for each block based on the measure of local signal strength and the selected compression curve. The gain is smoothed through the blocks and then multiplied with each block. Modified blocks are finally added with a tab to generate a modified y [n] beep.
[0039] It should be noted that although auditory scene analysis and digital DRC implementation, as described herein, divide the time domain audio signal into blocks for analysis and processing, DRC processing need not be performed using segmentation block. For example, auditory scene analysis can be performed using block segmentation and spectral analysis as described above, and the resulting auditory event locations and properties can be used to provide control information for a digital implementation of a traditional DRC implementation that typically operates on a sample to sample. Hence, however, the same block structure used for auditory scene analysis is used for DRC to simplify the description of their combination.
[0040] Continuing the description of the DRC implementation based on the block approach, the overlapping audio blocks can be represented as:
4 ". d = Mt «W» + tM! 2]<sub>for</sub> <n <M-1 (4) where M is the block length and the jump size is M / 2, in [n] is the window, n is the sample index within the block, and t is the block index (it should be noted that t is used here in the same way as for STDFT in equation 1; it represents time in discrete block units, not in, for example, seconds). In an ideal case, the window in [n] narrows to zero at both ends and adds up to one when it overlaps in half; for example, the commonly used sine wave window meets these criteria.
[0041] For each block, the mean square value (RMS) of power can then be calculated to generate a power measure P [t] in dB per block:
<img file="PL2011234T3_D0003.tif" />
(5) [0042] As mentioned earlier, this power measure can be smoothed by using a fast rake and slow slow down before processing with a compression curve, but as an alternative, the instantaneous power P [t] is processed and then the resulting gain is smoothed. This alternative approach has the advantage that a simple compression curve with sharp knees can be used, and the resulting gains are still smooth as power passes through the knee point. By presenting the compression curve as shown in Fig. 8c as a function of the signal level F that generates the gain, the block gain G [t] is given as:
G [r] = r {P [z]} (6) [0043] Assuming that the compression curve gives more damping as the signal level increases, the gain will fall when the signal is in "attack mode" and it will increase, when it will be in "release mode". Therefore, the smoothed gain G [t] can be calculated according to the expression:
GW = α [ί] · G [i -1} + (1 - a [<]) G [f] (7a) where:
<img file="PL2011234T3_D0004.tif" />
and
G [f] <G [/ —1] G [f] 2: G [fl] (7b) <sup>and</sup>release <sup>>> a</sup>attack [0044] The finally smoothed G [t] gain, which is expressed in dB, is applied to each signal block, and the modified blocks are added with an overlap to produce a modified sound:
y [n + tM 12] = (10<sup>C [</sup>'<sup>)/20</sup> ) * [«, T] + (lO<sup>0</sup>*'<sup>-11</sup>'<sup>20</sup> ) x [n + 4 // 2, / - 1] for 0 <n <M / 2 (8)
It should be noted that due to the fact that the blocks have been multiplied by the narrowing window, as shown in Equation 4, the synthesis with the addition of the overlap presented above effectively smooths the gain through the samples of the processed signal y [n]. Thus, the gain control signal receives smoothing beyond what is shown in Equation 7a. In a more traditional DRC implementation running sample-by-sample rather than block-by-block, you may need to use more advanced gain smoothing than the simple unipolar filter shown in Equation 7a to prevent audible distortion in the processed signal. Also, the use of block processing introduces an inherent delay of M / 2 sample length into the system, and as long as the decay time associated with anatarcia is close to this delay, the signal x [n] does not need to be further delayed before applying amplifications for overcurrent prevention purposes.
[0045] Figures 9a to 9c show the result of applying the described DRC treatment to the audio signal. For this particular implementation, a block length M = 512 with a sampling frequency of 44.1 kHz is used. A compression curve similar to the one in Fig. 8b was used: above -20 dB relative to full scale, the signal is attenuated with a 5: 1 ratio, and below -30 dB the signal is amplified with a 5: 1 ratio. The gain is smoothed with the anatear attack coefficient corresponding to a half decay time of 10 ms and the release release coefficient corresponding to a half decay time of 500 ms. The original sound signal shown in Fig. 9a consists of six successive piano chords, with a final chord located around the sample 1.7 x 10<sup>5</sup>fading to silence. By examining the G [t] gain graph in Fig. 9b, it can be seen that the gain remains close to 0 dB when these six chords are won.
[0046] This is due to the fact that the signal energy remains largely between -30dB and -20 dB, i.e. in an area where the DRC curve does not require modification. However, after the last chord hits, the signal energy drops below 30 dB and the gain begins to increase, eventually beyond 15 dB, as the chord disappears. Fig. 9c shows the resulting modified audio signal and it can be seen that the tail of the final chord is significantly amplified.
[0047] In terms of audibility, this natural amplification with a low fade sound produces an extremely unnatural result. The object of the present invention is to prevent problems of this kind that are associated with a traditional dynamic processor.
[0048] Figs. 10a to 10c show the results of applying exactly the same DRC system to another sound signal. In this case, the first half of the signal consists of a fast, high-level piece of music, followed by approximately a 10 x 10 sample<sup>4</sup> the signal goes to the second fast piece of music, but at a much lower level. By examining the gain in Fig. 6b, it can be seen that this signal is attenuated by approximately 10 dB during the first half, and then the gain increases back to 0 dB during the second half, when a softer portion is played. In this case, the reinforcement behaves as it is desired. It would be desirable for the second fragment to be strengthened relative to the first one and the gain should increase quickly after moving to the second fragment so that it is not audibly overlapping. You can see the gain behavior that is similar to the first signal in question, but here this behavior is desirable. Therefore, it is desirable to consolidate the first case without affecting the second. The use of auditory events to control the release time of this DRC system gives this kind of solution.
[0049] In the first signal that was tested in Fig. 9, the amplification of the last chord fading seems unnatural, since this chord and its fading are perceived as a single auditory event whose consistency appears to be maintained. In the latter case, however, there are many auditory events as the gain increases, which means that a small change is transmitted for each individual event. Therefore, a general change in gain is not so undesirable. Therefore, it can be argued that a change in gain should only be allowed close to the temporary surrounding of an auditory event boundary. This principle can be applied to gain, as long as it is neither attack mode nor release mode, but for most practical DRC implementations, the gain moves so fast in attack mode compared to the instant resolution of human perception that no control is needed. Therefore, you can use events to control smoothing of the DRC gain only when it is in release mode.
[0050] Appropriate release control behavior will now be described. In qualitative terms, if an event is detected, the gain is smoothed with a release time constant, as outlined in Equation 7a above. As time passes beyond the detected event and if no further events are detected, the release time constant increases steadily, so that the smoothed gain eventually becomes "frozen" in place. If another event is detected, then the smoothing time constant is reset to its initial value and the process repeats. In order to modulate the release time, it is possible to generate first a control signal based on the boundaries of the detected event.
[0051] As said before, event boundaries can be detected by looking for changes in successive spectra of the audio signal. In this particular implementation, a discrete Fourier DFT transform of each overlapping block x [n, t] can be calculated to generate STDFT audio signal x [n]:
<img file="PL2011234T3_D0005.tif" />
(9) [0052] Next, the difference between normalized logarithmic amplitude spectra of subsequent blocks can be calculated according to the expression:
<img file="PL2011234T3_D0006.tif" />
(10a) where
<img file="PL2011234T3_D0007.tif" />
(10b) [0053] Here, the maximum of | X [k, t] | is used for normalization among k packs, although other normalization factors can be used. For example, the average of | X [k, t] | among packages. If the difference D [t] exceeds the threshold Dmin, then the event is considered to have occurred. In addition, you can assign a force between zero and one to this event based on the size of D [t] compared to the maximum threshold Dmax. The resulting signal strength A [t] of an auditory event can be calculated as:
(11)
<img file="PL2011234T3_D0008.tif" />
ζ> [/] <£> "" "
- ° ππη <£> [*] <£> max w ».» [0054] By assigning force to an auditory event proportional to the magnitude of the spectral change associated with this event, greater control over dynamics processing is obtained compared to a binary decision. The inventors have noticed that larger gain changes are allowed during events of greater strength, while the signal in Equation 11 allows for this type of variable control.
[0055] Signal A [t] is an impulse signal with an impulse occurring at the location of the event boundary. For the purpose of controlling the release time, you can additionally smooth the signal A [t] so that it fades smoothly to zero when the event boundary is detected. The smoothed control signal A [t] of the event can be calculated from A [t] according to the expression:
<img file="PL2011234T3_D0009.tif" />
<img file="PL2011234T3_D0010.tif" />
(12) otherwise [0056] Here, the aevent parameter controls the decay time of the event control signal. Figures 9d and 10d show the control event signal A [t] for two corresponding audio signals, with a half-life of the smoother signal set to 250 ms.
In the first case, it can be seen that the event boundary is detected for each of the six piano chords, and that the event control signal fades smoothly to zero after each event. For the second signal many events are located very close together in time and therefore the event control signal never completely disappears to zero.
[0057] You can now use the event control signal A [t] to change the release time constant used to smooth the gain. When the control signal is equal to one, the smoothing factor a [t] of equation 7a equals release, as before, and when the control signal is zero, this factor equals one, so that the smoothed gain cannot change. The smoothing factor is interpolated between these two extreme values using a control signal in accordance with:
<img file="PL2011234T3_D0011.tif" />
(13) [0058] By interpolating the smoothing factor continuously as a function of the control signal of the event, the release time is adjusted to a value proportional to the strength of the event when the event is turned on, and then increases smoothly to infinite after the event occurs. The rate of this rise is dictated by the value of the aevent factor used to generate the smoothed event control signal.
[0059] Figures 9e and 10e show the effect of gain smoothing with the event-controlled factor of equation 13 as opposed to the non-event-controlled factor of equation 7b. In the first case, the event control signal drops to zero after the last piano chord, thus preventing gain gain. As a result, the corresponding modified sound of Fig. 9f does not experience the adverse effect of unnatural acceleration of gain drop. In the second case, the control gain signal never drops to zero, and therefore the smoothed gain signal is inhibited very little by using event control. The trajectory of the smoothed gain is almost identical to the non-event-controlled gain of Fig. 10b. This is exactly the desired effect.
AGC and DRC based on loudness [0060] As an alternative to traditional dynamics processing techniques in which signal modifications are a direct function of simple signal measurements such as peak power or mean square RMS, the international patent application number PCT / US2005 / 038579 discloses the use of the previously described a loudness model based on psychoacoustics as a skeleton, which performs dynamics processing. Several advantages are cited. First, measurements and modifications are determined in units of sons, which is a more accurate measure of loudness perception than more basic measures such as peak power or RMS. Secondly, the sound can be modified such that the perceived spectral balance of the original sound is maintained when the overall volume changes. As a result, changes in overall volume become perceptually less pronounced compared to a dynamic processor that uses, for example, broadband gain to modify the sound. Finally, the psychoacoustic model is naturally multi-band and therefore this system can easily be configured to perform multi-band dynamics processing to alleviate the well-known problems associated with inter-spectral pumping associated with a wide-band dynamics processor.
[0061] Although performing dynamic processing in this loudness domain already has several advantages over more traditional dynamic processing, this technique can be further refined by using auditory events to control various parameters. Consider the sound segment containing piano chords as shown in 27a and associated DRCs in Figures 10b and c. Similar DRCs can be made in the loudness domain, and in this case, when the decay loudness of the final piano chord is accelerated, this acceleration will be less pronounced because the spectral balance of the disappearing note will be maintained as it accelerates. However, a better solution is not acceleration of decay at all, and therefore the same principle of controlling attack and release times for auditory loudness events can be advantageously used as previously described for traditional DRC.
[0062] The loudness dynamic processing system that is now being described consists of AGC followed by DRC. The purpose of this combination is to ensure that all processed sound has approximately the same perceived loudness, while still maintaining at least some of the dynamics of the original sound. Figure 11 shows the appropriate set of AGC and DRC curves for this application. It should be noted that the input and output of both curves is represented in sons, since this processing is performed in the loudness domain. The AGC curve attempts to bring the output sound closer to a certain target level and, as mentioned earlier, does so with relatively slow time constants. It can be considered that AGC works so that the long-term sound volume is equal to the target, but in a short time the volume may be subject to significant fluctuations around this goal. Therefore, faster acting DRCs can be used to limit these fluctuations to a certain range considered acceptable for the application. Fig. 11 shows this type of DRC curve, where the AGC target falls within the DRC "zero band", the part of the curve that does not need modification. With this combination of curves, AGC locates the long-term sound volume within the zero band of the DRC curve, so that only minimal fast-acting DRC modifications need to be used. If the short-term volume is still fluctuating outside the zero band, DRC then operates to shift the volume of the sound towards that zero band. As a last general remark, it can be seen that it is possible to use slow-acting AGC, so that all bands of the loudness model receive the same volume of volume modification, thus maintaining the perceived spectral balance, and it is also possible to use fast-acting
DRC in a way that allows the volume modification to change in bands to alleviate inter-band pumping, which may otherwise result from fast-acting volume-independent volume modification.
[0063] Auditory events can be used to control the attack and release of both AGC and DRC. In the case of AGC, both attack and release times are large compared to the instant resolution of event perception, and therefore event control can be used advantageously in both cases. In the case of DRC, the attack is relatively short and therefore event control may only be needed to slow down, as in the case of the traditional DRC described above.
[0064] As said before, the specific loudness spectrum associated with the loudness model used for event detection may be used. Based on the specific loudness N [b, t] defined in equation 2, a differential signal D [t] can be calculated, similar to that in equations 10a and b, according to the following expression:
- ΣΙ ^ λκμϊλ / (¼ - N<sub>WELL</sub>rm [A t ~~ Π) b. (14a) where
<img file="PL2011234T3_D0012.tif" />
[0065] For normalization, the maximum value | N [b, t] | is used here in frequency bands, although other normalization factors may be used, e.g. mean | N [b, t] | in frequency bands. If the difference D [t] exceeds the threshold Dmin, then the event is considered to have occurred. The differential signal can then be processed in the same way as shown in equations 11 and 12 to generate the smooth control signal A [t] used to control the attack and release times.
[0066] The AGC curve shown in Fig. 11 can be represented as a function that takes as its loudness measure its input and generates the desired loudness output:
<img file="PL2011234T3_D0013.tif" />
(15a) [0067] The DRC curve can be similarly represented by:
<img file="PL2011234T3_D0014.tif" />
(15b) [0068] For AGC, input volume is a measure of long-term sound volume. You can calculate this type of measure by smoothing out the instantaneous volume L [t], defined in Equation 3, using relatively long time constants (on the order of a few seconds). It has been shown that when assessing the long-term loudness of sound segments, people give heavier weights to louder fragments than softer ones, and it is possible to use a faster attack than slowdown when smoothing to simulate this effect. In the event of event control being introduced for both attack and release, the long-term loudness used to determine the AGC modification can therefore be calculated according to the following expression:
<img file="PL2011234T3_D0015.tif" />
(16a) where <sup>and</sup>AGC Μ - ^ o<sub>ACCallaeh</sub> + (l - £ [/]) £ [r] <sup>></sup> - ^ AGcil Η
A ^ AGOtlcase + (1 "Yl) - AiceC *"!] (16b) [0069] In addition, the associated long-term specific loudness spectrum can be calculated, which will later be used for multi-band DRC:
<img file="PL2011234T3_D0016.tif" />
(16c) [0070] In practice, it is possible to select smoothing coefficients such that the rake time is approximately half the length of the release time. For a given long-term loudness measure, you can then calculate AGC-related loudness scaling as the ratio of output loudness to input loudness:
<img file="PL2011234T3_D0017.tif" />
(17) [0071] DRC modification can now be calculated based on loudness after applying AGC scaling. Instead of smooth measuring the loudness before applying the DRC curve, it is possible to apply the DRC curve to the instant loudness and then smooth the resulting modification. This is similar to the technique described earlier for smoothing the gain of traditional DRC. In addition, DRC can be used in multi-band mode, which means that DRC is a function of the specific loudness N [b, t] in each b band, not the overall loudness L [t]. However, in order to maintain the average spectral balance of the original sound, it is possible to apply DRC to each band so that the resulting modifications have the same average effect as would result from applying DRC to the overall loudness. This can be achieved by scaling each band by the ratio of long-term overall loudness (after applying AGC scaling) to long-term specific loudness and using this value as an argument in the DRC function. The result is then scaled by the inverse of this ratio to produce an output specific loudness. Thus, DRC scaling in each band can be calculated according to the expression:
<img file="PL2011234T3_D0018.tif" />
(18) [0072] AGC and DRC modifications can then be combined to form total volume scaling per band:
S<sub>TOT</sub> (¼ G <sup>=</sup> (19) [0073] This total scaling can then be smoothed in the time domain independently for each band with a fast attack and slow release and event control only applied to the release. In the ideal case, smoothing is performed on a logarithm of scaling similar to the traditional DRC smoothing in decibel representation, although this is not critical. To ensure that the smoothed total scaling moves synchronously with the specific loudness in each band, the attack and release modes can be determined by smoothing the specific loudness itself:
^ tot G <sup>—</sup> exp (oh <sub>ror</sub> t ^> ^ sC ^ ror IA> 1]) "* Ό <sup>α</sup>τοτ d) ^ sC ^ Tor IA> d)) (20a)
ΛφΜ] = a<sub>RO7</sub>. [fc, Z] TV [^ and ~ 1] + O -e<sub>in</sub>[and, (20b) where
<img file="PL2011234T3_D0019.tif" />
(20c) [0074] Finally, it is possible to calculate the target specific loudness based on the smoothed scaling applied to the original specific loudness,
<img file="PL2011234T3_D0020.tif" />
and then the solution for G [b, t] amplifications, which when applied to the original excitation result in a specific loudness equal to the target:
<img file="PL2011234T3_D0021.tif" />
(21) [0075] Gains can be applied to each band of the filter bank used to calculate the excitation, and the modified sound can then be generated by inverting the filter bank to produce a modified time domain sound signal.
Additional parameter control [0076] Although the above discussion focused on the parameters of AGC and DRC attack and release control via auditory scene analysis of the processed sound, other important parameters may also benefit from control via ASA results. For example, the control signal A [t] of equation 12 events can be used to dynamically adjust the sound gain. The ratio parameter, similarly to the parameters of the attack and release time, can significantly contribute to perceptual artifacts introduced via dynamic gain control.
Implementation [0077] The invention may be implemented on a hardware or software platform or as a combination of both of these platforms (for example, programmable logic matrices). Unless otherwise stated, the algorithms included as part of the invention do not naturally apply to any particular computer or apparatus. In particular, various general-purpose machines can be used along with programs written in accordance with the present disclosure, or it may be more convenient to construct more specialized apparatus (e.g. integrated circuits) designed to perform the required method steps. The invention may therefore be implemented in the form of one or more computer programs executed on one or more programmable computer systems, each of which contains at least one processor, at least one data storage system (including volatile and non-volatile memory and / or memory elements ), at least one input device or port, and at least one output device or port. The program code is applied to the input data to perform the functions described here and generate output information. The output information is applied to one or more output devices in a known manner.
[0078] Any such program can be implemented in any desired programming language (including machine language, assembly language or high-level procedural, logical or object-oriented language) for communication with a computer system. In any case, the language may be a compiled or interpreted language.
[0079] Any such computer program is preferably stored or downloaded to a storage medium or memory device (e.g. semiconductor memory, or also to a magnetic or optical medium) that can be read by a general or special purpose programmable computer for configuration and commissioning computer when this storage medium or device is read by the computer system to perform the procedures described herein. The system of the invention may also be implemented in the form of a computer-readable storage medium, configured using a computer program, said storage medium being configured to cause the computer system to operate in a predefined manner to perform the functions described herein.
[0080] Several embodiments of the present invention have been described herein. However, it should be understood that numerous modifications can be made without departing from the scope of the invention. For example, some of the steps described herein may be independent of the order and may be performed in a different order from that described.
[0081] It should be understood that the implementation of other variations and modifications of the invention and its various forms will become apparent to those skilled in the art, and the invention is not limited by those specific embodiments that have been described. Therefore, it should be considered that the invention covers each and all modifications, changes or equivalent solutions that fall within the scope of the basic principles disclosed and claimed herein.
[0082] The following patents, patent applications and publications disclose additional prior art documents:
Audio dynamics processing [0083] Audio Engineer's Reference Book, edited by Michael Talbot-Smith, edition
II. Limiters and Compressors, Alan Tutton, 2-1492-165. Focal Press, Reed Educational and Professional Publishing, Ltd., 1999.
Detection and use of auditory events [0084] US Patent Application No. 10/474387, "High Quality Time-Scaling and Pitch-Scaling of Audio Signals" to Brett Graham Crockett, published June 24, 2004 as US 2004/0122662 A1.
[0085] US Patent Application No. 10/478398, "Method for Time Aligning Audio Signals Using Characterizations Based on Auditory Events" to Brett G. Crockett et al., Published July 29, 2004 as US 2004/0148159 A1.
[0086] US Patent Application No. 10/478538, "Segmenting Audio Signals Into
Auditory Events "to Brett G. Crockett, published August 26, 2004 as US 2004/0165730 A1. Aspects of the present invention provide a method of detecting auditory events over what has been disclosed in said Crockett's application.
[0087] US Patent Application No. 10/478397, "Comparing Audio Using Characterizations Based on Auditory Events" to Brett G. Crockett et al., Published September 2, 2004 as US 2004/0172240 A1.
[0088] International application under PCT, number PCT / US 05/24630 filed on July 13, 2005, under the title "Method for Combining Audio Signals Using Auditory Scene Analysis," to Michael John Smithers, published March 9, 2006 as WO 2006/026161 .
[0089] International application under PCT, number PCT / US 2004/016964, filed May 27, 2004, under the title "Method, Apparatus and Computer Program for Calculating and Adjusting the Perceived Loudness of an Audio Signal" to Alan Jeffrey Seefeldt et al. , published on December 23, 2004 as WO 2004/111994 A2.
[0090] International application under PCT, number PCT / US2005 / 038579, filed
25 October 2005, titled "Calculating and Adjusting the Perceived Loudness and / or the Perceived Spectral Balance of an Audio Signal" by Alan Jeffrey Seefeldt and published as WO 2006/047600.
[0091] "A Method for Characterizing and Identifying Audio Based on Auditory Scene Analysis," Brett Crockett and Michael Smithers, Audio Engineering Society Convention
Paper 6416, 118th Convention, Barcelona, May 28-31, 2005.
[0092] "High Quality Multichannel Time Scaling and Pitch-Shifting using Auditory Scene Analysis," Brett Crockett, Audio Engineering Society Convention Paper 5948, New York, October 2003.
[0093] "A New Objective Measure of Perceived Loudness", Alan Seefeldt et al., Audio
Engineering Society Convention Paper 6236, San Francisco, October 28, 2004.
[0094] Handbook for Sound Engineers, The New Audio Cyclopedia, edited by Glen M. Ballou, 2nd edition. Dynamics, 850-851. Focal Press an imprint of Butterworth Heinemann, 1998.
[0095] Audio Engineer's Reference Book, edited by Michael Talbot-Smith, 2nd edition, part 2.9 ("Limiters and Compressors", Alan Tutton), pages 2149-2165, Focal Press, Reed Educational and Professional Publishing, Ltd., 1999.
121 members in 22 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 79580806 | United States of America | P | |
| 79580806 | United States of America | P | |
| 07754779 | European Patent Office (EPO) | A | |
| 2007008313 | United States of America | W | |
| 2007008313 | United States of America | W | |
| EP20070754779 | – | – | – |
| US20060795808P | – | – | – |
| WO2007US08313 | – | – | – |
Members121
| Document | Office | Kind | |
|---|---|---|---|
| AU2007243586A1 | Australia | A1 | |
| CA2648237A1 | Canada | A1 | |
| WO2007127023A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200803161A | Taiwan Province of China | A | |
| NO20084336L | Norway | L | |
| NO20161295A1 | Norway | A1 | |
| NO20161296A1 | Norway | A1 | |
| NO20161439A1 | Norway | A1 | |
| NO20180266A1 | Norway | A1 | |
| NO20180271A1 | Norway | A1 | |
| NO20180272A1 | Norway | A1 | |
| NO20190002A1 | Norway | A1 | |
| NO20190018A1 | Norway | A1 | |
| NO20190022A1 | Norway | A1 | |
| NO20190024A1 | Norway | A1 | |
| NO20190025A1 | Norway | A1 | |
| NO20191310A1 | Norway | A1 | |
| EP2011234A1 | European Patent Office (EPO) | A1 | |
| KR20090005225A | Republic of Korea | A | |
| MX2008013753A | Mexico | A | |
| CN101432965A | China | A | |
| IL194430A0 | Israel | A0 | |
| IL194430D0 | Israel | D0 | |
| US2009220109A1 | United States of America | A1 | |
| HK1126902A | Hong Kong, China | A | |
| HK1126902A1 | Hong Kong, China | A1 | |
| JP2009535897A | Japan | A | |
| MY141426A | Malaysia | A | |
| RU2008146747A | Russian Federation | A | |
| AU2007243586B2 | Australia | B2 | |
| EP2011234B1 | European Patent Office (EPO) | B1 | |
| AT493794T | Austria | T | |
| ATE493794T1 | Austria | T1 | |
| UA93243C2 | Ukraine | C2 | |
| DE602007011594D1 | Germany | D1 | |
| KR20110022058A | Republic of Korea | A | |
| DK2011234T3 | Denmark | T3 | |
| AU2011201348A1 | Australia | A1 | |
| RU2417514C2 | Russian Federation | C2 | |
| ES2359799T3 | Spain | T3 | |
| PL2011234T3This record | Poland | T3 | |
| KR101041665B1 | Republic of Korea | B1 | |
| JP2011151811A | Japan | A | |
| BRPI0711063A2 | Brazil | A2 | |
| US8144881B2 | United States of America | B2 | |
| US2012155659A1 | United States of America | A1 | |
| CN101432965B | China | B | |
| CN102684628A | China | A | |
| KR101200615B1 | Republic of Korea | B1 | |
| US2012321096A1 | United States of America | A1 | |
| JP5129806B2 | Japan | B2 | |
| CA2648237C | Canada | C | |
| AU2011201348B2 | Australia | B2 | |
| US8428270B2 | United States of America | B2 | |
| IL194430A | Israel | A | |
| HK1176177A | Hong Kong, China | A | |
| HK1176177A1 | Hong Kong, China | A1 | |
| JP5255663B2 | Japan | B2 | |
| US2013243222A1 | United States of America | A1 | |
| TWI455481B | Taiwan Province of China | B | |
| CN102684628B | China | B | |
| US9136810B2 | United States of America | B2 | |
| US9450551B2 | United States of America | B2 | |
| NO339346B1 | Norway | B1 | |
| US2016359465A1 | United States of America | A1 | |
| US9685924B2 | United States of America | B2 | |
| US2017179900A1 | United States of America | A1 | |
| US2017179901A1 | United States of America | A1 | |
| US2017179902A1 | United States of America | A1 | |
| US2017179903A1 | United States of America | A1 | |
| US2017179904A1 | United States of America | A1 | |
| US2017179905A1 | United States of America | A1 | |
| US2017179906A1 | United States of America | A1 | |
| US2017179907A1 | United States of America | A1 | |
| US2017179908A1 | United States of America | A1 | |
| US2017179909A1 | United States of America | A1 | |
| US9698744B1 | United States of America | B1 | |
| US9742372B2 | United States of America | B2 | |
| US9762196B2 | United States of America | B2 | |
| US9768749B2 | United States of America | B2 | |
| US9768750B2 | United States of America | B2 | |
| US9774309B2 | United States of America | B2 | |
| US9780751B2 | United States of America | B2 | |
| US9787268B2 | United States of America | B2 | |
| US9787269B2 | United States of America | B2 | |
| US9866191B2 | United States of America | B2 | |
| US2018069517A1 | United States of America | A1 | |
| NO342157B1 | Norway | B1 | |
| NO342160B1 | Norway | B1 | |
| NO342164B1 | Norway | B1 | |
| US10103700B2 | United States of America | B2 | |
| US2019013786A1 | United States of America | A1 | |
| US10284159B2 | United States of America | B2 | |
| NO343877B1 | Norway | B1 | |
| US2019222186A1 | United States of America | A1 | |
| NO344013B1 | Norway | B1 | |
| NO344361B1 | Norway | B1 | |
| NO344362B1 | Norway | B1 | |
| NO344363B1 | Norway | B1 | |
| NO344364B1 | Norway | B1 |
Numbers
- Publication, DOCDB
- 2011234
- Publication, EPODOC
- PL2011234T
- Application
- 754779
- Application, DOCDB
- 07754779
- Application, EPODOC
- PL20070754779T
Titles2
- English
- AUDIO GAIN CONTROL USING SPECIFIC-LOUDNESS-BASED AUDITORY EVENT DETECTION
- Polish
- Sposób sterowania wzmocnieniem dźwięku z wykorzystaniem detekcji zdarzeń słuchowych opartej na głośności właściwej
Classification
- CPC, 13
- H03G3/3089
- H03G3/30
- H03G3/3005
- H03G7/007
- H03G9/005
- H03G7/00
- G10L25/51
- H04R3/00
- H03G1/00
- H04R3/04
- H04R2430/03
- G10L21/038
- G10L25/21
- IPC, 4
- H03G3 30
- G10L21 034
- G10L21 0364
- H03G7 00