Method, apparatus and computer program for calculating and adjusting the perceived loudness of an audio signal
Abstract
One or a combination of two or more specific loudness model functions selected from a group of two or more of such functions are employed in calculating the perceptual loudness of an audio signal. The function or functions may be selected, for example, by a measure of the degree to which the audio signal is narrowband or wideband. Alternatively or with such a selection from a group of functions, a gain value G[t] is calculated, which gain, when applied to the audio signal, results in a perceived loudness substantially the same as a reference loudness. The gain calculating employs an iterative processing loop that includes the perceptual loudness calculation.
Term
Term ended
Projected expiry passed 27 May 2024, 2.3 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
21 claims: 17 independent, 4 dependent
- 1Zastrzeżenia patentowe 1. Sposób przetwaraania sygnału akustyzznego . oeejmujący tworzenie w odpowiedzi na sygnał akustyczne sygnału wzbudzenia i obliczanie percepcyjnej głośności teyo sygnału akustycznego w odpowiedzi na sygnał wzbudzenia oraz miary właściwości sygnału akustycznego, przy czym obliczanie takie wybiera z grupy co najmniej dwsch funkcji modelu głośności właściwej jedną lub kombinację dwsch lub więcej funkcji modelu głośności właściwej, ktsrych wybieranie jest regulowane poprzez miarę właściwości wejściowego sygnału akustycznego.
- 2Sposób . 1 . w h;S5^^m miara właściwość i łu akustycznego jest miara stopnia widmowego spłaszczenia sygnału wejściowego.
- 3Sposób według . 1 . w t^sr^m wymienione oblizzanie wybiera z dwsch lub łączy dwie funkcje modelu głośności właściwej, przy czym pierwsza funkcja modelu głośności jest wybierana przez miarę właściwości wynikających z sygnału wejściowego, ktSry nie jest widmowo płaski, druga funkcja modelu głośności jest wybierana przez miarę właściwości wynikających z wejściowego sygnału płaskiego widmowo, a połączenie pierwszej i drugiej funkcji modelu głośności jest wybierane przez miarę właściwości wynikających z sygnału wejściowego częściowo niepłaskiego, a częściowo płaskiego widmowo.
- 4Sposób według zat^z. 3 . w zarówno piewvsza ja. i druga funkcja modelu głośności rośnie monotonicznie powyżej progu ciszy z rosnącym wzbudzeniem według funkcji potęgowej, przy czym pierwszy - 40 funkcja modelu głośności rośnie szybciej niż druga funkcja modelu głośności.
- 5Sposób według zasrrz . 1 . w którym wymienione oblizzanie wybiera z grupy dwóch lub więcej modeli głośności właściwej jeden lub kombinację dwóch lub więcej wymienionych modeli głośności właściwej w każdym z odpowiednich pasm częskokliwości sygnału wzbudzenia.
- 6Sposób według zasto . 1 . w któiym wymienione oblizzanie wybiera z grupy dwóch lub więcej modeli głośności właściwej jeden lub kombinację dwóch lub więcej wymienionych modeli głośności właściwej w grupie odpowiednich pasm częskokliwości sygnału wzbudzenia.
- 7Sposób według zasEz. 6 . w kUrym odpowiednie pasma zzęskokliwościowe grupy są wszyskkie pasmami częskokliwościowymi sygnału wzbudzenia.
- 8Sposób wedługzaklTz. 1 . w k^óó^^m miara właściwość i sygnału akuskycznego jesk pobierana z sygnału wzbudzenia.
- 9Sposób według zaktre . 1 . w Μ0ιυγπ obliczanie zawiera obliczanie głośności właściwej w każdym z odpowiednich pasm częskokliwości sygnału wzbudzenia.
- 10Sposób według z^sózz . 9 . w któym oblizzanie obejmuje ponadko wybieranie głośności właściwej pasma częskokliwości, aby ukworzyć głośność percepcyjną lub łączenie głośności właściwej grupy pasm częskokliwości, aby ukworzyć głośność percepcyjną.
- 11Spoóób według zak^z . 1 . w Μ0ιυγπ wytwazzanie w odpowiedzi na sygnał akuskyczny sygnału wzbudzenia obejmuje:liniowe filkrowanie sygnału akuskycznego przez funkcję lub funkcje, kkóre symulują właściwości ludzkiego ucha zewnękrznego i środkowego, by ukworzyć liniowo filkrowany sygnał akuskyczny oraz dzielenie liniowo filkrowanego sygnału akuskycznego na pasma częskokliwości, kkóre symulują rozkład wzbudzenia wykwarzany wzdłuż błony podskawowej w uchu wewnękrznym, by wykworzyć sygnał wzbudzenia. - 41
- 12Sposób według któregokolwiek z poprzednich z^asir^eż^eń, obejmujący ponadto obliczanie, w odpowiedzi co najmniej na sygnał wzbudzenia, wartości wzmocnienia G[t] , przy czym obliczanie to obejmuje iteracyjną pętlę przetwarzania, która zawiera ustawianie wartości bezwzględnej sygnału wzbudzenia w odpowiedzi na wartość Gi iteracyjnego wzmocnienia tak, że regulowana wartość bezwzględna sygnału wzbudzenia zwiększa się ze zwiększającymi się wartościami Gi, a maleje z malejącymi wartościami Gi, porównywanie obliczonej odczuwalnej głośności sygnału akustycznego z odczuwalną głośnością odniesienia, by wytworzyć różnicę oraz regulowanie wartości wzmocnienia Gi w odpowiedzi na tę różnicę tak, aby zmniejszyć różnicę pomiędzy obliczoną odczuwalną głośnością a odczuwalną głośnością odniesienia.
- 13Sposób według zastrz. 12, w którym sygnał wzbudzenia jest wygładzany w czasie i/lub sposób ten zawiera ponadto wygładzanie w czasie wartości wzmocnienia G[t].
- 14Sposób według zastrz. 13, w którym sygnał wzbudzenia jest wygładzany w czasie liniowo.
- 15Sposób według zastrz. 13, który ponadto obejmuje wygładzanie wartości wzmocnienia G[t] wykorzystujące technikę histogramową.
- 16Sposób według 12, w którym pętla przetwarzania iteracyjnego według algorytmu minimalizowania odpowiednio reguluje wartość bezwzględną sygnału wzbudzenia, oblicza odbieraną głośność, porównuje obliczoną odbieraną głośność z odbieraną głośnością odniesienia i reguluje wartość wzmocnienia Gi na końcową wartość G[t].
- 17Sposób według zastrz. 16, w którym algorytm minimaiizacji jest zgodny z metodą minimalizowania ze spadkiem gradientu.
- 18Sposób według któregokolwiek z zastrz. 12 - 17, obejmujący ponadto - 42 regulowanie amplitudy wejściowego sygnału akustycznego ze wzmocnieniem G[t] tak, że wynikowa odbierana głośność wejściowego sygnału akustycznego jest zasadniczo taka sama jak głośność odniesienia.
- 19Sposób według któregokolwiek z zastrz. 12 - 18, w którym głośność odniesienia jest ustawiana przez użytkownika.
- 20Urządzenie zawierające środki dostosowane do przeprowadzania każdego z etapów sposobu według któregokolwiek z zastrz. 1 - 19.
- 21Program komputerowy, przechowywany na nośniku czytelnym dla komputera, do powodowania realizowania przez komputer każdego z etapów sposobu według któregokolwiek z zastrz. 1 - 19, kiedy program komputerowy jest wykonywany na komputerze. Tomasz Jabłkowski Rzecznik patentowy Sygnał wyjściowy o regulowanym wzmocnieniu -46FIG.4 Tłumienie Częstotliwość (Hz) FIG. 5 Częstotliwość (ERB)
Independent claims21
198 paragraphs, as filed
Technical field
This invention relates to loudness measurements of acoustic signals and to devices, methods and computer programs for controlling the volume of acoustic signals in response to such measurements.
State of the art
Loudness is a subjectively perceived attribute of the feeling of listeners, through which the sound can be set on a scale from quiet to loud. Because volume is a sensation perceived by the listener, it is not suitable for direct physical measurement, which makes it difficult to quantify. In addition, due to the perceptive component of loudness, different listeners with "normal" hearing may receive the same sound differently. The only way to reduce the changes introduced by individual perception and create an overall measure of loudness of the acoustic material is to gather a group of listeners and statistically determine or assess loudness. This is clearly an impractical approach to standard, daily volume measurements.
There have been many attempts to develop a satisfactory objective method of measuring loudness. Fletcher and Munson stated in 1933 that human hearing is less sensitive at low and high frequencies than at middle (or voice) frequencies. They also stated that
- The 2 relative change in sensitivity decreased as the sound level increased. The early volume meter consisted of a microphone, amplifier, meter and set of filters designed for rough simulation of hearing response at low, medium and high sound levels.
Although such devices provide a measurement of the volume of a single, isolated tone at a constant level, measurements of more complex sounds do not very well match the subjective sensations of volume. Sound level meters of this type have been standardized, but are only used for specific tasks such as monitoring and controlling industrial noise.
In the early 1950s, Zwicker and Stevens, among others, continued the work of Fletcher and Munson to develop a more realistic model of the loudness perception process. Stevens published the method "Calculation of the Loudness of Complex Noise" in the Journal of the Acoustical Society of America in 1956, and Zwicker published his article "Psychological and Methodical Basis of Loudness" in Acoustica in 1958. In 1959 Zwicker published a graphical procedure for calculating the volume, as well as several similar articles shortly thereafter. Stevens and Zwicker methods have been normalized as ISO 532, parts A and B respectively. Both methods include standard psychoacoustic phenomena such as division into critical bands, frequency masking and specific loudness. These methods are based on the division of complex sounds into components within the "critical bands" of frequency, which allows some components of the signal to mask other components and add specific loudness in each critical band to obtain the total sound volume.
Recent research indicates, as demonstrated by the Australian Broadcasting Authority (ABA) "Investigation into Loudness of Advertisements" (July 2002), that many ads (and some programs) are received
- 3 as too loud compared to other programs and therefore annoy the listeners. ABA research is only a recent attempt to approach the problem for years, in fact, all the material broadcast by broadcasters in all countries. These results show that the unpleasant feeling of listeners caused by inconsistent loudness of program material can be reduced or eliminated if reliable, consistent measurements of program volume can be developed and used to reduce annoying changes in volume.
The Barge scale is a unit of measurement used in the concept of critical bands. This critical band scale is based on the fact that human hearing analyzes a wide spectrum into parts that correspond to smaller critical sub-bands. Adding one critical band to the next in such a way that the upper limit of the lower critical band is the lower limit of the next higher critical band leads to the scale of the critical band measure. If the critical bands are added in this way, then each border point has a certain frequency. The first critical band covers the range from 0 to 100 Hz, the second from 100 Hz to 200 Hz, the third from 200 Hz to 300 Hz and so on up to 500 Hz, with the frequency range of each critical band being greater. The audible frequency range from 0 to 16 kHz can be further divided into 24 critical bands touching the boundaries, for which the bandwidth becomes larger as the frequency increases. These critical bands are numbered from 0 to 24, and they also have a "Bark" unit designating the Barge scale. This kind of relationship between the critical bandwidth transfer parameter and frequency is important for understanding many of the properties of the human ear. Please look at the example of "Psychoacoustics - Facts and Models" by E. Zwicker and H. Fasl, Springer-Verlag, Berlin 1990.
The ERB (Equivalent Rectangular Bandwidth) scale is a way of measuring frequency for human hearing, similar to the Barge scale. Developed by Moore, Glasberg and Baer, it is an improved variant of work on Zwicker's volume. Please look at Moore, Glasberg and Baer (BCJ Moore, B. Glasberg, T. Baer, "A Model for the Prediction of thresholds, Loudness, and Partial Loudness", Journal of the Audion Engineering Society, Volume 45, No. 4, April 1997 , pages 224240). Measurement of critical bands below 500 Hz is difficult because at such low frequencies, the efficiency and sensitivity of the human hearing aid drops sharply. Improved measurements of the audibility filter bandwidth led to the development of the ERB scale. For such measurements, slotted sound diaphragms were used to measure the bandwidth of the audibility filter. Generally, for the ERB scale, the audibility filter bandwidth (expressed in ERB units) is smaller than on the Barge scale. This difference is getting bigger for ever lower frequencies.
Frequency selectivity for the human hearing system can be approximately determined by subdividing the sound volume into parts corresponding to critical bands. This approximation leads to the emergence of the concept of critical band intensity. If instead of the infinitely steep slope of the hypothetical critical band filters, the actual slope for the human hearing is taken into account, then this procedure leads to an intermediate intensity value, known as agitation. In most cases, these values are not used as linear values, but as logarithmic values similar to the sound pressure level. Critical band and stimulation levels are relevant values that play an important role in many models as intermediate values. (see Psychoacoustics - Facts and Models, above).
The volume level can be measured in "phones". One fon is
- 5 defined as the received loudness of a pure 1 kHz sine wave reproduced at 1 dB sound pressure level (SPL), which corresponds to an effective pressure value of 2 * 10 Pa. N-phon is the received loudness of a 1 kHz tone reproduced at an SPL of N dB. By using this definition when comparing loudness of tones with frequencies other than 1 kHz with 1 kHz tone, you can determine the contour of equal volume for a given sound level. FIG. 7 presents the contours of the same volume level for frequencies from 20 Hz to 12.5 kHz and for phon levels from 4.2 phon (treated as hearing threshold) to 120 phon (ISO226: 1987 (E) "Acoustics - Normal Equal Loudness Level Contours") .
The volume level can also be measured in "sons". There is a simple relationship between phons and sons, as shown in Fig. 7. One son is defined as 40 dB (SPL) of pure 1 kHz sine wave and is equivalent to 40 phons. Sony are such units that a double increase in sons corresponds to a doubling of the received volume. For example, 4 sons are received twice as loud as 2 sons. Expressing loudness levels in sounds therefore provides more information.
Since a son is a measure of the loudness of an acoustic signal, the specific loudness is simply the volume per unit frequency. Therefore, when the Barka frequency scale is used, the specific loudness is measured in sons per shoulder, and similarly, when the ERB frequency scale is used, the units are sons on the ERB.
Hereinafter, terms such as "filter" or "filter assembly" herein include essentially any form of recursive or non-recursive filtering, such as filters or IIR transformations, and "filtered" information is the result of using such filters. The implementation examples described below use filter assemblies consisting of IIR filters and transforms.
- The essence of the invention
The object of the invention is to develop an objective loudness measurement technique that can be more accurately adapted to the subjective loudness results obtained by statistical loudness measurement with multiple listeners.
According to one aspect of the invention, the method of processing an audio signal includes creating an excitation signal in response to an audio signal and calculating the perceptual loudness of the audio signal in response to an excitation signal, and measuring the properties of an audio signal, wherein such calculation selects from at least two functions of the specific loudness model one or a combination of two or more specific loudness model functions, whose selection is regulated by a measure of the properties of the input audio signal.
According to a further aspect of the invention, a device and computer program have been developed according to claim 1. 20 and 21.
In an embodiment that uses aspects of the invention, the method or apparatus for processing the signal receives an input audio signal. This signal is linearly filtered by a filter or filter function that simulates the properties of the human outer and middle ear, and a set of filters or a function of a set of filters that divides the filtered signal into frequency bands that simulate the excitation model generated along the inner membrane of the inner ear. For each frequency band, the specific loudness is calculated using one or more specific loudness functions or models, the selection of which is controlled by the properties or characteristics derived from the input audio signal. The specific loudness for each frequency band is combined as loud as representing the broadband input audio signal. Single measure value
The loudness can be calculated for a certain finite time range of the input signal, or this loudness measure can be repeatedly calculated on time intervals or time blocks of the input audio signal.
In another embodiment that uses aspects of the invention, the signal processing method or device receives the input audio signal. This signal is linearly filtered by a filter or filter function that simulates the properties of the human outer and middle ear and by a set of filters or a filter set function that divides the filtered signal into frequency bands that simulate the excitation distribution generated along the inner membrane of the inner ear. For each frequency band, the specific loudness is calculated using one or more specific loudness functions or models, the selection of which is controlled by the properties or characteristics derived from the input audio signal. The specific loudness for each frequency band is combined as far as the loudness representative of the broadband input audio signal. This loudness measure is compared to a reference loudness value, and the difference is used to weigh or adjust the gain of the frequency-banded signals previously fed to the specific loudness input. Calculation of specific loudness, calculation of loudness and benchmarking are repeated until reference loudness and loudness become substantially equivalent. A single value of the loudness measure can be calculated for a certain finite range of the input signal, or the loudness measure can be repeatedly calculated in time intervals or blocks of the input acoustic signal. The recursive use of gain is beneficial due to the non-linear nature of the received loudness, as well as the structure of the loudness measurement process.
The various aspects of the present invention and its preferred embodiments can be better understood from the following description and the accompanying drawings, in which like reference numerals refer to like elements in several figures. The drawings, which show different devices or processes, show the more important elements that are helpful in understanding the present invention. For clarity, these drawings omit many other features that may be important in practical embodiments and are well known to those skilled in the art, but are not important for understanding the concept of the present invention. Signal processing for practicing the present invention can be carried out in a variety of ways, including programs implemented by microprocessors, digital signal processors, logic tables, and other forms of computing circuits.
Description of drawings
Fig. 1 is a block diagram of an embodiment of one aspect of the invention.
Fig. 2 is a block diagram of an embodiment of a further aspect of the invention.
Fig. 3 is a block diagram of an embodiment of yet another aspect of the invention.
Fig. 4 is the idealized characteristic of a linear P (z) filter suitable for a transmission filter in an embodiment of the present invention, where on the vertical axis there is attenuation in decibels (dB) and on the horizontal axis the frequency in hertz (Hz) on a logarithmic scale with a base 10.
Fig. 5 shows the relationship between the ERB frequency scale (vertical axis) and frequency in hertz (horizontal axis).
Fig. 6 shows a set of idealized hearing filter characteristics that approximate critical banding on the ERB scale. The horizontal scale is the frequency in hertz, and the vertical scale is the level in decibels.
Fig. 7 shows the contours of equal loudness according to ISO266. The horizontal scale is the frequency in hertz (log base 10 scale), and the vertical scale is the sound pressure level in decibels.
Fig. 8 shows contours of equal loudness according to ISO266 normalized by the transmission filter P (z). The horizontal scale is the frequency in hertz (log base 10 scale), and the vertical scale is the sound pressure level in decibels.
Fig. 9 (solid lines) shows loudness plots for both uniformly inducing noise and for a 1 kHz tone, where the solid lines are according to an embodiment of the present invention in which the parameters are selected to match the experimental data according to Zwicker (squares and circles) ). The vertical scale is loudness in sons (base 10 logarithmic), and the horizontal scale is the sound pressure level in decibels.
Fig. 10 is a block diagram of the operation of an embodiment according to a further aspect of the present invention.
Fig. 11 is a block diagram of the operation of an embodiment according to a still further aspect of the present invention.
Fig. 12 is a block diagram of the operation of an embodiment according to a further aspect of the present invention.
Fig. 13 is a block diagram of the operation of an embodiment according to a further aspect of the present invention.
The best ways to implement the invention
As described in more detail below, the embodiment of the first aspect of the present invention shown in Fig. 1 includes a regulator or specific volume control function ("control
- 10 specific volume '') 124, which analyzes and outputs the characteristics of the input audio signal. This acoustic characteristic is used to adjust the parameters in the specific loudness transducer or transducer function ('' specific loudness '') 120. By adjusting the specific loudness parameters using the signal characteristics, the present loudness measurement technique according to the present invention can be more accurately adapted to the results of subjective loudness by statistically measuring loudness using many listening people. Using the signal characteristics to adjust the volume parameters can also reduce the occurrence of incorrect measurements that make the signal volume feel nuisance to the listener.
As described in more detail below, the embodiment of the second aspect of the present invention shown in Fig. 2 adds an amplification device or function (`` iterative gain update '') 233 whose purpose is to iteratively regulate the gain of the time-averaged excitation signal from the input signal until the total volume at 223 in Fig. 2 is adjusted to the desired reference volume at 230 in Fig. 2. Since the objective measurement of perceived loudness requires a properly non-linear process, the iterative loop can be advantageously used to determine the appropriate gain to match the loudness of the input audio signal to the desired loudness level. However, an iterative gain loop surrounding the entire loudness measurement system, so that gain control is applied to the original audio input for each loudness iteration, would be expensive to implement due to the time integration needed to generate an accurate measure of long-term loudness. Usually in such a system, time integration requires a recalculation for each shift
- 11 enhancements in iteration. However, as explained further below, in the aspects of the invention shown in the embodiments of Fig. 2, as well as of Figs. 3 and 10-12, time integration can be implemented in linear processing paths that precede and / or follow a non-linear process, which forms part of an iterative gain loop. Linear processing paths need not be part of the iterative loop. Thus, for example, in the embodiment of Fig. 2 the loudness measurement path from input 201 to the specific loudness transducer or transducer function ('' specific loudness '') 220 may include time integration of the time averaging function ('' time averaging '') 206 and is linear. Consequently, iteration of the gain must be applied only to the reduced set of devices or functions for measuring loudness and does not need to include any time integration. In the embodiment of Fig. 2 transmission filter or transmission filter function ('' transmission filter ') 202, filter assembly or filter assembly function (' 'filter assembly' ') 204, time averaging unit or time averaging function (' 'time averaging' ') 206 and volume control proper or the specific volume control function ('' proper volume control '') 224 are not part of the iterative loop, which enables iterative gain control to be implemented in efficient and accurate real-time systems.
Returning to Fig. 1, a flowchart of an embodiment of a loudness meter or loudness measuring process 100 according to the first aspect of the present invention is shown. An acoustic signal for which the loudness measurement is to be determined is given to the loudness meter input 101 or the loudness measurement process 100. This input signal is fed on two paths - the first (main) path, which calculates the specific loudness in each of the many frequency bands that simulate the excitation distribution bands generated along the inner ear base membrane and the second (lateral) path having a regulator
- 12 specific loudness, which selects the specific loudness functions or models used in the main track.
In a preferred embodiment, the processing of the audio signal is carried out in the digital domain. The input audio signal is appropriately recorded by a discrete time sequence χ [η] which has been sampled from the audio source with a certain sampling frequency /<sub>χ</sub>. It is assumed that the sequence χ [η] has been weighed appropriately, so that the effective power of this sequence x [n] in decibels determined by
<img file="PL1629463T3_D0001.tif" />
is equal to the sound pressure level in dB at which the acoustic signal is heard by the listening person. In addition, to simplify exposure, it is assumed that the acoustic signal is monophonic. The implementation can, however, be adapted to the multi-channel audio signal as described later.
Transmission filter 102
In the main path, the input audio signal is fed to the transmission filter or the transmission filter function ("transmission filter") 102, whose output signal is a filtered version of the audio signal. This transmission filter 102 simulates the effect of the passage of the audio signal through the outer and middle ear using a linear P (z) filter. As shown in fig. 4, one frequency response P (z) of appropriate size is unity below 1 kHz, and above 1 kHz this response is consistent with the inverse of the hearing threshold specified in the ISO226 standard, with the normalization of the threshold to equal unity at 1 kHz. By using a transmission filter, the acoustic signal, which is processed by the volume measurement process, more closely resembles the acoustic signal received by human hearing, which improves the objective measure
- 13 volumes. Thus, the output signal of the transmission filter 102 is frequency dependent, the weighted version of the samples χ [η] of the time domain acoustic signal.
Filter assembly 104
The filtered audio signal is fed to the filter assembly or filter assembly function ("filter assembly") 104 (Fig. 1). The filter assembly 104 is intended to simulate the excitation distribution generated along the base membrane in the inner ear. The filter assembly 104 may include a set of linear filters whose bandwidth and spacing are constant on the Equivalent Rectangular Bandwith (ERB) frequency scale defined by Moore Glasberg and Baer (BCJ Moore, B. Glasberg, T. Baer, 'A Model for tha Prediction of Tresholds, Loudness, and Partial Loudness', above).
Although the ERB frequency scale is better suited to human perception and gives better results of objective loudness measurements matching subjective loudness results, the Barge frequency scale can be used with inferior results.
For the central frequency f in hertz, the width of one ERB band in hertz can be roughly expressed as:
ERB (f) = 24.7 (4.37f / 1000 +1) (1)
Based on this relationship, the warped frequency scale is defined such that at any point along such a warped scale, the corresponding equivalent rectangular bandwidth (ERB) in units of the warped scale is equal to one. The conversion function from the linear frequency in hertz to the ERB frequency scale is obtained by integrating the inverse of equation 1:
HzToERff) = f --- df = 214 Jog<sub>1D</sub>(4.37 /71000+1) <sup>J</sup> 24.7(4.37//1000+1) <sup>5,0</sup> ' (<sub>2a</sub>)
It is also useful to express the transformation from the ERB scale back to the linear frequency scale by solving equation 2a for
ERBToHz (e) = f = Ιθθθκ / "<sup>21</sup>·<sup>4-0</sup>
4.37 (2b) where e is in units of the ERB scale. Fig. 5 shows the relationship between the ERB scale and frequency in hertz.
The response of the auditory filters to the filter assembly 104 can be characterized and implemented using standard IIR filters. More specifically, individual auditory filters at the middle frequency fe in hertz, which are implemented in the filter assembly 104, can be defined by the twelfth order IIR transfer function:
-1. - 2_-2 (1 -z ~ ') (1-2 ^ (^ (2 ^<sup>/ J</sup> ) from ~ +<sup>g</sup>Z ~.) (1-2r <cosC<sup>2</sup>^<sup>/</sup>/ J) of 2 '<sup>1</sup> + r ^ -<sup>2</sup>)<sup>6</sup> (3) where
<img file="PL1629463T3_D0002.tif" />
(4a) (4b) (4c) (4d) (4e) fs is the sampling frequency in hertz, and G is the normalization factor to ensure that each filter has a unit peak gain in its frequency response; chosen so that {f<sub>e</sub>(e> ") |} = l max <2 / (4f)
The filter assembly 104 may comprise M of such hearing filters, called bands, with center frequencies fe [1] ... fe [M] distributed evenly along the ERB scale. Specifically
- 15 Λ1] = Λΰη fc [w] = f<sub>c</sub>\ m -1] + ERBToHz (HzToERB (f<sub>c</sub> \ m -1]) + Δ) m MM] <
(5a) <sup>2</sup>--<sup>M</sup> (5b) (5c) where Δ is the desired filter assembly spacing 104 and where / min and f<sub>max </sub>are the desired minimum and maximum frequencies, respectively. You can choose Δ = 1 and taking into account the frequency range to which the human ear is sensitive, you can set fmin = 50 Hz and fmax = 20,000 Hz. With such parameters, for example, the use of equations a - c gives M = 40 hearing filters. The sizes of these M hearing filters that approximate the critical banding on the ERB scale are shown in Figure 6.
Alternatively, filtering operations can be appropriately approximated using a finite length discrete Fourier transformation, commonly called fast discrete Fourier transformation (STDFT), because it is believed that the implementation of filter operation at the sampling frequency of the audio signal, treated as a full speed implementation, provides greater resolution time than necessary for accurate loudness measurements. By using STDFT instead of full-speed implementation, you can improve performance and reduce computational complexity.
The STDFT transformation of the input audio signal x [n] is defined as:
n ~ \ _ .fn
X \ k, d = Σ 14 « <sup>+</sup> tT] e <sup>N</sup><sub>n</sub>= o (6) where k is the frequency index, t is the index of the time block, N is the size of the DFT transformation, T is the value of the packet length measure, and in [n] is the length N normalized in the window, so that yi [«= in = Q (7)
It should be noted that the variable t in equation 6 is a discrete index representing the STDFT transformation time block as opposed to the measure of time in seconds. Each increase t represents a measure of the length of the packet T of samples along the signal x [n]. Further references to the t index take this definition. Although different parameter settings and window shapes can be used depending on the implementation details, for fs = 44100 Hz, excellent results are obtained by selecting N = 4096, T = 2048 and adopting the Hanning window as w [n]. The STDFT transformation described above may be more effective using Fast Fourier Transform (FFT).
In order to calculate the volume of the input audio signal, it is necessary to measure the energy of the audio signals in each filter of the filter set 104. The short-term output energy of each filter in the filter set 104 can be approximated by multiplying the frequency response of the frequency domain filter by the input power spectrum:
and*<sup>1</sup>* <sup>2</sup> Pi 12
A Λ = 0 <sup>m</sup> '(8) where m is the band number, t is the block number and P is the transmission filter. It should be noted that other modular response forms of acoustic filters than those given in equation 3 can be used in equation 8 to obtain similar results. For example, Moore and Glasberg suggest a filter shape described by an exponential function that behaves similarly to equation 3. In addition, with a slight reduction in efficiency, each filter can be approximated as a "wall" type band with a bandwidth of 1 ERB, and in further approximation the transmission filter P can be removed from the summation. In this case, Equation 8 is simplified to
- 17 £ lm, t] = Ι.Ιρ ^ ν.Μ'Λ) |<sup>2</sup> * 'k— k, = rwnd (ERBToHz (HzToEMtfJC [rm} -l / 2) N / /) k<sub>2</sub> = round (ERBToHz (HzToEIU3 (j<sub>c</sub> [m]) + l / 2) N / f<sub>s</sub>) (9a) (9b) (9c)
The excitation output of the filter assembly 104 is therefore a representation in the frequency domain of energy E in m of corresponding ERB bands at time t.
Multi-channel format
In the event that the input audio signal is in a multi-channel format for broadcasting through multiple loudspeakers, one for each channel, the excitation for each separate channel may first be calculated as described above. To then calculate the perceived loudness of all connected channels, individual excitations can be added together into one excitation to approximate the excitation reaching the listener's ears. All further processing is then carried out on this single, summed excitation.
Time averaging 106
Psychoacoustic studies and subjective loudness studies suggest that when comparing loudness between different acoustic signals, listeners perform some kind of temporal integration of short-term or "instantaneous" loudness of the signal to determine the value of long-term loudness to be used in such comparison. When building the loudness perception model, others suggested that this time integration would be performed after the non-linear transformation of the excitation into specific loudness. The inventors have found, however, that such time integration can be properly modeled using linear excitation smoothing before transforming it into specific loudness. By performing smoothing before calculating the specific loudness according to one aspect of the present invention, a significant advantage is obtained when calculating the gain to which a signal must be subjected to adjust its measured loudness in the prescribed manner. As explained further below, the gain can be calculated using an iterative loop that not only excludes the calculation of excitation, but preferably excludes such integration over time. In this way, the iterative loop can generate an advantage by calculations that depend only on the current time frame for which the gain is calculated, as opposed to calculations that depend on the entire time integration interval. As a result, you save both processing time and memory. Embodiments that calculate gain using an iterative loop include the embodiments described below in conjunction with Fig. 2, 3 and 10 - 12.
Returning to the description of Fig. 1, linear excitation smoothing can be carried out in various ways. For example, smoothing can be performed recursively using a device or time averaging function ("time averaging") 106 using the following equations:
Jm, i] = E [m, t -1] + <sup>1 (</sup>£ [m, /] - £ [m, t -1]) at ·! '".'] ' <sup>(</sup>and °<sup>and)</sup> σ [«7] = Λ" σ [ζη, ί-1] + 1 <sub>(10b)</sub> where the initial conditions are: E [m, -1] = 0 and σ [m, -1] = 0. The unique property of the smoothing filter is that by changing the smoothing parameter A<sub>m</sub> the smoothed energy E [m, t] can vary from the real time average E [m, t] to the average decay memory E [m, t], if Am = 1, then from equation (10b) it can be seen that σ [m, t] = t, and E [m, t] is then equal to the real time average E [m, t] for time blocks from 0 to t. If 0Am <1, then σ [m, t] -> 1 / (1 -Am) when t -> ~ and E [m, t] is simply the result of using a unipolar smoothing device against E [m, t], For applications where one number is needed
- 19 describing the long-term loudness of the finite-length acoustic segment, we can take Am = 1 for all m. For real-time applications where we want to track the long-term changing loudness of the continuous acoustic flux in real-time, you can set 0 <Am <1 and set Am to the same value for all m.
When calculating the time average E [m, t], it may be desirable to skip short-term segments that are considered "too quiet" and do not contribute to the volume received. To achieve this, the second threshold smoothener can work in parallel with the smoother of equation 10. This second smoothener maintains its current value if E [m, t] is small compared to E [m, t]:
<img file="PL1629463T3_D0003.tif" />
(11a)
<img file="PL1629463T3_D0004.tif" />
σ [ζη] /
1] + 1,
IN <sup>dB</sup> M m = l Λ »= ϊ othrrwisr (11b) where tdB is the relative threshold expressed in decibels. Although not critical for the invention, tdB = -24 gives good results. If there is no other smoothing in parallel, then E [<sup>m, t</sup>] = E [m, t].
Specific volume 120
There is still to be processed the time-divided excitation energy E [m, t] divided into bands into a single loudness measure in perceptual units, in this case in sons. In a specific loudness transducer or in a loudness processing function ("specific loudness") 120, each excitation band is converted to a certain specific loudness value, measured in sons on the ERB. In team
- volume combining or in the volume combining function ("loudness") 122 specific loudness values can be integrated or summed in bands to create total perceptual loudness.
Adjusting the specific volume 124 / specific volume 120
Multiple models
According to one aspect, the present invention uses multiple models in block 120 to convert the banded excitation into banded specific loudness. The control information obtained from the input audio signal by adjusting the specific loudness 124 in the lateral path selects the model or regulates the degree to which the model participates in the specific loudness. At block 124, certain properties or characteristics that are useful in selecting one or more specific loudness models from the available models are derived from the audio signal. Control signals that indicate which model or which model combinations should be used are generated from these derived properties or characteristics. Where it may be desirable to use more than one model, control information may also indicate how such models should be combined.
For example, the specific loudness of the N 'band [m, t] can be expressed as a linear combination of specific loudness of the band for each model N'q [m, t] as:
<img file="PL1629463T3_D0005.tif" />
(12) where Q is the total number of models and control information a<sub>q</sub>[m, t] represents the weight or contribution of each model. The sum of such weights may or may not be equal to one, depending on the models used.
Although the invention is not limited to them, two models have been found that give accurate results. One model works best when
- 21 when the audio signal has a narrow bandwidth and the second one works best when the audio signal has a wide bandwidth.
Initially, when calculating the specific loudness, the excitation level in each E [m, t] band can be converted into an equivalent excitation level at 1 kHz, as determined by the contours of the same loudness according to ISO 266 (Fig. 7) normalized by the transmission filter P (z) (Fig. . 8):
(13) where LikHz (EJ) is a function that generates a level at 1 kHz exactly as loud as an E level at frequency f. In practice, L<sub>1k</sub>H<sub>FROM</sub>(E, /) is implemented as an interpolation of an overview table of contours of equal loudness, normalized by a transmission filter. Transformation to equivalent levels at 1 kHz simplifies the following calculation of specific loudness.
Then the specific loudness in each band can be calculated as:
N '[m, f] = a [m, + (1- a [m' /]) N '<sub>IVB</sub> [w /] (14) where N'NB [m, t] and N'wB [m, t] are specific loudness values based on the narrowband and broadband signal model, respectively. The value a [m, t] is an interpolation coefficient in the range of 0 to 1, which is calculated from the audio signal, the details of which are described below.
The specific loudness values of the narrow band and wide band N'NB [m, t] and N'wB [m, t] can be estimated on the basis of the banded excitation using exponential functions:
f
<img file="PL1629463T3_D0006.tif" />
otherwise (15a)
- 22 /
<img file="PL1629463T3_D0007.tif" />
£<sub>l</sub>> A ", f]> 10 <sup>10</sup> otherwise (15b) where TQikHz is the level of excitation at the silence threshold for a tone of 1 kHz. With equal loudness contours (Figures 7 and 8), TQikhz equals 4.2 dB. It should be noted that both of these specific loudness functions are equal to zero when the excitation is equal to the threshold of silence. At excitations greater than the silence threshold, both functions increase monotonically according to the power function according to Stevens law of intensity. The exponent of the narrowband function is chosen as larger than the exponent of the broadband function, which results in a faster rise of the narrowband function than the broadband function. The specific selection of β exponents and G reinforcements for narrowband and broadband cases is discussed below.
Volume 122
Loudness 122 uses the banded specific loudness of Specific Loudness 120 to create one loudness measure for the audio signal, namely the output signal at terminal 123, which is the loudness value in perceptual units. This loudness measure can have any units, as long as a comparison of the loudness values of the different acoustic signals indicates which is louder and which is weaker.
Total loudness expressed in sons can be calculated as the sum of specific loudness for all frequency bands:
5 [z] = 4 £ jV '[<sub>m</sub>, z] (16) where Δ is the ERB distance defined in equation 5b. The Gnb and βΝΒ parameters in equation 15a are chosen such that when a [m, t] = 1, the S graph in sons as a function of SPL for a tone of 1 kHz is generally consistent with the corresponding
- 23 experimental data presented by Zwicker (circles in Figure 9) (Zwicker, H. Fastl, "Psychoacoustics - Facts and Models", above). The Gwb and 3wb parameters in equation 15b are chosen such that when a [m, t] = 0, the N chart in sons as a function of SPL for uniformly excited noise (noise with equal power in each ERB) is broadly consistent with analogous Zwicker results (squares in fig. 9). Matching the least squares to Zwicker data gives:
& νβ <sup>=</sup> θ · θ<sup>4</sup>θ<sup>4</sup> (17<sub>and</sub>) = <sup>0</sup>·<sup>279</sup> (17b) <sup>gy</sup> =0.<sup>05</sup>8 (17c) <sup>Ara</sup> =071<sup>2</sup> (17d)
Fig. 9 (solid lines) shows loudness plots for both uniformly excited noise and 1 kHz tone.
Adjusting the specific volume 124
As mentioned before, two specific loudness models (equations 15a and 15b) are used in practical implementation, one for narrowband signals and one for wideband signals. The specific volume control 124 in the side track calculates the measure a [m, t] of the degree to which the input signal is either narrowband or wideband in each band. in general, a [m, t] should be equal to one when the signal is narrowband near the center frequency / o [m] band, and zero when the signal is wideband near the center frequency / o [m] band. The regulation should change constantly between the two extremes, so that the mixtures of these characteristics change. For simplicity, the regulation a [m, t] can be selected as a constant in the bands, in which case a [m, t] is further treated as a [t] without the m band index. The a [t] control then represents a measure of narrowband signal in all bands. Although suitable
The method of obtaining such regulation is described below, this particular method is not critical and other suitable methods may be used.
Regulation a [t] can be calculated from the excitation E [m, t] at the output of the filter assembly 104 rather than by some other signal processing x [n]. E [m, t] can provide an appropriate reference from which "narrowband" and "broadband" x [n] are measured, and as a result a [t] can be generated with minor additional computational activities.
"Spectral flatness" is a property E [m, t] from which a [t] can be calculated. Spectral flatness, defined by Jayant and Noll (Digital Coding Of Waveforms, Prentice Hall, New Jersey, 1984), is the ratio of the geometric mean to the arithmetic mean, where the mean is determined by frequency (m index for E [m, t]). When E [m, t] is constant relative to m, the geometric mean is equal to the arithmetic mean and the spectral flatness is equal to one. This corresponds to the wide band case. If E [m, t] changes significantly relative to m, then the geometric mean is much smaller than the arithmetic mean and the spectral flatness approaches zero. This corresponds to a narrow band case. By calculating the difference of one minus spectral flatness, you can create a "narrowband" measure, where zero corresponds to a wide band and unity corresponds to a narrow band. In particular, the difference one minus the modified spectral flatness E [m, t] can be calculated:
and
<img file="PL1629463T3_D0008.tif" />
where P [m] is equal to the frequency response of the transmission filter P (z)
- 25 sampled at the frequency ω = 2nf<sub>c</sub>\ M] / f<sub>s</sub> . Normalizing E [m, t] through a transmission filter can provide better results because the use of a transmission filter introduces a "burst" in E [m, t] that tends to blow the "narrow band" measure. In addition, calculating spectral flatness for a subset of E [m, t] bands can give better results. The lower and upper summation limits in Equation 18, Ml [t] and Mu [t], define an area that may be smaller than the range of all M bands. It is desirable that Ml [t] and Mu [t] include the part E [m, t] that contains most of the energy, and that the range defined by Ml [t] and Mu [t] does not have a width greater than 24 units on the scale ERBIUM. In particular (and considering that fc [m] is the center frequency of the band in m Hz):
HzToERBJM ,, [r] D - HzToEIRBHffM, [/]]) = 24 as well as is required:
HzToERB (f<sub>e</sub>[M<sub>at</sub> R]])> C7R]> HzToEFWJCM, R]]) HzToEitffMM})> HzTtEERJJ}) HzToERB (f<sub>e</sub>[M ,, R] l) <HzToERB <f<sub>e</sub>[M]) (19b) (19c) (19d) where CT [t] is the spectral center E [m, t] measured on the ERB scale: Σ HzToERBJ Rm]) £ Rh, r] <a name="caption1"></a>CTR] = m = l (19e)
Ideally, the summation limits, Ml [t] and Mu [t], are centered around CT [t], measured on the ERB scale, but this is not always possible when CT [t] is near the lower or upper limit of its range.
Then NB [t] can be smoothed over time in a manner analogous to Equation 11a:
- 26 M '<sup>dR</sup> Μ ££ lm, r]> 10 '° £> («>, (] m = l otherwise
<img file="PL1629463T3_D0009.tif" />
(20) where σ [t] is equal to the maximum a [mt], according to equation 11 b, for all m.
Finally a [t] is calculated from NB [z] as follows:
<img file="PL1629463T3_D0010.tif" />
φ {χ} = 12.2568x<sup>3</sup> - 22.8320x<sup>2</sup> + 14.5869x- 2.9594 (21a) (21b)
Although the exact form Φ {x} is not critical, the polynomial in equation 21b can be determined by optimizing a [t] relative to the subjectively measured loudness of a large variety of acoustic material.
Fig. 2 is a flow chart of an embodiment of a loudness meter or loudness measuring process 200 according to the second aspect of the present invention. The devices or functions 202, 204, 206, 220, 222, 223 and 224 of Fig. 2 correspond to the corresponding devices or functions 102, 104, 106, 120, 122, 123 and 124 of Fig. 1.
According to a first aspect of the invention, the embodiment of which is shown in Figure 1, a loudness meter or loudness calculation produces a loudness value in perceptual units. To control the volume of the input signal, a useful measure is the gain G [t], which multiplied by the input signal x [t] (as in the embodiment of Fig. 3, described below, for example) makes its volume equal to the reference loudness level Sref. The reference loudness Sref can be given arbitrarily or measured by another device or operating process
According to a first aspect of the invention relative to a certain "known" audio reference signal. Assuming that ψ {.ν [//] ι represents all calculations performed on the signal x [n] to create the loudness S [t], we want to find such G [t] that
<img file="PL1629463T3_D0011.tif" />
= φί] = ψ {θ (ΦΜ ,,} (23)
Because part of the processing contained in Ψ {.} Is non-linear, there is no closed form for G [t], so you can use iterative technique instead to find an approximate solution. Let Gi represent the current rating G [t] at each iteration and in this process. For each iteration, Gi is updated so that the absolute error from the reference loudness decreases:
is · - ΨίσχΜ <| s "- ψ {σ<sub>Μ</sub>χ "], ί) (24)
There are many suitable techniques for upgrading Gi to obtain the above error reduction. One such method is to reduce the gradient (see Nonlinear Programming, Dimitri P. Bertseakas, Athena Scientific, Belmont, MA 1995), where Gi is updated with a value proportional to the error from the previous iteration:
G =<sup>Gw</sup> + / ¼-ψ {σ, _, χ [η], ί}) (25) where μ is the magnitude of the iteration stroke. The above iteration continues until the absolute error becomes less than a certain threshold, until the number of iterations reaches the set maximum limit value or until the set time has elapsed. From now on, G [t] is treated as equal to Gi.
Returning to equations 6-8, we notice that the excitation of the signal x [n] is obtained by linear operations on the square of the absolute value of the STDFT signal X [ł, t]<sup>2</sup> . It follows that the excitation obtained from the Gx [n] signal with modified gain is equal
- 28 excitation x [n] multiplied by β2. Furthermore, the time integration required to evaluate the long-term received loudness can be performed by linear averaging the excitation over time, and therefore the time-averaged excitation corresponding to Gx [n] is equal to the time-averaged excitation x [n] multiplied by G2. As a result, time averaging does not need to be calculated over the entire history of the input signal for each re-evaluation Ψ {G, .x [wJr in the iterative process described above. Instead, the time-average excitation E [m, t] can be calculated only once with x [n], and during iteration the updated loudness values can be calculated by applying the square of the updated loudness directly to E [m, t]. In particular, assuming that Ψ<sub>ε</sub>{ε [μ, ϊ] represents all processing performed on time-excited E [m, t] to produce S [t], for general multiplicative G gain:
By using this relationship, the iterative process can be simplified by
<img file="PL1629463T3_D0012.tif" />
it would be possible if the time integration needed to evaluate the long-term loudness received was performed after a non-linear transformation to the specific loudness.
The iterative calculation process G [t] is shown in Fig. 2. The output volume S [t] on terminal 223 can be subtracted in the subtraction device or in the subtraction function 231 from the reference volume Sref on terminal 230. The resulting error signal 232 is given to the device or an iterative gain update function ("iterative gain update") 233 that produces the next Gi gain in iteration. The square of this gain Gi<sup>2</sup> it is then returned to output 234 of the multiplicative connecting device 208, where Gi2 is multiplied by the time-average excitation signal from block 206.
- 29 The next value of S [t] in this iteration is then calculated from this version of the modified excitation gain averaged over time by blocks 220 and 222. The described loop iterates until the final conditions occur when the gain G [t] on terminal 235 is set as equal to the current value of Gi. The final G [t] value can be calculated by the iterative process described, for example, for each t frame of a fast Fourier transformation or only once at the end of an acoustic segment after averaging the excitation for the entire length of this segment.
If we want to calculate the signal volume without modifying the gain in connection with this iterative process, the Gi gain can be initialized to unity at the beginning of each iterative process for each time interval t. In this way the first value of the expression S [t] calculated in the loop represents the initial signal volume and can be written like that. However, if you do not want to save this value, Gi can be initialized with any value. If G [t] is calculated in subsequent time frames and you do not want to record the volume of the initial signal, it may be desirable to initialize Gi on the G [t] values from the previous time interval. In this way, if the signal does not change significantly from the previous time interval, it is likely that the G [t] value will remain substantially the same. Therefore, only a few iterations will be needed to achieve convergence to the correct value.
After the iteration, G [t] represents the gain applied to the input audio signal in 202 by some external device, so that the volume of the modified signal is matched to the reference volume. Fig. 3 shows one suitable arrangement in which the G [t] gain from iterative gain update 233 is applied to the control input of the device or signal level control function, e.g. voltage controlled amplifier (VCA) 236,
- 30 to create an output signal with adjustable gain.
VCA 234 in Fig. 3 can be replaced by an attendant who controls the gain control in response to the G [t] gain sensor indication on line 235. The sensor indication may be provided, for example, by a meter. G [t] reinforcement can be subjected to time smoothing (not shown).
For some signals, the smoothing alternative described in equations 10 and 11 may be desirable to calculate the long-term perceived loudness. Listeners tend to associate long-term signal volume with the loudest part of the signal. As a result, the smoothing presented in equations 10 and 11 may underestimate the received loudness of the signal containing long periods of relative silence interrupted by shorter segments of louder material. Such signals often occur in movie soundtracks with short dialogue segments surrounded by longer periods of background noise. Even when determining the threshold, shown in equation 11, the silent parts of such signals contribute too much to the time-averaged excitation E [m, t].
To deal with this problem, a statistical technique for calculating long-term loudness can be used in a further aspect of the present invention. First, the smoothing time constant in Equations 10 and 11 is set to be very small, and tdB is set to minus infinity, so that E [m, t] represents "momentary" excitation. In this case, the smoothing parameter A<sub>m</sub> you can choose to change in m bands to accurately model the way perception of instantaneous volume as a function of time changes. However, in practice, choosing the Am value as a constant in the function m provides results that are still acceptable. The rest of the previously described algorithm works unchanged, resulting in a signal of instantaneous loudness S [t], as given in equation 16. To some extent ti <t <t2 long-lasting loudness S<sub>p</sub>[T, t<sub>2</sub>] is then defined as a value that is greater than S [t] by p percent of the time value in this range, and less than S [t] for 100-p percent of the time value in this range. Experiments have shown that the setting of p equal to roughly 90% matches the subjectively perceived long-term loudness. With this setting, only 10% of the S [t] value must be significant to affect the long-term loudness. The remaining 90% of the value can be relatively quiet without reducing the long-term loudness measure.
The value of Sp [t1, t2] can be calculated by sorting in ascending order the value of S [t], t1 <t <t2 in the list Ssort {i}, 0 <and <t2-h, where i denotes the ith element of the sorted list . Long-term loudness is then determined by an element that is p percent of the way to the list:
<img file="PL1629463T3_D0013.tif" />
The above calculation, as such, is relatively simple. However, if you want to calculate the gain Gp [f, t2], which after multiplying by x [n] gives Sp [t1, t2] equal to a certain reference loudness Sref, the calculation becomes much more complex. As before, an iterative approach is needed, but now the measure of long-term loudness Sp [t1, t2] depends on a whole series of values of S [t], h <t <t2, each of which must be updated with each update of Gi in iteration. To calculate these updates, the signal E [m, t] must be stored in the entire range t1 <t <t2. In addition, since the relationship of S [t] to Gi is non-linear, the relative ordering of S [t], t1 <t <t2 may change at each iteration, and therefore Ssort {i} must also be recalculated. The need for re-sorting is clearly evident when considering short-term signal segments whose spectrum is just below the hearing threshold for a given gain in iteration. When the gain is increased, a significant portion of the spectrum of this segment may become audible, which may increase the total loudness of this segment relative to other narrowband signal segments that were previously audible. When the interval h <t <t2 becomes larger or if you want to calculate the gain Gp [h, t2] continuously as a function of the time shift window, the costs of calculating and remembering such an iterative process can be daunting.
Significant savings in calculations and memory are obtained by recognizing that S [t] is a monotonically increasing function of Gi. In other words, increasing Gi always increases the short-term volume at any time. With this knowledge, you can effectively calculate the needed matching gain Gp [ti, t<sub>2</sub>] in the following way. First, the previously defined matching gain G [t] with E [m, t] is calculated using the iteration described for all values in the interval t and <t <t2. It should be noted that for each value t, G [t] is calculated by iterating on a single value E [m, t]. Then the long-term matching gain Gp [ti, t2] is calculated by sorting in ascending order the value G [t] h <t <t2 to create a list of Gsortfi}, 0 <and <t2 - ti, and then setting
<img file="PL1629463T3_D0014.tif" />
We will now justify that Gp [ti, t2] is equal to the gain, which after multiplying by x [n] results in Sp [ti, t2] equal to the desired reference loudness Sref. From equation 28 it follows that G [t] <Gp [ti, t<sub>2</sub>] for 100-p percent of the time value in the interval ti <t <t2 and that G [t]> Gp [ti, t2] for the remaining p percent. For G [t] values such that G [t] <Gp [ti, t2], it can be seen that if Gp [ti, t2] would be applied to the corresponding values of E [m, t] and not to Gt, then the resulting S [t] would be greater than the reference loudness needed. This is correct because S [t] is a monotonically increasing function of gain. Similarly, if Gp [ti, t<sub>2</sub>] would be applied to E [m, t] values corresponding to G [t] such that G [t]> Gp [ti, t2],
- then the resulting values of S [t] would be less than the reference loudness needed. Therefore, applying Gp [ti, t2] to all values E \ m, t] in the interval ti <t <t2 causes that S [t] is greater than the desired reference value in 100 - p percent of the time, and smaller than the reference value wp percentage of time. In other words, Sp [ti, t2] is equal to the desired reference value.
This alternative method of calculating the matching gain avoids storing E [m, t] and S [t] during the time interval ti <t <t2. Only G [t] needs to be stored. In addition, for each calculated value of Gp [ti, t2], sorting G [t] in the time interval ti <t <t2 must be carried out only once, unlike the previous approach, at which S [t] must be re-sorted at each iteration. In the event that Gp [ti, t2] is to be calculated continuously in a certain sliding window of length T (i.e. ti = t - T, t2 = t), the list Gsort {i} can be maintained efficiently by simply removing and adding a single value in a sorted list for each new moment of time. When the interval ti <t <t2 becomes too large (for example, the length of the entire song or movie), then the memory needed to store G [t] can still be inhibiting. In this case, Gp [ti, t2] can be approximated from a discretized histogram G [t]. Practically, this histogram is created from G [t] in decibels. This histogram can be calculated as
H [i] = the number of samples in the interval ti <t <t2 such that
<img file="PL1629463T3_D0015.tif" />
where Ajb is the histogram resolution and dBmin is the minimum histogram. The matching gain is then approximated as
<img file="PL1629463T3_D0016.tif" />
where (30a)
- 34 100-22- = ρ
Σ * Μ "<sup>ο</sup> (30b) and I is the maximum histogram index. When using a discretized histogram, only the I values need to be stored, and Gp [h, t2] is easily updated with each new G [t] value.
Other ways to approximate Gp [h, t2] from G [t] are conceivable, and the invention is intended to include such techniques. A key aspect of this part of the invention is to perform some kind of smoothing on the matching G [t] gain to generate long-term matching Gp [h, t2] instead of processing the instantaneous loudness S [t] to generate the long-lasting loudness Sp [h, t2], from which, then, iterative process estimates Gp [t1, t<sub>2</sub>].
Figs. 10 and 11 show systems similar to Figs. 2 and 3, respectively, but smoothing (device or function 237) of the matching gain G [t] is used to generate the signal Gp [h, t2] of the smoothed gain (signal 238).
The reference loudness at the input 230 (Figures 2, 3, 10, 11) may be "constant" or "variable", and the reference loudness source may be internal or external in a system implementing aspects of the invention. For example, the reference volume can be set by the user, in which case its source is external and may remain "constant" for a certain period of time until it is set by the user. Alternatively, the reference loudness may be a measure of the loudness of another audio source, derived from a loudness process or apparatus according to the present invention, such as the system shown e.g. in Fig. 1. The normal loudness adjustment of the acoustic signal generating device may be replaced by the process or device according to aspects of the invention, as in the examples of Fig. 3 or
- Fig. 11. In this case, the user-operated knob, slider etc., volume control would be used to adjust the reference volume in 230 in Fig. 3 or in Fig. 11, and as a consequence this device generating the acoustic signal will have a volume in accordance with the regulation of the strength user voice.
An example of a variable reference is shown in Figure 12, where the reference loudness Sref is replaced by the variable reference Sref [t], which is calculated, for example, from the loudness signal S [t] by the device or the variable reference loudness function ("Variable Reference Volume") 239. In this system, at the beginning of each iteration in each time interval t, the variable reference Sref [t] can be calculated from the unmodified loudness S [t] before any gain is applied to the excitation in 208. Relationship S<sub>re</sub>f [t] and S [t] through the variable reference loudspeaker function 239 can take different forms to obtain different results. For example, such a function can simply scale S [t] to generate a certain reference that is in a constant relation to the original loudness. Alternatively, this function could produce a reference greater than S [t] when S [t] is below a certain threshold and less than S [t] when S [t] is above a certain threshold, thereby reducing the dynamic range of the received signal volume acoustic. Regardless of the form of this function, the iteration described previously is carried out to calculate G [t] so that
<img file="PL1629463T3_D0017.tif" />
(31)
The matching G [t] gain can then be smoothed as described above or some other suitable method to achieve the desired perceptual result. Finally, a delay 240 between the acoustic signal 201 and the VCA block 236 can be introduced to compensate for any delay in calculating the smoothed gain. Such a delay can also be applied to the systems of Figures 3 and 11.
The gain control signal G [t] of the system of Fig. 3 and the control signal smooth smoothening Gp [t1, t2] of the system of Fig. 11 can be useful in many different applications, including, for example, television or satellite radio, where the received volume changes in different channels. In such environments, the device and method of the present invention may compare the audio signal from each channel with a reference loudness level (or with a reference signal loudness). An operator or automatic device can use this gain to adjust the volume of each channel. All channels would therefore have essentially the same received volume. FIG. 13 shows an example of such a system in which audio signals from multiple TV or audio channels, 1-N, are fed to the respective inputs 201 of processes or devices 250, 252, each in accordance with aspects of the invention as shown in Figures 3 or 11. The same the reference loudness level is given to each of the processes or devices 250, 252, which causes the adjustment of the volume of the acoustic signal in the channels from the first to the Nth on each output 236.
This gain measurement and adjustment technique can also be used in a real-time measuring device that monitors input audio material, performs a process that identifies audio content essentially containing human speech signals, and calculates gain so that speech signals generally conform to a predetermined reference level . Suitable methods for identifying speech in an acoustic material are disclosed in US Patent Application SN 10 / 233.073 of August 30, 2002 and published as published US patent application US 2004/0044525 A1, publication of March 4, 2004. Because irritation of listeners with a loud acoustic signal tends to focus on parts of speech in the material
- 37 program, the method of measuring and adjusting gain can significantly reduce the difference in the level of irritation in relation to the acoustic signal commonly used in television, film and music material.
Implementation
The invention may be implemented in hardware and / or software (e.g., programmable logic tables). Unless otherwise specified, the algorithms contained in part of the invention are not permanently associated with a particular computer or other device. In particular, various general-purpose machines can be used with programs written in accordance with the principles given here, or it may be more convenient to construct more specialized devices (e.g. integrated circuits) to perform the necessary method steps. The invention may therefore be implemented in one or more computer programs executed on at least one programmable computer system, each of which includes at least one processor, at least one data memory system (including volatile and non-volatile memory and / or memory elements), which at least one input device or input port and at least one device or output port. The input data is subjected to program code to perform the functions described in it and generate output information. This output information is fed to at least one output device in a known manner.
Any such program can be implemented in any computer language you need (including machine language, assembly language or higher-level procedural, logical or object-oriented programming languages) for communication with a computer system. In any case, the language may be a compiled or translated language.
Each such computer program is preferably stored
- 38 on or downloaded to a storage medium or device (e.g. semiconductor storage medium or memory, or magnetic or optical storage medium) read by a general or special purpose programmable computer to configure or control this computer when the storage medium or device is read by a computer system to carry out its procedures. The system according to the invention can also be considered as implemented in the form of a computer readable storage medium configured with a computer program, where the storage medium so configured causes the computer system to operate in a specific and predefined manner to perform the functions described therein.
A number of embodiments of the invention have been described. However, it is understood that various modifications can be made without departing from the scope of the invention. For example, some of the steps described above may be independent of the order and thus may be carried out in a different order than described. Other embodiments are included within the scope of the following claims, respectively. The scope of the invention is therefore limited only by the appended claims.
Tomasz Jabłkowski Patent Attorney
34 members in 19 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 47407703 | United States of America | P | |
| 47407703 | United States of America | P | |
| 04776174 | European Patent Office (EPO) | A | |
| 2004016964 | United States of America | W | |
| 2004016964 | United States of America | W | |
| EP20040776174 | – | – | – |
| US20030474077P | – | – | – |
| WO2004US16964 | – | – | – |
Members34
| Document | Office | Kind | |
|---|---|---|---|
| AU2004248544A1 | Australia | A1 | |
| CA2525942A1 | Canada | A1 | |
| WO2004111994A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2004111994A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20060013400A | Republic of Korea | A | |
| MXPA05012785A | Mexico | A | |
| EP1629463A2 | European Patent Office (EPO) | A2 | |
| BRPI0410740A | Brazil | A | |
| CN1795490A | China | A | |
| HK1083918A1 | Hong Kong, China | A1 | |
| JP2007503796A | Japan | A | |
| US2007092089A1 | United States of America | A1 | |
| EP1629463B1 | European Patent Office (EPO) | B1 | |
| AT371246T | Austria | T | |
| ATE371246T1 | Austria | T1 | |
| EP1835487A2 | European Patent Office (EPO) | A2 | |
| DE602004008455D1 | Germany | D1 | |
| DK1629463T3 | Denmark | T3 | |
| PL1629463T3This record | Poland | T3 | |
| ES2290764T3 | Spain | T3 | |
| HK1105711A1 | Hong Kong, China | A1 | |
| EP1835487A3 | European Patent Office (EPO) | A3 | |
| DE602004008455T2 | Germany | T2 | |
| AU2004248544B2 | Australia | B2 | |
| JP4486646B2 | Japan | B2 | |
| CN101819771A | China | A | |
| IL172108A | Israel | A | |
| CN101819771B | China | B | |
| KR101164937B1 | Republic of Korea | B1 | |
| SG185134A1 | Singapore | A1 | |
| US8437482B2 | United States of America | B2 | |
| EP1835487B1 | European Patent Office (EPO) | B1 | |
| CA2525942C | Canada | C | |
| IN2913KON2010A | India | A |
Numbers
- Publication, DOCDB
- 1629463
- Publication, EPODOC
- PL1629463T
- Application
- 776174
- Application, DOCDB
- 04776174
- Application, EPODOC
- PL20040776174T
Titles2
- English
- METHOD, APPARATUS AND COMPUTER PROGRAM FOR CALCULATING AND ADJUSTING THE PERCEIVED LOUDNESS OF AN AUDIO SIGNAL
- Polish
- Sposób, urządzenie i program komputerowy do obliczania i regulowania odczuwalnej głośności sygnału akustycznego
Classification
- CPC, 7
- H03G9/005
- G10L25/27
- G10L25/48
- H03G5/00
- H03G5/005
- H03G9/025
- G10L19/02
- IPC, 2
- G10L11 00
- H03G9 02