Method for determining intensity parameters of background nose in speech pauses of voice signals
Claim Score by NHIP
Abstract
Known methods for determining intensity parameters are based on the evaluation of short signal segments and their direct allocation to speech pauses or speech activity. In order to distinguish speech from speech pauses, intensity thresholds are often used. When the undisturbed source signal is used to mark, speech pauses, a variably occurring time lag between source voice signal and disturbed voice signal often impedes exact transfer of the marking. Intensity parameters of background noises in speech pauses can be determined from the frequency distribution of the intensity values for short signal segments using the method disclosed in the invention. In order to assign intensity values, the fraction of speech pauses in the entire signal is calculated from the undisturbed source signal and defined as frequency threshold. Intensity values below the frequency threshold are assigned to the speech pauses. The arithmetic mean value of said intensity value is determined as intensity parameter for the background noise in the speech pauses. Percentile parameters for background noises in speech pauses can also be calculated with the inventive method.

Term
Term ended
Projected expiry passed 13 September 2024, 2 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
4 claims: 1 independent, 3 dependent
- 1Broadest claimClaim Score 45, average(NHIP)A method for determining intensity characteristics of background noise during speech pauses of speech signals, the undisturbed source speech signal and the disturbed speech signal of which being available in recorded form, and the proportion of speech pauses in the overall signal being determined from the undisturbed source speech signal according to known methods, and the disturbed speech signal being divided into short successive signal elements, and an intensity value being determined for each signal element. wherein the cumulative relative frequency distribution (1) is formed from the intensity values of the individual signal elements of the disturbed speech signal;the determined proportion of speech pauses in the source speech signal is defined as the frequency threshold, and the frequency threshold is applied to the disturbed speech signal;the intensity threshold value (3) which corresponds to the defined frequency threshold (2) is determined from the frequency distribution of the intensity values of the signal segments;all signal segments having a smaller intensity value than that of the intensity threshold value are assessed as belonging to the speech pauses;the distribution function for the intensity values of the signal segments in the region below the intensity threshold value represents the frequency distribution for the intensity values during the speech pauses (4);and this region of the distribution function is able to be used for determining intensity characteristics of the background noise during the speech pauses.
74 paragraphs in 8 sections, as filed
SPECIFICATION
PRELIMINARY REMARKS
[0001] The present invention relates to a method for assessing background noise during speech pauses of recorded or transmitted speech signals.
[0002] The perceived speech quality, for example, in telephone connections or radio transmissions, is chiefly determined by speech-simultaneous interference, that is, by interference during speech activity. However, noise during the speech pauses goes into the quality decision as well, in particular in the case of high-quality speech reproduction.
[0003] The intensity of the background noise during the speech pauses can be used as a supplementary characteristic for determining the speech quality.
[0004] Speech quality evaluations of speech signals are generally carried out by listening (“subjective”) tests with test subjects.
[0005] On the other hand, the goal of instrumental (“objective”) methods for determining speech quality is to determine characteristics which describe the speech quality of the speech signal from properties of the speech signal to be assessed, using suitable calculation methods without having to draw on the judgements of test subjects.
[0006] A reliable quality assessment is provided by instrumental methods which are based on a comparison of the undisturbed reference speech signal (source speech signal) and the disturbed speech signal at the end of the transmission chain. There are many such methods, which are mostly employed in so-called “test connection systems”. In this context, the undisturbed source speech signal is injected at the source and recorded after transmission.
RELATED ART AND DISADVANTAGES OF KNOWN METHODS
[0007] Known methods for determining the intensity of background noise usually start from the disturbed signal itself and use a determined intensity threshold to distinguish active speech and speech pauses (FIG. 1). In the simplest case, this threshold is set to be constant in the method, but can also be adapted on the basis of the signal pattern (for example, a defined distance from the signal peak value). The goal is a reliable distinction between speech and speech pause. If the distinction is achieved, the sought intensity characteristics of the background noise can be determined from the signal segments that have been identified as a speech pause. To this end, the signal segments that have been identified as a speech pause are generally further divided into shorter segments (typically 8 . . . 40 ms) and the intensity calculations (for example, effective value or loudness) are carried out for these shorter segments. Then, intensity characteristics can be determined from the results.
[0008] Given low noise intensities during speech pauses and, at the same time, high speech intensity (high speech-to-noise ratio), these methods yield reliable measured values because a reliable distinction can be made between speech and speech pause (FIG. 1).
[0009] In the case of increasing noise intensities during speech pauses (decreasing speech-to-noise ratio), increasingly uncertainties arise in the distinction between speech and speech pauses. Here, it is difficult to fix the threshold value in such a manner that, on one hand, no noise segments with higher intensities than speech are detected (threshold too low) and, on the other hand, no speech segments of lower intensity are judged as a speech pause (threshold too high) (FIG. 2).
[0010] If the intensity of the noise during the speech pauses reaches or even exceeds the intensity of the active speech, no intensity threshold can be found that would permit a distinction between speech and speech pause.
[0011] Solutions to the described problems are possible if, for example, speech and background noise have different spectral characteristics. By appropriately prefiltering the signal or via spectral analysis and evaluation of selected frequency bands, it is possible here to achieve a higher speech-to-background noise ratio in the observed frequency bands, making a reliable distinction between speech and speech pause possible again.
[0012] Other solutions make use of certain parameters, which are determined in speech coding, and use them to distinguish between speech and segments containing background noise. In this context, the goal is to derive from the parameters whether the observed signal segment has typical properties of speech (for example, voiced portions). An example of this is the “Voice-Activity Detector” (ETSI Recommendation GSM 06.92, Valboune, 1989).
[0013] In the case of low speech-to-noise ratios, these methods work more ruggedly and are primarily used to suppress the transmission of speech pauses, for example, in mobile radiocommunications. However, the methods show uncertainties when the background noise itself contains speech or is similar to speech. Such segments are then classified as speech although they are perceived by a listener as disturbing background noise.
[0014] Instrumental speech quality measurement methods are usually based on the principle of signal comparison of the undisturbed reference speech signal and the disturbed signal to be assessed. Examples of this include the publications:
[0015] “A perceptual speech-quality measure based on a psychacoustic sound representation” (Beerends. J. G.: Stemerdink, J. A., J. Audio Eng. Soc. 42 (1994) 3, p. 115-123).
[0016] “Auditory distortion measure for speech coding” (Wang, S; Sekey, A.; Gersho, A.: IEEE Proc. Int. Conf. acoust., speech and signal processing (1991), p. 493-496).
[0017] Such a method is also described in the ITU-T standard P.861 currently in force: “Objective quality measurement of telephone-band speech codecs” (ITU-T Rec. P.861, Geneva 1996).
[0018] Such measurement methods are employed in so-called “test connection systems”, in which a knot, reference speech signal (source speech signal) is injected at the source, transmitted, for example, via a telephone connection, and recorded at the sink. Subsequent to recording the speech signal, its properties are compared to those of the undisturbed source speech signal to assess the speech quality of the possibly disturbed speech signal.
[0019] If the undisturbed source speech signal is available to determine the background noise during speech pauses, then this signal can be used to determine the transition moments from speech to speech pause or from speech pause to speech, respectively. To this end, for example, a method with threshold value determination, as described above, is applied to the source speech signal. The method provides reliable distinctions between speech and speech pause because the speech-to-noise ratio in the undisturbed source speech signal is sufficiently high (FIG. 3<i>a</i>). The moments of threshold passage, that is, beginning and end of speech activity can now be transferred to the disturbed speech signal (FIG. 3<i>b</i>).
[0020] Such a method can be modified without problems if a constant time lag (for example, a delay due to signal transmission) occurs between the source speech signal and the disturbed signal. However, the condition is that this time lag can be reliably determined in advance and that it is then used to correct the end or beginning points of speech activity. This is mostly possible in the case of time-invariant systems because these have a constant delay (FIG. 3<i>c</i>)
[0021] In principle such a method works also if the time offset between the two signals is not constant for the entire signal length but is variable. These time-invariant systems include, in particular, packet-based transmission systems where marked fluctuations in the system delay can occur due to different packet transit times and a corresponding starting points management in the receiver. To prevent losses due to packets that arrive late, sometimes speech pauses are extended and later ones are shortened in the receiver. Starting or end points of speech activity can then only be transmitted if the current delay at these points is known. The adaptive determination of the time offset is computing-time intensive and frequently only inadequately achieved, especially in the case of reduced speech-to-noise ratios. If the adaptive determination of the time offset is not achieved reliably then the beginning and the end of speech pauses cannot be determined exactly or not at all. Because of this, the intensity characteristics of noise during pauses cannot or only unreliably be determined.
OBJECTIVE
[0022] As described, it is difficult or sometimes impossible to determine background noise during speech pauses even if the undisturbed source speech signal is known, especially when
[0023] a low speech-to-background noise ratio exists,
[0024] the background noise contains speech or is similar to speech itself,
[0025] the time offset between the undisturbed source speech signal and the disturbed speech signal is not constant over the entire signal length.
[0026] The intention is to present a method which ensures reliable and rapid determination of intensity characteristics of the background noise during speech pauses even under the conditions mentioned. The condition is that both the source speech signal and the disturbed speech signal are available completely recorded.
PRINCIPLE OF SOLUTION
[0027] The known methods are based on determining the starting and end points of a speech pause as accurately as possible. As a result, the signal of the pause segments is then available for further evaluation. The intensity characteristics are determined from these separated pause segments
[0028] Using the present method, intensity characteristics of background noise during speech pauses can be determined without having to determine the exact starting or end points of a pause segment. Moreover, it is not necessary to separate the speech pause signal for the evaluation.
[0029] The method for determining intensity characteristics of background noise during speech pauses of speech signals here described is based on the cumulative frequency distribution of the intensity values of the signal segments into which the speech signal is previously divided. These short-time signal intensities refer to signal segments having a duration of, for example. 8 ms or 16 ms. The frequency distribution indicates the magnitude of the fraction of short-time intensities below a defined threshold value.
[0030] To calculate the frequency distribution, the speech signal to be analyzed is divided into short successive signal segments and the intensity value (for example, loudness or effective value) is determined for each signal segment.
[0031]FIG. 4 shows a typical curve shape for speech signals containing stationary background noise (speech-to-noise ratio: approximately 10 dB). The cumulative frequency distribution is depicted by the example of short-time loudnesses (loudnesses calculated in accordance with ISO532). 2000 segments having a length of 16 ms were evaluated. It can be seen that none of the segments has a lower value than 30 sone (P=0%) and none of the segments reaches a higher value than 80 sone either since here the value P=100% is already reached. The steep rise of the function at about 30 sone suggests a low fluctuation of the signal intensity over large ranges (almost 70%) of the signal. The signal used here was a speech signal with additive white noise.
[0032] Such a distribution function is now intended to be used to determine intensity characteristics of background noise during the speech pauses. To this end, it is necessary to know the proportion of speech pauses in the overall signal. This proportion can be determined from the undisturbed source speech signal (FIG. 3<i>a</i>).
Total length of the speech pauses=(<i>t</i>1<i>−t</i>0)+(<i>t</i>3<i>−t</i>2)
Total length of the signal segment=(<i>t</i>4<i>−t</i>0) <maths id="MATH-US-00001" num="1"><math overflow="scroll"><mrow><mrow><mi>Proportion</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>of</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>speech</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>pauses</mi></mrow><mo>=</mo><mfrac><mrow><mi>total</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>length</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>of</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>the</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>speech</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>pauses</mi></mrow><mrow><mi>total</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>length</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>of</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>the</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>signal</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>segment</mi></mrow></mfrac></mrow></math><img file="US20030191633A1-20031009-M00001.TIF" id="EMI-M00001" he="18.96615" wi="216.027" img-format="tif" img-content="mf" /><attachments><attachment idref="MATHEMATICA-00001" attachment-type="nb" file="US20030191633A1-20031009-M00001.NB" /></attachments></maths>
[0033] When assuming that the ratio of active speech to speech pauses remains substantially constant during the transmission, this value can also be applied to the disturbed signal.
[0034] If the proportion of speech pauses of the overall speech signal is known and if this proportion is defined as the frequency threshold, then the intensity threshold value which corresponds to the frequency threshold can be determined from the frequency distribution of the short-time intensities.
[0035] In FIG. 4, a proportion of speech pauses of 58% is plotted as an example. This frequency threshold P<sub>z</sub>=0.58 corresponds to an intensity threshold value of N=34.5 sone, which means that 58% of the signal segments do not exceed the intensity value (loudness) of 34.5 sone.
[0036] The region below the intensity threshold value shows the frequency distribution for intensity values of signal segments during the speech pauses and can be used to determine intensity characteristics of the background noise during the speech pauses.
[0037] It is assumed that no speech pause segment has a higher intensity value than a speech segment so that the intensity threshold value can be regarded as the maximum value for the background noise during speech pauses.
DETERMINATION OF THE ARITHMETIC MEAN OF INTENSITIES
[0038] The arithmetic mean of all segments whose intensities are below a previously determined frequency threshold can also be derived from the cumulative distribution function. To this end, initially, the cumulative distribution function P(x) has to be differentiated to a distribution density function p(x).
[0039] The arithmetic mean of all evaluated intensities X of the overall signal is calculated in known manner from the integral of the distribution density function p(x): <maths id="MATH-US-00002" num="2"><math overflow="scroll"><mrow><mover><mi>X</mi><mi>_</mi></mover><mo>=</mo><mrow><msubsup><mo>∫</mo><mrow><mo>-</mo><mi>∞</mi></mrow><mi>∞</mi></msubsup><mo></mo><mrow><mi>x</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><mi>p</mi><mo></mo><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow></mrow><mo></mo><mrow><mo></mo><mi>x</mi></mrow></mrow></mrow></mrow></math><img file="US20030191633A1-20031009-M00002.TIF" id="EMI-M00002" he="18.96615" wi="216.027" img-format="tif" img-content="mf" /><attachments><attachment idref="MATHEMATICA-00002" attachment-type="nb" file="US20030191633A1-20031009-M00002.NB" /></attachments></maths>
[0040] By limiting the integration at a certain value x<sub>G</sub>, it becomes possible to determine the arithmetic mean over all values X below this limiting value. In this context, however, the result has to be weighted with frequency P(x<sub>G</sub>). This frequency corresponds to the integral over p(x) up to value x<sub>G</sub>.
[0041] Intensity threshold value x<sub>G </sub>can be derived from distribution function P(x). In the example according to FIG. 4, frequency threshold value P(x<sub>G</sub>) is the proportion of speech pauses in overall signal P<sub>z</sub>=0.58 with which is associated the intensity threshold value x<sub>G</sub>=34.5 sone. The arithmetic mean of all segments having ant intensity which is smaller than x<sub>G </sub>is calculated according to equation 2, where x<sub>G</sub>=34.5 sone. Here, the frequency of 58% corresponds to the weighting value P(x<sub>G</sub>=34.5)=0.58. This procedure is graphically shown in FIG. 5. <maths id="MATH-US-00003" num="3"><math overflow="scroll"><mrow><mover><mi>X</mi><mi>_</mi></mover><mo>=</mo><mrow><mrow><msubsup><mo>∫</mo><mrow><mo>-</mo><mi>∞</mi></mrow><msub><mi>x</mi><mi>G</mi></msub></msubsup><mo></mo><mrow><mi>x</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><mi>p</mi><mo></mo><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow></mrow><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><mrow><mo></mo><mi>x</mi></mrow><mo>/</mo><mrow><msubsup><mo>∫</mo><mrow><mo>-</mo><mi>∞</mi></mrow><msub><mi>x</mi><mi>G</mi></msub></msubsup><mo></mo><mrow><mrow><mi>p</mi><mo></mo><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow></mrow><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><mo></mo><mi>x</mi></mrow></mrow></mrow></mrow></mrow></mrow><mo>=</mo><mrow><msubsup><mo>∫</mo><mrow><mo>-</mo><mi>∞</mi></mrow><msub><mi>x</mi><mi>G</mi></msub></msubsup><mo></mo><mrow><mi>x</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><mi>p</mi><mo></mo><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow></mrow><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><mrow><mo></mo><mi>x</mi></mrow><mo>/</mo><mrow><mi>P</mi><mo></mo><mrow><mo>(</mo><msub><mi>x</mi><mi>G</mi></msub><mo>)</mo></mrow></mrow></mrow></mrow></mrow></mrow></mrow></math><img file="US20030191633A1-20031009-M00003.TIF" id="EMI-M00003" he="18.96615" wi="216.027" img-format="tif" img-content="mf" /><attachments><attachment idref="MATHEMATICA-00003" attachment-type="nb" file="US20030191633A1-20031009-M00003.NB" /></attachments></maths>
[0042] If now, again, it is assumed that the intensities of segments during speech pauses do not exceed the intensities of speech segments or that the background noise has only weak temporal fluctuations, the calculated arithmetic mean can be regarded as the mean of the intensities during speech pauses.
SIMPLIFIED METHOD FOR DETERMINING THE ARITHMETIC MEAN
[0043] A simplified method for determining the mean over all X starts from the assumption that the relative frequency distribution of the intensity values of the signal segments in the region p(x)=0 up to the frequency threshold value of speech pauses P<sub>z </sub>can be approximated by a weighted normal distribution G(x, μ, σ). The value for the distribution function G(x, μ, σ) for x →∞ is 1. As is known, value x for which G(x, μ, σ)=0.5 corresponds to the arithmetic mean over all individual values X.
[0044] If an approximation of relative frequency distribution P(x) in the region of P(x)=0 to P<sub>z </sub>is achieved with a weighted normal distribution κP<sub>z </sub>G(x, μ, σ), then the arithmetic mean over X for the weighted normal distribution corresponds to value x for which G(x, μ, σ)=0.5 κP<sub>z</sub>. Due to the assumption that κP<sub>z </sub>G(x, μ, σ) approximates distribution P(x) in the region of P(x)=0 to P<sub>z </sub>to a good degree and κ≧1, the arithmetic mean sought corresponds to value x<sub>A </sub>for which P(x<sub>A</sub>)=0.5 κP<sub>z</sub>.
[0045] For the application case of speech with additive background noise observed here, values for κ=1 . . . 1.3 show good approximation results. An example of the approximation through weighted normal distributions is shown in FIG. 6. In this context, a value κ=1.1 was selected. The diagram shows speech as background nose and features a proportion of speech pauses of 58%. The strong temporal fluctuation of the speech background can be clearly seen as a flat gradient in the region N=0 . . . 40 sone. The arithmetic mean derived from the normal distribution function with P(x<sub>A</sub>)=0.5 κP<sub>z</sub>=0.32 is 20 sone.
[0046] The advantage of this simplified method is the smaller computing intensity because the calculation of the distribution density and the integration thereof can be dispensed with. Likewise, it is not necessary to accurately determine the normal distribution function κP<sub>z </sub>G(x, μ, σ), it is already sufficient to define κ. Since P<sub>z </sub>is known, the mean is determined over all X<x<sub>G </sub>as a value x<sub>A </sub>for which P(x<sub>A</sub>)=0.5 κP<sub>z</sub>. Thus, the arithmetic mean over all X up to x<sub>G </sub>corresponds to the intensity value that corresponds to a frequency value of 0.5 *κ* proportion of the speech pauses of the overall signal, that is, the intensity which is not exceeded by a proportion of segments of 0.5 *κ* proportions of the speech pauses.
DETERMINATION OF FURTHER STATISTICAL CHARACTERISTICS
[0047] Using this method, other statistical intensity characteristics can be determined as well. In FIG. 7, it is demonstrated by the example from FIG. 4, how the intensity value which is only exceeded by 20% of the speech pause segments (20% percentile loudness) can be determined from the function.
[0048] In the given example, the intensity value is sought which is not reached by 80% of the segments during speech pauses, that is, the abscissa value is sought which applies to ordinate value P=0.58 * 0.8=0.46. Due to the low-fluctuation disturbing noise selected in the example, the value is only slightly smaller than the maximum value.
EXEMPLARY EMBODIMENT OF THE DETERMINATION OF THE ARITHMETIC MEAN FOR THE DISTRIBUTION DENSITY FUNCTION
[0049] The exemplary embodiment of the method or determining the intensity of background noise presented here determines the arithmetic mean of all loudnesses of the segments below a certain frequency threshold. This frequency threshold corresponds to the proportion of speech pauses in the signal, and the calculated arithmetic mean is regarded as the mean loudness during speech pauses. In this exemplary embodiment, the distribution density function is used for that purpose.
[0050] The prerequisite is that both signals, i.e., the undisturbed source speech signal and the disturbed signal to be assessed are available completely recorded.
[0051] Initially, the proportion of speech pauses P<sub>z </sub>in this signal is determined on the basis the source speech signal using a suitable threshold.
[0052] The second step is the calculation of the desired intensity values for successive short signal segments of the speech signal to be assessed. In this exemplary embodiment, the loudnesses are calculated according to ISO532 in successive signal segments having a length of 16 ms. The distribution function is approximated by a series of single values (discrete relative frequency distribution). These single values are denoted by successive indices m. The series of single values is limited at a maximum value M (for example: P<sub>0 </sub>. . . P<sub>200</sub>). During evaluation, each single value P<sub>m </sub>whose index exceeds the determined intensity X of the evaluated signal segment is increased by the numerator 1. Upon evaluation of the entire signal, all single values are divided by the number of all evaluated signal segments. Then, each single value P<sub>m </sub>contains the relative frequency of the signal segments that have a loudness which is smaller than the value of the index.
[0053] On the basis of the previously determined proportion of speech pauses P<sub>z</sub>, the frequency value P<sub>s </sub>is determined which has the smallest absolute difference from P<sub>z</sub>. Index S of this single value P<sub>s </sub>indicates the corresponding loudness, that is, the loudness which is not exceeded by a proportion P<sub>s </sub>of all segments. Next, to determine the arithmetic mean of the loudnesses of all segments whose loudnesses are below the predetermined frequency threshold P<sub>s</sub>, the discrete frequency distribution P<sub>0 </sub>. . . P<sub>M </sub>has to be converted to a discrete frequency density (strip frequency) P<sub>0 </sub>. . . P<sub>M</sub><sub><sup>−</sup></sub><sub>1</sub>. To this end, the differences of two successive single values are generated and stored as set of values P<sub>0 </sub>. . . P<sub>N</sub><sub><sup>−</sup></sub><sub>1</sub>.
P<sub>m</sub>=p<sub>m+1</sub>−p<sub>m </sub>for all m=0 . . . M−1
[0054] Value p<sub>m </sub>the contains the relative frequency of the segments whose loudness is between m and m−1. The arithmetic mean sought corresponds to the weighted sum over the strip frequency P<sub>m </sub>up to m=S, that is, to the loudness which is not exceeded by a proportion P<sub>s </sub>of all segments: <maths id="MATH-US-00004" num="4"><math overflow="scroll"><mrow><msub><mover><mi>N</mi><mo>~</mo></mover><mi>av</mi></msub><mo>=</mo><mrow><mrow><munderover><mo>∑</mo><mrow><mi>m</mi><mo>=</mo><mn>0</mn></mrow><mi>S</mi></munderover><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><mrow><mo>(</mo><mrow><mi>m</mi><mo>+</mo><mfrac><mn>1</mn><mn>2</mn></mfrac></mrow><mo>)</mo></mrow><mo></mo><mrow><msub><mi>p</mi><mi>m</mi></msub><mo>/</mo><mrow><munderover><mo>∑</mo><mrow><mi>m</mi><mo>=</mo><mn>0</mn></mrow><mi>S</mi></munderover><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><msub><mi>p</mi><mi>m</mi></msub></mrow></mrow></mrow></mrow><mo>=</mo><mrow><munderover><mo>∑</mo><mrow><mi>m</mi><mo>=</mo><mn>0</mn></mrow><mi>S</mi></munderover><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><mrow><mo>(</mo><mrow><mi>m</mi><mo>+</mo><mfrac><mn>1</mn><mn>2</mn></mfrac></mrow><mo>)</mo></mrow><mo></mo><mrow><msub><mi>p</mi><mi>m</mi></msub><mo>/</mo><msub><mi>P</mi><mi>s</mi></msub></mrow></mrow></mrow></mrow></mrow></math><img file="US20030191633A1-20031009-M00004.TIF" id="EMI-M00004" he="29.9943" wi="216.027" img-format="tif" img-content="mf" /><attachments><attachment idref="MATHEMATICA-00004" attachment-type="nb" file="US20030191633A1-20031009-M00004.NB" /></attachments></maths>
[0055] The correction value ½ corresponds to half the distance of two successive indices. Value p<sub>m </sub>contains the relative frequency of segments whose loudnesses are between m and m+1. Assuming uniform distribution of the loudnesses from m . . . m−1, the expected value of all loudnesses determined here is therefore m+0.5.
[0056] As described in the application case, the method yields a discrete frequency distribution with a resolution of 1 sone since index m is integral and the loudness values are directly associated with the corresponding indices. To achieve other, higher or reduced resolutions if desired, the loudness value has to be multiplied by corresponding factors prior to calculating the relative frequency distribution.
[0057] To demonstrate the measuring accuracy of the presented method, measured values for different signals and background noises are listed in Table 1. Speech signals having a length of 32 s and different proportions of speech pauses (35%, 58% and 91%) were each mixed with different noises. Initially, white noise having different speech-to-noise ratios was used as noise. Moreover, continuously spoken speech and two noises from real acoustic environments (street and office) were used.
[0058] Prior to calculating the frequency distribution, all loudness values are multiplied by a factor 2 to increase the resolution of the representation when using integral indices. This then corresponds to a loudness grading of 0.5 sone integral indices. With the frequency distribution function being limited at P<sub>200</sub>, it is thus possible to image loudnesses of 0 . . . 100 sone in steps of 0.5 sone. However, it should be observed that this factor is must be applied to all results as a divisor for correction. In the exemplary embodiment selected here, this means that the calculate arithmetic mean has to be divided by 2.
[0059] Explanations on Table 1: The speech-to-noise ratio serves only for information purposes; the basis is formed by the distance of the mean effective level during speech activity from the mean effective level of the background noise. The mean loudness value (target value) was determined in a reference measurement in which the speech pauses were manually marked and evaluated in segments of 16 ms. The calculated standard deviations refer to the reference loudnesses measured in this manner and provide information on the magnitude of the occurring fluctuations. The measured values in column 5 were determined using the method described in this exemplary embodiment. <tables id="TABLE-US-00001" num="1"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="42PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="center" /><colspec colname="5" colwidth="35PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><thead><row><entry namest="1" nameend="6" align="center">TABLE 1</entry></row><row><entry /></row><row><entry namest="1" nameend="6" align="center" rowsep="1" /></row><row><entry /><entry /><entry /><entry /><entry>Mean</entry><entry /></row><row><entry /><entry /><entry /><entry /><entry>loudness</entry></row><row><entry /><entry /><entry>Mean</entry><entry>Standard</entry><entry>(sone)</entry></row><row><entry /><entry /><entry>loudness</entry><entry>deviation</entry><entry>measured</entry><entry>Deviation</entry></row><row><entry /><entry /><entry>(sone)</entry><entry>of the</entry><entry>with the</entry><entry>(measuring</entry></row><row><entry /><entry /><entry>target</entry><entry>segment</entry><entry>described</entry><entry>error)</entry></row><row><entry>Noise</entry><entry>SNR</entry><entry>value</entry><entry>loudnesses</entry><entry>method</entry><entry>abs./rel.</entry></row><row><entry namest="1" nameend="6" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><tbody valign="top"><row><entry>Proportion of pauses of the speech signal 91%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="42PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="char" char="." /><colspec colname="5" colwidth="35PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><tbody valign="top"><row><entry>White noise</entry><entry> 6 dB</entry><entry>41.4</entry><entry>1.55</entry><entry>42.0</entry><entry>0.6/1.4%</entry></row><row><entry>White noise</entry><entry>10 dB</entry><entry>32.3</entry><entry>1.22</entry><entry>32.6</entry><entry>0.3/0.9%</entry></row><row><entry>White noise</entry><entry>16 dB</entry><entry>22.2</entry><entry>0.87</entry><entry>22.3</entry><entry>0.1/0.4%</entry></row><row><entry>Speech</entry><entry> 6 dB</entry><entry>21.3</entry><entry>11.7</entry><entry>20.6</entry><entry>−0.7/−3.3%</entry></row><row><entry>Speech</entry><entry>10 dB</entry><entry>16.5</entry><entry>9.16</entry><entry>16.2</entry><entry>−0.3/−1.8%</entry></row><row><entry>Speech</entry><entry>16 dB</entry><entry>11.2</entry><entry>6.21</entry><entry>11.3</entry><entry>0.1/0.9%</entry></row><row><entry>Street noise</entry><entry>10 dB</entry><entry>26.0</entry><entry>3.22</entry><entry>26.2</entry><entry>0.2/0.8%</entry></row><row><entry>Office noise</entry><entry>10 dB</entry><entry>26.3</entry><entry>2.78</entry><entry>26.6</entry><entry>0.3/1.1%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><tbody valign="top"><row><entry>Proportion of pauses of the speech signal: 58%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="42PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="char" char="." /><colspec colname="5" colwidth="35PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><tbody valign="top"><row><entry>White noise</entry><entry> 6 dB</entry><entry>41.3</entry><entry>1.55</entry><entry>44.8</entry><entry>3.5/8.5%</entry></row><row><entry>White noise</entry><entry>10 dB</entry><entry>32.3</entry><entry>1.22</entry><entry>34.2</entry><entry>1.9/6.0%</entry></row><row><entry>White noise</entry><entry>16 dB</entry><entry>22.1</entry><entry>0.87</entry><entry>22.6</entry><entry>0.5/2.2%</entry></row><row><entry>Speech</entry><entry> 6 dB</entry><entry>20.7</entry><entry>11.7</entry><entry>19.0</entry><entry>−1.7/−8.2%</entry></row><row><entry>Speech</entry><entry>10 dB</entry><entry>16.0</entry><entry>9.16</entry><entry>15.4</entry><entry>−0.6/−3.8%</entry></row><row><entry>Speech</entry><entry>16 dB</entry><entry>10.7</entry><entry>6.21</entry><entry>10.8</entry><entry>0.1/0.9%</entry></row><row><entry>Street noise</entry><entry>10 dB</entry><entry>26.1</entry><entry>3.22</entry><entry>27.0</entry><entry>0.9/3.4%</entry></row><row><entry>Office noise</entry><entry>10 dB</entry><entry>26.3</entry><entry>2.78</entry><entry>27.3</entry><entry>1.0/3.8%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><tbody valign="top"><row><entry>Proportion of the speech signal 35%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="42PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="char" char="." /><colspec colname="5" colwidth="35PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><tbody valign="top"><row><entry>White noise</entry><entry> 6 dB</entry><entry>41.3</entry><entry>1.55</entry><entry>46.1</entry><entry> 4.8/11.6%</entry></row><row><entry>White noise</entry><entry>10 dB</entry><entry>32.3</entry><entry>1.22</entry><entry>35.6</entry><entry> 3.3/10.2%</entry></row><row><entry>White noise</entry><entry>16 dB</entry><entry>22.1</entry><entry>0.87</entry><entry>23.3</entry><entry>1.2/5.4%</entry></row><row><entry>Speech</entry><entry> 6 dB</entry><entry>20.0</entry><entry>11.22</entry><entry>17.6</entry><entry>−2.4/−12% </entry></row><row><entry>Speech</entry><entry>10 dB</entry><entry>15.6</entry><entry>8.7</entry><entry>15.0</entry><entry>−0.6−3.8%</entry></row><row><entry>Speech</entry><entry>16 dB</entry><entry>10.9</entry><entry>5.93</entry><entry>11.8</entry><entry>0.9/8.3%</entry></row><row><entry>Street noise</entry><entry>10 dB</entry><entry>26.1</entry><entry>3.22</entry><entry>27.3</entry><entry>1.2/4.6%</entry></row><row><entry>Office noise</entry><entry>10 dB</entry><entry>26.3</entry><entry>2.78</entry><entry>27.9</entry><entry>1.6/6.1%</entry></row><row><entry namest="1" nameend="6" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
[0060] First of all, it can be established that the measuring accuracy increases as the proportion of pauses in the signal to be assessed increases. An increase in measuring accuracy can also be established in the case of a decrease in the noise intensity or a reduced temporal fluctuation of the background noise. Starting from a typical proportion of speech pauses in a telephone communication of P<sub>z</sub>>50%. the measured values achieved by the presented method are satisfactory even in the case of stronger fluctuations in the background noise (for example, speech).
EXEMPLARY EMBODIMENT OF THE DETERMINATION OF THE ARITHMETIC MEAN A SIMPLIFIED METHOD
[0061] This particular exemplary embodiment shows an application of the described simplified method for determining the arithmetic mean, using a weighted normal distribution.
[0062] The simplified method dispenses with the calculation of the strip frequency and derives an estimate for the arithmetic mean of the loudnesses of all segments whose loudnesses are below predetermined frequency threshold P<sub>z </sub>directly from relative frequency distribution P<sub>m</sub>. As described, only value k has to be defined for the estimation.
[0063] In this exemplary embodiment, the definition is done with k=1.1. The estimate then corresponds to the loudness value which is not exceeded by a proportion of 0.5 *1.1 * P<sub>z </sub>of all evaluated segments. In the exemplary embodiment, this estimate of the arithmetic mean of the loudnesses corresponds to the index m of the frequency value which has the lowest absolute difference from 0.55 P<sub>z</sub>. The measured values which have been obtained by this simplified method are listed in Table 2. Here too, all loudness values were multiplied by a factor 2 and the results were corrected accordingly to increase the resolution to 0.5 sone. <tables id="TABLE-US-00002" num="2"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="42PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="center" /><colspec colname="5" colwidth="35PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><thead><row><entry namest="1" nameend="6" align="center">TABLE 2</entry></row><row><entry /></row><row><entry namest="1" nameend="6" align="center" rowsep="1" /></row><row><entry /><entry /><entry /><entry /><entry>Mean</entry><entry /></row><row><entry /><entry /><entry /><entry /><entry>loudness</entry></row><row><entry /><entry /><entry>Mean</entry><entry>Standard</entry><entry>(sone)</entry></row><row><entry /><entry /><entry>loudness</entry><entry>deviation</entry><entry>measured</entry><entry>Deviation</entry></row><row><entry /><entry /><entry>(sone)</entry><entry>of the</entry><entry>with the</entry><entry>(measuring</entry></row><row><entry /><entry /><entry>target</entry><entry>segment</entry><entry>simplified</entry><entry>error)</entry></row><row><entry>Noise</entry><entry>SNR</entry><entry>value</entry><entry>loudnesses</entry><entry>method</entry><entry>abs./rel.</entry></row><row><entry namest="1" nameend="6" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><tbody valign="top"><row><entry>Proportion of pauses of the speech signal 91%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="42PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="char" char="." /><colspec colname="5" colwidth="35PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><tbody valign="top"><row><entry>White noise</entry><entry> 6 dB</entry><entry>41.4</entry><entry>1.55</entry><entry>41.5</entry><entry>0.1/0.2%</entry></row><row><entry>White noise</entry><entry>10 dB</entry><entry>32.3</entry><entry>1.22</entry><entry>32.5</entry><entry>0.2/0.6%</entry></row><row><entry>White noise</entry><entry>16 dB</entry><entry>22.2</entry><entry>0.87</entry><entry>22.5</entry><entry>0.3/1.3%</entry></row><row><entry>Speech</entry><entry> 6 dB</entry><entry>21.3</entry><entry>11.7</entry><entry>20.5</entry><entry>−0.8/−3.8%</entry></row><row><entry>Speech</entry><entry>10 dB</entry><entry>16.5</entry><entry>9.16</entry><entry>16.5</entry><entry>0.0/0.0%</entry></row><row><entry>Speech</entry><entry>16 dB</entry><entry>11.2</entry><entry>6.21</entry><entry>11.0</entry><entry>−0.2/1.8% </entry></row><row><entry>Street noise</entry><entry>10 dB</entry><entry>26.0</entry><entry>3.22</entry><entry>26.0</entry><entry>0.0/0.0%</entry></row><row><entry>Office noise</entry><entry>10 dB</entry><entry>26.3</entry><entry>2.78</entry><entry>26.5</entry><entry>0.2/0.6%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><tbody valign="top"><row><entry>Proportion of pauses of the speech signal 58%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="42PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="char" char="." /><colspec colname="5" colwidth="35PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><tbody valign="top"><row><entry>White noise</entry><entry> 6 dB</entry><entry> 41.30</entry><entry>1.55</entry><entry>41.5</entry><entry>0.2/0.5%</entry></row><row><entry>White noise</entry><entry>10 dB</entry><entry>32.3</entry><entry>1.22</entry><entry>32.5</entry><entry>0.2/0.6%</entry></row><row><entry>White noise</entry><entry>16 dB</entry><entry>22.1</entry><entry>0.87</entry><entry>22.5</entry><entry>0.4/1.8%</entry></row><row><entry>Speech</entry><entry> 6 dB</entry><entry>20.7</entry><entry>11.7</entry><entry>20.0</entry><entry>−0.7/−3.4%</entry></row><row><entry>Speech</entry><entry>10 dB</entry><entry>16.0</entry><entry>9.16</entry><entry>16.0</entry><entry>0.0/0.0%</entry></row><row><entry>Speech</entry><entry>16 dB</entry><entry>10.7</entry><entry>6.21</entry><entry>11.0</entry><entry>0.3/2.8%</entry></row><row><entry>Street noise</entry><entry>10 dB</entry><entry>26.1</entry><entry>3.22</entry><entry>26.0</entry><entry>−0.1/−0.4%</entry></row><row><entry>Office noise</entry><entry>10 dB</entry><entry>26.3</entry><entry>2.78</entry><entry>26.5</entry><entry>0.2/0.8%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><tbody valign="top"><row><entry>Proportion of pauses of the speech signal 35%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="42PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="char" char="." /><colspec colname="5" colwidth="35PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><tbody valign="top"><row><entry>White noise</entry><entry> 6 dB</entry><entry>41.3</entry><entry>1.55</entry><entry>41.0</entry><entry>−0.3/0.7% </entry></row><row><entry>White noise</entry><entry>10 dB</entry><entry>32.3</entry><entry>1.22</entry><entry>32.5</entry><entry>0.2/0.6%</entry></row><row><entry>White noise</entry><entry>16 dB</entry><entry>22.1</entry><entry>0.87</entry><entry>22.5</entry><entry>0.4/1.8%</entry></row><row><entry>Speech</entry><entry> 6 dB</entry><entry>20.0</entry><entry>11.22</entry><entry>19.0</entry><entry>−1.0/−5% </entry></row><row><entry>Speech</entry><entry>10 dB</entry><entry>15.6</entry><entry>8.7</entry><entry>15.5</entry><entry>−0.1/−0.6%</entry></row><row><entry>Speech</entry><entry>16 dB</entry><entry>10.9</entry><entry>5.93</entry><entry>11.5</entry><entry>0.6/5.5%</entry></row><row><entry>Street noise</entry><entry>10 dB</entry><entry>26.1</entry><entry>3.22</entry><entry>25.5</entry><entry>−0.6/−1.4%</entry></row><row><entry>Office noise</entry><entry>10 dB</entry><entry>26.3</entry><entry>2.78</entry><entry>26.5</entry><entry>0.2/0.8%</entry></row><row><entry namest="1" nameend="6" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
[0064] The simplified method not only saves computing time, but also yields measured values with a markedly higher accuracy in the evaluated examples compared to the values from Table 1. Since index m is directly used as the estimate, the accuracy of the estimation is limited to the resolution of the relative discrete frequency distribution (here: 0.5 sone).
[0065] Using the simplified measurement method described, good measured values are attained even in the case of noises with stronger fluctuation. For the selected speech-to-noise ratios of 6 dB, moreover, it can no longer be assumed that all loudnesses during speech pauses have a smaller loudness than speech segments. Nevertheless, the measured values were hardly corrupted. The simplified method described is also suitable for signals having a smaller proportion of pauses.
EXEMPLARY EMBODIMENT OF THE DETERMINATION OF PERCENTILE LOUDNESSES FROM THE RELATIVE FREQUENCY DISTRIBUTION
[0066] The percentile loudness of all segments below a certain frequency threshold P<sub>z</sub>, can be carried out by multiplying this relative frequency P<sub>z </sub>by a value 1-percentile value (for example, 10% percentile loudness: P<sub>z10%</sub>=0.9 * P<sub>z</sub>). The integral index m of frequency value P<sub>m </sub>value which has the lowest absolute difference from P<sub>S10% </sub>yields the percentile loudness value sought.
[0067] The 10% percentile loudnesses for the examples already listed in Tables 1 and 2 are given in Table 3 and compared to a manually determined reference value. <tables id="TABLE-US-00003" num="3"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="28PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="center" /><colspec colname="5" colwidth="49PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><thead><row><entry namest="1" nameend="6" align="center">TABLE 3</entry></row><row><entry /></row><row><entry namest="1" nameend="6" align="center" rowsep="1" /></row><row><entry /><entry /><entry>10%</entry><entry /><entry /><entry /></row><row><entry /><entry /><entry>percentile</entry><entry>Standard</entry><entry>10% percentile</entry><entry /></row><row><entry /><entry /><entry>loudness</entry><entry>deviation</entry><entry>loudness (sone)</entry><entry>Deviation</entry></row><row><entry /><entry /><entry>(sone)</entry><entry>of the</entry><entry>measured</entry><entry>(measuring</entry></row><row><entry /><entry /><entry>target</entry><entry>segment</entry><entry>over frequency</entry><entry>error)</entry></row><row><entry>Noise</entry><entry>SNR</entry><entry>value</entry><entry>loudnesses</entry><entry>distribution</entry><entry>abs./rel.</entry></row><row><entry namest="1" nameend="6" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><tbody valign="top"><row><entry>Proportion of pauses of the speech signal 91%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="28PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="char" char="." /><colspec colname="5" colwidth="49PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><tbody valign="top"><row><entry>White</entry><entry> 6 dB</entry><entry>42.5</entry><entry>1.55</entry><entry>43.0</entry><entry> 0.5/1.2%</entry></row><row><entry>noise</entry></row><row><entry>White</entry><entry>10 dB</entry><entry>33.0</entry><entry>1.22</entry><entry>34.0</entry><entry> 1.0/3.0%</entry></row><row><entry>noise</entry></row><row><entry>White</entry><entry>16 dB</entry><entry>22.5</entry><entry>0.87</entry><entry>23.5</entry><entry> 1.0/4.4%</entry></row><row><entry>noise</entry></row><row><entry>Speech</entry><entry> 6 dB</entry><entry>37.0</entry><entry>11.7</entry><entry>34.5</entry><entry> −2.5/−6.8%</entry></row><row><entry>Speech</entry><entry>10 dB</entry><entry>28.5</entry><entry>9.16</entry><entry>27.5</entry><entry> −1.0/−3.5%</entry></row><row><entry>Speech</entry><entry>16 dB</entry><entry>19.0</entry><entry>6.21</entry><entry>19.5</entry><entry> 0.5/2.6%</entry></row><row><entry>Street</entry><entry>10 dB</entry><entry>29.5</entry><entry>3.22</entry><entry>30.0</entry><entry> 0.5/1.7%</entry></row><row><entry>noise</entry></row><row><entry>Office</entry><entry>10 dB</entry><entry>29.0</entry><entry>2.78</entry><entry>29.5</entry><entry> 0.5/1.7%</entry></row><row><entry>noise</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><tbody valign="top"><row><entry>Proportion of pauses of the speech signal 58%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="28PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="char" char="." /><colspec colname="5" colwidth="49PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><tbody valign="top"><row><entry>White</entry><entry> 6 dB</entry><entry>42.5</entry><entry>1.55</entry><entry>42.5</entry><entry> 0.0/0.0%</entry></row><row><entry>noise</entry></row><row><entry>White</entry><entry>10 dB</entry><entry>33.0</entry><entry>1.22</entry><entry>33.5</entry><entry> 0.5/1.5%</entry></row><row><entry>noise</entry></row><row><entry>White</entry><entry>16 dB</entry><entry>22.5</entry><entry>0.87</entry><entry>23.0</entry><entry> 0.5/2.2%</entry></row><row><entry>noise</entry></row><row><entry>Speech</entry><entry> 6 dB</entry><entry>36.0</entry><entry>11.7</entry><entry>29.0</entry><entry> −7.0/−19% </entry></row><row><entry>Speech</entry><entry>10 dB</entry><entry>28.5</entry><entry>9.16</entry><entry>24.5</entry><entry> −4.0/−14% </entry></row><row><entry>Speech</entry><entry>16 dB</entry><entry>19.0</entry><entry>6.21</entry><entry>18.0</entry><entry> −1.0/−5.3%</entry></row><row><entry>Street</entry><entry>10 dB</entry><entry>30.0</entry><entry>3.22</entry><entry>29.0</entry><entry> −1.0/−3.3%</entry></row><row><entry>noise</entry></row><row><entry>Office</entry><entry>10 dB</entry><entry>29.0</entry><entry>2.78</entry><entry>28.5</entry><entry> −0.5/−1.6%</entry></row><row><entry>noise</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><tbody valign="top"><row><entry>Proportion of pauses of the speech signal 35%</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="6"><colspec colname="1" colwidth="28PT" align="left" /><colspec colname="2" colwidth="28PT" align="center" /><colspec colname="3" colwidth="35PT" align="center" /><colspec colname="4" colwidth="35PT" align="char" char="." /><colspec colname="5" colwidth="49PT" align="center" /><colspec colname="6" colwidth="42PT" align="center" /><tbody valign="top"><row><entry>White</entry><entry> 6 dB</entry><entry>42.5</entry><entry>1.55</entry><entry>42.5</entry><entry> 0.0/0.0%</entry></row><row><entry>noise</entry></row><row><entry>White</entry><entry>10 dB</entry><entry>33.0</entry><entry>1.22</entry><entry>33.5</entry><entry> 0.5/1.5%</entry></row><row><entry>noise</entry></row><row><entry>White</entry><entry>16 dB</entry><entry>22.5</entry><entry>0.87</entry><entry>23.5</entry><entry> 1.0/2.2%</entry></row><row><entry>noise</entry></row><row><entry>Speech</entry><entry> 6 dB</entry><entry>35.5</entry><entry>11.22</entry><entry>24.0</entry><entry>−11.5/−33% </entry></row><row><entry>Speech</entry><entry>10 dB</entry><entry>27.5</entry><entry>8.7</entry><entry>21.0</entry><entry> −6.5/−24% </entry></row><row><entry>Speech</entry><entry>16 dB</entry><entry>19.0</entry><entry>5.93</entry><entry>17.5</entry><entry> −1.5/−7.9%</entry></row><row><entry>Street</entry><entry>10 dB</entry><entry>29.5</entry><entry>3.22</entry><entry>28.0</entry><entry> −1.5/−4.8%</entry></row><row><entry>noise</entry></row><row><entry>Office</entry><entry>10 dB</entry><entry>29.0</entry><entry>2. 78</entry><entry>28.5</entry><entry> −0.5/−1.6%</entry></row><row><entry>noise</entry></row><row><entry namest="1" nameend="6" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
[0068] The measured values show a good estimation of the percentile loudness for background noises with weak fluctuation. For speech, only inadequate accuracies are attained, above all in the case of a small proportion of pauses. Only in the case of higher speech-to-noise ratios, the results are serviceable to good.
Contents8
10 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US7664733B2 | Cited by | United States of America | Search report |
| US2009180697A1 | Cited by | United States of America | Pre-grant |
| US2015156329A1 | Cited by | United States of America | Pre-grant |
| US2007288523A1 | Cited by | United States of America | Pre-grant |
| US7643705B1 | Cited by | United States of America | Applicant |
| US7616840B2 | Cited by | United States of America | Applicant |
| US2004205041A1 | Cited by | United States of America | Pre-grant |
| US8281230B2 | Cited by | United States of America | Applicant |
| US2007204229A1 | Cited by | United States of America | Pre-grant |
| US7698646B2 | Cited by | United States of America | Applicant |
| US2003156633A1 | Cites | United States of America | Pre-grant |
| US4481593A | Cites | United States of America | Pre-grant |
| US4811404A | Cites | United States of America | Pre-grant |
| US5598466A | Cites | United States of America | Pre-grant |
| US6031915A | Cites | United States of America | Pre-grant |
| US6044342A | Cites | United States of America | Pre-grant |
19 members in 8 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 10120168 | Germany | A | |
| 10120168 | Germany | A | |
| 0201200 | Germany | W | |
| 0201200 | Germany | W | |
| 101201680 | – | – | – |
| DE2001120168 | – | – | – |
| WO2002DE01200 | – | – | – |
Members19
| Document | Office | Kind | |
|---|---|---|---|
| US4792538A | United States of America | A | |
| EP0311553A1 | European Patent Office (EPO) | A1 | |
| BR8804996A | Brazil | A | |
| BR8804996A | Brazil | A | |
| JPH0238361A | Japan | A | |
| EP0311553B1 | European Patent Office (EPO) | B1 | |
| DE3881816D1 | Germany | D1 | |
| CA1321406C | Canada | C | |
| DE3881816T2 | Germany | T2 | |
| JP2525039B2 | Japan | B2 | |
| DE10120168A1 | Germany | A1 | |
| WO02084644A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2003191633A1 | United States of America | A1 | |
| EP1382034A1 | European Patent Office (EPO) | A1 | |
| EP1382034B1 | European Patent Office (EPO) | B1 | |
| AT289442T | Austria | T | |
| ATE289442T1 | Austria | T1 | |
| DE50202281D1 | Germany | D1 | |
| US7277847B2 | United States of America | B2 |
43 transactions on the USPTO file
Allowed after 1 non-final rejection and 1 final rejection.
- Non-final rejections
- 1
- Final rejections
- 1
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | |
|---|---|
| Payment of Maintenance Fee, 12th Year, Large Entity | |
| Correspondence Address Change | |
| Post Issue Communication - Certificate of Correction | |
| Recordation of Patent Grant Mailed | |
| Patent Issue Date Used in PTA CalculationAllowed | |
| Issue Notification MailedAllowed | |
| Dispatch to FDC | |
| Application Is Considered Ready for Issue | |
| Issue Fee Payment Verified | |
| Issue Fee Payment Received | |
| Mail Notice of AllowanceAllowed | |
| Notice of Allowance Data Verification CompletedAllowed | |
| Date Forwarded to Examiner | |
| Response after Final Action | |
| Mail Final Rejection (PTOL - 326)Final rejection | |
| Final RejectionFinal rejection | |
| Date Forwarded to Examiner | |
| Information Disclosure Statement considered | |
| Information Disclosure Statement (IDS) Filed | |
| Information Disclosure Statement (IDS) Filed | |
| Response after Non-Final Action | |
| Request for Extension of Time - Granted | |
| Correspondence Address Change | |
| Change in Power of Attorney (May Include Associate POA) | |
| Mail Non-Final RejectionNon-final rejection | |
| Non-Final RejectionNon-final rejection | |
| Case Docketed to Examiner in GAU | |
| Case Docketed to Examiner in GAU | |
| Case Docketed to Examiner in GAU | |
| IFW TSS Processing by Tech Center Complete | |
| Case Docketed to Examiner in GAU | |
| Case Docketed to Examiner in GAU | |
| Case Docketed to Examiner in GAU | |
| Case Docketed to Examiner in GAU | |
| Cleared by OIPE CSR | |
| Application Dispatched from OIPE | |
| Notice of DO/EO Acceptance Mailed | |
| Information Disclosure Statement considered | |
| Preliminary Amendment | |
| Reference capture on IDS | |
| Information Disclosure Statement (IDS) Filed | |
| Information Disclosure Statement (IDS) Filed | |
| Initial Exam Team nn |
6 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Maintenance fee paymentMAFP | MAFP | |
| Fee paymentFPAY | FPAY | |
| Fee paymentFPAY | FPAY | |
| Certificate of correctionCC | CC | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| AssignmentAS | AS |
Numbers
- Publication, DOCDB
- 2003191633
- Publication, EPODOC
- US2003191633
- Application
- 10311487
- Application, DOCDB
- 31148702
- Application, EPODOC
- US20020311487
Titles
- English
- Method for determining intensity parameters of background nose in speech pauses of voice signals
Patent term adjustment
- A delay
- +923 daysthe office missed an examination deadline
- Applicant delay
- −29 days
- Net adjustment
- 894 days
Classification
- CPC, 4
- G10L25/69
- G10L25/78
- G10L2021/02168
- G10L2025/786
- IPC, 5
- G10L11 02
- G10L19 00
- G10L21 0216
- G10L25 69
- G10L25 78
- USPC, 3
- 704205000
- 704E11003
- 704E19002